Industry

Turning Unstructured Content into a Reliable AI Data Foundation

AI deployments often fail to scale because unstructured content—documents, transcripts, images, and more—is not prepared, governed, and linked to structured enterprise data.

Turning Unstructured Content into a Reliable AI Data Foundation

AI can make enterprise data appear simpler than it is: contracts become summaries, call transcripts become recommended actions, policies become direct answers. Those outputs don’t emerge spontaneously—AI systems repeatedly decompose and recombine data across documents, systems, prompts, and workflows. To scale AI, organizations must treat those transformed data artifacts as authoritative, traceable, and governed; otherwise agents cannot act responsibly and users cannot trust results.

This data bottleneck helps explain why only 7 percent of companies have fully scaled AI across their organizations. Addressing unstructured data alone is insufficient: AI data readiness requires integrating structured and unstructured data into a governed, traceable, and reusable foundation.

A financial-services case: rebuilding unstructured data pipelines

A financial-services company rebuilt its unstructured-data pipelines with the same rigor traditionally applied to structured enterprise data. The initiative targeted documents, images, audio files and other inputs that needed extraction, quality checks, enrichment and preparation for AI consumption.

The company treated a source artifact as more than a single static object. For example, a PDF could yield extracted text, parsed tables, images, image summaries, metadata, sensitivity tags, quality scores and other intermediate artifacts. Those outputs had to remain linked to the original source and to one another so that meaning, lineage and control were preserved as artifacts moved through the pipeline.

The goal was not merely to load documents into a model but to make the right content discoverable, retrievable and usable by AI applications. To that end the company created curated unstructured data products accessible via text, metadata and vector search, and through APIs. Applications could retrieve the correct documents, passages, tables, images or entities before sending context to a model.

To speed adoption, the company developed a common, extensible pipeline pattern. Teams reused foundational mechanics for ingestion, extraction, quality checks, metadata, lineage, indexing and exposure, configuring only the steps required for each curated dataset. A video-focused use case required different processing than an image-and-table-heavy document use case, and teams could add rules at defined extension points rather than building a new pipeline per use case.

The result was a more repeatable way to deliver the last mile of data access for AI: business and application teams consumed governed content through standard interfaces instead of receiving sanitized data and having to determine retrieval, ranking and usage themselves.

Why AI-era data challenges differ from traditional data practices

Generative and agentic AI systems depend heavily on large volumes of unstructured content—documents, emails, call transcripts, video. Each source file can expand into multiple representations (text, tables, images), increasing both data volume and management complexity. AI systems also reuse outputs across applications, so small data issues can quickly scale into large problems. As data is transformed and recombined, it becomes harder to trace, validate and defend outputs.

Many companies try to address the problem by digitizing content and making it searchable. For AI, searchability does not equal usability. Reliable AI requires data with clear versioning, structure and context, including links between unstructured content and structured enterprise data that defines customers, products, contracts, assets, policies and transactions.

Others invest in tools—vector databases, model gateways and retrieval pipelines—but still struggle to explain outputs, trace answers to source documents or prevent sensitive information from surfacing. Tools alone, including retrieval-augmented approaches, are not a silver bullet.

Accordingly, more than two-thirds of high-performing companies identify data as the primary obstacle to enabling AI (Exhibit 1 in the original analysis).

Turning data from a roadblock into a core asset: six disciplines

AI leaders should ensure data is reliable, well-understood, traceable and reusable so outputs can be produced consistently and trusted across applications. Structure and governance must encompass structured enterprise data, unstructured artifacts, derived artifacts and the tools and skills that let models and agents use data consistently. The authors outline six disciplines CDOs should apply:

  • Observability: make data ingestion, transformation and the assembly of context observable end-to-end so teams can detect issues early, including stale content influencing answers, retrieval or orchestration failures, and drift between generated responses and source material.
  • Data quality: ensure completeness, correctness and freshness not only at ingestion but across extraction, chunking, embedding, retrieval and generation; include semantic integrity so superseded content does not contaminate outputs.
  • Metadata: use metadata as the control layer for unstructured artifacts—ownership, sensitivity, intent and permitted usage must be explicit for extracted objects as well as for source files; enterprise graphs should link unstructured artifacts to core entities (customers, contracts, products, assets, policies, transactions).
  • Lineage: capture artifact-level provenance—what document version was indexed, how it was segmented, which chunks were retrieved, how prompts were constructed and which tools were invoked—and the links between unstructured artefacts and structured records.
  • Governance: extend controls from storage to runtime so policies apply during retrieval, embedding usage, prompt assembly, memory layers and generated outputs, preventing fragments of sensitive content from surfacing even when document-level access controls exist.
  • Platform and tooling: standardize how unstructured content becomes AI-ready via reusable extraction pipelines, shared embedding and indexing infrastructure, common retrieval layers and standardized guardrails to avoid duplication and inconsistent outputs across teams.

These disciplines are complementary: observability identifies where quality breaks down; metadata and lineage provide context and traceability; governance enforces runtime policies; platformization makes practices repeatable and cost-effective.

Organizational implications: the CDO mandate and operating model

The mandate of the chief data officer (CDO) is expanding. CDOs must ensure data can be reused, traced and governed consistently wherever AI systems operate, and increasingly must link structured and unstructured data and the reusable tools and skills that enable deterministic behavior where required. Collaboration with product and engineering teams becomes essential, and talent demands shift as data, engineering, product and governance roles blend.

Traditional architectures designed for stable datasets and predictable workflows must evolve; core data disciplines should be rewired—not discarded—to support AI systems that assemble information dynamically and operate in real time.

Conclusion

Scaling AI requires treating data as a strategic enterprise asset. Organizations that establish standards, controls, observability, artifact-level lineage, semantics-aware quality practices, and reusable platforms can scale AI with greater consistency, safety and speed. Without that foundation, each step toward broader AI use risks increasing errors, exposing sensitive information and eroding trust in outputs.

Authors: Asin Tavakoli (McKinsey Düsseldorf), Brian Goodman (McKinsey New York), Kayvaun Rowshankish (senior partner, McKinsey New York), Camila Cartafina (consultant, McKinsey New York), Stephen Reddin (McKinsey Toronto), Maheshwar Muralidharan (consultant, McKinsey Boston). Edited by Kristi Essick (executive editor, Bay Area).