Why does data fragmentation prevent enterprises from scaling agentic AI?

|
Share

Enterprises have spent the last two years racing to adopt AI. They’ve signed contracts with model providers, stood up pilots, and trained employees. Yet most are still struggling to move from experimentation to enterprise-wide impact.

According to Box’s State of AI 2026 report, which surveyed 1,640 IT decision-makers across the US, UK, France, and Japan, 83% of organizations are now running AI agents. But only 19% are doing so autonomously at scale. There’s a significant gap between “running AI” and “scaling AI” so that it makes a real impact in the organization.

83% of organizations are now running AI agents. But only 19% are doing so autonomously at scale.

Enterprise organizations have a common barrier: AI agents are only as good as the content they can access, and most enterprise content is fragmented, scattered across technologies, siloed within teams, and inaccessible to those who need it most.

“A lot of challenges with AI strategies are actually data strategy challenges in disguise,” says Aaron Levie, CEO of Box. “This is why there’s such a significant premium on getting structured and unstructured data environments set up properly so agents can work with information effectively.”

Key insights:

  • Data fragmentation across disconnected legacy systems and siloed tools acts as a complete barrier to AI, causing agents to fail to find the right content, draw from outdated sources, or produce hallucinations
  • Scattered information directly hinders high-value AI workflows, including automated structured data extraction, workflow automation (due to conflicting permission models), and enterprise-wide scaling
  • Without a unified governance and permission-aware content layer, scaling AI leads to severe security risks, with 49% of organizations already experiencing an AI-related data exposure incident

Download report

What is data fragmentation and why does it matter for AI?

Data fragmentation is what happens when an organization’s information lives in dozens of disconnected systems: document sharing services, network file shares, legacy ECM platforms, FTP sites, point solutions, and cloud storage tools that were never designed to work together.

For human workers, fragmentation is an inconvenience. For AI agents, it’s an existential blocker.

AI agents need to retrieve, reason over, and act on information in real time. When that information is scattered across systems with inconsistent formats, conflicting versions, and mismatched permissions, agents can’t do their jobs. They either fail to find the right content, draw from outdated sources, or in the worst-case scenario, confidently produce incorrect answers.

More than two-thirds of respondents called legacy and on-premises systems a moderate or major barrier to AI progress

“For most enterprises that were started before a world of agents, the context that agents need for doing their work is often not in a clean format readily available to agents,” Levie explains. “Most enterprises are still entrenched with many legacy and fragmented systems. We see this all the time at Box, where companies still have large unstructured data repositories sitting inside of systems that will never play well with agents.”

The State of AI 2026 report confirms this problem. When asked about the top barriers to AI adoption, more than two-thirds of respondents called legacy and on-premises systems a moderate or major barrier to AI progress. Specifically, they struggle with:

  • Data fragmented across systems (25%)
  • Difficulty integrating AI into existing systems (24%)
  • Missing permissions and access controls (21%)
  • Content that isn’t well organized or classified (18%)
  • Poor or outdated content quality (16%)

While organizations deal with different specific AI barriers under the umbrella of enterprise content, data fragmentation is an across-the-board culprit. 

Agents that can’t access context can’t deliver value

In plain terms, 96% of organizations say AI agents need access to company-specific content to be effective. Yet only 36% have actually connected agents to trusted content across many use cases.

That gap — 96% need it, 36% have it — is where AI ROI goes to die.

When agents can’t access the right content, they can’t complete workflows. They can’t extract structured data from contracts, invoices, or reports. They can’t answer questions grounded in your organization’s actual policies, products, or history. They hallucinate, stall, or require constant human intervention in order to deliver correct answers.

With too much information or conflicting sources, the agent can easily draw from the data and produce the wrong result.

Aaron Levie, CEO of Box

“Too much information or conflicting sources, and the agent can easily draw from the data and produce the wrong result,” Levie notes. “Conflicting sources of truth for documents, data sources that haven’t been kept up to date, knowledge management systems that rely on tribal knowledge to navigate.”

The result is a compounding problem: the more fragmented your data, the more your agents underperform, and the harder it becomes to justify further AI investment.

Data fragmentation blocks three critical AI workflows

This chain reaction plays out in three profound ways within the enterprise.

1. Data extraction is much harder if data is fragmented

Automated structured data extraction (pulling key fields from unstructured data like contracts, invoices, forms, and reports) is one of the highest-value AI use cases in the enterprise. But it requires clean, accessible, well-organized content.

The State of AI 2026 report found that 49% of all organizations call automated structured data extraction “very important.” Among leading-edge AI organizations (the ones most advanced with AI at this point), that number jumps to 70%, compared to just 36% among early-stage organizations. This data point suggests that as an enterprise moves along the AI maturity curve, data extraction becomes even more important — or, conversely, that data extraction is a big factor in AI maturity to begin with.

70% of leading-edge organizations call automated structured data extraction "very important".

Infrastructure is key to data extraction, which is why leading-edge organizations have invested in content platforms that make their unstructured data legible to AI. Early-stage organizations are more likely to still be fighting fragmentation.

2. Fragmented data hinders workflow automation

Enterprises often have decades of legacy infrastructure containing valuable context for AI agents, but for it to matter, your agents must be able to talk to your data securely across systems. As Levie says, “Most companies have orgs that have always operated in silos. But agents are most effective when they’re tied to a process, which often cuts across these silos. The big question is how to deploy centrally managed agents that can work across organizational boundaries.”

AI agents don’t just answer questions, they execute multi-step workflows: reviewing documents, routing approvals, triggering actions, updating records. But agentic workflows require agents to move fluidly across systems, accessing the right content at the right time with the right permissions.

Fragmented data environments break this fluidity. When content lives in 12 different systems with 12 different permission models, agents can’t complete workflows without constant human intervention, and the promise of automation evaporates.

3. Scaling agentic AI relies upon data extraction

The most advanced AI use cases — autonomous agents that operate across the enterprise, make decisions, and take actions — require something even more demanding: a unified, permission-aware, governance-ready content layer.

Without it, agents either operate in a dangerously permissive environment (accessing content they shouldn’t) or a frustratingly restrictive one (unable to access content they need). Neither scales.

The State of AI 2026 report found that only 34% of organizations have formal standards for how agents access company data. Twenty-seven percent still describe their AI governance as “ad hoc.” And 49% have already experienced an AI-related data exposure incident.

Only 34% of organizations have formal standards for how agents access company data

“Data fragmentation remains a major issue for most organizations,” Levie wrote in a recent LinkedIn post. “As long as data remains highly fragmented and not in standard formats, or data is not available to the right people and agents, enterprises are dealing with issues around being able to get answers from agents that are accurate or that conform to their business practices.”

The competitive divide: leading-edge vs. early-stage

The State of AI 2026 data reveals a stark divergence between organizations that have solved the data problem and those that haven’t.

Just one year ago, only 8% of organizations surveyed qualified as “advanced” or “leading-edge” AI adopters. Today that number is 64%. The early stage cohort collapsed from 53% to just 9%.

But the gap in outcomes is even more striking:

  • 63% of leading-edge organizations describe their unstructured data as an active competitive advantage, compared to just 26% of early stage organizations
  • 50% of leading-edge organizations report ROI of more than 25% from AI, compared to roughly 1 in 9 early stage organizations
  • 80% of all organizations report at least 10% improvement from AI, but the ceiling is much higher for those who’ve solved the data layer

The organizations winning the AI race aren’t the ones with the best models. They’re the ones whose content is organized, accessible, and ready for agents to use.

Your content is your competitive moat

At Box, we’ve seen firsthand what separates the organizations that scale AI from those that stall. The answer is almost always the same: the content layer. Box is built to be that content layer. 

Box connects enterprise content (contracts, reports, invoices, policies, presentations, and more) to AI agents with the right access controls, the right data governance, and the right security guardrails in place.

That means:

  • Box Extract automates structured data extraction from unstructured documents, pulling key fields from contracts, invoices, and forms at scale, without manual data entry
  • Box Automate enables AI-powered workflow automation, routing documents and triggering actions based on content so agents can complete multi-step processes end to end
  • Box AI gives agents access to your company’s content with full context, grounded in your actual documents, not generic training data
  • Box MCP Server acts as a secure, compliant bridge between your content and any AI agent or model, enabling agents to access Box content through a standardized protocol with enterprise-grade security
  • Box Shield and Box Governance ensure that as agents access and act on content, they do so within the boundaries your organization has defined with audit trails, legal holds, and access controls that extend to agentic sessions

And for developers, Box provides a filesystem for AI agents to securely connect your critical enterprise content to your agents — both those within Box and those outside of it. The result: AI agents that can actually do their jobs because they have access to the content they need, with the permissions and governance your organization requires.

How to solve data fragmentation: a practical framework

Solving data fragmentation isn’t a one-time project. It’s an ongoing discipline. Here’s where to start:

1. Audit your content landscape: Map where your enterprise content actually lives. Identify the systems, the formats, and the gaps. You can’t fix what you can’t see.

2. Consolidate around a unified content platform: The goal isn’t to eliminate every system but to create a content layer that makes your data accessible to AI agents regardless of where it originated. A platform like Box can serve as that layer, connecting content from across your organization to any AI agent or model.

3. Establish permission-aware retrieval: AI agents should access content based on the same permissions as the humans they’re working with. Agents that can access everything produce worse results than agents that access the right things.

4. Invest in content quality and classification: Outdated documents, conflicting versions, and unclassified content all degrade agent performance. Invest in content hygiene as a prerequisite for AI scale.

5. Build governance before you need it: 49% of organizations have already experienced an AI-related data exposure incident. Don’t wait for an incident to build your governance framework. Establish formal standards for how agents access, use, and act on company data — before you scale.

From data fragmentation to AI readiness

Data fragmentation isn’t a technical problem that will solve itself as models improve; it’s an organizational and infrastructure challenge that requires deliberate investment right now. The organizations that address data fragmentation at an infrastructure level are pulling ahead fast. 

Box customer Samsung has reduced vendor assessment time by 90% using agentic AI and data extraction. The US Air Force used Box AI Studio to create a custom proposal reviewer agent that saves teams 30 to 40 hours on roughly 20 proposals a month. Cisco takes advantage of Box’s integration with Webex to give the customer success team a single source of truth that helps contractors work with consistent quality across regions.

These leading-edge AI adopters are reporting dramatically better outcomes: more ROI, more automation, more competitive advantage from their content. 

How to become a leading-edge AI adopter? Download the ebook The enterprise AI agent advantage.

Get the handbook

Box’s State of AI 2026 report surveyed 1,640 IT decision-makers across the US, UK, France, and Japan. Download the full report at box.com.

Frequently Asked Questions (FAQ)

What is data fragmentation and why is it a major barrier to scaling AI? 

Data fragmentation occurs when an organization’s information is scattered across dozens of disconnected, legacy, or siloed systems (such as document sharing services, network file shares, and cloud storage tools) that were never designed to work together. While human workers find this inconvenient, it acts as a complete blocker for AI agents. Because agents require real-time access to company-specific content to retrieve, reason, and act, fragmented data environments cause them to fail to find the right content, draw from outdated sources, or produce confidently wrong answers (hallucinations).

What does the State of AI 2026 report reveal about the gap between AI adoption and scaling? 

According to Box's State of AI 2026 report, which surveyed 1,640 IT decision-makers, there’s a massive gap between experimentation and actual scale: while 83% of organizations are currently running AI agents, only 19% are doing so autonomously at scale. This gap exists because 96% of organizations state that AI agents need access to company-specific content to be effective, yet only 36% have actually connected their agents to trusted content across multiple use cases.

How does data fragmentation specifically impact automated data extraction?

Automated structured data extraction (pulling key fields from unstructured documents like contracts, invoices, and forms) is one of the highest-value AI use cases, with 70% of leading-edge adopters calling it “very important.” However, data fragmentation makes this process incredibly difficult. Without a centralized, clean, and organized content platform to make unstructured data legible to AI, organizations struggle to extract accurate data, which directly stalls their progression along the AI maturity curve.

Why do fragmented data environments break workflow automation?

AI agents are designed to execute multi-step workflows, such as reviewing documents, routing approvals, and updating records across organizational boundaries. For these workflows to run fluidly, agents must be able to securely access the right content at the right time with the correct permissions. When data is fragmented across multiple systems with conflicting permission models, agents cannot complete these multi-step processes without constant human intervention, which ultimately eliminates the automation benefits.

What are the primary risks associated with scaling AI without a unified governance layer?

Scaling AI without a unified, permission-aware content layer leads to severe security and quality issues. Agents either operate in a dangerously permissive environment where they access sensitive content they shouldn’t, or a highly restrictive one where they lack the context to do their jobs. The State of AI 2026 report highlights this risk, showing that only 34% of organizations have formal standards for agent data access, 27% describe their AI governance as ad hoc, and 49% have already experienced an AI-related data exposure incident.

Download report