Two sponsored studies were published one day apart. Cloudera, with Wakefield Research, surveyed 1,500 enterprise architects, cloud infrastructure leads, and data architects across nine markets in June 2026. MIT Technology Review Insights, sponsored by Google Cloud, surveyed 300 senior data and technology executives at mostly $500M+ organizations in February and March.
Different samples, different scopes, different sponsors. Read together, the agreements are more credible than either report alone — and the disagreements are more useful.
Viewed together, the studies reveal three important lessons. First, AI performance depends as much on data quality, system integration, and governance as it does on the model itself. Second, while the reports diagnose a similar problem, their proposed solutions reflect the commercial interests of their sponsors. Finally, evaluating AI costs requires a broader view than token pricing alone. Organizations must include the cost of accessing, transferring, and validating data, and establish clear accountability for inaccurate outputs.
Where they converge
The bottleneck sits below the model. MIT reports that 55% call their current data platform the primary bottleneck to scaling agents, and 52% say legacy latency prevents decisions at speed. Cloudera reports that 95% delayed or canceled AI projects over governance, compliance, or regulatory issues, and 72% say their architecture needs a significant overhaul. These figures are not competing estimates: the questions differ in scope, event definition, and response construction. But neither report's headline findings locate the principal scaling barrier in model capability.
Cost is infrastructure-shaped, not token-shaped. Cloudera finds 84% reporting higher infrastructure costs from AI workloads. MIT finds 50% saying legacy data systems significantly damage agent ROI. Both reports name the same trap almost verbatim: Cloudera writes that "token consumption becomes another variable in the cost together with compute or storage units," and Google Cloud's Andi Gutmans describes teams struggling to demonstrate ROI given "various costs, such as tokens." Neither frames token cost as a sufficient unit of analysis — the more defensible cost object is the full infrastructure required to produce a useful answer.
Unstructured data and missing context are the access problem. MIT's top platform constraints are entrenched silos (50%), unstructured data (40%), real-time access (34%), and missing business context (32%). Cloudera's top driver of architecture change is security, governance, and compliance (42%), followed by latency and real-time capability (35% each). The reports describe the same failure surface from two angles.
Where they split
Access versus control. MIT's diagnosis is scarcity: agents reach only 45% of enterprise data on average, and its preferred remedy is to open more of it. Cloudera emphasizes mobility: 97% move data between environments monthly and 31% daily, increasing the burden of maintaining governance across locations. One report emphasizes gates that are too closed; the other emphasizes control as data crosses them. Both conditions can be real, but they imply different first moves.
Consolidation versus distribution. Google Cloud argues that a patchwork architecture "will lead to spiraling costs" and presents a single AI-native platform as the remedy. Cloudera argues that workloads belong wherever the data lives, highlighting the 66% who moved at least some AI workloads from public cloud to private cloud or on-premises. Each sponsor's preferred architecture appears in its diagnosis. Neither survey tests whether consolidation or distribution produces lower total cost or better governance.
Broad adoption is not the same as scaled use. MIT reports that 83% are "using agents," but only 10% report widespread use; the remaining 73% describe limited use. Cloudera reports that 77% use AI "in some form," while only 26% characterize their organizations as highly AI-driven. In both reports, the large adoption headline becomes much smaller when use at scale or organizational depth is examined.
The most useful number is the one against interest
Cloudera sells hybrid and private data platforms. Its respondents ranked private cloud the hardest environment to govern for AI workloads (29%), above hybrid (24%) and public cloud (20%), with on-premises the easiest (9%). They rated public cloud the best-performing environment for AI workloads (30%), and named greater cloud spend the most common two-year plan (29%) — ahead of hybrid-first (25%) and more on-premises spend (24%).
That is a report foregrounding repatriation while its own respondents rank private cloud hardest to govern and plan to spend more on public cloud. The 66% headline is real, but it is a binary question that does not reveal how many workloads moved, whether the moves were permanent, or the net direction of workload and spending change. Read the tables, not the pull quotes.
MIT's marquee finding deserves the same treatment. The report defines 25 of its 300 respondents as "data leaders" because AI can access more than 70% of their enterprise data. All 25 report trusting their agents' decisions to be mostly or consistently accurate and relevant. That is a striking association, but it is a small, sponsor-defined subgroup reporting its own trust — not an observed performance test and not evidence that broad access caused accuracy. Broad access may instead indicate that the expensive work of classification, lineage, semantics, access policy, and evaluation is already done.
What survives both readings
Strip out each sponsor's preferred remedy and a narrower claim remains: retrieval-grounded agents can turn data access from a mostly administrative control into a recurring runtime operation. An over-broad permission, an unclassified store, or a retrieval step that returns twenty documents when three would suffice is no longer an occasional design flaw. It can be repeated thousands of times a day.
That makes retrieval a metered, governed operation, which is an awkward thing to own. It creates costs at multiple layers: ingestion and indexing, query compute, data movement or egress, and finally the input tokens used to process whatever the retrieval system returns. Architecture controls classification coverage, evidence size per request, and where inference sits relative to data. FinOps needs to connect data-platform spend with model spend through a common cost model. Otherwise, two parts of the same answer-production process may be optimized against each other. Governance has to answer what an agent may retrieve, on whose authority, and what record survives.
Neither study establishes that its sponsor's platform is required. Both establish that the constraint is real. Test the placement question against your own egress, latency, and sovereignty data.
Neither report directly measures the cost of a verified, decision-ready answer. That may be the more useful unit for organizations to develop themselves.
The question is not what percentage of enterprise data an agent can reach. It is what one trustworthy answer costs to produce — and who is accountable when it is wrong.
References
- Cloudera — The Great AI Re-Architecture (with Wakefield Research; 1,500 respondents, fielded June 5–22, 2026)
- MIT Technology Review Insights / Google Cloud — Scaling AI agents with trustworthy data (300 executives, surveyed February–March 2026)
