Every argument for unifying a firm’s data so agents can use it is also an argument for a larger blast radius, and the people making the first argument do not price the second: this dimension is where usefulness and exposure are the same lever.

Why this matters

The consolidation case is real, and the prescription that follows from it, making the sources of truth machine-readable, is sound. An agent with access to unified data can draw cross-domain connections nobody would attempt by hand, and the constraint is a storage-layer problem rather than a presentation one: institutional knowledge lives in PDFs, decks and formatted documents built as outputs for humans rather than as sources of truth for machines.

The same essay omits the other half: the consolidation plan and the access-scoping plan are the same project, funded together, or the firm has built a very efficient way to lose everything at once. Cross-domain aggregation is exactly what converts one prompt injection into firm-wide exfiltration. The lethal trifecta’s middle leg is access to sensitive data, and a unification project is a program to grow that leg on purpose, which argues for funding the scoping alongside the consolidation rather than against consolidating at all.

Readiness is low and the honest numbers are unflattering. Only 15% of 603 respondents said their data and systems were fully ready to support agents in core processes (12% said the same of risk and governance controls); a separate survey of 623 found 13% calling their data architecture well-equipped. Both are Harvard Business Review Analytic Services panels fielded in July 2025 and vendor-sponsored, so they count as one measurement rather than two.

Two things to disbelieve. The circulating figure that 80% of enterprise knowledge is trapped in unstructured formats has no primary source: it traces to a 1998 analyst note with no methodology, and vendors have been reprinting it since. The second is the practitioner assertion that “data quality is foundational, agents amplify existing data problems”, which carries no evidence however plausible it sounds; cite it as a position rather than a finding.

The cost shape does have evidence behind it, and the recurring bill for agents at scale is data plumbing rather than model tokens. Just over half of firms with unified data architectures name storage, movement and duplication as their largest ongoing AI expense, rising to roughly two-thirds among less integrated firms, more than double the share naming compute (n=1,221 CIO/CTO/CDAO respondents; the sponsor sells the unified architecture the finding recommends, a conflict worth factoring in).

For a fund, the sensitivity gradient is the part that generalizes. Positions, counterparties, client identifiers and MNPI do not become less regulated because a retrieval step put them in a context window. The trace store inherits every classification of everything the agent read, a fact most firms discover after the traces exist.

Where you stand

LevelLooks likeCheapest next move
CrawlNo classification an agent can act on. Agents read whatever their builder’s account can reach.Enumerate what the three most-used agents can actually reach, rather than what they should. The gap is the finding.
WalkSensitive stores identified. Agent access still inherits the human’s entitlements wholesale.Give one agent its own scoped read credential and see what breaks; that is the migration estimate.
RunAccess scoped per tier, retrieval logged, classification enforced at the retrieval point rather than the document.Sample traces monthly for data the agent had no business seeing.
FlyMachine-readable sources of truth with recorded provenance. Access reviewed on a cadence, and revocation tested.Re-run the enumeration and compare against last quarter’s; drift is the metric.

Concerns this dimension covers

Controls that answer them

  • Data access governance — scoping reads to the task rather than to the owner. The ceiling on every exfiltration story in this wiki.
  • Egress control — where the data actually stops leaving.
  • Logging and audit — because the trace is now a record of sensitive data, whether or not that was the plan.

Who sells it

Sequencing and where this is checked

Open questions

  • Nobody has published a working production implementation of memory-store validation. One would settle the open item in memory poisoning.
  • There is no measurement anywhere in the sources here of the relative payoff between making sources machine-readable and scoping access. The order in the table above is this page’s own judgment call.