Open questions
What nobody in the field knows yet. These are durable unknowns, not homework: each one names what would settle it, so a firm can tell when the answer arrives, and several of them a firm could settle for itself faster than the literature will.
Distinct from this build’s verification register in OPEN-QUESTIONS.md, which empties as sources
get chased. This page is permanent.
The incident record is empty, and that cuts both ways
No published incident describes an end-to-end lethal-trifecta exploitation at a financial firm. Every instance in our corpus is a lab study, a reference scenario, or a definition. There are real zero-click exfiltration incidents in other settings and a documented agent-orchestrated intrusion campaign that named financial institutions among its targets, but the specific thing this wiki warns about, a citizen-built agent at a fund exploited end to end, has no public case.
Treat that honestly in both directions. It is not evidence of safety: agent incidents are hard to attribute, funds publish nothing, and the deployment base is two years old. It is also not something to hide from a skeptical CTO. What would settle it: a published post-incident writeup, or an enforcement action naming an agent. Until then, the argument for controls rests on mechanism and on adjacent-sector incidents, and should say so.
Nobody knows whether amnesty programs surface hidden agents
Three independent sources recommend discovery-and-graduation over suppression, and the mechanism is plausible: bans make disclosure punishable, so disclosure stops. No one has published a before/after or matched-firm study of whether an amnesty window actually surfaces anything. What would settle it: registration curves from firms that ran one, ideally with the post-window tail, which any firm running an amnesty window can produce for itself in two quarters; see shadow agents.
There is no basis for sizing a reviewer pool
No reviews-per-FTE figure, no throughput limit, no defect-escape rate exists anywhere in the literature for agent review. Every staffing recommendation in this wiki, including ours, is derived from arithmetic rather than measured. What would settle it: any firm publishing review volume, time per review, and the rate at which problems reach production anyway. See the reviewer pool and oversight decay.
Cross-session persistence is the unmonitored zone
Monitoring is scoped to sessions; the highest-dwell-time attacks span them. Sub-agent detection is explicitly unsolved in the visibility literature rather than merely unimplemented, and there is no benchmark for cross-session detection, so nobody can say what monitoring it would cost or catch. What would settle it: a benchmark that defines the unit of detection. Someone has to decide what a “session” even is for an agent with persistent memory first. See cross-session delayed detonation.
Agent identity has no settled architecture
Centralized provider schemes with formal proofs and decentralized identifier designs are opposite answers to the same question, and neither is deployed at scale. An agent acting for a user across a trust boundary is governed by no ratified standard, no adopted working-group draft, no registry and no maintained implementation, while every major cloud ships an incompatible in-stack answer. What would settle it: a ratified cross-boundary delegation standard with two independent implementations; the current product field sits at identity platforms for agents.
Nobody can govern what compliant agents compose into
Two agents that each pass every policy can produce a joint outcome that violates one. No single-agent view catches it, no practical multi-agent control exists, and the EU AI Act is silent on multi-agent risk. What would settle it: a control that reasons over the composition rather than the components, demonstrated on a system nobody designed for the demonstration; see multi-agent cascades.
Four tool categories still lack a credible leader
Agent runtime security, MCP gateways, tool identity, and AI governance platforms. Money is arriving; products are immature; the absence of an anchor is the finding rather than a gap for a vendor to fill. The open question is which of the four consolidates and which gets absorbed into platforms; the M&A record so far favours absorption, which argues against long contracts with point tools. What would settle it: eighteen months of the same M&A record, or a break in it.
No framework treats citizen developers as an actor class
ISO, NIST, OWASP and CSA all model the professional developer, the provider, or the deployer. The business analyst who built something on a low-code platform over a weekend has no role and no pathway to apply the controls. CSA’s own suggested fix is to extend its controls matrix with the role rather than mint another framework; nobody has done it. What would settle it: a published actor role with a reduced, applicable control set, mapped against the standards at frameworks.
No regulator has addressed agent-to-agent records
No SEC or FINRA guidance treats agent-initiated or agent-to-agent communication as a distinct recordable category. The existing rules are category-based and reach agent output on their own terms, so this is a question about clarity rather than about exposure. What would settle it: a no-action letter, an FAQ, or an examination finding that names agent traffic. See recordkeeping and compliance gaps.
Nobody knows whether allocators will ask
The claim that institutional buyers will demand auditable AI governance is asserted in the practitioner literature with no evidence behind it. For this audience it would matter more than any regulatory forecast: “your LPs will ask” moves budgets that “the SEC might” does not. What would settle it: one operational-due-diligence questionnaire from a large allocator or consultant containing AI-governance questions; a firm that receives one knows before the field does.
Nobody has measured any of this at your size
Every adoption, governance and monitoring statistic in this wiki comes from firms one to three orders of magnitude larger than a 20-to-500-person fund. The smallest tier in the best unsponsored survey is 1,000+ employees; the consultant surveys start at $1B in revenue. There is no published base rate for anything at fund scale: not agents per employee, not registry coverage, not retirement rate. What would settle it: a survey with a small-firm cut, or an industry body collecting it. In the meantime a firm’s own numbers beat anyone’s published ones, the argument the problem makes for measuring rather than citing.
See also
- The problem — why the measurement gap is itself the operating condition.
- Frameworks — where the actor-class gap sits in each standard.
- Bibliography — the sources that get closest to these questions.