Recordkeeping obligations did not pause for agents. The rules are old and the floors specific, while the default configuration of the most widely deployed citizen-agent platform destroys transcripts after 30 days, shorter than every period that could apply to a regulated firm.
The mechanism
The binding rules are category-based, so they never mention AI and never need to. For an SEC-registered adviser, Rule 204-2(a)(3) covers trade memoranda and (a)(7) covers communications concerning recommendations, advice, orders and the movement of funds; 204-2(e)(1) sets five years from the fiscal year of last entry, the first two in an appropriate office. For a broker-dealer, 17a-4(b) sets a three-year tier and FINRA Rule 4511 defaults to six. The question is never “is this an agent record” but “does this content fall in a listed category.” An agent drafting a recommendation to a client has produced an (a)(7) record.
There is no SEC or FINRA rule requiring an AI-specific policy. The obligation is 206(4)-7 and 3110: stronger and less exotic than an AI rule, and harder to wave away than one that doesn’t exist. Advisers Act Rule 206(4)-7 requires a compliance program reasonably designed for a firm’s actual business, FINRA Rule 3110 requires supervision, and the SEC’s FY2026 exam priorities say staff will assess whether firms adequately supervise their AI use.
Copilot Studio conversation transcripts land in Dataverse with a preconfigured job that permanently deletes anything older than 30 days. Three aggravators: admins can disable transcript saving altogether; where SharePoint is the knowledge source the stored transcript marks the agent’s own answer REDACTED, so the record omits what the agent said; and no transcript is written at all for Dataverse for Teams, developer environments, or Microsoft 365 Copilot agents.
Copilot Studio does not “violate 17a-4.” The loose version of the claim is wrong and the precise one is narrower. What is true: the default configuration auto-destroys content that would be a required record where it falls in a required-record category, and a plain Dataverse table fails to qualify as a 17a-4(f) system at any retention horizon, being rewriteable, with no audit-trail recreation, no (f)(3)(v) undertakings and no WORM. Advisers under the flexible 204-2(g) regime can get there with extended retention, indexing, prompt production and a duplicate copy. No SEC or FINRA action has named agent-transcript retention; the exposure is inference from rule text plus the off-channel enforcement record, and that inference is ours.
The tooling to fix it exists and is off by default. Purview retention policies now carry dedicated locations for Copilot experiences and enterprise AI apps, preserving content in a hidden Exchange mailbox folder subject to eDiscovery and litigation hold, retained in inactive mailboxes after someone leaves. Interactions are searchable through the mailbox, not the Dataverse transcript table. All of it depends on Purview licensing and a collection policy someone configured.
Meanwhile the market treats this as undecided. Deloitte’s survey of 3,235 leaders frames retention of system-behaviour records as an open design question. For an SEC-registered adviser or FINRA member the question is closed: the floors exist, and FINRA has said since Regulatory Notice 17-18 that a firm must be able to retain records before adopting a technology. A consultancy calling settled law unresolved is the clearest signal that the compliance framing has not reached the people buying tools.
Two failure modes sit next to retention and are governed separately. Whether the agent announces itself, and whether a firm’s statements to clients and regulators about its AI are accurate, is agent disclosure. Whether an agent optimising a profit objective discovers conduct that looks like manipulation is emergent market abuse. Both used to live on this page and were split out because they answer different questions than this one.
The cost is already visible in the sector. More than a quarter of financial executives name regulation as their biggest barrier to autonomy, and about three in five financial firms cite human review and compliance as a top ongoing cost against roughly half across industries (n=1,221, fielded November 2025 to January 2026). The CIO of an SEC-registered adviser puts it plainly: a regulated manager has to account for its significant decisions, and an autonomous system that cannot explain itself creates an audit gap.
What to do
Decide which agent outputs fall in a required-record category, then archive those to the system already used for email and chat (comms archiving). A mapping exercise rather than a technology project, and the first thing to do.
Turn off the 30-day default before deploying anything. Configure Purview retention for Copilot and AI apps, confirm transcript saving is enabled per environment, and check whether the knowledge-source configuration redacts the agent’s answers out of the record retained.
Do not treat the platform’s storage as the archive. A Dataverse table fails to qualify as a 17a-4(f) system; the archive is where the firm’s existing books-and-records controls run.
Watch the inverse failure too. SEC rules set retention floors, not ceilings: over-retention exposure comes from privacy commitments, contracts, erasure obligations and discovery. Data sent to a vendor for training or testing must not outlive the firm’s permission to hold it, structured or unstructured.
Write the AI policy anyway, filed under 206(4)-7 rather than an imaginary AI rule. Then put agents on the surveillance map: an agent touching MNPI that nobody listed is the gap that turns a technology problem into an enforcement one (data access governance).
Decide what the recordkeeping obligation actually is before buying anything to satisfy it. An adviser-only firm is under 204-2 with no format regime; a broker-dealer is under 17a-4(f) and must choose WORM or an audit trail. Buying a product marketed against a rule that binds someone else is a poor opening to an examination.
How you’d know it’s working
Pick an agent interaction from four months ago and produce it with its prompt, output and authorization chain. Four months is deliberate: past the 30-day default, short of any real floor.
The surveillance map lists agents and someone can say which touch MNPI. If it lists only humans and channels, agents are outside supervision by construction.
Compliance can name which agent outputs are books-and-records without consulting IT. If nobody owns that distinction the answer defaults to “none,” which is wrong.
What this doesn’t solve
The EU AI Act deadline for funds is widely misreported as August 2026. Regulation (EU) 2026/1744 moved standalone Annex III high-risk obligations to 2 December 2027 and embedded Annex I systems to 2 August 2028; 2 August 2026 survives for everything else, including Article 50 transparency. More to the point, a fund’s trading, research, portfolio and operations agents fall outside Annex III high-risk entirely: the creditworthiness category covers natural persons, and algorithmic trading is absent from the listed categories, so any sentence implying a high-risk cliff for a hedge fund this year is wrong twice.
Archiving is not supervision. Retaining every prompt and output satisfies a records obligation and says nothing about whether anyone reviewed them; 3110 and 206(4)-7 want a program, not a bucket. See oversight decay for why the review half rots first.
Per-agent compliance monitoring cannot see composed violations. Two agents that each stay in policy can jointly move restricted information across a barrier, and parameter-keyed monitoring reports both compliant (multi-agent cascades).
Explainability has a ceiling nobody has raised. A reasoning trace is an artifact the agent produced, not a faithful account of the computation, so an audit trail built on chain-of-thought is evidence rather than testimony. “The AI did it” fails as an answer to a regulator, and so does a plausible reconstruction the model wrote afterwards.
Finally, none of these numbers describe firms this small: the compliance-cost figures come from a survey skewed to large enterprises, the disclosure statistics cover 30 flagship agents, and the market-abuse results are laboratory work rather than enforcement actions. The rules fit exactly, and they are what to work from rather than the survey data.
See also
- Logging and audit — the substrate every examiner response is built on, and the 17a-4 audit-trail-or-WORM choice.
- Data access governance — keeping MNPI away from agents nobody put on the surveillance map.
- Comms archiving and surveillance — the vendor category built for this, and whether it covers agents yet.
- Multi-agent cascades — the barrier breach that per-agent compliance monitoring reports as compliant.
- Shadow agents — the population outside supervision because nobody knows it exists.
- Oversight decay — why the supervision half erodes faster than the retention half.
- Regulatory exposure — what a compliance-program failure costs, and why the absence of an AI-specific rule offers no cover.
- Agent disclosure and AI-washing — the transparency half that used to live on this page.
- Emergent market abuse — the trading-conduct half that used to live on this page.
- Frameworks — where the SEC/FINRA/EU AI Act mapping lives.