Archiving is not supervision. Retaining every prompt and output satisfies a records obligation and says nothing about whether anyone read them. A record kept and never reviewed is the worst of the three available positions, because it proves the firm had the evidence.

The checklist

1. Put agents on the surveillance map, by name. The map most firms run lists humans and channels. Agents missing from it fall outside supervision by construction, and the gap is structural rather than an oversight. Compliance advisers writing for private funds now say plainly to “use surveillance to monitor employees’ use of AI systems, in addition to communications and trading indicators of MNPI misuse.” That is a third surveillance domain alongside comms and trading, and almost nobody has built it.

Where the surveillance itself runs on a model at a member firm, FINRA has named what the policy has to cover: technology governance, model risk management, data privacy and integrity, and the reliability and accuracy of the model (Regulatory Notice 24-09, 2024-06-27, applying Rule 3110 to the review of electronic correspondence). Four elements, from the regulator, for the recursive case where the supervisor is itself the thing that can hallucinate.

2. Decide what gets reviewed before deciding what gets logged. This is where the sources openly conflict, and the conflict is real. One side says surveil agent use as a supervised channel. The other warns that retained conversational records are themselves exposure, because an agent output that “does not accurately reflect the intent of the users” still reads to an examiner as evidence of what the firm believed. Both are right, and neither acknowledges the other. Our position: logging less is the wrong resolution. Retention is largely not optional, and under-logging fails a records obligation to avoid an interpretive risk. The resolution is that the review has to be real. Logged and unreviewed is the position that carries both costs at once.

3. Sample properly, and say how. FINRA’s recurring finding against communications review is worth stealing whole: reviewing “without selecting adequate samples or using targeted key word searches,” and failing to review communications in languages the firm actually does business in. A sample of “whatever the queue surfaced” is the finding waiting to be written. Document the sampling method, the size, and why they are adequate for the population.

4. Surveil for absence, not only presence. The cleverest supervisory practice in the FINRA report is monitoring for a decrease or absence of activity on approved channels as an indicator that someone moved off-channel. The agent analogue is direct: a builder whose sanctioned-platform usage drops has probably not stopped building. Pair it with shadow agents.

5. Revise the keywords as the vocabulary moves. FINRA expects firms to revise surveillance terms frequently and tailor them to the business. Agent vocabulary turns over faster than trading slang. A keyword list written at deployment is stale within two quarters.

6. Log the three artifacts a regulator has actually named. Prompt and output logs for accountability, which model version was used and when, and human-in-the-loop review of outputs including regular checks for errors and bias. Model-version-and-date is the one teams skip and the one that makes a reconstruction possible at all: see logging and audit.

7. Treat information barriers as a design constraint, not a policy sentence. Section 204A requires advisers to establish, maintain and enforce written policies reasonably designed to prevent misuse of material nonpublic information, and examiners treat weak barriers as a serious red flag. An agent that can traverse research, portfolio and trading data crosses a barrier by default unless someone stopped it. Prohibit MNPI in external or public AI tools explicitly, and check the retrieval scope rather than the policy: data access governance.

8. Review for degradation on a calendar. Most IOSCO members require periodic review to detect model degradation, framed through existing risk-based monitoring or model-governance obligations, and the supervisory question is simply what processes review and document AI outputs periodically. An agent that was accurate at deployment and drifted is the ordinary case, not the exotic one.

The conflict this page does not resolve cleanly

MNPI can be inferred as well as ingested. An agent with access to several individually unremarkable datasets can compose something material out of them, and access controls govern reach rather than inference, so composition escapes them. No source in this material offers a control for it, and neither do we. It is a live gap, it belongs in a firm’s risk assessment as an accepted unknown rather than a solved problem, and it is logged in open questions.

How you’d know it’s working

The surveillance map lists agents, and someone can say which ones touch MNPI. If it lists only humans and channels, this page has not started.

The firm can produce the sampling methodology for agent-output review, with a number in it. “We look at them” is the answer that becomes a finding.

Someone has been surprised by a review at least once. A supervisory process that has never found anything is either running on a clean population or not running.

Reviewed volume is reported alongside logged volume. The ratio is the honest measure of whether this is supervision or storage.

What this doesn’t solve

This is the recurring review process only. What must be retained and for how long is recordkeeping and compliance gaps; the substrate is logging and audit.

Per-agent review cannot see composed violations. Two agents that each stay in policy can jointly move restricted information across a barrier, and parameter-keyed monitoring reports both compliant: multi-agent cascades.

Nobody has published a review-rate benchmark for agent output. There is no equivalent of a communications-surveillance sampling standard, so the sampling design is the firm’s to defend and the defence is currently unaided by any published norm.

Scope: FINRA Rule 3110 and the communications findings quoted here bind member broker-dealers. The adviser equivalents are Rule 206(4)-7 for the programme and §204A for MNPI, and they demand a programme without specifying its mechanics. The FINRA material is the best available description of what a working programme looks like, and for an adviser it is a description rather than an applicable rule.

See also