For a 20–500 person fund the decision is ChatGPT vs Microsoft Copilot vs Claude, and neither chat quality nor the certificate list decides it: on the published enterprise bar the three are at parity, so it comes down to where the firm’s data already lives and what the assistant is allowed to touch.

What this category solves (and what you did before it)

It is the sanctioned alternative to the personal accounts staff are already using: the same capability with retention terms, an audit trail, and a tenant boundary. Before it, shadow use with neither; see shadow agents.

The enterprise bar does not discriminate

This is the finding, and it should save a procurement cycle. All three publish SOC 2 Type 2 and ISO 27001; two of the three also publish ISO/IEC 42001. All three offer SSO, provisioning, RBAC, admin audit and compliance log export, tenant isolation, no-training-by-default on business tiers, and configurable retention. Nobody offers on-premises. There is no acquisition risk to weigh: the M&A churn in this space sits in the layers around the assistants rather than in the assistants.

Two precision points that survive the parity finding:

No-training and zero retention are different levers. No-training applies to the business tiers by default; zero data retention is a separate, negotiated configuration. Ask for both explicitly. Consumer tiers are typically opt-out rather than off by default: one provider changed its consumer default to training-with-opt-out in August 2025 while excluding its Work, Team, Enterprise and API products. That is the concrete version of a point legal scholarship makes in the abstract: a vendor agent owes the buyer no fiduciary loyalty, so loyalty has to be a contract term.

Deployment mode changes the certification answer. One vendor’s model hosted inside another vendor’s cloud carries a different, later certification status than the same model direct: one such combination is listed as in-process rather than certified. A firm planning to consume a model through a hyperscaler should verify that path’s certifications, not the vendor’s headline ones.

What actually differentiates them

Where the compliance surface lives. Microsoft’s genuine advantage is that Copilot inherits the tenant’s identity model, permissions and sensitivity labels, with retention policies and audit search in the same estate as Exchange and SharePoint, the estate an examiner asks about. The caveat that vendors omit: Purview capability varies by subscription tier. A firm on Business Premium lacks an E5 firm’s Purview surface, so governance coverage should not be assumed with a Copilot seat. Its “enterprise data protection” is a contractual tier under the data protection addendum, not a private model, and the “private LLM” framing that circulates should not be repeated. Web search queries leave that boundary and are handled by the search provider as an independent controller.

Connector reach into the systems a fund actually runs. This is where a fund’s decision is usually made. One vendor ships a financial-services configuration with pre-built connectors to the data providers a fund already licenses: Box, Daloopa, Databricks, FactSet, Morningstar, Palantir, PitchBook, S&P Global and Snowflake. No price is published for it; a six-figure minimum circulates among practitioners with no primary source, so treat it as sales-quoted hearsay (see data access governance).

The agent store’s gate is the tenant admin, not the vendor. Microsoft’s store gating is real and documented: makers submit, agents are invisible to users until an admin approves, deployment is scoped to individuals, groups or the tenant, and admins can block or remove. But certification is optional and publisher attestation is a self-assessment. Presence in the store implies nothing about vetting; the tenant’s own approval is the vetting, and that approval is a job the firm has to staff; see supply-chain vetting.

Pricing, as of August 2026

Quoted list prices, which move: ChatGPT Business is $20 per user per month billed annually ($25 monthly), from two users; ChatGPT Enterprise is custom-quoted with no published list price or seat minimum. Microsoft 365 Copilot is $30 per user per month annually, with a Business add-on tier around $21 and Copilot now bundled into Business Standard and Business Premium, so a forty-person fund on Business Premium is buying it inside the suite price rather than as a $30 add-on. The 300-seat purchase minimum was removed in January 2024. Claude Team Standard is $20 per seat annually ($25 monthly) from two seats, with Enterprise priced as seat plus metered API usage.

That last one is a structural change worth planning for: hybrid seat-plus-consumption pricing breaks per-seat budgeting, and it is the shape the category is drifting toward. See cost controls.

What the usage evidence actually shows

Three findings worth more than any feature comparison.

The analytical app is the least used. In a randomized field experiment across 66 large firms and 7,137 workers, average weekly Copilot usage in months four to six ran Teams 66%, Word 54%, Outlook 49%, chat 42%, PowerPoint 25%, and Excel 8.6%. The application closest to analytical work at a fund was used least by a wide margin. The dominant use was summarizing dispersed information in Teams, not generating text. Date and scope it hard: the experiment ran September 2023 to October 2024, the firms were very large and Microsoft-heavy, and it predates agents entirely.

Standardizing on one assistant is a losing assumption. Among 95 engineers tracked over six months, 82% changed their tool combination and the mean number of tools per person rose from 1.9 to 2.9, with general-purpose chat share falling and specialised coding tools rising. The sample is small, geographically skewed and self-reported, but the direction is unambiguous.

Incumbency decides more purchases than capability does. Organizational tool use tracks the primary cloud provider, and procurement leaders in regulated workflows say plainly that they would rather wait for an existing partner to add AI than gamble on a startup. The stated buyer criteria are data boundaries, minimal disruption, workflow fit and vendor trust, rather than model quality.

The client matters as much as the platform

When researchers tested seven MCP-capable assistant and IDE clients at pinned November 2025 builds, the spread in defensive posture was wide. One popular coding client had no static validation, low parameter visibility, no injection detection, no user warnings and no audit logging; it executed a hidden-parameter file read that exfiltrated its own MCP config and an SSH secret while showing the user only an innocuous request. The best-behaved client was safe on three of four attacks and partial on the fourth. A widely-quoted “0% to 100% attack success” framing is absent from the published version of that paper; don’t repeat it. The safe clients were safe mostly by model-level refusal rather than sandboxing, which is a runtime check that a different model or a jailbreak removes. Every one of these clients has shipped many releases since; re-test rather than inheriting the verdict.

Tier

Day 1. A sanctioned assistant is the cheapest visibility win available and the fastest way to shrink the shadow population.

M&A state and category maturity

Stable at the vendor level, churning at the feature level. The consolidation to watch is horizontal: a single business seat now carries coding agents, workspace agents, deep research, connectors to the firm’s document stores, and spreadsheet and presentation extensions. That crowds the vertical specialists and the low-code builders from above. Data vendors are responding by distributing through the assistant platforms rather than building their own, which is partnering rather than surrender and leaves the question of who owns the customer relationship unresolved.

Where the category is immature

Nobody publishes comparative accuracy on financial-document tasks. The evidence base for enterprise deployment is thin enough that the best-attested named asset-manager deployment in this corpus is a vendor’s own case study of its customer, with a productivity figure drawn from a self-selected survey of about 400 of its 4,200 users, and that thinness is itself the finding.

See also