For a mid-size fund, buy the platform and build the agents on top of it, and treat the contract terms (zero retention, no-training clauses, audit rights, exit path) as the real deliverable, because the feature matrix will be obsolete before the contract is.
Status is contested, and the honest summary is that the market has not settled. The most-cited evidence for buying supports a direction rather than a ratio. It comes from a study of 52 organizations reporting that externally sourced tools reached deployment roughly twice as often as internal builds, and three caveats travel with it: the same 66/33 pair appears in that report both as a rate and as a composition of deployments, which are different quantities; the authors state the causation caveat themselves (organizational capability confounds the comparison); and they run a protocol project whose thesis the finding supports.
Watch the whipsaw before adopting anyone’s default, because none of these surveys establishes a direction of travel. One consultancy survey has 67% calling “buy existing capability” the quickest and preferred route: a stated preference, from a pulse survey rather than the report’s own research, among global executives at billion-dollar firms. A different survey measuring revealed budget found off-the-shelf’s share of AI investment falling from 38% to 32% in H1 2025 while in-house build rose. Those two findings are compatible: one is stated preference, the other revealed. Then the same tracker reversed in mid-2026, with off-the-shelf focus rising to 66% even as 76% of respondents said off-the-shelf no longer met their needs.
The checklist
1. Decide the platform before the tools proliferate, and decide it separately from the agents. The pattern that survives the survey noise is a split: very few firms want to build the underlying platform, while most intend to build or co-build the agents that run on it. In one survey 9% would build the platform in-house against 80% choosing a hyperscaler, a specialised platform or a combination; at the agent level the same respondents split 38% in-house, 23% hybrid, 14% buy. (That survey’s own summary sentence, “nearly three quarters” building, fails to follow from its chart, and its population is 110 respondents at mostly multi-billion-dollar firms in one country, which any citation of it should state.)
2. Score the platform on the enterprise bar, not the demo. Deployment mode, SSO and SCIM, RBAC, environment separation, exportable audit logs, zero retention, no-training terms, audit rights, viability and exit path. See supply-chain vetting. Selection criteria have already moved this way in survey data: scalability first, security of sensitive data second, with cost falling to seventh from first in 2023. That ordering rests on a small base and on new answer options that mechanically shift ranks, so read the ordering rather than the movement.
3. Reserve the build lane for defensible reasons rather than preferences. Two hold up. Sovereignty: data residency and extraterritorial subpoena exposure are legitimate build triggers. Quality floor nobody clears: one practitioner in a 16-person interview study trains models from scratch for internal proprietary languages, having found retrieval, LoRA and prompt-stuffing all failed his quality bar. That is an existence proof of the far build end, not a prevalence claim.
4. Price lock-in as a switching cost with a clock on it. Systems that learn a firm’s workflows create compounding switching costs, and some vendor relationships lock in within a few quarters, a finding derived from seventeen procurement and sourcing leaders plus public procurement disclosures showing two-to-eighteen-month RFP-to-implementation cycles. This directly contradicts the more optimistic claim that natural-language integration makes systems modular enough to avoid lock-in, which is asserted with no evidence behind it; for anything trained on a firm’s own data, the switching-cost finding is the better bet.
There is a subtler version at the protocol layer: early platform and protocol choices create network-effect lock-in that is hard to reverse, with BGP as the cautionary precedent: a routing protocol whose insecurity has persisted for decades because nobody can unilaterally move. That is an argument by analogy rather than evidence about agent protocols. And note the irony in the usual remedy: an abstraction layer sold as preventing model lock-in becomes the lock-in point itself, and the vendors selling them stop short of claiming code-free provider swapping in their own documentation.
5. Let cost inform the decision rather than settle it. The commoditization is real: inference cost for GPT-3.5-level performance fell from about $20 to about $0.07 per million tokens between late 2022 and late 2024, per an independent index rather than a vendor. But the same index reports per-task price declines ranging from 9× to 900× a year, so no single multiple works as a planning number, and reasoning models bill for reasoning tokens at a multiple of their non-reasoning equivalents. Claims that open-weight models match proprietary quality at a fraction of serving cost are task-specific vendor benchmarks on vendor-built evaluations, with optimization cost amortized over volume; they support no general claim about open-weight cost parity.
6. Enable experimentation rather than mandating one tool. Among 95 professional engineers tracked over six months, 82% changed their tool combinations and the mean number of tools per person rose from 1.9 to 2.9. Mandating a single tool in a market churning that fast buys compliance theatre and a shadow toolchain.
The disagreement worth understanding
One consultancy argues that off-the-shelf agents streamline routine work but rarely create advantage, so differentiating processes need custom agents. That firm sells the build, its report carries a foreword by a model vendor’s chief executive, and the same document concedes that roughly 90% of vertical use cases remain stuck in pilots. Read it against the field data pointing the other way, and notice that both sides are arguing about differentiation while a fund’s actual question is usually risk.
Vendor viability is a live and unusually noisy input right now. One advisory firm forecasts up to $500B of enterprise software revenue at risk as agents displace knowledge-worker tooling, a scenario with no stated model. In early February 2026 a two-day software selloff wiped roughly $285B of market value, triggered by a model vendor’s agentic desktop release; that figure was immediately superseded by larger estimates over the following weeks, alongside a visible counter-narrative that the thesis was wrong. A major chip vendor’s chief executive called the selloff illogical on the grounds that agents will use software rather than reinvent it. Everyone quoted has a position. For a fund the practical consequence has nothing to do with market timing. It is that exit path, data portability and source-code escrow deserve more contract attention than they did two years ago.
One term to negotiate carefully: outcome-based pricing, where the buyer pays for resolutions rather than seats. The mechanism is real and shipping. The governance problem is worth naming: outcome pricing makes results observable and leaves process unobservable, the wrong shape for vendor-oversight and books-and-records expectations. See recordkeeping and compliance gaps.
How you’d know it’s working
Time-to-deployed, and a signed contract file that passes the enterprise bar. Not a feature matrix. If internal builds outnumber deployed vendor tools two years in, the build lane has become a hobby.
The platform decision is stable while the agent portfolio churns. That is the shape to aim for: one durable procurement, many cheap, disposable builds on top of it.
What this doesn’t solve
Buying transfers construction, not accountability. Vendor-oversight obligations, exit paths, and the governance of everything employees build on top of the bought platform all remain the firm’s problem. Governance capability specifically cannot be bought and installed: it is policies, people and process, and the vendor sells none of those.
It also leaves unanswered what happens to the build lane when the buy option improves underneath it. That is the genuine uncertainty in this decision, and nobody in this corpus has measured it.
See also
- Enterprise AI assistant platforms — the category comparison this decision draws on.
- Low-code agent platforms — the compose option between build and buy.
- Supply-chain vetting — the vetting the buy lane requires.
- Dimension 11: External partnerships & vendors — the hub for vendor oversight.
- Day 3 sequencing — where the platform decision falls in the sequence.
- Domain-specific finance AI platforms — the buy side for fund-specific workflows.