An agent inherits every over-broad share and mislabeled file in the estate on day one, and DSPM is how a firm finds out what that agent can reach before the agent demonstrates it.

What this category solves (and what you did before it)

Discovery and classification of sensitive data across stores, so that access scoping has something to scope against. Before it, permission sprawl was invisible because no human ever traversed a document estate the way a retrieval-backed agent does on its first afternoon.

The most useful framing of the job comes from a pre-agentic cloud framework, and it still holds: apply detection both to the stores known to hold sensitive data and to the stores that by design should not, then monitor for unintended disclosure. The second half, misplacement detection, is the whole DSPM pitch. The same guidance adds three things worth stealing: alert on sensitivity escalation, align tags to the firm’s classification policy, and monitor the sensitivity level of model output, redacting or quarantining a response whose classification rises. Cite it as lineage; its publisher now marks it historical-reference-only and it predates agents, MCP and tool use entirely.

What actually differentiates products

Coverage of the stores a firm actually has. The four products most commonly listed together for this job have scopes that barely overlap, and treating them as substitutes is the standard procurement error.

  • Amazon Macie covers Amazon S3 and nothing else: managed data identifiers, custom regex, allow lists, bucket posture. It cannot see an agent’s prompts at all. It belongs to the question of where the grounding corpus lives.
  • Nightfall is SaaS, email and API-centric, with no documented endpoint DLP. Its generative-AI story is a redaction API the calling application has to invoke, not passive coverage.
  • BigID is the only standalone discovery-and-classification pure-play of the four, spanning discovery, risk, privacy automation, AI governance and access governance.
  • Microsoft Purview is Microsoft 365-centric, and reaches third-party SaaS and cloud partly by integrating with partner products, including BigID. In Microsoft’s own architecture these two are complements, not alternatives.

Whether findings drive enforcement or produce a report. This is the same question as on every governance-tooling page in this section, and it is the one that separates a control from a PDF.

Classification honesty. Purview classifies three ways: manual, pattern-based, and trainable ML classifiers. On a direct count of Microsoft’s own published definition lists in August 2026 there are over three hundred built-in pattern classifiers and roughly eighty pretrained ML classifiers. No Microsoft documentation page states a total; the widely-quoted “315 classifiers” is a marketing snapshot of a continuously growing list, so use the shape rather than the number. Two limits the count hides and vendors rarely lead with: classifiers only work on unencrypted items, and custom trainable classifiers are English-only.

Read the coverage claims before you rely on them

Two examples from the best-documented product in the category, because they generalize.

The default oversharing assessment is a sample, not a sweep. It genuinely runs weekly with no activation needed; its scope is the top 100 SharePoint sites by usage in the organization, selected by usage rank with no agent-usage criterion anywhere in the documentation. The commonly repeated claim that it covers all SharePoint sites used by agents is wrong. The hard limits matter too: a maximum of 200,000 items per location with file counts unreliable above 100,000, item-level scanning restricted to Microsoft 365 and currently capped at ten SharePoint sites, OneDrive unsupported for item-level scanning, a four-day delay before the first default assessment, and custom assessment results that take 48 hours and expire after 30 days. A reader who takes the marketing at face value will believe they have tenant-wide weekly coverage of a top-100 sample.

The risky-AI-usage detection has preconditions. The mechanism is real: insider-risk templates score prompts and responses containing sensitive information across assistants and agents, correlate with departures, and flag agents generating sensitive responses or exceeding their own baseline. But the departing-employee scenario requires the HR connector or an Entra deletion signal, and the risky-AI-usage template requires the insider-risk browser extension installed on user devices plus at least one browsing indicator selected. Without those the template stays silent. And note the scope of “risky agents”: it means sanctioned agents drifting rather than unsanctioned ones appearing, and shadow agents remain a different problem.

One vocabulary correction for anyone reading vendor material: the documentation exposes files referenced and sensitive files referenced, not “grounding data.” Prompt and response visibility is permission-gated to specific admin roles rather than available to any data security admin.

A Cloudflare citation in this context has to name the right product. Its AI Gateway is an observability, caching and rate-limiting proxy with no documented DLP, so it belongs on the gateways page. Its prompt-PII capability is a separate WAF detection, Firewall for AI, which inspects incoming prompts for phone numbers, emails, national identifiers and card numbers, unsafe topics and injection attempts. It is generally available, and the AI detection fields require an enterprise plan plus a paid add-on.

Tier

Day 2 in general, pulled hard toward Day 1 the moment an assistant is connected to internal documents. The oversharing lesson from the first wave of Copilot deployments generalizes to every agent with retrieval.

The enterprise bar

The usual list, plus two category-specific additions. First, what is the actual scanned scope under the licence tier on offer, asked for in items and locations rather than in adjectives. Second, can classification results be exported to drive access changes in systems beyond the vendor’s own. See data access governance.

M&A state and category maturity

Heavily consolidated, and still consolidating. Of the four products above, only one is a standalone pure-play; the rest are features of larger platforms. In the wider DSPM market most of the recognizable independents of 2023–24 have been bought by cloud, backup and security platform vendors, with several large deals closing across 2025 and 2026 and at least one still pending as of late July 2026. Evaluate the acquirer’s roadmap rather than the acquired brand, and ask what happens to the classification engine when it becomes a feature of something else.

Where the category is immature

No independent party we could find has measured classification accuracy on the categories a fund actually cares about: material non-public information, deal-team walls, client identifiers. Vendors publish classifier counts, not precision and recall. And nothing in this category yet answers the agent-specific question: not “who can reach this file” but “what did the agent retrieve and does that combination of retrieved items constitute a disclosure that no individual item would.”

See also