This post explains why a Microsoft Purview deployment built before generative AI is structurally incomplete once Microsoft 365 Copilot is running, rather than simply needing a few new settings. It covers the threat model shift that makes existing controls insufficient, the Purview capabilities that exist only because of AI and when they became generally available, the reasons your current data loss prevention policies do not cover Copilot even though you have data loss prevention, the long-standing controls whose behavior quietly changed once Copilot arrived, why the decisions involved are no longer compliance decisions, and where Purview genuinely stops. If you are being asked whether your data protection is ready for Copilot, the honest answer depends on which of these has been looked at.
Copilot Did Not Add a Risk. It Changed the Threat Model
This is the part that matters most, and it is the reason a checklist of new settings misses the point.
Information protection as most organizations built it was designed to stop sensitive data leaving. Data loss prevention rules watch the exits: email, endpoints, external sharing, removable media. Classification exists so those rules know what to stop. The whole apparatus is tuned for exfiltration, and it was tuned that way for good reasons, because that was the dominant risk.
Copilot’s dominant risk points the other way. It makes content that was already inside the organization, already accessible, and already forgotten, trivially findable by people who technically always had access but would never have stumbled across it.
Nothing leaves. That is precisely why an exfiltration-shaped control set does not catch it.
Microsoft’s own documentation acknowledges the amplification directly, stating that because of the power and speed of AI it can proactively surface content that might be obsolete, over-permissioned, or lacking governance controls, and that generative AI amplifies the problem of oversharing data.
So the question to ask of an existing Purview deployment is not whether it works. It is what it was built to prevent. A policy set tuned entirely for egress does not address latent-access discovery, and most deployments were tuned entirely for egress.
Why “We Already Have DLP and Labels” Does Not Hold
This is the most common thing we hear, and it fails for three separate reasons. Two of them are mechanical rather than a matter of degree.
Your Existing Policies Do Not Cover Copilot, and Cannot
Microsoft 365 Copilot is a distinct data loss prevention policy location. Microsoft’s documentation states that the Microsoft 365 Copilot policy location is only available in the Custom policy template, and that when you select it, all other locations for that policy are disabled.
That is a scoping mechanic with a blunt consequence. You cannot add Copilot coverage to your existing policy. It has to be its own policy. Which means an organization with a mature, well-tuned data loss prevention deployment covering Exchange, SharePoint, OneDrive, Teams, and endpoints has exactly zero coverage over Copilot interactions until somebody builds a separate policy for it.
Label-Based Controls Are Blind to Unlabeled Content
Microsoft’s security blog puts it plainly: unlabeled sensitive data is high-risk data, and if a file has not been classified it will not be labeled or protected, which increases the risk of accidental exposure.
Every control that keys on a sensitivity label does nothing on content that does not carry one. In most estates the unlabeled tail is the majority of the content, and it is disproportionately the old, forgotten material that Copilot is best at surfacing.
On how large that tail typically is: Microsoft has not published a figure, and we are not going to borrow one from a vendor selling classification tooling. Measure your own estate. That number is knowable and it is the one that matters.
Retention Was Not Built to Cover Copilot Either
There is a dedicated retention location for Copilot interactions, and Microsoft states it is a new, dedicated policy location that is not included in existing policies and must be selected when creating or editing a policy.
Same mechanic as data loss prevention, same consequence. A pre-AI retention policy set does not cover Copilot prompts and responses until the location is explicitly added.
And there is a second-order effect worth thinking through before it arrives unannounced. Copilot prompts and responses are stored as compliance copies in a hidden folder in the user’s mailbox, which makes them discoverable through eDiscovery. Your employees’ conversations with Copilot are now potential evidence. Broad mailbox searches will capture them unless deliberately filtered, and legal holds need to account for them.
Does your DLP policy actually have a dedicated Copilot location, or did someone add it to an existing one?
That second scenario doesn’t add coverage — it silently disables the policy’s other protections. We can check which one you actually have.
Controls You Already Had That Now Behave Differently
Copilot did not only add capabilities alongside the old ones. It changed what some long-standing settings actually do.
The EXTRACT Usage Right Became a Gate
Where a sensitivity label applies encryption, Microsoft states that users must have the EXTRACT usage right as well as VIEW for AI apps to return the data.
This right has existed for years and was, for most organizations, a detail nobody thought about. It is now the switch that decides whether Copilot can read a file. A label granting VIEW but not EXTRACT means Copilot cannot summarize a document that the person asking can open and read themselves.
Which produces a specific and confusing failure. Users report that Copilot cannot see documents they are looking at. The cause is a labeling decision made years earlier for a completely different reason. This is the single most common thing we unpick.
The related trap is labels configured with the option that lets users assign permissions themselves. Users grant VIEW without EXTRACT, because why would they think about it, and Copilot silently stops working on that content.
Label Priority Order Now Has Functional Consequences
Copilot in Word, PowerPoint, and Outlook inherits the source file’s label onto new content, and where multiple files are used, Microsoft states the sensitivity label with the highest priority is used for inheritance.
Label priority was ordering in a list. It now determines what protection Copilot-generated content receives when it draws on several sources at once. Organizations that set that order casually, which is most of them, should look at it again.
One limit to be clear about: Copilot does not apply or generate its own sensitivity label. If there is no label on the content it works from, the output carries none either. Inheritance is real but it only inherits what exists, and it applies in specific applications rather than everywhere.
Purview Is Not a Compliance Tool Anymore
The belief that Purview belongs to the compliance team survives right up until Copilot is deployed, and then it fails for a structural reason rather than a cultural one.
Once classification governs what an AI assistant can retrieve, the classification decisions stop being compliance configuration and become operating decisions about the business.
Microsoft’s own architecture guidance says there is no universal standard for classification or for defining a taxonomy, and that it is driven by an organization’s motivation for protecting data. It tells workload owners to rely on the organization for a well-defined taxonomy rather than defining their own. Its adoption guidance says to get buy-in on the taxonomy from compliance and business stakeholders, and its labeling guidance recommends a working virtual team that identifies and manages the business and technical requirements.
Read what that actually requires. Someone has to decide what counts as sensitive, and different functions answer differently. Someone has to arbitrate when legal and a business unit disagree. Someone has to own the unlabeled majority of the estate and decide what happens to it. And now those decisions determine whether an expensive AI tool is useful or useless, because over-classification breaks Copilot as surely as under-classification exposes content.
None of those are compliance-desk toggles. They are business decisions with a compliance implementation.
Worth knowing what Microsoft does not provide here. It gives framework-level advice to involve stakeholders. It does not publish an operating model, a dispute-arbitration process, or a responsibility matrix that separates business decisions from technical ones. That gap is real, and it is where most deployments stall: not on configuration, but on nobody owning the decision the configuration implements.
What Purview Delivers for AI, and Where It Stops
There is a genuine AI-specific layer in Purview now, most of it built or rebuilt recently. Data Security Posture Management for AI runs data risk assessments and gives visibility into AI activity. The Copilot data loss prevention location can stop Copilot using labeled or sensitive content at the moment of interaction rather than only when data moves. Prompts and responses are audited. Insider risk gained a template for risky AI usage.
The limits are what determine your plan.
A Copilot policy location that blocks labeled or sensitive content at interaction time
It disables every other location in the same policy, so it needs its own. Custom agents are governed separately
Prompt and response auditing in standard auditing, retained 180 days by default
The audit record shows that an interaction happened and which files it touched, not the prompt or response text. Reading content needs eDiscovery or the activity explorer
A recurring data risk assessment that runs without activation
The default assessment covers only the top SharePoint sites by usage, not the estate
Custom assessments with item-level scanning and remediation
Real caps on items per location and on how many sites can be scanned at item level, and OneDrive is not supported for item-level scanning
Auto-labeling and trainable classifiers to classify at scale
Classification is probabilistic. It needs simulation, produces false positives and negatives, and cannot recognize context-dependent sensitivity a person would spot immediately
Label inheritance onto Copilot output in Word, PowerPoint, and Outlook
Unlabeled input produces unlabeled output, because Copilot does not create labels. Inheritance applies in specific apps, not everywhere
Encrypting labels that travel with content and gate Copilot through EXTRACT
Nothing retroactive. Content already overshared and unlabeled stays exposed until it is remediated
Site-level controls that remove content from Copilot and tenant search
They do not change permissions. Authorized users still open the content directly. This buys time, it does not fix anything
Two patterns run through the right column and both are worth saying out loud to whoever is approving the budget.
The first is that Purview protects content going forward and does almost nothing about the mess already in place. Labeling new content from creation is the right long-term move and it does not touch the decade of material sitting in sites nobody owns.
The second is that the full experience sits behind Microsoft 365 E5 or the Purview suite, with a reduced version available at lower tiers. Which capability you can reach is a licensing question, and it is worth answering before the readiness project is scoped rather than after.
Want to know how much of your estate is actually unlabeled?
A data protection assessment measures your own estate rather than borrowing someone else’s average, and shows exactly which content Copilot can and can’t see.
What Goes Wrong When Organizations Do This Themselves
The failures repeat, and none of them are exotic.
Taxonomies nobody can use. A label for every business unit and every exception produces something users do not understand, automation cannot reliably apply, and support teams inherit. Simpler taxonomies survive contact with actual users.
Policies at the wrong scope. Covered above, and worth repeating because it fails silently. Adding the Copilot location to an existing policy disables that policy’s other coverage rather than extending it.
Auto-labeling straight to enforcement. Without a meaningful simulation period you get mislabeled content at scale, user complaints, and a loss of trust that is far harder to recover than the labels are to fix.
Encryption applied without tracing the consequences. This breaks external collaboration and, through the EXTRACT gate, silently breaks Copilot on content people expect it to read.
Over-restriction that defeats the purpose. Label enough content as highly confidential and Copilot cannot retrieve enough to be useful. The rollout stalls, and the conclusion drawn is usually that the AI does not work rather than that the classification is too blunt.
Treating labels as a permissions fix. Labels classify and can encrypt. They do not repair oversharing. Overshared sites, broken inheritance, and old sharing links are still there afterwards. That is a separate problem with a separate fix, and it is covered in our companion post.
Where to Start
Run a data risk assessment first and read it as a discovery risk report rather than an oversharing report. The question it answers is how much sensitive content is unlabeled and how much is reachable by people who never go looking. Check whether the default scope is representative of your estate, because for larger organizations it will not be.
Then settle ownership before you build anything. Who owns the taxonomy, who defines sensitive, who arbitrates disputes, who owns the unlabeled majority. A taxonomy built without business agreement is the single largest source of rework in this work, and it is cheaper to have the argument before the labels exist.
Then close the AI-specific gaps: a dedicated Copilot data loss prevention policy, the Copilot retention location added, and auditing confirmed with a retention period that matches your regulatory position rather than the default.
Then reconcile encryption with Copilot. Audit your encrypting labels for the VIEW-without-EXTRACT problem and check that label priority produces the inheritance you actually want. Validate with test prompts, because the failure mode here is silent and users will report it as Copilot being broken.
WME runs Copilot data protection assessments covering classification coverage, policy gap analysis, encryption and label reconciliation, and the ownership model that has to sit underneath all of it.
Over-classification breaks Copilot as surely as under-classification exposes content.
WME runs Copilot data protection assessments covering classification coverage, policy gap analysis, encryption and label reconciliation, and the ownership model underneath it all.