Sensitive data discovery only finds the copies you already knew about
.jpg)
Short answer: Sensitive data discovery scans the repositories you already know you own and finds the PII, PCI, PHI and IP sitting inside them. That is useful and incomplete, because the copies that create real exposure are made the moment data leaves a device for a tool nobody registered. A discovery program is only as complete as its view of data in motion, which is why dope.security pairs CASB Neural for data at rest in OneDrive and Google Drive with Dopamine DLP inspecting uploads and AI prompts on the endpoint.
Every sensitive data discovery project starts the same way. Point the scanner at the file shares, the SharePoint tenant, the S3 buckets, the databases. Wait. Get a report that says you have 41,000 files containing something that looks like a Social Security number, and a heat map that makes the finance folder glow.
Then someone asks the obvious question and the room goes quiet. Is that all of it?
It is not, and the reason is structural rather than a tuning problem. Scanners enumerate locations. You supply the list of locations. The data you are most exposed by is, by definition, in a location you did not put on the list.
What sensitive data discovery does well
Let us be fair to the category before taking it apart. Data discovery earns its budget.
It gives you a defensible inventory for audit and regulatory work. It finds the genuinely forgotten material, the 2017 export sitting in a departed employee's OneDrive, the CSV of customer records attached to a closed ticket. It tells you which repositories carry concentrated risk so you can prioritize access reviews. And it establishes a classification baseline, which every downstream control needs.
Modern tooling does this well. Classifiers moved past regex-only matching into context and semantics. Coverage extends across cloud object storage, SaaS collaboration suites, and structured databases. If your question is "what sensitive data is in the systems we manage," discovery answers it.
That is precisely the sentence to look at closely. In the systems we manage.
Discovery is a census of the places you already listed
A scanner has to be told where to look. It needs credentials, an API connection, a network path. Every one of those is a registration step, and registration is the exact thing shadow SaaS and shadow AI skip.
Consider what a normal Tuesday produces. Somebody pastes a paragraph of a contract into a personal ChatGPT account to get a plain-English summary. Somebody uploads a customer spreadsheet to a free PDF converter because the corporate tool needs a ticket. Somebody connects a personal Google Drive to an unvetted OAuth app that requests broad file scopes. Somebody drops a database export into a project management tool their team adopted without telling IT.
None of that appears in a discovery scan. There is no connector for a destination you never approved. The copy exists, it contains regulated data, and it lives somewhere your inventory has no row for. Our breakdown of CASB versus DLP covers why data at rest and data in motion need different instruments, and this is the sharpest version of that argument.
The falsifiable claim is straightforward: run a discovery scan and an egress inventory in the same month and compare the destination lists. If the egress list contains destinations the scan never touched, your discovery coverage is a subset, and you now know the size of the gap.
Three blind spots that no scanner closes on its own
The unregistered destination
This is the primary one. Discovery covers approved systems. Data flows to unapproved ones. The gap is not a percentage of files, it is a whole category of location, and it grows every time a team adopts a tool in an afternoon. Our shadow IT discovery playbook covers how to build the destination list you are missing.
The prompt
An AI prompt is sensitive data in transit that never becomes a file anywhere you can scan. Someone pastes patient notes, source code or a compensation table into a chat box. There is no object created in your tenant. A repository scanner has nothing to find, and it never will, because the artifact only ever existed as a request body.
This is now a large share of real-world exposure and it is invisible to the entire at-rest tooling category. It is only catchable at the moment of egress, which means on the device or in the path.
The lag
Scans are periodic. Even continuous monitoring is a polling loop against known locations. Between one scan and the next, data is created, copied and moved. Discovery tells you what was true at scan time in the places you scanned. It is a photograph, and an incident is a video.
How the approaches compare
Four things get sold as sensitive data discovery. They answer different questions, and it is worth being precise about which.
Repository scanners and DSPM
These connect to storage and SaaS platforms and classify what is inside. Strong for audit evidence, access review prioritization and finding forgotten data. Blind to any destination without a connector and to anything that never lands as an object. Answers: what sensitive data sits in systems we manage.
Network DLP
Sits inline and inspects traffic crossing a chokepoint. Sees data in motion, which is the right layer. Depends on the traffic passing the chokepoint, which stopped being reliable when work left the building, and in a cloud-proxy model it means every request detours to a point of presence and back before reaching its destination. Answers: what sensitive data crossed our network.
Cloud proxy DLP add-ons
The same inspection, delivered from a vendor PoP, usually as a separately licensed module. Zscaler gates prompt DLP behind its Data Protection add-on and licenses AI Guard and AI Scanning separately. Netskope's AI Guardrails, which are genuinely capable, ship in the higher Max Advantage tier with CASB API as a separate SKU. Forcepoint's GenAI Security is assembled from multiple SKUs, and its classification brain was external until April 2025. Broadcom and Symantec's documented GenAI capability still dates to April 2023 with no tenant-aware corporate versus personal control. Menlo's browser DLP is regex and dictionary based, 380-plus dictionaries rather than semantic, and bound to the browser, so API, IDE and desktop AI traffic is out of scope. Cisco Umbrella's DNS-layer base cannot read payloads at all, and Cisco's own doc 225162 confirms tenant-aware AI control requires the intelligent proxy with SSL decryption and a root certificate.
On-device inspection (dope.security)
Inspection runs on the endpoint, where the data leaves. Every browser, every thick client, every CLI and every AI desktop app goes through the same point, because the point is the device rather than a network location. Dopamine DLP intercepts file uploads and AI prompts and classifies them through zero-retention APIs, in Block, Monitor or Off mode (US Patent 12,464,023). Answers: what sensitive data actually left, to where, from which process, regardless of network.
Stated as pairs, the difference is concrete:
- Coverage: a repository scanner covers locations you registered; dope.security covers every destination a managed device reaches, registered or not.
- AI prompts: at-rest scanners cannot see a prompt because no object is created; Dopamine DLP classifies prompt content in motion before it leaves.
- Timing: a scan reports what was true at scan time; on-device inspection makes the decision at the moment of egress, so Block actually prevents the copy.
- Network dependence: network DLP needs traffic to cross a chokepoint and cloud DLP adds a PoP round trip to every request; dope.security inspects locally and the traffic flies direct.
- Licensing: AI and prompt inspection are commonly a separate module or a higher tier on legacy platforms; dope.security ships DLP, Shadow IT discovery and Cloud Application Control in one console.
- Privacy: cloud inspection means decrypted content exists in third-party infrastructure; on-device inspection keeps the plaintext on the device the data came from and uses zero-retention classification APIs.
Every vendor named above is competent and most score well. The claim is narrow and testable: an architecture that inspects at a registered location or a network chokepoint cannot see data going somewhere unregistered from a device that is not on your network.
What a complete discovery program looks like
The fix is not to abandon scanning. It is to run discovery in two directions and reconcile them.
Start with the at-rest census, because you need the audit artifact and the classification baseline. Scan the approved repositories, classify, and record where concentrated risk lives. CASB Neural does this with LLM-powered classification across OneDrive and Google Drive, flagging publicly or externally shared files containing PII, PCI, PHI or IP, with one-click remediation.
Then build the egress inventory. Watch what leaves managed devices and to where. This produces the destination list your scanner does not have, including personal SaaS accounts, unvetted AI tools and one-off utilities. Shadow IT discovery on the endpoint is how you get it, and it is also how you find out that the destination list is longer than anyone guessed.
Now reconcile. Any destination in the egress list that is not in the at-rest scope is either a repository to bring into scope or a flow to block. That single comparison converts a coverage question into a work queue.
Finally, enforce at the moment of egress, not after the fact. Discovery tells you what happened. A Block decision on the upload is the only control that changes the outcome. Our guide to endpoint DLP and data in motion covers how that decision gets made without a network detour, and the complete DLP guide covers the policy model.
The test that shows the gap in an afternoon
You do not need a project to prove this. Take a managed laptop. Paste a fake but realistic record, a full name with a plausible SSN and account number, into a personal AI account in the browser. Then attach a spreadsheet with the same data to a free online converter. Then upload it to a personal cloud drive.
Now go ask your discovery tooling to find those three copies. It will find none of them, because none of them landed in a scanned repository and one of them never became a file at all.
Run the same three actions with on-device inspection enabled. Each one is classified as it leaves, attributed to the process and the destination, and blocked or logged according to policy. That is the difference between an inventory and a control. For the wider category picture, our write-up on SaaS sprawl and shadow SaaS discovery sets out where each layer belongs.
Where dope.security fits
dope.security runs a single lightweight agent, the dope.endpoint, under 100 MB of RAM. Our Fly Direct secure web gateway performs SSL inspection on the device, which is what makes on-device DLP possible at all: you cannot classify what you cannot decrypt, and decrypting locally means no traffic detour to a point of presence. Dopamine DLP handles data in motion. CASB Neural handles data at rest. Cloud Application Control keeps people on approved SaaS tenants so the corporate copy is the only copy. Shadow IT discovery builds the destination list. All of it in one console, built from scratch rather than assembled through acquisitions.
It scales the boring way. A Fortune 100 customer went from 900 devices to over 18,000 in weeks, roughly 3,000 per week, pushed silently through Intune. A healthcare organization across 34 offices secured 99% of devices in one week and cut web access tickets 70% in 90 days. Greylock Partners moved off Cisco Umbrella in 27 days, because DNS-only filtering could not see inside HTTPS.
Sensitive data discovery answers a question worth answering. Just be clear which question it is. It tells you what is in the buildings on your map. The exposure that ends up in a breach notification is usually in a building nobody drew.
Want to see your own egress destination list? Book a 20-minute demo and we will run discovery and on-device DLP side by side.
Frequently Asked Questions
What is sensitive data discovery?
Sensitive data discovery is the process of scanning storage locations and SaaS platforms to find and classify regulated or confidential data such as PII, PCI, PHI and intellectual property. It produces an inventory of what sensitive data exists and where. It covers the repositories you connect it to, which means unregistered destinations and data that never lands as a file fall outside its scope.
Why can't a discovery scan find data pasted into ChatGPT?
Because no object is ever created in a system you can scan. A prompt is a request body sent to a third party, so there is no file, row or blob for a scanner to classify. The only place to catch it is at the moment it leaves the device. Dopamine DLP inspects prompt content on the endpoint and applies Block, Monitor or Off policy before the request goes out.
How is sensitive data discovery different from DLP?
Discovery finds and classifies data at rest in known locations and reports on it. DLP makes a decision about data in motion at the moment it moves and can prevent the movement. Discovery tells you what happened and what you hold; DLP changes the outcome. Most programs need both, plus a reconciliation step where egress destinations get compared against scan scope.
Do I still need at-rest scanning if I have on-device DLP?
Yes. At-rest scanning gives you the audit inventory, the classification baseline, and the forgotten historical data that predates any egress control. dope.security covers it with CASB Neural across OneDrive and Google Drive, including publicly or externally shared files, with one-click remediation. On-device DLP covers what leaves from here on.
Does prompt and upload inspection require a separate license on most platforms?
Usually. Zscaler gates prompt DLP behind its Data Protection add-on and licenses AI Guard and AI Scanning separately. Netskope ships AI Guardrails in the higher Max Advantage tier with CASB API as a separate SKU. Forcepoint's GenAI Security is a multi-SKU assembly. dope.security includes Dopamine DLP, Shadow IT discovery and Cloud Application Control in one platform and one console.
Will on-device inspection slow users down or drain their laptops?
No detour is added, which is the main source of latency in cloud DLP: inspection happens where the request originates instead of at a point of presence, so traffic flies direct. The dope agent runs under 100 MB of RAM, and dope.security measures up to 4x performance versus legacy proxy secure web gateways.
How do I find the destinations my discovery tool is missing?
Build an egress inventory from managed devices and diff it against your scan scope. Shadow IT discovery on the endpoint shows which SaaS and AI tools people actually reach, and whether they reached them with a corporate or personal account. Any destination on that list without a corresponding connector in your discovery tooling is a coverage gap you can act on immediately.


.jpg)
.jpg)
.jpg)

