Employees are uploading files to tools you've never heard of. RAG pipelines are surfacing content users were never supposed to see. Agents are acting on data no human in that role would be authorized to touch, all within approved tools, all completely outside your governance framework.
Confidencial's AI data governance operates at the content layer, applying protection before data enters any AI workflow, so your governance isn't just a policy document.
One knowledge base, two outcomes. When sensitive spans are unprotected, a RAG pipeline surfaces what it can reach, not what you intended. With Confidencial, the same pipeline delivers answers with sensitive content encrypted at the span level, regardless of who's asking or which agent is running the query.
From classification to continuous enforcement. No re-architecture. Protection follows the data, not the system.
Know what AI can reach
Confidencial scans OneDrive, SharePoint, S3, Google Drive, and on-prem repositories, including every source your AI systems connect to. Sensitive content surfaced at the span level: PII, PHI, IP, regulated data.
Know what AI can reach
Confidencial scans OneDrive, SharePoint, S3, Google Drive, and on-prem repositories, including every source your AI systems connect to. Sensitive content surfaced at the span level: PII, PHI, IP, regulated data.
Enforce least-privilege for humans and agents
Per-user and per-agent access policies define exactly which spans each identity is authorized to see. This is enforced at query time, not just at configuration. Scope never assumed. Always explicit.
See what AI is actually touching
Every retrieval, every inference, every agent action generates a span-level audit trail, with identity, timestamp, data accessed, and workflow step. Not file-level logs. A cryptographic record of what AI did with your data.
Demonstrate governance that actually works
Span-level evidence for HIPAA, GDPR, ISO 42001, NIST AI RMF, and EU AI Act- the kind of proof policy documents and folder-level logs cannot produce. When a regulator asks, the answer is specific, verifiable, and complete.
Most organizations govern AI from the outside: policies, acceptable use guidelines, approved tool lists. The data moves anyway. Once a file enters a pipeline, a prompt, or a knowledge base, the governance you built for documents no longer applies. Nothing flags it because, technically, nothing went wrong.
Employee uploads a client file to an AI tool using a personal account. The source system logs a legitimate access event. Nothing flags. The data is now in an environment you don't run, can't audit, and can't revoke.
AI Guard redacts or encrypts sensitive spans before upload. Policy applies regardless of which tool the employee uses, sanctioned or not. Protection is in the content, not the channel.
RAG pipeline with appropriate read permissions retrieves content a user was never intended to see. Backend access controls don't consistently carry through to the retrieval layer. The document was restricted for the human. The model accessed it anyway.
Protection persists through chunking, embedding, and retrieval. Encrypted spans stay encrypted and the model reads around them. Sensitive content remains inaccessible even when the pipeline reaches the source document.
Agent provisioned with broader access than necessary. Access rarely revisited. The agent acts on data no human in that role would be authorized to see, and no alert fires because the access was technically granted.
Per-agent policies enforced at the span level continuously and not just at setup. Agents access only what they're authorized for. If an agent's scope doesn't include compensation data, it cannot surface it, regardless of what system it's connected to.
Standard logs track who opened a file. They don't capture what a RAG pipeline retrieved, what a copilot surfaced, or what an agent acted on across a multi-step workflow. The regulator asks for AI-specific data flow evidence. You have file access logs.
Span-level tamper-proof audit trail logs every AI interaction, such as retrieval, inference, and agent steps, by identity, time, and workflow. The specific, verifiable chain regulators require for HIPAA, GDPR, ISO 42001, and EU AI Act.
Publicly reported · Vercel / Context.ai · April 2026
A Vercel employee signed up for an AI productivity tool, Context.ai, using their corporate Google Workspace account and granted it “Allow All” OAuth permissions. Context.ai was subsequently compromised. The attacker used the stolen OAuth tokens to take over the employee's account and move laterally into Vercel's internal systems, accessing employee records, API keys, and internal credentials.
The entry point was an AI tool an employee had integrated into their workflow with permissions that were never reviewed, scoped, or revoked. The attacker had the same access the AI tool had, which was access to everything.
“No exploit. No zero-day. Just an unsanctioned AI tool, an overpermissioned OAuth grant. Your employees are doing the same things on their machines right now. The question is whether you know about it.”
David Lindner, CISO, Contrast Security · Dark Reading, April 2026Access controls govern what users can reach. They don't govern what AI tools do with that access once it's granted. Confidencial enforces data-layer protection so that even a fully compromised AI tool can only surface encrypted spans, and not readable sensitive content. The OAuth token gets stolen. The data inside stays locked.
DSPMs classify. DLP watches uploads. AI firewalls filter outputs. None of them protect what the model sees, what the agent retrieves, or what ends up in a response. Confidencial operates at the data layer, which is the only layer that travels with the content through every AI workflow.
Authorization to access a file and authorization for an AI pipeline to retrieve specific content from it are different controls, and most systems don't enforce the second one. A user restricted from opening a document directly can still have a RAG assistant retrieve and surface content from it, because backend access controls don't consistently carry through to the retrieval layer.
A policy that says “use only approved tools” is not a technical control. 89% of AI usage in organizations is invisible to security, happening outside any visibility or audit trail, despite policies. Confidencial protects the data itself, so a file uploaded to an unsanctioned tool arrives with sensitive spans already encrypted. The tool doesn't matter. The protection is in the content.
AI agents are non-human identities and should be governed as such, with explicitly scoped, continuously enforced access rather than inherited permissions provisioned once at configuration. Confidencial applies per-agent policies at the span level, enforced at query time on every workflow step. If an agent's scope doesn't include compensation data, it cannot surface that content, regardless of what system it's connected to.
DLP watches uploads and known file-transfer paths. It doesn't see what a RAG pipeline retrieves, what a copilot surfaces in a response, what an agent acts on across a multi-step workflow, or what ends up in a prompt history or vector database. Confidencial operates inside the content itself, so protection persists through every stage of an AI workflow, not just at the upload boundary.
Span-level, tamper-proof audit trails that record every retrieval, inference, and agent action by identity, timestamp, workflow step, and data accessed. The specific, verifiable chain of what AI systems did with sensitive data, sufficient for HIPAA, GDPR, ISO 42001, NIST AI RMF, EU AI Act, and CCPA. The standard logs your existing stack generates don't capture AI-specific data flows at this level.
The exposure is already happening. 77% of employees regularly paste company data into generative AI tools, and 82% of those interactions happen through personal accounts outside your visibility. The gap between what your policies say and what's actually happening widens with every new tool, agent, and workflow you add. Governance at the data layer scales with adoption. Governance by policy doesn't.
Clinical data. Client records. Compensation files. IP. Most organizations can't answer that question for their AI workflows. The AI Data Exposure Assessment maps it, including what your pipelines can access, what your agents are authorized for, and where governance stops applying.