Govern AI

Your AI policy isn't your AI control.

Employees are uploading files to tools you've never heard of. RAG pipelines are surfacing content users were never supposed to see. Agents are acting on data no human in that role would be authorized to touch, all within approved tools, all completely outside your governance framework.
Confidencial's AI data governance operates at the content layer, applying protection before data enters any AI workflow, so your governance isn't just a policy document.

CONFIDENCIALAI GUARD ACTIVE
AI Knowledge Base Ingestion Log
Source documentMatter files · privilege log
Patient recordsENCRYPTED FIELD
Formulation dataENCRYPTED FIELD
Compensation dataENCRYPTED FIELD
Agent identityAnalytics-Agent-04
Access scope3 of 11 spans authorized
Protection enforced at ingestion — before the model sees it
KB-2026-0291 · RAG Pipeline · Legal Affairs
89%
of AI usage in organizations is invisible to security
LayerX Enterprise GenAI Security Report 2025
47%
of CISOs have already observed AI agents exhibit unintended or unauthorized behavior
Cybersecurity Insiders 2026 CISO Survey
79%
of organizations deploying agentic AI lack a mature governance model
Deloitte, 2026
$670K
average additional cost of a shadow AI breach compared to a standard breach
IBM CDB, 2025
97%
of organizations that reported an AI-related breach  lacked proper AI access controls
IBM CDB, 2025
See it in action

The model gets context. Not your data.

One knowledge base, two outcomes. When sensitive spans are unprotected, a RAG pipeline surfaces what it can reach, not what you intended. With Confidencial, the same pipeline delivers answers with sensitive content encrypted at the span level, regardless of who's asking or which agent is running the query.

AI Guard Demo
Query the knowledge base
How it works

Five steps from exposed to governed.

From classification to continuous enforcement. No re-architecture. Protection follows the data, not the system.

01  classify

Know what AI can reach

Confidencial scans OneDrive, SharePoint, S3, Google Drive, and on-prem repositories, including every source your AI systems connect to. Sensitive content surfaced at the span level: PII, PHI, IP, regulated data.

02  protect

Know what AI can reach

Confidencial scans OneDrive, SharePoint, S3, Google Drive, and on-prem repositories, including every source your AI systems connect to. Sensitive content surfaced at the span level: PII, PHI, IP, regulated data.

03  s

Enforce least-privilege for humans and agents

Per-user and per-agent access policies define exactly which spans each identity is authorized to see. This is enforced at query time, not just at configuration. Scope never assumed. Always explicit.

04  m

See what AI is actually touching

Every retrieval, every inference, every agent action generates a span-level audit trail, with identity, timestamp, data accessed, and workflow step. Not file-level logs. A cryptographic record of what AI did with your data.

05  p

Demonstrate governance that actually works

Span-level evidence for HIPAA, GDPR, ISO 42001, NIST AI RMF, and EU AI Act- the kind of proof policy documents and folder-level logs cannot produce. When a regulator asks, the answer is specific, verifiable, and complete.

The AI exposure lifecycle

AI governance fails at the data layer.

Most organizations govern AI from the outside: policies, acceptable use guidelines, approved tool lists. The data moves anyway. Once a file enters a pipeline, a prompt, or a knowledge base, the governance you built for documents no longer applies. Nothing flags it because, technically, nothing went wrong.

Exposure vector
Without protection
With Confidencial
Shadow upload
Ungoverned

Employee uploads a client file to an AI tool using a personal account. The source system logs a legitimate access event. Nothing flags. The data is now in an environment you don't run, can't audit, and can't revoke.

PROTECTED

AI Guard redacts or encrypts sensitive spans before upload. Policy applies regardless of which tool the employee uses, sanctioned or not. Protection is in the content, not the channel.

RAG retrieval gap
Retrieval gap

RAG pipeline with appropriate read permissions retrieves content a user was never intended to see. Backend access controls don't consistently carry through to the retrieval layer. The document was restricted for the human. The model accessed it anyway.

Enforced

Protection persists through chunking, embedding, and retrieval. Encrypted spans stay encrypted and the model reads around them. Sensitive content remains inaccessible even when the pipeline reaches the source document.

Agent permission creep
Over-scoped

Agent provisioned with broader access than necessary. Access rarely revisited. The agent acts on data no human in that role would be authorized to see, and no alert fires because the access was technically granted.

Scoped

Per-agent policies enforced at the span level continuously and not just at setup. Agents access only what they're authorized for. If an agent's scope doesn't include compensation data, it cannot surface it, regardless of what system it's connected to.

Compliance audit
No lineage

Standard logs track who opened a file. They don't capture what a RAG pipeline retrieved, what a copilot surfaced, or what an agent acted on across a multi-step workflow. The regulator asks for AI-specific data flow evidence. You have file access logs.

Full lineage

Span-level tamper-proof audit trail logs every AI interaction, such as retrieval, inference, and agent steps, by identity, time, and workflow. The specific, verifiable chain regulators require for HIPAA, GDPR, ISO 42001, and EU AI Act.

Case in point

Publicly reported · Vercel / Context.ai · April 2026

The AI tool was approved. The access it granted wasn't.

01 — What happened

A Vercel employee signed up for an AI productivity tool, Context.ai, using their corporate Google Workspace account and granted it “Allow All” OAuth permissions. Context.ai was subsequently compromised. The attacker used the stolen OAuth tokens to take over the employee's account and move laterally into Vercel's internal systems, accessing employee records, API keys, and internal credentials.

02 — The gap

The entry point was an AI tool an employee had integrated into their workflow with permissions that were never reviewed, scoped, or revoked. The attacker had the same access the AI tool had, which was access to everything.

“No exploit. No zero-day. Just an unsanctioned AI tool, an overpermissioned OAuth grant. Your employees are doing the same things on their machines right now. The question is whether you know about it.”

David Lindner, CISO, Contrast Security · Dark Reading, April 2026
580
employee records accessed in the breach, including names, account statuses, and internal credentials
$2M
ransom demanded by the attacker for the stolen data
0
technical controls on what the AI tool could access once it had been granted permissions

Access controls govern what users can reach. They don't govern what AI tools do with that access once it's granted. Confidencial enforces data-layer protection so that even a fully compromised AI tool can only surface encrypted spans, and not readable sensitive content. The OAuth token gets stolen. The data inside stays locked.

Why existing tools fall short

Visibility isn't the same as control.

DSPMs classify. DLP watches uploads. AI firewalls filter outputs. None of them protect what the model sees, what the agent retrieves, or what ends up in a response. Confidencial operates at the data layer, which is the only layer that travels with the content through every AI workflow.

ScenarioDSPMDLP / AI FirewallVector DB vendorsConfidencial
RAG pipeline surfaces content a user couldn't access directly Classifies source files. No enforcement at the retrieval layer.~ May flag outputs after the fact. Can't prevent retrieval.~ Protects embeddings partially. No span-level access policy. Encrypted spans persist through chunking and retrieval. The model can't surface what it can't read.
Employee uploads a client file to an unsanctioned AI tool~ Classifies the file at rest. No control after upload.~ May block known tools. Can't intercept personal accounts or new apps. No coverage outside the vector store. Sensitive spans encrypted at the source. Protection travels with the file regardless of tool or account.
AI agent accesses data beyond its authorized scope No agent identity governance. No visibility into agent-level data access. No per-agent access policy. Per-agent least-privilege enforced at the span level — continuously, not just at setup.
Prove to a regulator what data AI workflows accessed and when~ Posture reports. File-level classification logs.~ Upload and block events. No AI-specific data flow tracking.~ Limited vector access logs. No span-level detail. Span-level, tamper-proof audit trail by identity, workflow, and timestamp. Regulator-ready.
DSPMs get you to classification. DLP watches the upload. AI firewalls review the output. None of them govern the data between ingestion and response. Confidencial is the only layer that protects the data itself, so what the model retrieves, what the agent surfaces, and what ends up in a response is ciphertext, not sensitive content.
Common questions

Hard questions. Direct answers.

01

Our RAG pipeline only connects to files employees are authorized to access. Isn't that enough?

Authorization to access a file and authorization for an AI pipeline to retrieve specific content from it are different controls, and most systems don't enforce the second one. A user restricted from opening a document directly can still have a RAG assistant retrieve and surface content from it, because backend access controls don't consistently carry through to the retrieval layer.

02

We have an approved AI tool list. Doesn't that handle shadow AI?

A policy that says “use only approved tools” is not a technical control. 89% of AI usage in organizations is invisible to security, happening outside any visibility or audit trail, despite policies. Confidencial protects the data itself, so a file uploaded to an unsanctioned tool arrives with sensitive spans already encrypted. The tool doesn't matter. The protection is in the content.

03

How do we govern AI agents that act autonomously across multiple systems?

AI agents are non-human identities and should be governed as such, with explicitly scoped, continuously enforced access rather than inherited permissions provisioned once at configuration. Confidencial applies per-agent policies at the span level, enforced at query time on every workflow step. If an agent's scope doesn't include compensation data, it cannot surface that content, regardless of what system it's connected to.

04

We already have DLP. Why isn't that sufficient?

DLP watches uploads and known file-transfer paths. It doesn't see what a RAG pipeline retrieves, what a copilot surfaces in a response, what an agent acts on across a multi-step workflow, or what ends up in a prompt history or vector database. Confidencial operates inside the content itself, so protection persists through every stage of an AI workflow, not just at the upload boundary.

05

What does AI-specific governance evidence look like for a regulatory examination?

Span-level, tamper-proof audit trails that record every retrieval, inference, and agent action by identity, timestamp, workflow step, and data accessed. The specific, verifiable chain of what AI systems did with sensitive data, sufficient for HIPAA, GDPR, ISO 42001, NIST AI RMF, EU AI Act, and CCPA. The standard logs your existing stack generates don't capture AI-specific data flows at this level.

06

We're still in early AI adoption. Is this relevant now?

The exposure is already happening. 77% of employees regularly paste company data into generative AI tools, and 82% of those interactions happen through personal accounts outside your visibility. The gap between what your policies say and what's actually happening widens with every new tool, agent, and workflow you add. Governance at the data layer scales with adoption. Governance by policy doesn't.

Map your AI exposure

Which files in your knowledge bases are readable to any model that can reach them?

Clinical data. Client records. Compensation files. IP. Most organizations can't answer that question for their AI workflows. The AI Data Exposure Assessment maps it, including what your pipelines can access, what your agents are authorized for, and where governance stops applying.