Skip to content

PII Sanitization ​

Automatically detect and remove personally identifiable information (PII) from documents before they reach an AI model — and restore original values in the response.

Overview ​

The PII Sanitization module sits between your application and the LLM provider. It scans documents for sensitive data (names, emails, phone numbers, financial identifiers, credentials), replaces each value with a surrogate token like [PERSON_1], and sends the cleaned version to the AI. When the AI responds, the module restores the original values.

Why it matters:

  • Prevents PII from being sent to third-party AI providers
  • Meets GDPR, HIPAA, and SOC 2 requirements for data minimization
  • Works transparently — the AI still produces useful results because token format is preserved
  • Sessions are encrypted at rest, have a configurable TTL, and can be explicitly purged when no longer needed

How it works ​

┌──────────┐    ┌───────────┐    ┌───────────┐    ┌─────────┐
│ Document │───>│  Detect   │───>│ Tokenize  │───>│   LLM   │
│          │    │ (Presidio │    │ [PERSON_1]│    │ Provider│
│          │    │  + spaCy) │    │ [EMAIL_1] │    │         │
└──────────┘    └───────────┘    └───────────┘    └────┬────┘
                                                       │
┌──────────┐    ┌───────────┐                          │
│ Original │<───│  Restore  │<─────────────────────────┘
│ Response │    │ (variant  │    LLM response with tokens
│          │    │  matching)│
└──────────┘    └───────────┘
  1. Detect — Presidio and spaCy NER models scan the document for PII entities
  2. Tokenize — Each detected value is replaced with a format-preserving surrogate token ([PERSON_1], [EMAIL_ADDRESS_2])
  3. Store — The mapping between tokens and original values is encrypted with AES-256-GCM and stored in an isolated session
  4. Send — The sanitized document is forwarded to the LLM provider
  5. Restore — When the LLM responds, surrogate tokens are matched (including variants the LLM may have introduced) and replaced with original values
  6. Purge — Session data and encryption keys are destroyed on explicit purge; expired sessions are cleaned up automatically by backend retention jobs

Layered detection ​

Shield combines automation with precise human control across three layers. Together they balance broad coverage with the judgement only you have about your own data:

  1. Built-in detection (automatic). Presidio + spaCy detect common PII — names, emails, phones, addresses, IBANs, German social-security numbers, credit cards, API keys (incl. Google AIza… and Anthropic sk-ant-…).
  2. AI-assisted detection (automatic, context-aware). An internal model (GLiNER) finds context-dependent sensitive data the rules miss, from a few examples you provide. See Example-based entities.
  3. Your own rules (manual, precise). Define the formats and columns specific to your business — see Custom entities and Spreadsheet handling. You can add as many patterns and sensitive columns as you need.

Why a manual layer? Conventions like customer-number or staff-ID formats vary per company and depend on context no general model should guess at. Defining them explicitly guarantees they are always protected.

Coverage & limitations ​

Sanitization only works on readable text. Every layer above — Presidio, spaCy, GLiNER, and your custom rules — operates on the text layer of a document. The sanitizer never sees pixels or image content. As a result:

  • Images (PNG, JPG, screenshots, photos) are not analyzed. Any personal or confidential information visible in the picture is not detected and may reach the AI provider unprotected.
  • Scanned or non-OCR PDFs — PDFs that are really just images of pages, with no embedded text layer — yield no extractable text, so they are not analyzed either.

For a document to be protected it must contain a machine-readable text layer (for example, a native digital document or an OCR-processed PDF). If you only have a scanned document, run it through OCR first so the text becomes machine-readable.

Unreadable attachment policy ​

Administrators choose what happens when an attachment cannot be analyzed, in Shield → Settings → Unreadable attachment policy:

PolicyBehavior
Warn onlyThe attachment is processed and a coverage warning is attached. The send is never held.
Require acknowledgement (default)The send is held until you acknowledge that the attachment is sent unprotected.
BlockUnreadable attachments are never allowed through. There is no override.

The company-wide default can be overridden per scope on a Data Protection Profile (Company → Project → User). In chat, the policy surfaces as a warning banner with a Send anyway option (when acknowledgement is allowed). On the external API there is no interactive prompt: Warn only returns the warning in the response and the X-Shield-Coverage-Warning header; Require acknowledgement returns 422 unless you pass acknowledge_unprotected_attachments: true; Block always returns 422. See Shield → Coverage & limitations for the full workflow.

Sensitivity profiles ​

Profiles define which categories of PII the sanitizer detects. You can select one or more profiles per request.

ProfileDetectsUse case
personal_informationNames, emails, phone numbers, addresses, dates of birthGeneral PII protection
financial_dataCredit cards, IBANs, SSNs, tax IDs, bank account numbersFinancial document processing
health_informationMedical license numbers, national registration numbersHealthcare and insurance
credentialsAPI keys, connection strings, passwordsDeveloper documents, config files
corporate_confidentialOrganization names, IP addresses, internal URLsInternal memos, infrastructure docs

Profiles are additive — selecting personal_information and financial_data detects entity types from both profiles.

Replacement style: tokens vs realistic surrogates ​

When Shield masks non-credential PII, you choose how the masked value looks in the text sent to the AI. There are two styles:

StyleExample outputLooks like
Placeholder tokens (default)[PERSON_1], [EMAIL_2]An opaque bracketed label
Realistic surrogatesJon Doe_1, iris.dale@example.org_1, Northgate_1, Globex_2, Project AtlasA natural-looking fake of the same type

Both styles are fully reversible — Shield restores the real values in the model's response either way. Neither style ever sends a real value.

Before / after example ​

Original text:

text
Sarah Thompson met Michael Chen at the Munich office.

With placeholder tokens:

text
[PERSON_1] met [PERSON_2] at the [LOCATION_1] office.

With realistic surrogates:

text
Nora Wren_1 met Olivia Knox_2 at the Northgate_1 office.

In both cases, when the AI replies, Shield swaps the masked values back so you see Sarah Thompson, Michael Chen and Munich in the final result.

Why realistic surrogates exist ​

Large language models reason and write better over natural-looking text than over bracket tokens. A sentence full of [PERSON_1] and [LOCATION_1] reads like a fill-in-the-blank form, which can degrade the quality of summaries, drafts, and analysis. Realistic surrogates keep the prose readable, so the model produces better output — while the real data still never leaves your side.

Realistic surrogates are safe by design:

  • Zero leak — fakes are drawn from a curated pool, indexed by a non-invertible SHA-256 hash of the original value. The fake reveals nothing about the real one.
  • Reserved fake domains — surrogate emails use RFC 2606 reserved domains (e.g. example.org), so they can never reach a real inbox.
  • Deterministic and consistent — the same input always maps to the same surrogate within and across a conversation. The trailing _<n> id (e.g. Nora Wren_1) keeps coreference stable, so the model knows two mentions are the same person.

When to use realistic vs tokens ​

  • Use realistic surrogates when output quality matters on natural-language tasks — summarization, drafting emails or reports, rewriting, narrative analysis. The model handles Nora Wren_1 far more naturally than [PERSON_1].
  • Use placeholder tokens when you want the masking to be unmistakable to a human reviewer, when you are auditing exactly what was redacted, or for highly structured/extraction tasks where a clearly-marked slot is preferable.

Where to set it ​

Shield resolves the replacement style with the same most-specific-wins logic as everything else:

  1. Company-wide default — Shield → Settings → Replacement style card. Pick Placeholder tokens or Realistic surrogates for the whole organization.
  2. Per-category override — a chat category can pin a style for everything routed through it.
  3. Per-prompt choice — a saved prompt can carry its own style.

If nothing more specific is set, the company-wide default applies.

Sanitization Profiles & Per-User / Per-Scope Configuration ​

You rarely have to configure sanitization for every message. Shield resolves an effective profile for each request by layering several levels of configuration, from broad organizational defaults down to the exact toggles you pick for one message. This means different users and different projects can run completely different sanitization profiles — all without anyone re-entering the same rules over and over.

The precedence chain (highest precedence first) ​

#LevelWhere it's setTypical owner
1Per-request selectionsThe toggles, profiles, column rules and custom recognizers you choose in the chat for this messageAny user, ad-hoc
2Chat / session configSticky settings that stay applied for the whole conversationThe user, per conversation
3Prompt-level configA sanitization profile attached to a saved promptPrompt author
4Project-level configA default profile for everything in a projectProject owner
5Company-level policyThe organization-wide default profileAdmin / CISO

How the layers combine. Lower levels are inherited; higher levels override or augment them. The merge is additive:

  • Profiles are unioned — a project that adds financial_data is combined with, not replaced by, the company's personal_information.
  • Column rules merge by column, with the more specific level winning on a conflict (e.g. a per-request rule for a column beats the project's rule for the same column).
  • Custom recognizers are concatenated and de-duplicated, so the same pattern defined at two levels is only applied once.

On top of everything sits an admin redlist (always-mask floor) enforced inside the sanitizer itself — values on the redlist are always protected regardless of which profile resolves.

Worked example ​

Imagine three levels stacked for one chat message:

  1. Company default masks PERSON and EMAIL for every request in the organization.
  2. The Marketing project this prompt belongs to adds LOCATION to the mix — so anything run inside that project masks names, emails and locations.
  3. For one message, a user pastes a planning doc and adds a per-request custom rule for an internal project code:
json
{
  "label": "Internal Project Code",
  "type": "pattern",
  "pattern": "PROJ-\\d{4}-\\d{3}"
}

The effective profile for that single message is the union of all three: PERSON + EMAIL (company) + LOCATION (project) + the ad-hoc PROJ-2024-001 style codes (per-request) — plus anything on the company redlist. A colleague in a different project, or the same user on their next message without the ad-hoc rule, gets a different effective profile automatically.

Per-request rule-sets ​

Anything you can persist in a profile can also be supplied for a single request, layered on top of the saved configuration:

  • Ad-hoc column rules — designate spreadsheet/CSV columns to tokenize just for this upload (see Spreadsheet handling).
  • Custom recognizers — add a regex pattern or a few example values for an identifier specific to this document (see Custom entities).

These ad-hoc rules don't change the saved profiles; they apply once, on top of whatever the persisted Company → Project → Prompt chain resolved.

Attachments & multi-part prompts ​

Prompts that include file attachments — a CV, a PDF, or a spreadsheet alongside your typed message — are sanitized together with the message as a single request. Every part is tokenized consistently (the same value gets the same token across the message text and each attachment), and when the AI responds, every part is detokenized back correctly. You don't need to sanitize attachments separately or worry about tokens from one segment leaking unrestored into another.

Custom entities ​

When built-in profiles do not cover your data, define custom entity patterns. You can define multiple patterns in one request — in the UI, open Custom sensitive data → Add pattern for each format you want to protect.

Pattern-based (regex) ​

Define a regex pattern to match application-specific identifiers:

json
{
  "label": "Employee ID",
  "type": "pattern",
  "pattern": "EMP-\\d{6}",
  "description": "Internal six-digit employee identifier"
}

The sanitizer replaces matches with [EMPLOYEE_ID_1], [EMPLOYEE_ID_2], etc. To protect several formats at once, send several entries:

json
{
  "custom_entities": [
    { "label": "Customer Number", "type": "pattern", "pattern": "CUST-\\d{3,}" },
    { "label": "Staff ID",        "type": "pattern", "pattern": "EMP-\\d{3,}" },
    { "label": "Invoice",         "type": "pattern", "pattern": "INV-\\d{3}-\\d{3}" }
  ]
}

Example-based (GLiNER — Phase 2) ​

Provide 2–5 example values and let the system infer the pattern using zero-shot NER:

json
{
  "label": "Internal Project Code",
  "type": "examples",
  "examples": ["PROJ-2024-001", "PROJ-2024-042", "PROJ-2025-100"],
  "description": "Internal project tracking codes"
}

Example-based entities require GLiNER (Phase 2 deployment). If GLiNER is not available, the request returns a 400 error with guidance.

Spreadsheet handling ​

For CSV and XLSX files, the sanitizer supports two scanning strategies:

StrategyHow it worksBest for
column_headerInfers entity types from column headers, and tokenizes every value in a designated columnClean spreadsheets with descriptive headers
cell_scanScans every cell individually with NERUnstructured or poorly labeled data

Real-world exports often begin with banner rows (a company name, a report title, a reporting period) before the real column header. Shield auto-detects the header row beneath such a preamble, so column designation and header inference work without manual cleanup. Multi-sheet workbooks are fully processed — every sheet is sanitized, and the same value gets the same token across sheets.

  1. Upload your spreadsheet
  2. Click Detect columns to list the headers with automatic type suggestions
  3. Mark the sensitive columns — you can mark as many as you need (e.g. Customer No. → Customer number, Staff ID → Staff/employee ID)
  4. Confirm to sanitize with the finalized column rules

Every value in a designated column is tokenized — even formats that automatic detection alone would miss. You can also set column rules directly via the API to mark multiple columns:

json
{
  "options": {
    "column_rules": {
      "Customer No.": "customer_id",
      "Customer Name": "organization",
      "Order No.": "order_number",
      "Notes": "skip"
    }
  }
}

Use "skip" to exclude a column from scanning entirely.

Pasted CSV ​

You can paste CSV directly instead of uploading a file. Under the paste box, set Treat as → CSV (Shield also detects CSV-shaped text and offers a one-click switch). Pasted CSV then gets the full tabular pipeline — header-row detection, Detect columns, and multi-column designation — exactly like an uploaded .csv. For API calls, send document_type: "csv".

Review before sending ​

Shield shows you exactly what will be shared before anything leaves your side:

  • A detection summary (counts per category).
  • A masked diff — each detected value shown partially masked (e.g. S██████l) next to its token, so you can confirm a secret was caught without re-exposing it.
  • The full sanitized text that will be sent.

In Proxy (BYOK) the flow is two-step: Preview redactions → review → Confirm & send to AI. Nothing is sent to the model until you confirm.

Correcting a miss (Layer 4) ​

If the three detection layers miss something, or over-redact, you can fix it on the spot from the result view:

  • Redact something we missed — type the exact text and pick a type; every occurrence is tokenized consistently.
  • Un-redact — click any token (e.g. [ORGANIZATION_3]) to restore its original value when a non-sensitive term was caught.

For values that recur across many documents, add them to the redlist (always redact) so you never have to correct them again.

Diff preview ​

Before the sanitized document is sent to the LLM, you can review a diff preview showing exactly what will be replaced:

RowOriginal (masked)Replacement
1M██████g[PERSON_1]
1m██████@acme.com[EMAIL_ADDRESS_1]
3+49 170 ███████[PHONE_NUMBER_1]

Original values are partially masked in the preview for security — you can see enough to confirm the detection is correct without exposing the full value.

Credential detection ​

The sanitizer automatically detects credentials and secrets in your documents:

  • API keys — AWS, Azure, GCP, Stripe, GitHub, and generic key patterns
  • Connection strings — Database URLs, Redis URIs
  • Passwords — Plain-text passwords near common labels (password:, pwd=)

Detected credentials are replaced with format-preserving fakes that maintain the same structure (prefix, length, checksum) so the LLM can still reason about the document structure. For example, an AWS access key AKIA1234567890ABCD56 becomes a structurally-valid fake such as AKIA8652176294KHZH36_FAKE — the AKIA prefix and character classes are preserved, an _FAKE audit suffix is appended, and the suffix is stripped automatically on restore.

This format-preserving credential masking now applies on the chat and Guru paths too, not only in the standalone Sanitize tool — so secrets pasted into a conversation get the same protection as those uploaded to Shield.

When credentials are found, the response includes a credential_warning field alerting you to the presence of secrets.

Session lifecycle ​

Each sanitization creates an encrypted session:

active  ──(restore)──>  restored  ──(purge/TTL)──>  purged
  │                                                    ▲
  └──────────(purge/TTL)───────────────────────────────┘
StatusDescription
ActiveSession created, sanitized document ready, awaiting LLM response
RestoredLLM response received and de-tokenized at least once
PurgedSession data and encryption keys destroyed (irreversible)
  • TTL: Sessions expire automatically (default: 1 hour, configurable)
  • Manual purge: Delete a session immediately when you no longer need it
  • Encryption: All PII mappings use AES-256-GCM with per-session keys

After TTL expiry, restore should be treated as unavailable. Expired session rows are then removed by the sanitizer cleanup process.

Policy and package controls ​

The package gate is the single on/off switch ​

Whether Shield sanitization is available at all is controlled by one thing: the PII / Shield Sanitization package feature on your subscription.

  • When the feature is ON, sanitization is active everywhere it can run — chat, Guru, and batch.
  • When the feature is OFF, sanitization is unavailable, regardless of any other setting.

A routing policy can no longer silently switch sanitization off when the package feature is on. If the package is enabled and a policy leaves sanitization unset (or set to disabled), Shield defaults to automatic rather than going dark. This removes the confusing "Shield is not active" state that used to appear in Guru even though the package was enabled. A deliberate manual or automatic policy setting is still honored — the gate only ever errs toward masking more, never less.

Sanitization mode hierarchy: org → category → per-prompt ​

Beyond the on/off gate, you control how sanitization runs through a mode, resolved most-specific-wins:

text
per-prompt choice  ->  chat category setting  ->  org-wide default

The available modes are:

ModeBehavior
Per-promptDefer to the choice made on the individual prompt
StandardApply the standard sanitization profile automatically
SmartContext-aware sanitization
DisabledDo not sanitize at this level

Where to set each level:

  • Org-wide default — Admin → Compliance Config, in the new Shield Protection section (which also surfaces HIPAA recommendations).
  • Per chat category — Admin → Chat Categories, on each category.
  • Per prompt — the prompt's own sanitization choice.

Trigger sources ​

In manual mode, a routing policy can restrict which trigger sources are accepted:

  • prompt
  • api
  • user_action

UI walkthrough ​

  1. Enable sanitization — Toggle PII sanitization on from prompt execution, prompt testing, prompt comparison reruns, Guru Run It, the chat terminal, or analytics live-comparison runs
  2. Select profiles — Choose which sensitivity profiles to apply (e.g., Personal Information + Financial Data)
  3. Add custom rules — Optionally define regex patterns or example-based entities for application-specific data
  4. Review preview — Inspect the diff preview to confirm detections are accurate
  5. Confirm & send — The sanitized document is sent to the LLM; you see the cleaned version
  6. Restore — When the LLM responds, tokens are automatically replaced with original values
  7. Session cleanup — Use explicit purge when you want immediate irreversible deletion; expired sessions are removed automatically after the restore window closes

API quick reference ​

MethodEndpointPurpose
POST/api/v1/sanitizeSanitize a document
POST/api/v1/inspect-headersAuto-detect column entity types in spreadsheets
POST/api/v1/desanitizeStateless de-tokenization
GET/api/v1/sessions/{token}Get session info and sanitized content
POST/api/v1/sessions/{token}/restorePush LLM response for de-tokenization
DELETE/api/v1/sessions/{token}Purge a session

See the full PII Sanitization API Reference for request/response schemas, code examples, and error handling.

Troubleshooting ​

Session expired ​

Symptom: 404 error when trying to restore or access a session.

Cause: The session TTL has elapsed. The session is no longer restorable and may already have been removed by cleanup.

Fix: Re-sanitize the document to create a new session. To avoid this, increase TOKEN_TTL_SECONDS or restore the session promptly after receiving the LLM response.

No entities detected ​

Symptom: total_entities_found is 0 in the response.

Possible causes:

  • The selected profiles do not cover the entity types in your document
  • The document language does not match the configured language (en or de)
  • The document format is corrupted or empty

Fix: Try adding more profiles, check the language setting, or verify the document content with detect-only mode.

GLiNER not available ​

Symptom: 400 error when using example-based custom entities.

Cause: The sanitizer is running in Phase 1 mode (GLiNER disabled).

Fix: Deploy Phase 2 with GLINER_ENABLED=true and BUILD_PHASE=2. Phase 2 requires ~2-3 GB RAM.

Unresolved tokens in restored content ​

Symptom: unresolved_tokens array is non-empty after restore.

Cause: The LLM restructured the tokens in a way that could not be matched (e.g., split across lines, heavily abbreviated).

Fix: The variant matcher handles most LLM modifications (case changes, spacing, partial matches). If tokens remain unresolved, check the confidence field — "low" indicates significant restructuring. You may need to adjust your prompt to instruct the LLM to preserve token formatting.