Appearance
Data Protection Controls: Sanitization, Blacklist, Redlist & Whitelist
VeriPrompt gives administrators four complementary controls for keeping sensitive data out of AI providers. They are easy to confuse because they sound alike, so this guide explains what each one does, when to use it, who manages it, and how — and, critically, how they combine.
At a glance:
| Control | One-line job | Effect on the message |
|---|---|---|
| PII Sanitization | Detect & mask sensitive values automatically | Values are hidden, the message still goes through and is restored in the reply |
| Blacklist | Forbid specific content outright | The message is blocked — it never reaches the model |
| Redlist | Always mask specific terms | The term is always hidden (a floor users can't switch off) |
| Whitelist | Never mask specific terms | The term is never hidden — removes false positives |
A useful mental model: the blacklist is a gate ("you can't say this at all"), while the redlist and whitelist tune the masking ("always hide this" / "never hide this") on top of the automatic detection.
1. PII Sanitization (the foundation)
PII Sanitization sits between your users and the LLM provider. For every prompt it:
- Detects sensitive values — names, emails, phone numbers, addresses, financial identifiers, credentials, and more.
- Masks each value with a surrogate token such as
[PERSON_1]or[EMAIL_2]. - Sends the cleaned text to the AI provider — the real values never leave the platform.
- Restores the original values in the model's response before the user sees it.
Sessions are encrypted at rest, carry a configurable TTL, and can be purged on demand.
How detection works
- Tier 1 (always on): Microsoft Presidio + spaCy NER for pattern- and language-based detection, plus dedicated recognizers for credentials/secrets.
- Tier 2 (optional): example-based zero-shot detection (GLiNER) for organization-specific entity types.
- Credentials are always detected and replaced with format-preserving fakes (a fake that keeps the shape of a key/token but contains no real secret).
Detection sensitivity
Automatic detection is tunable with a single company-wide Detection sensitivity slider in Shield → Settings. It sets the minimum confidence a detection needs before a value is masked:
- Lower → catches more potential PII, but flags more false positives (over-redaction).
- Higher → masks only higher-confidence matches — cleaner output, fewer false positives.
- Default: 0.4, which suits most teams.
The slider is capped so that names, emails and credentials stay masked at every setting — only weak-signal values (e.g. loosely-formatted phone digits) are dropped at the strict end. Use the slider first when masking feels too aggressive; reach for the whitelist only for specific recurring false positives.
The sensitivity setting is company-wide and admin-controlled. See also the Redlist (force more masking) and Whitelist (remove specific false positives) below.
2. Blacklist — block the message outright
The chat blacklist is a content gate. When a user's message contains a blacklisted term, the message is rejected before it ever reaches the model.
| Property | Value |
|---|---|
| Action | Blocks the entire message |
| Runs | First, in the chat path (pre-LLM) |
| Matching | Case-insensitive substring (plain keywords/phrases, not regex) |
| Scope | Per company, optionally narrowed to specific classifications |
| On a match | Returns a "blocked" response, writes a BLACKLIST_VIOLATION audit entry, and (optionally) notifies admins |
| Managed in | /admin/chat-blacklists |
| Who | SUPER_ADMIN, SAAS_ADMIN, ACCOUNT_OWNER, ADMIN |
How to manage it — two ways:
- Online form: name the list, add keyword tags, choose which classifications it applies to (empty = all), and toggle "notify admin on detection."
- Bulk import: download a CSV or XLSX template, fill it in, and upload with one of three modes — Create-only, Append, or Replace. A violations dashboard lets you filter by date and export / archive the records.
Use the blacklist when content must never be processed at all — forbidden topics, competitor confidential terms, compliance no-go phrases. It protects against the request itself, not just the data inside it.
3. Redlist — always mask (the floor)
The redlist forces specific terms to always be masked, on top of automatic detection. It is a floor: individual users cannot switch it off — their own rules can only add more masking. The message still goes through; the term is simply always hidden and restored afterward.
| Property | Value |
|---|---|
| Action | Always masks the term (tokenized, then restored in the reply) |
| Runs | Inside the sanitizer, as part of detection |
| Matching | exact term, pattern (regex), or column_header (redacts a whole spreadsheet column) |
| Scope | Company (workspace-wide floor) or a specific project |
| Stored in | The sanitizer service (not the main database) |
| Managed in | Shield → Settings |
| Who | SUPER_ADMIN, SAAS_ADMIN, ACCOUNT_OWNER, ACCOUNT_ADMIN, ADMIN |
How to manage it: an inline add form (choose type → value → label → company/project scope → Add), one entry at a time.
Use the redlist when automatic detection might miss a term you must always protect — internal customer IDs, project codenames, contract numbers, or any organization-specific identifier that isn't standard PII.
4. Whitelist — never mask (the false-positive killer)
The whitelist is the exact inverse of the redlist. Any detected value that matches a whitelisted term is dropped from masking and passes through to the model in the clear.
| Property | Value |
|---|---|
| Action | Never masks the term (excluded from redaction) |
| Runs | Last — after automatic detection and the redlist, so it overrides detection |
| Matching | Exact, case-insensitive term |
| Scope | Company or project |
| Stored in | The sanitizer service |
| Managed in | Shield → Settings |
| Who | Same admin roles as the redlist |
How to manage it: a simple add form (value → Add).
Use the whitelist sparingly — only for specific, recurring false positives. The classic case: your own company or product name keeps getting flagged as an organization and redacted. Whitelist it so it stops being masked. For broad over-masking, prefer the Detection sensitivity slider instead.
How they combine (order of operations)
For a chat message, the controls apply in this order:
1. BLACKLIST → if any term matches, BLOCK the message. Stop. (Nothing below runs.)
2. SANITIZATION → detect PII (tuned by the sensitivity slider)
+ REDLIST always-mask terms
+ any per-user / per-request custom rules (additive)
3. WHITELIST → remove known false positives (applied last)
4. Send the sanitized prompt to the model → restore originals in the replyKey implications:
- The blacklist short-circuits everything — a blocked message is never sanitized or sent.
- The redlist and automatic detection are additive — they can only add masking.
- The whitelist is applied last and overrides detection — which is why it should be used sparingly. (As a safety improvement, redlist terms are being made exempt from whitelist removal, so the floor can never be accidentally un-masked.)
Which control should I use?
| Goal | Use |
|---|---|
| This content must never be processed at all | Blacklist |
| This specific term must always be hidden from the model | Redlist |
| This term keeps getting wrongly redacted | Whitelist |
| Masking is generally too aggressive (or not aggressive enough) | Detection sensitivity slider |
| I want standard PII (names, emails, etc.) handled automatically | PII Sanitization (on by policy) |
On the roadmap
These improvements are in development and not yet released — listed here so administrators know what's coming:
- Redlist import/export as Excel — bringing the same CSV/XLSX template + Create/Append/Replace import (and export) the blacklist already has to the redlist, so long lists of always-mask terms can be loaded in bulk.
- Replacement terms & realistic surrogates — instead of an opaque
[PERSON_1]token, masking a value with a realistic, type-faithful stand-in that carries a stable ID (e.g.Jon Doe_1). This keeps the prompt natural and answerable for the model while remaining uniquely recognizable and fully reversible, and protects exactly as well as a token (no real value ever leaves the platform). Redlist entries will be able to specify their own replacement; everything else falls back to a clear typed placeholder.
Related documentation
- PII Sanitization — the detection & masking engine in depth
- Shield — the Rotate / Proxy / Gateway product flows
- CISO & DPO Guide — security-leadership control points
