Skip to content

Data Protection Controls: Sanitization, Blacklist, Redlist & Whitelist ​

VeriPrompt gives administrators four complementary controls for keeping sensitive data out of AI providers. They are easy to confuse because they sound alike, so this guide explains what each one does, when to use it, who manages it, and how — and, critically, how they combine.

At a glance:

ControlOne-line jobEffect on the message
PII SanitizationDetect & mask sensitive values automaticallyValues are hidden, the message still goes through and is restored in the reply
BlacklistForbid specific content outrightThe message is blocked — it never reaches the model
RedlistAlways mask specific termsThe term is always hidden (a floor users can't switch off)
WhitelistNever mask specific termsThe term is never hidden — removes false positives

A useful mental model: the blacklist is a gate ("you can't say this at all"), while the redlist and whitelist tune the masking ("always hide this" / "never hide this") on top of the automatic detection.


1. PII Sanitization (the foundation) ​

PII Sanitization sits between your users and the LLM provider. For every prompt it:

  1. Detects sensitive values — names, emails, phone numbers, addresses, financial identifiers, credentials, and more.
  2. Masks each value with a surrogate token such as [PERSON_1] or [EMAIL_2].
  3. Sends the cleaned text to the AI provider — the real values never leave the platform.
  4. Restores the original values in the model's response before the user sees it.

Sessions are encrypted at rest, carry a configurable TTL, and can be purged on demand.

How detection works ​

  • Tier 1 (always on): Microsoft Presidio + spaCy NER for pattern- and language-based detection, plus dedicated recognizers for credentials/secrets.
  • Tier 2 (optional): example-based zero-shot detection (GLiNER) for organization-specific entity types.
  • Credentials are always detected and replaced with format-preserving fakes (a fake that keeps the shape of a key/token but contains no real secret).

Detection sensitivity ​

Automatic detection is tunable with a single company-wide Detection sensitivity slider in Shield → Settings. It sets the minimum confidence a detection needs before a value is masked:

  • Lower → catches more potential PII, but flags more false positives (over-redaction).
  • Higher → masks only higher-confidence matches — cleaner output, fewer false positives.
  • Default: 0.4, which suits most teams.

The slider is capped so that names, emails and credentials stay masked at every setting — only weak-signal values (e.g. loosely-formatted phone digits) are dropped at the strict end. Use the slider first when masking feels too aggressive; reach for the whitelist only for specific recurring false positives.

The sensitivity setting is company-wide and admin-controlled. See also the Redlist (force more masking) and Whitelist (remove specific false positives) below.


2. Blacklist — block the message outright ​

The chat blacklist is a content gate. When a user's message contains a blacklisted term, the message is rejected before it ever reaches the model.

PropertyValue
ActionBlocks the entire message
RunsFirst, in the chat path (pre-LLM)
MatchingCase-insensitive substring (plain keywords/phrases, not regex)
ScopePer company, optionally narrowed to specific classifications
On a matchReturns a "blocked" response, writes a BLACKLIST_VIOLATION audit entry, and (optionally) notifies admins
Managed in/admin/chat-blacklists
WhoSUPER_ADMIN, SAAS_ADMIN, ACCOUNT_OWNER, ADMIN

How to manage it — two ways:

  • Online form: name the list, add keyword tags, choose which classifications it applies to (empty = all), and toggle "notify admin on detection."
  • Bulk import: download a CSV or XLSX template, fill it in, and upload with one of three modes — Create-only, Append, or Replace. A violations dashboard lets you filter by date and export / archive the records.

Use the blacklist when content must never be processed at all — forbidden topics, competitor confidential terms, compliance no-go phrases. It protects against the request itself, not just the data inside it.


3. Redlist — always mask (the floor) ​

The redlist forces specific terms to always be masked, on top of automatic detection. It is a floor: individual users cannot switch it off — their own rules can only add more masking. The message still goes through; the term is simply always hidden and restored afterward.

PropertyValue
ActionAlways masks the term (tokenized, then restored in the reply)
RunsInside the sanitizer, as part of detection
Matchingexact term, pattern (regex), or column_header (redacts a whole spreadsheet column)
ScopeCompany (workspace-wide floor) or a specific project
Stored inThe sanitizer service (not the main database)
Managed inShield → Settings
WhoSUPER_ADMIN, SAAS_ADMIN, ACCOUNT_OWNER, ACCOUNT_ADMIN, ADMIN

How to manage it: an inline add form (choose type → value → label → company/project scope → Add), one entry at a time.

Use the redlist when automatic detection might miss a term you must always protect — internal customer IDs, project codenames, contract numbers, or any organization-specific identifier that isn't standard PII.


4. Whitelist — never mask (the false-positive killer) ​

The whitelist is the exact inverse of the redlist. Any detected value that matches a whitelisted term is dropped from masking and passes through to the model in the clear.

PropertyValue
ActionNever masks the term (excluded from redaction)
RunsLast — after automatic detection and the redlist, so it overrides detection
MatchingExact, case-insensitive term
ScopeCompany or project
Stored inThe sanitizer service
Managed inShield → Settings
WhoSame admin roles as the redlist

How to manage it: a simple add form (value → Add).

Use the whitelist sparingly — only for specific, recurring false positives. The classic case: your own company or product name keeps getting flagged as an organization and redacted. Whitelist it so it stops being masked. For broad over-masking, prefer the Detection sensitivity slider instead.


How they combine (order of operations) ​

For a chat message, the controls apply in this order:

1. BLACKLIST    → if any term matches, BLOCK the message. Stop. (Nothing below runs.)
2. SANITIZATION → detect PII (tuned by the sensitivity slider)
                  + REDLIST always-mask terms
                  + any per-user / per-request custom rules (additive)
3. WHITELIST    → remove known false positives (applied last)
4. Send the sanitized prompt to the model → restore originals in the reply

Key implications:

  • The blacklist short-circuits everything — a blocked message is never sanitized or sent.
  • The redlist and automatic detection are additive — they can only add masking.
  • The whitelist is applied last and overrides detection — which is why it should be used sparingly. (As a safety improvement, redlist terms are being made exempt from whitelist removal, so the floor can never be accidentally un-masked.)

Which control should I use? ​

GoalUse
This content must never be processed at allBlacklist
This specific term must always be hidden from the modelRedlist
This term keeps getting wrongly redactedWhitelist
Masking is generally too aggressive (or not aggressive enough)Detection sensitivity slider
I want standard PII (names, emails, etc.) handled automaticallyPII Sanitization (on by policy)

On the roadmap ​

These improvements are in development and not yet released — listed here so administrators know what's coming:

  • Redlist import/export as Excel — bringing the same CSV/XLSX template + Create/Append/Replace import (and export) the blacklist already has to the redlist, so long lists of always-mask terms can be loaded in bulk.
  • Replacement terms & realistic surrogates — instead of an opaque [PERSON_1] token, masking a value with a realistic, type-faithful stand-in that carries a stable ID (e.g. Jon Doe_1). This keeps the prompt natural and answerable for the model while remaining uniquely recognizable and fully reversible, and protects exactly as well as a token (no real value ever leaves the platform). Redlist entries will be able to specify their own replacement; everything else falls back to a clear typed placeholder.