Skip to content

PII Sanitization API ​

The Sanitizer service detects, tokenizes, and de-tokenizes personally identifiable information in documents before they are sent to LLM providers.

  • Base URL: https://<your-domain>:8100
  • Content-Type: application/json
  • Interactive docs: GET /docs (Swagger UI)

Authentication & sessions ​

The sanitizer uses session tokens to track sanitization state. Tokens are generated when a document is sanitized and required for subsequent de-tokenization.

  • Format: stk_ followed by 32 hex characters (e.g., stk_a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6)
  • TTL: Configurable via TOKEN_TTL_SECONDS (default: 3600s)
  • Encryption: AES-256-GCM per-session keys, destroyed on purge

TTL expiry and explicit purge are different lifecycle events:

  • TTL expiry ends guaranteed restore availability and the session is later removed by cleanup
  • Explicit purge immediately destroys the encrypted mapping, encrypted sanitized document, and session key

POST /api/v1/sanitize ​

Detect and replace PII in a document with surrogate tokens.

Request body ​

FieldTypeRequiredDefaultDescription
documentstringYes—Base64-encoded document content
document_typestringYes—"docx", "pdf", "csv", or "xlsx"
profilesstring[]No["personal_information"]Sensitivity profiles to apply
custom_entitiesCustomEntityDef[]No[]User-defined entity patterns
optionsSanitizeOptionsNodefaultsProcessing options

SanitizeOptions:

FieldTypeDefaultDescription
modestring"sanitize""sanitize" or "detect_only"
spreadsheet_strategystring"cell_scan""column_header" or "cell_scan"
column_rulesobject{}Manual column-to-entity mappings
preview_rowsinteger25Rows to include in diff preview
languagestring"en""en" or "de"

CustomEntityDef:

FieldTypeRequiredDescription
labelstringYesDisplay label (e.g., "Employee ID")
typestringYes"pattern" (regex) or "examples" (Phase 2 GLiNER)
patternstringConditionalRegex pattern (required if type is "pattern")
examplesstring[]ConditionalExample values (required if type is "examples")
descriptionstringNoHuman-readable description

Response (200) ​

FieldTypeDescription
session_tokenstring | nullSession token (null if detect_only)
expires_atstring | nullISO 8601 expiration timestamp
sanitized_documentstring | nullBase64-encoded sanitized document
protected_documentstring | nullBase64 provider transport with nonce-scoped protected-reference wrappers
protected_contextobjectProvider-safe wrapped versions of sanitized context segments
protection_manifestarraySafe replacement ids, values, types, and refs; never contains originals
detection_summaryobject{ total_entities_found, by_category }
diff_previewobject | nullPreview of changes
credential_warningstring | nullAlert if credentials were detected
extractionobject | nullPDFs only: what extraction left out. hidden_chars (invisible, white, tiny, off_page), images_removed, pages_without_text, ocr_pages, annotations, form_fields, embedded_files, javascript, metadata_fields, page_count. Counts only.

PDFs are extracted to Markdown, one block per page: headings, lists and tables (each cell once). Hidden text (invisible, white, tiny or off-page) is removed and replaced by [Hidden text removed], images by [Image removed], and pages without a text layer by a marker. A PDF with no text layer at all yields an empty document. On a scanned page with an OCR layer (at least 80% of its text is invisible text over an image) that text is kept and sanitized, under the marker [Scanned page: this text comes from its OCR layer and was not checked against the image], and the page is listed in ocr_pages.

Errors ​

CodeCondition
400Invalid document_type, malformed base64, empty document
422Missing required fields
500Internal processing error

cURL example ​

bash
curl -X POST https://sanitizer.veriprompt.tech/api/v1/sanitize \
  -H "Content-Type: application/json" \
  -d '{
    "document": "'$(base64 < employees.csv)'",
    "document_type": "csv",
    "profiles": ["personal_information", "financial_data"],
    "options": {
      "spreadsheet_strategy": "column_header",
      "column_rules": {"SSN": "US_SSN"},
      "preview_rows": 10
    }
  }'

TypeScript example ​

typescript
const response = await fetch('https://sanitizer.veriprompt.tech/api/v1/sanitize', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    document: btoa(csvContent),
    document_type: 'csv',
    profiles: ['personal_information'],
    options: { spreadsheet_strategy: 'column_header' }
  })
});

const { session_token, protected_document, protection_manifest, detection_summary } = await response.json();
console.log(`Detected ${detection_summary.total_entities_found} entities`);
// Send protected_document—not the readable preview—to an external model.

POST /api/v1/inspect-headers ​

Auto-detect column entity types from spreadsheet headers. Use this before sanitization to let users confirm or override column mappings.

Request body ​

FieldTypeRequiredDescription
documentstringYesBase64-encoded spreadsheet
document_typestringYes"csv" or "xlsx"

Response (200) ​

json
{
  "columns": [
    {"column_name": "Full Name", "column_index": "A", "suggested_entity_type": "PERSON", "confidence": "high"},
    {"column_name": "Email", "column_index": "B", "suggested_entity_type": "EMAIL_ADDRESS", "confidence": "high"},
    {"column_name": "Department", "column_index": "C", "suggested_entity_type": null, "confidence": null}
  ],
  "sheet_names": []
}

cURL example ​

bash
curl -X POST https://sanitizer.veriprompt.tech/api/v1/inspect-headers \
  -H "Content-Type: application/json" \
  -d '{
    "document": "'$(base64 < employees.csv)'",
    "document_type": "csv"
  }'

POST /api/v1/desanitize ​

Stateless de-tokenization — replace surrogate tokens with original values using an existing session.

Request body ​

FieldTypeRequiredDefaultDescription
contentstringYes—Text containing surrogate tokens
session_tokenstringYes—Session token from the sanitize call
delete_afterbooleanNofalsePurge session after de-tokenization

Response (200) ​

json
{
  "restored_content": "The report for Max Berg shows max.berg@example.com as the primary contact.",
  "tokens_restored": 2,
  "unresolved_tokens": []
}

cURL example ​

bash
curl -X POST https://sanitizer.veriprompt.tech/api/v1/desanitize \
  -H "Content-Type: application/json" \
  -d '{
    "content": "The report for [PERSON_1] shows [EMAIL_1] as the primary contact.",
    "session_token": "stk_a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6"
  }'

GET /api/v1/sessions/ ​

Retrieve session metadata and sanitized content.

Response (200) ​

FieldTypeDescription
session_tokenstringThe session token
statusstring"active" or "restored"
sanitized_documentstringBase64-encoded sanitized content
protected_documentstring | nullBase64 provider transport with protected refs
protected_contextobjectWrapped provider context segments
protection_manifestarraySafe protected-reference manifest
document_typestringOriginal document type
detection_summaryobjectEntity detection summary
profiles_appliedstring[]Profiles used
created_atstringISO 8601 creation timestamp
expires_atstringISO 8601 expiration
ttl_remaining_secondsintegerSeconds until expiry

Errors ​

CodeCondition
404Session not found or purged

POST /api/v1/sessions/{token}/restore ​

Push an LLM response back through the session for restoration. Use strict mode for provider output; it restores only complete nonce-bound wrappers and never guesses from ordinary prose.

Request body ​

FieldTypeRequiredDefaultDescription
contentstringYes—LLM response text or base64 content
content_typestringNo"text""text" or "base64"
delete_afterbooleanNofalsePurge session after restore
restoration_modestringNo"fuzzy""strict" for provider output; "fuzzy" is legacy sidecar compatibility only

Response (200) ​

FieldTypeDescription
session_tokenstringThe session token
restored_contentstringDe-tokenized content
tokens_restoredintegerTotal replacements made
tokens_foundarrayDetails of each match with variant_found, canonical_token, match_category
unresolved_tokensstring[]Tokens not found in mapping
safe_unrestoredstring[]Known synthetic values left bare and therefore not restored in strict mode
restoration_statusstringrestored, not_applicable, safe_unrestored, or incomplete
confidencestring"high", "medium", or "low"

Variant match categories:

CategoryDescriptionExample
exactToken matches exactly[PERSON_1] → [PERSON_1]
caseDiffers in letter case[person_1] or [Person_1]
spacingUnderscores replaced with spaces/hyphens[PERSON 1] or [PERSON-1]
partialPartial match without bracketsPerson 1

cURL example ​

bash
curl -X POST https://sanitizer.veriprompt.tech/api/v1/sessions/stk_a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6/restore \
  -H "Content-Type: application/json" \
  -d '{
    "content": "Based on the data, [PERSON_1] has the highest rating. Contact [EMAIL_1] for details.",
    "content_type": "text"
  }'

TypeScript example ​

typescript
const restoreResponse = await fetch(
  `https://sanitizer.veriprompt.tech/api/v1/sessions/${sessionToken}/restore`,
  {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      content: llmResponse,
      content_type: 'text',
      delete_after: true  // purge session after restore
    })
  }
);

const { restored_content, confidence, unresolved_tokens } = await restoreResponse.json();

DELETE /api/v1/sessions/ ​

Permanently destroy a session and its encryption keys. Irreversible.

This endpoint explicitly deletes the sensitive session layer:

  • encrypted token-to-original mapping
  • encrypted sanitized document
  • per-session encryption key

The sanitizer preserves only minimal session metadata needed for operational audit within the sanitizer service itself.

Response (200) ​

json
{
  "session_token": "stk_a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6",
  "purged": true,
  "message": "Session purged successfully"
}

Credential exposure log ​

Track and audit credential detections across sanitization sessions.

POST /api/v1/admin/credential-exposure/record ​

Record a credential exposure event.

GET /api/v1/admin/credential-exposure/list ​

List credential exposure events with optional filters (session_token, credential_type, date range).

GET /api/v1/admin/credential-exposure/csv ​

Download credential exposure events as CSV.

POST /api/v1/admin/credential-exposure/clear ​

Clear the credential exposure log (marks entries as cleared, append-only).


Redlist management ​

Manage custom terms that should always be detected, regardless of profile settings.

POST /api/v1/admin/redlist ​

Add a redlist entry.

GET /api/v1/admin/redlist ​

List all redlist entries.

DELETE /api/v1/admin/redlist/ ​

Remove a redlist entry.


Sensitivity profiles reference ​

ProfileEntity types
personal_informationPERSON, EMAIL_ADDRESS, PHONE_NUMBER, LOCATION, DATE_OF_BIRTH, ADDRESS
financial_dataCREDIT_CARD, IBAN, US_SSN, TAX_ID, US_BANK_NUMBER
health_informationMEDICAL_LICENSE, NRP
credentialsAPI_KEY, CONNECTION_STRING, PASSWORD
corporate_confidentialORGANIZATION, IP_ADDRESS, URL

Configuration reference ​

Environment variableDefaultDescription
ENCRYPTION_KEY(required)Base64-encoded 32-byte AES-256 master key
DB_PATH/data/sanitizer.dbSQLite database path
SUPPORTED_LANGUAGESen,deComma-separated language codes
TOKEN_TTL_SECONDS3600Session TTL in seconds
GLINER_ENABLEDfalseEnable GLiNER (Phase 2)
GLINER_MODEL_NAMEurchade/gliner_multi_pii-v1GLiNER model identifier
GLINER_THRESHOLD0.5Minimum GLiNER confidence score
GLINER_IDLE_TIMEOUT300GLiNER idle unload timeout (seconds)
LOG_LEVELINFOLogging level
HOST0.0.0.0Server bind host
PORT8100Server bind port