Skip to content

Watermark Avoidance ​

Can the AI provider's response be traced back to you? Shield's watermark avoidance check lets you refuse a request before it ever reaches a provider known to watermark its output.

Overview ​

PII sanitization already strips identifiable content out of what you send to a provider. But some providers fingerprint their own output with a watermark, and that watermark can in principle be used to trace a generated response back to whoever it was generated for — even when the input was fully sanitized. Watermark avoidance closes that second half of the traceability problem: it checks the resolved provider/model against an internal watermarking verdict, and if that provider is known to watermark, it refuses the request before any document content is sent.

Who this is for ​

This matters most to regulated and high-sensitivity customers who need a hard guarantee that a request cannot be traced back to them, including:

  • Law firms handling privileged or client-confidential material
  • Defense contractors working under strict data-handling and non-attribution requirements
  • Healthcare organizations subject to confidentiality obligations around patient data

If your organization's threat model includes "the provider could later attribute this output to us," this feature is for you.

How it works ​

Watermark avoidance is a request-level option on the Shield document-proxy endpoint — the same POST route that already sanitizes a document before calling the LLM.

Set avoid_marked_providers: true on the request. When set, Shield checks the resolved provider/model against an internal watermarking verdict before sending anything. If that provider/model is flagged, the request is refused immediately — no tokens spent, no document sent to the provider. It is optional and defaults to false (off), so existing integrations keep their current behavior unless they opt in.

Watermark verdicts come from an internal probe process; no further detail on that process is documented here.

Not every flagged provider is handled the same way, because not every watermarking technique can be cleaned up after the fact:

  • A provider known to embed a statistical, word-choice-level watermark — the kind baked into the text itself with no reliable way to strip it back out — is always refused outright, the same as before.
  • A provider that might embed hidden or invisible characters in its response is still used, but the response is automatically scanned and cleaned of those characters before it is returned to you.

Either way, you get an answer back — you just never get one that could still carry a hidden marker. This distinction only affects which providers get a hard refusal versus a cleaned response; it does not change the fact that avoid_marked_providers is opt-in and off by default.

Example: request with watermark avoidance on ​

json
{
  "document": "<base64-encoded document>",
  "document_type": "pdf",
  "prompt": "Summarize this document",
  "provider": {
    "type": "example-provider",
    "apiKey": "sk-...",
    "model": "example-model"
  },
  "avoid_marked_providers": true
}

Shield does the sanitizing itself as part of this same call — you send the raw document, not a pre-sanitized one.

Example: refusal response ​

If the resolved provider/model is known to watermark its output, the request is refused before the document is sent anywhere, with an HTTP 502 and a plain-text message:

json
{
  "error": "Provider EXAMPLE-PROVIDER/example-model is known to watermark its output. The request was refused rather than sent to the provider."
}

There is currently no separate machine-readable error code in the response for this case — you can only match on the message text. This is a known limitation; a distinct error_code (the way the coverage-gate error already has one) is a reasonable future improvement.

The exact wording depends on which condition triggered the refusal — "is known to watermark its output" for a flagged provider, or language about an uncleanable statistical watermarking mechanism when that's the reason — so match on the response status code and the general shape of the message rather than an exact string.

If the resolved provider/model is not flagged, the request proceeds through the normal sanitize -> LLM -> restore flow with no change in behavior.

Coverage & limitations ​

  • Not exhaustive. Watermark avoidance only catches providers with a known verdict — either watermarking detected via active testing, or watermarking the provider documents themselves. It does not catch every possible watermarking scheme, including ones no one has tested for yet.
  • Off by default. Because this is new behavior that could otherwise unexpectedly block an existing integration, avoid_marked_providers defaults to false. You must explicitly set it to true per request.
  • No UI toggle yet. There is currently no dedicated settings toggle for this in the Shield UI — it can only be set via the API parameter above. If you use the Shield UI directly rather than calling the API yourself, you cannot turn this on yet. A UI toggle is a possible future addition, but does not exist today.
  • Hidden-character cleaning is text-only. Shield's document-proxy responses are plain text, so the automatic cleaning pass covers hidden/invisible characters in that text. It does not apply to file-container-level watermarking (for example, metadata embedded in a generated DOCX or PDF file) — that scenario doesn't arise on this text-only endpoint.