Skip to content

Document Parsing ​

Endpoints for the PDF parsing sub-service: read and set the delivery defaults, measure what parsing saves on a real file, and inspect a PDF's tables so Shield can protect their columns.

See Document Parsing for the concepts.

All endpoints require an authenticated session and a company context.


GET /api/documents/parsing/settings ​

Returns the caller's entitlement plus the stored default at every scope, so a UI can show "inherited from company" without a second call.

Query parameters

NameRequiredDescription
projectIdnoProject.id or Project.projectId. Include the project's stored override.
promptIdnoStoredPrompt.id or StoredPrompt.promptId. Include the prompt's stored override.

Response 200

json
{
  "entitlement": {
    "allowedModes": ["NATIVE", "TEXT", "STRUCTURED"],
    "structuredAllowed": true,
    "upgradeMessage": "",
    "maxPages": 300
  },
  "modes": ["NATIVE", "TEXT", "STRUCTURED"],
  "company": { "mode": "TEXT" },
  "project": { "id": "...", "projectId": "proj_x", "name": "Invoices", "mode": "NATIVE" },
  "prompt": null,
  "effectiveMode": "NATIVE",
  "canManage": true
}

effectiveMode is what would apply right now, before any per-message override. project and prompt are null when not requested or not found; a mode of null on either means "inherit".


PUT /api/documents/parsing/settings ​

Sets — or clears — the default at one scope.

json
{ "scope": "project", "mode": "STRUCTURED", "targetId": "proj_x" }
FieldRequiredDescription
scopeyescompany, project, or prompt.
modeyesNATIVE, TEXT, STRUCTURED, or null to clear the override and inherit again. The company scope cannot be null.
targetIdfor project / promptThe project or prompt identifier.

Responses

StatusMeaning
200{ "success": true, "scope": "project", "mode": "STRUCTURED" }
400Invalid body, missing targetId, or an attempt to clear the company default.
403mode: "STRUCTURED" without the entitlement, or a non-admin changing a company/project default.
404Project or prompt not found in your company.

Entitlement is enforced on write as well as at parse time — storing a mode your plan cannot run would leave a setting that silently never applies.


POST /api/documents/parsing/preview ​

Parses a PDF and returns both the result and its token-savings figures. Nothing is sent to an AI provider, and no DocumentParseLog row is written — the ledger counts documents actually sent to a model, not UI experiments.

Request — multipart/form-data

FieldRequiredDescription
fileyesThe PDF. Max 15MB.
modenoTEXT or STRUCTURED. NATIVE is treated as TEXT (there is nothing to preview; the cost figure is in the metrics either way).
includeImagesno"true" to extract embedded images (STRUCTURED only).

Response 200

json
{
  "success": true,
  "fileName": "quarterly.pdf",
  "mode": "STRUCTURED",
  "contentType": "text/markdown",
  "content": "# Quarterly Report\n\n| Region | Revenue |\n| --- | --- |\n…",
  "contentTruncatedForPreview": false,
  "tables": [{ "page": 1, "header": ["Region", "Revenue"], "rows": [["EMEA", "1,240"]] }],
  "images": [{ "page": 2, "ref": "p2-img1", "width": 800, "height": 600, "byteSize": 48213 }],
  "metrics": { "pageCount": 12, "pagesParsed": 12, "tablesFound": 3, "embeddedImages": 2, "truncated": false },
  "savings": { "parsedTokens": 4210, "nativeTokens": 21600, "tokensSaved": 17390, "percentSaved": 80 },
  "extraction": { "pageCount": 12, "hiddenTextRemoved": 47, "imagesRemoved": 2, "pagesWithoutText": [],
                  "ocrPages": [], "activeContentRemoved": 1, "metadataRemoved": 2 },
  "extractionSource": "standard",
  "entitlement": { "structuredAllowed": true, "maxPages": 300 }
}

extraction says what the PDF extraction standard left out of content (counts only, never text): hiddenTextRemoved (characters a reader would not see: invisible, white, under 1.5 pt, off the page; replaced by [Hidden text removed]), imagesRemoved, pagesWithoutText (scans, left out), ocrPages (scans whose OCR text was kept; it was not checked against the image), activeContentRemoved (links, form fields, embedded files, scripts) and metadataRemoved. extractionSource is standard, or fallback when the protection service was unreachable and the PDF was read without hidden-text removal (extraction is then null).

The same two fields are returned by POST /api/v1/files/parse for PDFs, under parsing:

bash
curl -s https://<host>/api/v1/files/parse -b cookies.txt \
  -F "file=@offer.pdf" -F "mode=TEXT" | jq '.parsing | {mode, extractionSource, extraction}'

Image bytes are omitted from this response — the preview reports what was found; the full payload goes to the model, not to a settings pane.

StatusMeaning
400No file, empty file, or not a PDF.
413Over 15MB.
500The PDF could not be parsed (encrypted, corrupt, or a scan with no text layer).

POST /api/documents/parsing/inspect-tables ​

Returns the headers of the tables in a PDF with a suggested entity type, so they can be turned into Shield column rules. A column rule protects a whole column rather than whichever cells the detector happened to match.

Request — multipart/form-data with file (PDF, max 15MB).

Response 200

json
{
  "success": true,
  "fileName": "payroll.pdf",
  "mode": "STRUCTURED",
  "structuredAvailable": true,
  "tablesFound": 2,
  "columns": [
    { "header": "Mitarbeiter", "suggestedEntityType": "PERSON", "page": 1, "sampleValues": ["M. Schneider"] },
    { "header": "IBAN", "suggestedEntityType": "IBAN_CODE", "page": 1, "sampleValues": ["DE89…"] },
    { "header": "Abteilung", "suggestedEntityType": null, "page": 1, "sampleValues": ["Vertrieb"] }
  ],
  "upgradeMessage": null
}

Notes:

  • Table recovery requires structured parsing. Without the entitlement, structuredAvailable is false, tablesFound is 0 and upgradeMessage explains why — rather than an empty list implying the PDF has no tables.
  • suggestedEntityType: null means nothing matched. Pick a type yourself, or leave the column unruled.
  • Tables whose header row could not be confirmed are skipped, not guessed. A rule keyed to a data value would match nothing and silently protect no one.
  • Suggestions come from the sanitizer service where it can classify them, and from local German/English heuristics when it cannot — a sidecar outage degrades the suggestions, not the endpoint.

Feed the confirmed pairs back as sanitization.columnRules on your next gateway execute call.


Sending documents through the gateway ​

POST /api/gateway/execute accepts two document-related attachment types.

json
{
  "prompt": "Summarize the attached report.",
  "attachments": [
    { "type": "document", "fileName": "report.pdf", "mimeType": "application/pdf",
      "data": "<base64>", "content": "<extracted text>" },
    { "type": "image", "fileName": "report.pdf#p2-img1", "mimeType": "image/webp",
      "data": "<base64>" }
  ],
  "acknowledgeUnprotectedDocument": false
}
FieldMeaning
type: "document"Send as-is. data is the file; content is your extracted text, which is still scanned for prompt injection even though the provider reads the bytes. Always send content — omitting it removes injection coverage for that file.
type: "image"An image for a vision-capable provider.
acknowledgeUnprotectedDocumentInformed consent that Shield cannot protect this content.

Behaviour when PII sanitization is active (which is the default once your plan includes it, regardless of what you send in sanitization.enabled):

  • Without acknowledgeUnprotectedDocument: documents are converted to their extracted text so Shield can mask them, and images are dropped with an explanatory note. This is the fail-closed default.
  • With it: the bytes go through, and the acceptance is written to your audit log as SHIELD_UNPROTECTED_DOCUMENT_DELIVERY.

The flag is per request and is never remembered — send it on each call where you mean it. Providers without native document support receive the extracted text instead, so a failover never drops your document.