Appearance
Document Parsing
Endpoints for the PDF parsing sub-service: read and set the delivery defaults, measure what parsing saves on a real file, and inspect a PDF's tables so Shield can protect their columns.
See Document Parsing for the concepts.
All endpoints require an authenticated session and a company context.
GET /api/documents/parsing/settings
Returns the caller's entitlement plus the stored default at every scope, so a UI can show "inherited from company" without a second call.
Query parameters
| Name | Required | Description |
|---|---|---|
projectId | no | Project.id or Project.projectId. Include the project's stored override. |
promptId | no | StoredPrompt.id or StoredPrompt.promptId. Include the prompt's stored override. |
Response 200
json
{
"entitlement": {
"allowedModes": ["NATIVE", "TEXT", "STRUCTURED"],
"structuredAllowed": true,
"upgradeMessage": "",
"maxPages": 300
},
"modes": ["NATIVE", "TEXT", "STRUCTURED"],
"company": { "mode": "TEXT" },
"project": { "id": "...", "projectId": "proj_x", "name": "Invoices", "mode": "NATIVE" },
"prompt": null,
"effectiveMode": "NATIVE",
"canManage": true
}effectiveMode is what would apply right now, before any per-message override. project and prompt are null when not requested or not found; a mode of null on either means "inherit".
PUT /api/documents/parsing/settings
Sets — or clears — the default at one scope.
json
{ "scope": "project", "mode": "STRUCTURED", "targetId": "proj_x" }| Field | Required | Description |
|---|---|---|
scope | yes | company, project, or prompt. |
mode | yes | NATIVE, TEXT, STRUCTURED, or null to clear the override and inherit again. The company scope cannot be null. |
targetId | for project / prompt | The project or prompt identifier. |
Responses
| Status | Meaning |
|---|---|
200 | { "success": true, "scope": "project", "mode": "STRUCTURED" } |
400 | Invalid body, missing targetId, or an attempt to clear the company default. |
403 | mode: "STRUCTURED" without the entitlement, or a non-admin changing a company/project default. |
404 | Project or prompt not found in your company. |
Entitlement is enforced on write as well as at parse time — storing a mode your plan cannot run would leave a setting that silently never applies.
POST /api/documents/parsing/preview
Parses a PDF and returns both the result and its token-savings figures. Nothing is sent to an AI provider, and no DocumentParseLog row is written — the ledger counts documents actually sent to a model, not UI experiments.
Request — multipart/form-data
| Field | Required | Description |
|---|---|---|
file | yes | The PDF. Max 15MB. |
mode | no | TEXT or STRUCTURED. NATIVE is treated as TEXT (there is nothing to preview; the cost figure is in the metrics either way). |
includeImages | no | "true" to extract embedded images (STRUCTURED only). |
Response 200
json
{
"success": true,
"fileName": "quarterly.pdf",
"mode": "STRUCTURED",
"contentType": "text/markdown",
"content": "# Quarterly Report\n\n| Region | Revenue |\n| --- | --- |\n…",
"contentTruncatedForPreview": false,
"tables": [{ "page": 1, "header": ["Region", "Revenue"], "rows": [["EMEA", "1,240"]] }],
"images": [{ "page": 2, "ref": "p2-img1", "width": 800, "height": 600, "byteSize": 48213 }],
"metrics": { "pageCount": 12, "pagesParsed": 12, "tablesFound": 3, "embeddedImages": 2, "truncated": false },
"savings": { "parsedTokens": 4210, "nativeTokens": 21600, "tokensSaved": 17390, "percentSaved": 80 },
"extraction": { "pageCount": 12, "hiddenTextRemoved": 47, "imagesRemoved": 2, "pagesWithoutText": [],
"ocrPages": [], "activeContentRemoved": 1, "metadataRemoved": 2 },
"extractionSource": "standard",
"entitlement": { "structuredAllowed": true, "maxPages": 300 }
}extraction says what the PDF extraction standard left out of content (counts only, never text): hiddenTextRemoved (characters a reader would not see: invisible, white, under 1.5 pt, off the page; replaced by [Hidden text removed]), imagesRemoved, pagesWithoutText (scans, left out), ocrPages (scans whose OCR text was kept; it was not checked against the image), activeContentRemoved (links, form fields, embedded files, scripts) and metadataRemoved. extractionSource is standard, or fallback when the protection service was unreachable and the PDF was read without hidden-text removal (extraction is then null).
The same two fields are returned by POST /api/v1/files/parse for PDFs, under parsing:
bash
curl -s https://<host>/api/v1/files/parse -b cookies.txt \
-F "file=@offer.pdf" -F "mode=TEXT" | jq '.parsing | {mode, extractionSource, extraction}'Image bytes are omitted from this response — the preview reports what was found; the full payload goes to the model, not to a settings pane.
| Status | Meaning |
|---|---|
400 | No file, empty file, or not a PDF. |
413 | Over 15MB. |
500 | The PDF could not be parsed (encrypted, corrupt, or a scan with no text layer). |
POST /api/documents/parsing/inspect-tables
Returns the headers of the tables in a PDF with a suggested entity type, so they can be turned into Shield column rules. A column rule protects a whole column rather than whichever cells the detector happened to match.
Request — multipart/form-data with file (PDF, max 15MB).
Response 200
json
{
"success": true,
"fileName": "payroll.pdf",
"mode": "STRUCTURED",
"structuredAvailable": true,
"tablesFound": 2,
"columns": [
{ "header": "Mitarbeiter", "suggestedEntityType": "PERSON", "page": 1, "sampleValues": ["M. Schneider"] },
{ "header": "IBAN", "suggestedEntityType": "IBAN_CODE", "page": 1, "sampleValues": ["DE89…"] },
{ "header": "Abteilung", "suggestedEntityType": null, "page": 1, "sampleValues": ["Vertrieb"] }
],
"upgradeMessage": null
}Notes:
- Table recovery requires structured parsing. Without the entitlement,
structuredAvailableisfalse,tablesFoundis0andupgradeMessageexplains why — rather than an empty list implying the PDF has no tables. suggestedEntityType: nullmeans nothing matched. Pick a type yourself, or leave the column unruled.- Tables whose header row could not be confirmed are skipped, not guessed. A rule keyed to a data value would match nothing and silently protect no one.
- Suggestions come from the sanitizer service where it can classify them, and from local German/English heuristics when it cannot — a sidecar outage degrades the suggestions, not the endpoint.
Feed the confirmed pairs back as sanitization.columnRules on your next gateway execute call.
Sending documents through the gateway
POST /api/gateway/execute accepts two document-related attachment types.
json
{
"prompt": "Summarize the attached report.",
"attachments": [
{ "type": "document", "fileName": "report.pdf", "mimeType": "application/pdf",
"data": "<base64>", "content": "<extracted text>" },
{ "type": "image", "fileName": "report.pdf#p2-img1", "mimeType": "image/webp",
"data": "<base64>" }
],
"acknowledgeUnprotectedDocument": false
}| Field | Meaning |
|---|---|
type: "document" | Send as-is. data is the file; content is your extracted text, which is still scanned for prompt injection even though the provider reads the bytes. Always send content — omitting it removes injection coverage for that file. |
type: "image" | An image for a vision-capable provider. |
acknowledgeUnprotectedDocument | Informed consent that Shield cannot protect this content. |
Behaviour when PII sanitization is active (which is the default once your plan includes it, regardless of what you send in sanitization.enabled):
- Without
acknowledgeUnprotectedDocument: documents are converted to their extracted text so Shield can mask them, and images are dropped with an explanatory note. This is the fail-closed default. - With it: the bytes go through, and the acceptance is written to your audit log as
SHIELD_UNPROTECTED_DOCUMENT_DELIVERY.
The flag is per request and is never remembered — send it on each call where you mean it. Providers without native document support receive the extracted text instead, so a failover never drops your document.
