Skip to content

Endpoint: Execution Modes & Holds ​

Control how the AI Gateway handles prompt execution using three modes: AUTONOMOUS (default), DRY_RUN (plan-only), and SUPERVISED (human-in-the-loop). The Execution Holds API lets reviewers approve or reject held responses from supervised executions.

Endpoints ​

MethodPathAuthDescription
POST/api/gateway/executeAPI Key or SessionExecute a prompt with an executionMode parameter
GET/api/v1/execution-holdsNextAuth sessionList pending holds for the company
GET/api/v1/execution-holds/:idNextAuth sessionGet hold detail
PATCH/api/v1/execution-holds/:idNextAuth sessionApprove or reject a hold

Authentication ​

  • Gateway Execute accepts a Bearer token (Gateway API Key) or a NextAuth session cookie.
  • Execution Holds endpoints require an authenticated NextAuth session (cookie-based).

Execution Modes Overview ​

The executionMode parameter is an optional field on the standard Gateway Execute request body.

ModeBehaviour
AUTONOMOUS(Default) Full execution. The prompt is routed, sent to the provider, and the response is returned immediately.
DRY_RUNPlan only. Returns the routing plan, selected provider, policy evaluation, and cost estimates without calling the AI provider. No tokens are consumed.
SUPERVISEDExecute and hold. The prompt is sent to the provider and the response is captured, but instead of returning the content it creates an Execution Hold that must be approved or rejected by a human reviewer.

Mode Flow Diagram ​

                        ┌──────────────┐
   executionMode?       │   Gateway    │
   ─────────────────>   │   Execute    │
                        └──────┬───────┘
                               │
              ┌────────────────┼────────────────┐
              │                │                │
         DRY_RUN          SUPERVISED        AUTONOMOUS
              │                │                │
      ┌───────▼──────┐  ┌─────▼──────┐  ┌──────▼───────┐
      │ Return plan  │  │ Call provider│  │ Call provider │
      │ (no tokens)  │  │ Hold result │  │ Return result │
      └──────────────┘  └─────┬──────┘  └──────────────┘
                              │
                     ┌────────▼────────┐
                     │  PENDING_REVIEW │
                     └────────┬────────┘
                        ┌─────┴─────┐
                   approve       reject
                        │           │
                   ┌────▼───┐  ┌───▼─────┐
                   │APPROVED│  │REJECTED │
                   └────────┘  └─────────┘

POST /api/gateway/execute (with executionMode) ​

The executionMode field is added to the standard gateway execute request body. All other fields documented in the Gateway Execute endpoint remain unchanged.

Additional Request Parameter ​

FieldTypeRequiredDefaultDescription
executionModestringNoAUTONOMOUSOne of AUTONOMOUS, DRY_RUN, SUPERVISED

DRY_RUN Response ​

When executionMode is set to DRY_RUN, the gateway evaluates routing policies, security checks, and reference library injection but does not call the AI provider. No tokens are consumed and no execution record is billed.

json
{
  "success": true,
  "mode": "DRY_RUN",
  "routingPlan": {
    "provider": "openai",
    "model": "gpt-4o-mini",
    "policyId": "cm1policy...",
    "fallbackProviders": [],
    "geofencingApplied": false,
    "geofencingRegion": null,
    "securityPassed": true,
    "protectivePromptApplied": true,
    "documentsInjected": 3,
    "tokensInjected": 1240,
    "truncated": false
  },
  "timeline": { "..." : "..." }
}

DRY_RUN Response Fields ​

FieldTypeDescription
modestringAlways DRY_RUN
routingPlan.providerstringProvider that would be selected (or auto)
routingPlan.modelstringModel that would be used (or auto)
routingPlan.policyIdstringEffective routing policy ID
routingPlan.fallbackProvidersstring[]Ordered list of fallback providers
routingPlan.geofencingAppliedbooleanWhether geofencing rules were evaluated
routingPlan.securityPassedbooleanWhether security checks passed
routingPlan.protectivePromptAppliedbooleanWhether a protective system prompt was injected
routingPlan.documentsInjectednumberCount of reference library documents that would be injected
routingPlan.tokensInjectednumberEstimated token count from injected documents
routingPlan.truncatedbooleanWhether reference context was truncated to fit limits
timelineobjectStep-by-step timing breakdown

DRY_RUN Examples ​

bash
curl -X POST https://app.veriprompt.tech/api/gateway/execute \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Summarize our Q4 earnings report.",
    "executionMode": "DRY_RUN",
    "policyId": "cm1policy..."
  }'
javascript
const response = await fetch('/api/gateway/execute', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer YOUR_API_KEY',
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({
    prompt: 'Summarize our Q4 earnings report.',
    executionMode: 'DRY_RUN',
    policyId: 'cm1policy...'
  })
});
const plan = await response.json();
console.log('Selected provider:', plan.routingPlan.provider);
console.log('Documents injected:', plan.routingPlan.documentsInjected);
python
import requests

response = requests.post(
    'https://app.veriprompt.tech/api/gateway/execute',
    headers={
        'Authorization': 'Bearer YOUR_API_KEY',
        'Content-Type': 'application/json'
    },
    json={
        'prompt': 'Summarize our Q4 earnings report.',
        'executionMode': 'DRY_RUN',
        'policyId': 'cm1policy...'
    }
)
plan = response.json()
print(f"Selected provider: {plan['routingPlan']['provider']}")
print(f"Documents injected: {plan['routingPlan']['documentsInjected']}")

SUPERVISED Response ​

When executionMode is set to SUPERVISED, the gateway executes the prompt against the AI provider but does not return the response content. Instead, it creates an Execution Hold and returns the hold ID. A human reviewer must then approve or reject the hold via the Execution Holds API.

json
{
  "success": true,
  "mode": "SUPERVISED",
  "holdId": "cm1hold...",
  "message": "Execution held for approval",
  "timeline": { "..." : "..." }
}

SUPERVISED Response Fields ​

FieldTypeDescription
modestringAlways SUPERVISED
holdIdstringID of the created execution hold
messagestringConfirmation message
timelineobjectStep-by-step timing breakdown

SUPERVISED Examples ​

bash
curl -X POST https://app.veriprompt.tech/api/gateway/execute \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Draft a press release announcing the merger.",
    "executionMode": "SUPERVISED",
    "model": "gpt-4o"
  }'
javascript
const response = await fetch('/api/gateway/execute', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer YOUR_API_KEY',
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({
    prompt: 'Draft a press release announcing the merger.',
    executionMode: 'SUPERVISED',
    model: 'gpt-4o'
  })
});
const { holdId } = await response.json();
console.log('Hold created:', holdId);
// Now approve or reject via the Execution Holds API
python
import requests

response = requests.post(
    'https://app.veriprompt.tech/api/gateway/execute',
    headers={
        'Authorization': 'Bearer YOUR_API_KEY',
        'Content-Type': 'application/json'
    },
    json={
        'prompt': 'Draft a press release announcing the merger.',
        'executionMode': 'SUPERVISED',
        'model': 'gpt-4o'
    }
)
hold_id = response.json()['holdId']
print(f"Hold created: {hold_id}")
# Now approve or reject via the Execution Holds API

GET /api/v1/execution-holds ​

List execution holds for the authenticated user's company.

Query Parameters ​

ParameterTypeDefaultDescription
statusstringPENDING_REVIEWFilter by hold status: PENDING_REVIEW, APPROVED, REJECTED, EXPIRED, or ALL
limitnumber50Results per page (max 100)
offsetnumber0Pagination offset

Success Response ​

json
{
  "holds": [
    {
      "id": "cm1hold...",
      "companyId": "cm1company...",
      "executionId": "cm1exec...",
      "executionMode": "SUPERVISED",
      "status": "PENDING_REVIEW",
      "requestPayload": {
        "prompt": "Draft a press release announcing the merger.",
        "systemPrompt": null,
        "model": "gpt-4o",
        "provider": "openai"
      },
      "routingDecision": {
        "provider": "openai",
        "model": "gpt-4o",
        "policyId": "cm1policy..."
      },
      "responsePayload": {
        "content": "FOR IMMEDIATE RELEASE...",
        "usage": { "promptTokens": 45, "completionTokens": 312, "totalTokens": 357 },
        "model": "gpt-4o-2024-08-06"
      },
      "costEstimate": 0.0042,
      "expiresAt": "2026-02-15T11:00:00.000Z",
      "resolvedBy": null,
      "resolvedAt": null,
      "resolution": null,
      "resolver": null,
      "createdAt": "2026-02-15T10:00:00.000Z",
      "updatedAt": "2026-02-15T10:00:00.000Z"
    }
  ],
  "pagination": {
    "total": 3,
    "limit": 50,
    "offset": 0,
    "hasMore": false
  }
}

Hold Status Values ​

StatusMeaning
PENDING_REVIEWAwaiting human approval or rejection
APPROVEDReviewer approved the response
REJECTEDReviewer rejected the response
EXPIREDHold expired without action (default: 60 minutes)

Examples ​

bash
curl "https://app.veriprompt.tech/api/v1/execution-holds?status=PENDING_REVIEW&limit=10" \
  -H "Cookie: next-auth.session-token=YOUR_SESSION_TOKEN"
javascript
const response = await fetch(
  '/api/v1/execution-holds?status=PENDING_REVIEW&limit=10',
  { credentials: 'include' }
);
const { holds, pagination } = await response.json();
console.log(`${pagination.total} holds pending review`);
python
response = session.get(
    'https://app.veriprompt.tech/api/v1/execution-holds',
    params={'status': 'PENDING_REVIEW', 'limit': 10}
)
data = response.json()
print(f"{data['pagination']['total']} holds pending review")

GET /api/v1/execution-holds/:id ​

Get full details for a single execution hold.

Path Parameters ​

ParameterTypeDescription
idstringExecution hold ID

Success Response ​

Returns the full hold object (same shape as items in the list response above).


PATCH /api/v1/execution-holds/:id ​

Approve or reject a pending execution hold.

Path Parameters ​

ParameterTypeDescription
idstringExecution hold ID

Request Body ​

FieldTypeRequiredDescription
actionstringYesapprove or reject
reasonstringNoRejection reason (only used when action is reject)

Success Response — Approve ​

json
{
  "hold": {
    "id": "cm1hold...",
    "status": "APPROVED",
    "resolvedBy": "cm1user...",
    "resolvedAt": "2026-02-15T10:05:00.000Z",
    "resolution": "APPROVED",
    "..." : "..."
  },
  "responsePayload": {
    "content": "FOR IMMEDIATE RELEASE...",
    "usage": { "promptTokens": 45, "completionTokens": 312, "totalTokens": 357 },
    "model": "gpt-4o-2024-08-06"
  }
}

When approving, the responsePayload is returned so the caller can retrieve the AI-generated content that was held.

Success Response — Reject ​

json
{
  "id": "cm1hold...",
  "status": "REJECTED",
  "resolvedBy": "cm1user...",
  "resolvedAt": "2026-02-15T10:05:00.000Z",
  "resolution": "Content not appropriate for public release",
  "..." : "..."
}

Examples ​

bash
# Approve a hold
curl -X PATCH "https://app.veriprompt.tech/api/v1/execution-holds/cm1hold123" \
  -H "Content-Type: application/json" \
  -H "Cookie: next-auth.session-token=YOUR_SESSION_TOKEN" \
  -d '{ "action": "approve" }'

# Reject a hold with reason
curl -X PATCH "https://app.veriprompt.tech/api/v1/execution-holds/cm1hold456" \
  -H "Content-Type: application/json" \
  -H "Cookie: next-auth.session-token=YOUR_SESSION_TOKEN" \
  -d '{ "action": "reject", "reason": "Content not appropriate for public release" }'
javascript
// Approve
const approveRes = await fetch(`/api/v1/execution-holds/${holdId}`, {
  method: 'PATCH',
  headers: { 'Content-Type': 'application/json' },
  credentials: 'include',
  body: JSON.stringify({ action: 'approve' })
});
const { responsePayload } = await approveRes.json();
console.log('Approved content:', responsePayload.content);

// Reject
const rejectRes = await fetch(`/api/v1/execution-holds/${holdId}`, {
  method: 'PATCH',
  headers: { 'Content-Type': 'application/json' },
  credentials: 'include',
  body: JSON.stringify({
    action: 'reject',
    reason: 'Content not appropriate for public release'
  })
});
python
# Approve
response = session.patch(
    f'https://app.veriprompt.tech/api/v1/execution-holds/{hold_id}',
    json={'action': 'approve'}
)
payload = response.json()['responsePayload']
print(f"Approved content: {payload['content'][:100]}...")

# Reject
response = session.patch(
    f'https://app.veriprompt.tech/api/v1/execution-holds/{hold_id}',
    json={
        'action': 'reject',
        'reason': 'Content not appropriate for public release'
    }
)
print(f"Status: {response.json()['status']}")

Error Responses ​

StatusCondition
400Invalid request body, hold already resolved, or invalid execution mode value
401Unauthorized — missing API key or session
403Policy or geofence block (gateway execute)
404Hold not found or does not belong to the user's company
429Concurrency guardrail exceeded (gateway execute)
500Internal server error