Skip to content

Prompt Chaining & Execution Modes API Documentation ​

Version: 1.1.0 Last Updated: 2026-02-10

Overview ​

The Prompt Chaining system allows prompts to reference and embed other prompts using the {{prompt:promptID}} syntax. Execution modes control how referenced prompts are executed:

  • TEMPLATE (default): Fast text injection, no routing preservation
  • EXECUTE: Full execution via internal API, preserves routing and settings
  • AUTO: Smart mode selection based on routing configuration

Key Features ​

1. Embeddability Classification ​

Prompts are automatically analyzed for embeddability when created or updated:

  • ✅ Embeddable: Static prompts without attachments or dynamic variables
  • ❌ Not Embeddable: Prompts with attachments, {{variables}}, or [data] markers

2. Execution Modes ​

ModeSpeedCostRouting PreservedUse Case
TEMPLATE<1msFree❌ NoStatic templates, guidelines
EXECUTE500-2000msCharged✅ YesRouting-critical prompts
AUTODynamicVariable✅ SmartRecommended default

3. Caching ​

EXECUTE mode results are cached for 5 minutes to improve performance:

  • Cache key: prompt:execute:{promptId}:{hash}:depth{N}
  • In-memory fallback (Redis-ready)
  • Automatic cache invalidation on prompt update

4. Cost Attribution ​

Internal executions are tracked for transparent billing:

  • isInternal: true - Marks prompt chaining executions
  • parentExecutionId - Links to parent execution
  • chainDepth - Tracks nesting level (max: 3)
  • referencedPromptId - ID of referenced prompt

API Endpoints ​

GET /api/v1/prompts ​

List prompts with embeddability info

Response:

json
{
  "prompts": [
    {
      "id": "...",
      "promptId": "prompt_123",
      "name": "Legal Analysis Prompt",
      "userPrompt": "Analyze the following: {{prompt:guidelines}}",
      "embeddability": {
        "isEmbeddable": true,
        "executionMode": "TEMPLATE",
        "requiresAttachment": false,
        "requiresDynamicData": false,
        "blockedReason": null
      },
      ...
    }
  ]
}

POST /api/v1/prompts ​

Create a new prompt with auto-classification

Request:

json
{
  "name": "Contract Review",
  "userPrompt": "Review this contract: {{prompt:legal_guidelines}}",
  "systemPrompt": "You are a legal expert",
  "temperature": 0.3,
  "attachment": null
}

Response:

json
{
  "success": true,
  "prompt": {
    "id": "...",
    "promptId": "prompt_456",
    "isEmbeddable": true,
    "executionMode": "TEMPLATE",
    "requiresAttachment": false,
    "requiresDynamicData": false,
    ...
  }
}

Auto-Classification Logic:

  • ✅ No attachment → embeddable
  • ✅ No {{variables}} → embeddable
  • ✅ No dynamic markers ([file], [url], etc.) → embeddable
  • Execution mode = EXECUTE if routing configured, else TEMPLATE

PUT /api/v1/prompts/[promptId] ​

Update prompt with re-classification

Request:

json
{
  "userPrompt": "Updated prompt with {{variable}}",
  "attachment": {
    "type": "file",
    "fileId": "file_123"
  }
}

Response:

json
{
  "success": true,
  "prompt": {
    "isEmbeddable": false,
    "embeddingBlockedReason": "Requires 1 runtime variable(s): variable; Prompt has file/URL attachment configured",
    "requiresAttachment": true,
    "requiresDynamicData": true,
    "executionMode": "TEMPLATE",
    ...
  }
}

POST /api/gateway/execute ​

Execute prompt with reference resolution

Request:

json
{
  "prompt": "Generate report using {{prompt:template_123}}",
  "sourcePromptId": "prompt_parent",
  "temperature": 0.7
}

Internal Flow:

  1. Extract prompt references: [template_123] or {{prompt:template_123}}
  2. For each reference:
    • Check embeddability
    • Determine execution mode (TEMPLATE/EXECUTE/AUTO)
    • If EXECUTE: Check cache → Execute via internal API → Cache result
    • If TEMPLATE: Inject text directly
  3. Replace references with resolved content
  4. Execute final prompt

Response:

json
{
  "response": "...",
  "executionChain": [
    {
      "promptId": "template_123",
      "promptName": "Report Template",
      "depth": 1,
      "executionMode": "TEMPLATE",
      "response": "..."
    }
  ],
  "referenceErrors": []
}

POST /api/v1/chains ​

Create a stored prompt chain that runs multiple prompts sequentially. Requires an API key with the execute:chains permission.

Request:

json
{
  "name": "Document Processing Chain",
  "description": "Intake → Policy Review → Summary",
  "steps": [
    { "order": 1, "promptId": "chain.document-intake", "outputVariable": "intake_summary" },
    { "order": 2, "promptId": "chain.policy-check", "variables": { "summary": "{{intake_summary}}" } },
    { "order": 3, "promptId": "chain.summary" }
  ],
  "tags": ["chain", "documents"],
  "stopOnError": true,
  "maxConcurrency": 1,
  "metadata": { "department": "compliance" }
}

Response:

json
{
  "chainId": "chain_b3d4f0c9...",
  "name": "Document Processing Chain",
  "stepsCount": 3,
  "settings": {
    "stopOnError": true,
    "maxConcurrency": 1
  },
  "usage": {
    "totalExecutions": 0,
    "lastExecutedAt": null
  }
}

POST /api/v1/chains/{chainId}/execute ​

Execute a stored chain end-to-end. Each step inherits the referenced prompt’s routing policy and feeds its output variable back into the shared context.

Request:

json
{
  "inputs": {
    "company_name": "VeriPharma",
    "risk_tier": "HIGH"
  },
  "attachments": [
    { "type": "url", "url": "https://example.com/docs/kyc/veripharm.pdf" }
  ],
  "options": {
    "stopOnError": true
  }
}

Response:

json
{
  "chainId": "chain_b3d4f0c9...",
  "chainExecutionId": "b7bde270-4a6c-4e0d-a43f-7ab79fca6371",
  "status": "SUCCESS",
  "steps": [
    {
      "order": 1,
      "promptId": "chain.document-intake",
      "status": "SUCCESS",
      "providerName": "openai",
      "providerModel": "gpt-4o-mini",
      "response": "### Intake Summary ...",
      "usage": { "promptTokens": 812, "completionTokens": 350 }
    },
    {
      "order": 2,
      "promptId": "chain.policy-check",
      "status": "SUCCESS",
      "providerName": "anthropic",
      "providerModel": "claude-3-sonnet",
      "response": "### Compliance Decision ...",
      "usage": { "promptTokens": 540, "completionTokens": 210 }
    }
  ],
  "context": {
    "intake_summary": "### Intake Summary ..."
  },
  "tokens": {
    "promptTokens": 1352,
    "completionTokens": 560,
    "totalTokens": 1912
  },
  "errors": [],
  "elapsedMs": 8420
}

Notes

  • Guardrails enforce max concurrency, rate limit, and token ceilings per chainId. Violations return HTTP 429 with Chain guardrail violation.
  • Attachments are injected into every step; per-step attachment overrides can be stored in the chain metadata (via the admin UI).
  • Each chain step forwards a W3C traceparent header to /api/gateway/execute.
  • Each chain step sets telemetry: { mode: "full", force: true } so prompt-side telemetry is always captured for chain runs.
  • Step metadata includes chainExecutionId, chainStep, traceId, and spanId for cross-step correlation.
  • Ideal for the sample Document Processing Chain exercised by scripts/test-external-agent.js.

Embeddability Analysis ​

Service: lib/prompt-embeddability.ts ​

Function: analyzePromptEmbeddability(prompt)

Returns:

typescript
{
  isEmbeddable: boolean;
  reason?: string;
  requiresAttachment: boolean;
  requiresDynamicData: boolean;
  suggestedMode: 'TEMPLATE' | 'EXECUTE';
  warnings: string[];
}

Blocking Conditions:

  • Has attachment field populated
  • Contains {{variable}} placeholders (excluding {{prompt:*}})
  • Contains markers: [file], [url], [data], [attachment], [input]
  • Explicitly marked requiresAttachment=true or requiresDynamicData=true

Execution Mode Selection:

  • Has routingPolicyId → EXECUTE
  • Has customGatewayId → EXECUTE
  • Temperature < 0.3 or > 1.2 → Warning (may need EXECUTE)
  • Otherwise → TEMPLATE

Internal Execution ​

Service: lib/internal-execution.ts ​

Function: executePromptInternally(options)

Options:

typescript
{
  promptId: string;
  parentExecutionId?: string;
  chainDepth: number;          // Max: 3
  companyId: string;
  userId: string;
}

Flow:

  1. Validate prompt exists and is embeddable
  2. Check cache: generateCacheKey(promptId, hash, depth)
  3. If cached → return immediately (0ms latency)
  4. If not cached:
    • Execute via gateway API (TODO: implement actual API call)
    • Record in AnonymousExecution table with isInternal=true
    • Cache result for 5 minutes
  5. Return response with execution metadata

Caching ​

Service: lib/redis-cache.ts ​

Current Implementation: In-memory fallback Ready For: Redis integration (see TODOs in file)

Functions:

  • getCachedResult(key) - Retrieve cached execution
  • setCachedResult(key, value, {ttl: 300}) - Cache with 5min TTL
  • clearPromptCache(promptId) - Clear on prompt update
  • getCacheStats() - Monitor cache performance

Cache Keys:

prompt:execute:{promptId}:{hash}:depth{N}

Example:

prompt:execute:prompt_123:dGVzdA==:depth1

Security & Limits ​

Rate Limiting ​

Chain Depth: Maximum 3 levels

typescript
if (chainDepth > 3) {
  return { success: false, error: 'Maximum chain depth exceeded' };
}

Circular Reference Detection: Automatically detected and blocked during prompt save/update.

Access Control ​

All prompt references validated via:

  • Company-level access (prompts must be in same company)
  • Project/repository-level access (if PROJECT scope)
  • User permissions (via checkUserPromptAccess)

Error Handling ​

Embeddability Errors ​

json
{
  "error": "Invalid prompt references",
  "details": [
    {
      "promptId": "template_999",
      "reason": "Prompt is not embeddable: Requires 2 runtime variable(s): name, date",
      "position": 45
    }
  ]
}

Execution Errors ​

json
{
  "success": false,
  "error": "Failed to execute [template_123]: Maximum chain depth exceeded"
}

Migration Guide ​

Updating to Execution Modes ​

Before:

typescript
const result = await executePromptWithReferences(
  prompt,
  promptId,
  userId,
  3 // maxDepth
);

After:

typescript
const result = await executePromptWithReferences(
  prompt,
  promptId,
  userId,
  {
    maxDepth: 3,
    parentExecutionId: executionId,
    companyId: user.companyId
  }
);

Monitoring & Analytics ​

Track Internal Executions ​

sql
SELECT
  COUNT(*) as internal_count,
  AVG(executionTimeMs) as avg_latency,
  SUM(CASE WHEN selectedProvider = 'cache' THEN 1 ELSE 0 END) as cache_hits
FROM anonymous_executions
WHERE isInternal = true
  AND executedAt > NOW() - INTERVAL '24 hours';

Cost Attribution Query ​

sql
SELECT
  parentExecutionId,
  COUNT(*) as child_executions,
  SUM(executionTimeMs) as total_latency
FROM anonymous_executions
WHERE isInternal = true
GROUP BY parentExecutionId;

Future Enhancements ​

  1. Variable Passing: {{prompt:id|var1=value}}
  2. Parallel Execution: Execute multiple EXECUTE-mode prompts concurrently
  3. Cost Dashboard: Real-time chaining cost analytics
  4. Redis Integration: Replace in-memory cache with Redis
  5. Gateway Integration: Implement actual API call in executeViaGateway

Troubleshooting ​

Prompt Not Embeddable ​

Error: "Prompt is not embeddable: Requires runtime variables"

Solution:

  1. Remove {{variables}} from prompt
  2. Remove [file], [url] markers
  3. Remove attachments
  4. Or use EXECUTE mode in calling code (not through embedding)

Cache Not Working ​

Check:

  1. Cache stats: getCacheStats()
  2. Verify cache key format
  3. Check TTL expiration
  4. Monitor memory usage

Performance Issues ​

Solutions:

  1. Use TEMPLATE mode for static content
  2. Enable caching for EXECUTE mode
  3. Reduce chain depth
  4. Monitor with executionChain metadata

For more information, see:

  • IMPLEMENTATION_TODOS.md - Full implementation details
  • lib/prompt-embeddability.ts - Classification logic
  • lib/internal-execution.ts - Execution service
  • lib/redis-cache.ts - Caching layer