Skip to content

Intelligent Routing ​

Intelligent routing lets you control how Veriprompt selects AI providers for your requests. Define routing policies to optimize for cost, speed, quality, or compliance requirements.

Overview ​

Routing policies define rules for provider selection. When a request comes in, Veriprompt evaluates available providers against your policy and selects the best match based on:

  • Strategy - Overall optimization goal (cost, latency, quality, balanced)
  • Weights - Fine-tuned preference scores for different factors
  • Geofencing - Geographic and compliance restrictions
  • Provider Groups - Limit to specific provider sets
  • Hard Constraints - Minimum requirements providers must meet

Access and Permissions ​

Required roles: Account Owner, Admin, Security Custodian, or Developer

UI Path: Admin > Chat Settings > 1. Routing Policies

Routing Strategies ​

StrategyDescriptionBest For
BalancedEqual weight to all factorsGeneral use
Lowest CostPrioritize cheapest providersHigh-volume, non-critical tasks
Lowest LatencyPrioritize fastest responseReal-time applications
Highest QualityPrioritize best output qualityCritical or customer-facing tasks
Round RobinRotate between providersLoad distribution, testing

Task-Aware Routing ​

A policy can also rank models by how well they handle the task of each request (coding, contract review, translation) instead of one general quality number. It only reorders the models your policy already allows. See Task-Aware Routing.

Preference Weights ​

Fine-tune provider selection with weighted preferences (0-100% each):

WeightDescription
LatencyResponse time importance
CostPrice per token importance
QualityOutput quality importance
StabilityProvider uptime importance

Example: For a cost-sensitive batch job, set Cost to 60% and reduce Quality to 20%.

Geofencing Rules ​

Restrict providers by geographic and compliance requirements:

Allowed Regions/Countries (Whitelist) ​

Only use providers in specified locations:

  • Regions: US, EU, UK, APAC, LATAM, ME, AF
  • Countries: US, CA, UK, DE, FR, NL, IE, JP, AU, SG, IN, BR

Blocked Regions (Blacklist) ​

Exclude providers from specific regions.

Blocked Countries (Blacklist) ​

Exclude providers from specific countries.

Required Compliance Standards ​

Require providers to meet specific certifications:

  • GDPR - EU Data Protection
  • SOC 2 Type II - Security controls
  • HIPAA - Healthcare data
  • ISO 27001 - Information security
  • PCI DSS - Payment card data

Geofencing Input Rules (Strict Mode) ​

Routing policy geofencing values are validated against your company's active provider configuration baseline:

  • Country values must match ProviderModelConfig.locationCountry
  • Region values must match ProviderModelConfig.locationNetwork
  • Baseline includes imported provider rows and active BYOK provider entries

UI Behavior ​

In both Admin and non-admin policy editors:

  • Country/region fields are autocomplete selectors
  • You can only add values from the suggestion list
  • Pressing Enter adds the first match only (no free-text insertion)

Save Behavior on Invalid Values ​

If policy payloads include country/region values not present in your active provider baseline, save/update is rejected with 400 Bad Request.

This prevents silent typos like GERMANY when baseline values are DE.

Provider Groups ​

Restrict routing to specific provider groups you've configured. Leave empty to allow all providers.

Provider groups let you:

  • Create geographic clusters (e.g., "EU Premium Providers")
  • Define quality tiers (e.g., "High Quality", "Budget")
  • Organize by use case (e.g., "Code Generation", "Summarization")

Hard Constraints ​

Set minimum requirements providers must meet:

ConstraintDescriptionExample
Min AvailabilityMinimum uptime percentage0.95 (95%)
Max LatencyMaximum response time (ms)5000 (5 seconds)
Max Cost/TokenMaximum price per token0.0001

Providers not meeting constraints are excluded before scoring.

Hard constraints are absolute: if none of your providers satisfy them, the request does not route at all. When a requirement is ranked rather than absolute — "prefer a German provider, but accept a Swiss one" — express it with routing preference tiers instead, which sit on top of the hard constraints and degrade in a recorded, auditable way rather than blocking.

Fallback Configuration ​

Configure behavior when the primary provider fails:

SettingDescriptionDefault
Enable FallbackTry next-best provider on failureYes
Max RetriesNumber of retry attempts3
Retry DelayDelay between retries (ms)1000

How to Create a Routing Policy ​

  1. Navigate to Admin > Chat Settings > 1. Routing Policies
  2. Click Create Policy
  3. Enter a Name and optional Description
  4. Select Scope (Company, User, or Prompt level)
  5. Choose a Routing Strategy
  6. Adjust Preference Weights using the sliders
  7. (Optional) Expand Geofencing Rules and select regions/compliance requirements
  8. (Optional) Expand Advanced Settings for provider groups and hard constraints
  9. Click Create Policy

Policy Scope ​

ScopeDescription
CompanyApplies to all users in your company
UserApplies to a specific user
PromptApplies to a specific prompt

Lower-level scopes override higher-level scopes (Prompt > User > Company).

Sanitization Controls in Routing Policies ​

Routing policies can now also decide how PII sanitization behaves for executions that use the gateway.

Available modes ​

  • Disabled: never sanitize for this policy
  • Automatic: always sanitize before the request reaches the provider
  • Manual: sanitize only when a prompt run, API call, or user action explicitly enables it

Manual trigger rules ​

If a policy uses Manual mode, admins can choose which trigger sources are allowed:

  • prompt
  • api
  • user_action

This lets one policy enforce automatic sanitization for regulated flows while another policy allows selective sanitization for lower-risk work.

Example Policies ​

US-Only HIPAA Compliant ​

json
{
  "strategy": "HIGHEST_QUALITY",
  "geoFenceRules": {
    "allowedCountries": ["US"],
    "requireCompliance": ["HIPAA", "SOC2"]
  },
  "hardConstraints": {
    "minAvailability": 0.99
  }
}

Cost-Optimized Batch Processing ​

json
{
  "strategy": "LOWEST_COST",
  "preferences": {
    "weights": {
      "cost": 0.6,
      "latency": 0.1,
      "quality": 0.2,
      "stability": 0.1
    }
  },
  "fallback": {
    "enabled": true,
    "maxRetries": 5
  }
}

Low-Latency Real-Time ​

json
{
  "strategy": "LOWEST_LATENCY",
  "hardConstraints": {
    "maxLatencyMs": 2000
  },
  "preferences": {
    "weights": {
      "latency": 0.7,
      "quality": 0.2,
      "stability": 0.1
    }
  }
}

Simulating Policies ​

Test your policy before applying it:

  1. Open the policy in the list view
  2. Click Simulate
  3. Review the Primary Candidate and Fallback Candidates
  4. Check scores for latency, cost, quality, and availability

Using Policies with Chat Categories ​

Routing policies are assigned to Chat Categories:

  1. Create a routing policy
  2. Go to Admin > Chat Settings > 3. Categories
  3. Create or edit a category
  4. Select your routing policy from the dropdown

API Reference ​

Internal Service Policy Enforcement ​

Routing policies are automatically enforced for all provider calls, including internal platform services such as Guru, Prompt Optimizer, Architect, and MCP Executor. This means your company-level routing policy applies even when these services call an AI provider on your behalf.

How It Works ​

When an internal service resolves a provider, VeriPrompt checks the provider against your company's default routing policy before executing the request. The check validates:

  1. Provider allow/deny lists — is the provider type permitted?
  2. Model deny list — is the specific model blocked?
  3. Geo enforcement — does the provider's registered location comply with your geo rules?

If the provider fails any check, the request is blocked and you receive an actionable error message.

Policy Violation Error ​

When a service is blocked by your routing policy, the response looks like this:

json
{
  "success": false,
  "error": "[Policy Violation] \"Prompt Optimizer\" cannot use provider deepseek/deepseek-chat: Provider country CN is denied. Your company routing policy \"EU-Only Compliance\" restricts this provider. To enable this function, add a compliant provider via BYOK credentials (Settings → Credentials) or ask your admin to update the service configuration (Admin → Service Config).",
  "provider": "deepseek",
  "model": "deepseek-chat",
  "policyChecked": true,
  "policyName": "EU-Only Compliance"
}

Resolving Policy Violations ​

ResolutionSteps
Add a compliant provider (BYOK)Go to Settings → Credentials, add API keys for a provider in an allowed region
Update the service configurationGo to Admin → Service Config, change the provider/model assigned to the service
Adjust the routing policyGo to Admin → Chat Settings → Routing Policies, modify geo or provider restrictions

Metadata on Successful Execution ​

When a policy is evaluated and the provider passes, the response includes policy metadata:

json
{
  "success": true,
  "policyChecked": true,
  "policyName": "EU-Only Compliance"
}

Capability-Based Routing & Automatic Failover ​

Veriprompt resolves providers by capability, not by a hardcoded vendor name. This makes your routing resilient to model upgrades, provider outages, and catalog changes — usually with zero configuration changes on your part.

No hardcoded providers ​

Internal platform services — Prompt Optimizer, Guru (chat, explain, refine), Architect, and the MCP Executor — never pin a specific vendor or model in code. Neither does routing itself. Every provider is resolved at request time from your live Provider Management catalog plus any BYOK credentials you've added.

This means: when you change which providers are available, the services follow automatically. There is no code path that secretly defaults to a particular vendor.

Quality tiers ​

Instead of naming a specific model, a service or policy can reference a capability tier. Veriprompt then picks a matching-tier model from your live catalog:

TierIntentTypical use
cheapLowest costHigh-volume, low-stakes work
fastLowest latencyReal-time / interactive flows
qualityBest outputCritical or customer-facing tasks
securePrivacy-sensitiveRegulated or confidential data

Example defaults already in use:

  • Prompt Optimizer defaults to the cheap tier
  • Architect defaults to the quality tier

If you later add a stronger or cheaper model to the catalog in that tier, the service starts using it without any config edit.

Provider selectors ​

A dynamic provider group (or a service config) can be defined by a selector — a set of filters evaluated against the live catalog at request time. All clauses are combined with AND; empty clauses match everything.

FieldMeaning
providersProvider types to include (e.g. anthropic, openai)
tiersQuality tiers (cheap, fast, quality, secure)
regionsISO country codes for data residency (e.g. DE, FR)
excludeModelsSpecific model IDs to exclude
privacyTiersOptional privacy-tier filter

Example selector — "any quality-tier model from Anthropic or OpenAI, hosted in Germany or France":

json
{
  "providers": ["anthropic", "openai"],
  "tiers": ["quality"],
  "regions": ["DE", "FR"]
}

The selector resolves to a live, ordered pool of real models — never a frozen list.

Automatic cross-provider failover ​

At execution time, Veriprompt walks an ordered pool of candidate providers (not just models on one vendor). For each candidate it:

  1. Skips providers whose health circuit-breaker is open — a provider that has been failing recently (>50% error rate over its recent window) is automatically bypassed during its cooldown.
  2. Checks routing-policy compliance — geo rules and provider allow/deny lists are enforced; non-compliant candidates are skipped.
  3. Fails over on any runtime failure — a network error, 401 revoked key, 429 rate limit, 5xx, timeout, or model-not-found is recorded, and Veriprompt moves to the next candidate — which may be an entirely different provider.

Because the pool spans multiple providers, an entire provider going down (not just one of its models) is handled by switching to another provider automatically. No user action is needed.

Self-healing on catalog updates ​

When you update or re-import providers and models — for example, when a new model generation ships — routes and services automatically begin using the newest models that match their tier or selector.

Stored references to a model that no longer exists are validated against the live catalog and silently skipped, never attempted. A reseed or catalog refresh needs zero config edits: stale pins fall away, fresh models take over.

When to use tiers vs. selectors vs. explicit models ​

Reach for the least specific option that meets your need. Use a quality tier when you care about the outcome (cheap, fast, quality, secure) but not the vendor — it's the most resilient choice and upgrades itself as your catalog grows. Use a selector when you need to constrain which providers or regions are eligible (for data residency or vendor preference) while still letting Veriprompt pick the best match and fail over within that set. Pin an explicit model only when a workflow genuinely depends on one specific model's behavior — and remember that a pin is still validated against the live catalog, so it is skipped (not errored) if that model is retired.

Troubleshooting ​

No providers match policy ​

  • Check if geofencing rules are too restrictive
  • Verify providers have required compliance certifications
  • Lower hard constraint thresholds

Internal service blocked by policy ​

  • Check the error message for which provider and geo rule caused the violation
  • Add a compliant provider via BYOK (Settings → Credentials)
  • Or update the service's provider assignment (Admin → Service Config)

Unexpected provider selection ​

  • Review preference weights
  • Check if provider groups are restricting selection
  • Verify provider health status

Fallback not working ​

  • Ensure fallback is enabled in policy
  • Check max retries setting
  • Verify fallback providers meet constraints

Learn More ​