Appearance
Routing Advisory API
Overview
The Routing Advisory API provides intelligent routing recommendations for AI requests based on cost, performance, quality, and compliance requirements. It analyzes available providers and models to recommend the optimal routing strategy for each request.
Endpoint
POST /api/gateway/routing-advisoryAuthentication
Requires valid session authentication. Only authenticated users can access this endpoint.
Rate Limits
- Free Tier: 500 requests/hour
- Standard Tier: 5,000 requests/hour
- Professional Tier: 50,000 requests/hour
- Enterprise Tier: Unlimited
Request Format
Headers
http
Content-Type: application/json
Authorization: Bearer <session-token>Request Body
json
{
"requestId": "string (required)",
"prompt": {
"text": "string (required)",
"tokens": number,
"type": "completion" | "chat" | "embedding" | "image" | "code" | "analysis",
"expectedResponseTokens": number
},
"requirements": {
"maxLatency": number,
"maxCost": number,
"minQuality": number,
"preferredProviders": ["string"],
"excludeProviders": ["string"],
"features": ["streaming", "function_calling", "vision", "code_execution"]
},
"context": {
"userId": "string (required)",
"companyId": "string (required)",
"tier": "enterprise" | "professional" | "standard" | "free",
"priority": "realtime" | "high" | "normal" | "low" | "batch",
"region": "string",
"industry": "string"
}
}Field Descriptions
| Field | Type | Required | Description |
|---|---|---|---|
requestId | string | ✅ | Unique identifier for this routing request |
prompt.text | string | ✅ | The prompt text to be processed |
prompt.tokens | number | ❌ | Estimated input token count |
prompt.type | enum | ✅ | Type of AI task being performed |
prompt.expectedResponseTokens | number | ❌ | Expected output token count |
requirements.maxLatency | number | ❌ | Maximum acceptable latency (ms) |
requirements.maxCost | number | ❌ | Maximum cost per request (USD) |
requirements.minQuality | number | ❌ | Minimum quality score (0-100) |
requirements.preferredProviders | array | ❌ | Preferred provider names |
requirements.excludeProviders | array | ❌ | Providers to exclude |
requirements.features | array | ❌ | Required model features |
context.userId | string | ✅ | User making the request |
context.companyId | string | ✅ | Company/organization ID |
context.tier | enum | ✅ | Service tier |
context.priority | enum | ❌ | Request priority level |
context.region | string | ❌ | Preferred geographic region |
context.industry | string | ❌ | Industry context for compliance |
Response Format
Success Response (200 OK)
json
{
"requestId": "string",
"recommendation": {
"primary": {
"provider": "string",
"model": "string",
"endpoint": "string",
"estimatedLatency": number,
"estimatedCost": number,
"qualityScore": number,
"availability": number,
"reasoning": "string"
},
"fallback": [
{
"provider": "string",
"model": "string",
"endpoint": "string",
"triggerCondition": "string"
}
],
"loadBalancing": {
"strategy": "round_robin" | "weighted" | "least_latency" | "cost_optimized",
"distribution": {
"provider_name": number
}
}
},
"costAnalysis": {
"estimatedCost": number,
"costBreakdown": {
"inputTokens": number,
"outputTokens": number,
"apiCalls": number,
"totalCost": number
},
"savings": {
"comparedToDefault": number,
"optimizationApplied": ["string"]
}
},
"performanceMetrics": {
"expectedLatency": {
"p50": number,
"p95": number,
"p99": number
},
"throughput": number,
"successRate": number
},
"compliance": {
"dataResidency": boolean,
"gdprCompliant": boolean,
"hipaaCompliant": boolean,
"sox": boolean
},
"warnings": ["string"],
"metadata": {
"decisionTime": number,
"factorsConsidered": ["string"],
"cacheHit": boolean
}
}Response Field Descriptions
| Field | Type | Description |
|---|---|---|
requestId | string | Echo of request ID |
recommendation.primary | object | Primary routing recommendation |
recommendation.primary.provider | string | Recommended provider (e.g., "openai") |
recommendation.primary.model | string | Recommended model (e.g., "gpt-4-turbo") |
recommendation.primary.endpoint | string | API endpoint URL |
recommendation.primary.estimatedLatency | number | Expected latency (ms) |
recommendation.primary.estimatedCost | number | Expected cost (USD) |
recommendation.primary.qualityScore | number | Quality score (0-100) |
recommendation.primary.availability | number | Availability percentage (0-1) |
recommendation.primary.reasoning | string | Explanation of recommendation |
recommendation.fallback | array | Fallback options if primary fails |
recommendation.loadBalancing | object | Load balancing configuration (enterprise) |
costAnalysis | object | Detailed cost breakdown |
performanceMetrics | object | Expected performance characteristics |
compliance | object | Compliance certifications |
warnings | array | Warnings about the recommendation |
metadata | object | Processing metadata |
Load Balancing Strategies
| Strategy | Description | Use Case |
|---|---|---|
round_robin | Distribute requests evenly | Even load distribution |
weighted | Distribute based on provider capacity | Optimized resource usage |
least_latency | Route to fastest available provider | Latency-critical applications |
cost_optimized | Route to most cost-effective provider | Budget-conscious workloads |
Trigger Conditions
| Condition | Description |
|---|---|
primary_timeout | Primary provider response timeout |
primary_error | Primary provider returns error |
rate_limit | Primary provider rate limit exceeded |
availability | Primary provider availability below threshold |
quality_degradation | Primary provider quality scores declining |
Error Responses
400 Bad Request
json
{
"error": "Missing required fields",
"details": {
"missingFields": ["requestId", "prompt.text"]
}
}401 Unauthorized
json
{
"error": "Unauthorized",
"message": "Valid authentication required"
}404 Not Found
json
{
"error": "No suitable providers found",
"message": "No providers match the specified requirements",
"suggestions": [
"Relax quality requirements",
"Increase budget constraints",
"Remove provider exclusions"
]
}429 Too Many Requests
json
{
"error": "Rate limit exceeded",
"retryAfter": 3600,
"limit": {
"requests": 500,
"window": "hour",
"tier": "free"
}
}500 Internal Server Error
json
{
"error": "Failed to generate routing recommendation",
"requestId": "string"
}Provider Coverage
Supported Providers
| Provider | Models | Features | Regions |
|---|---|---|---|
| OpenAI | GPT-4 Turbo, GPT-3.5 Turbo | Streaming, Function Calling, Vision | US, EU |
| Anthropic | Claude-3 Opus, Claude-3 Sonnet | Streaming, Vision, Long Context | US, EU |
| DeepSeek | DeepSeek-Coder, DeepSeek-Chat | Streaming, Code Execution | US, Asia |
| Gemini Pro, Gemini Ultra | Streaming, Multimodal | Global | |
| Azure OpenAI | GPT-4, GPT-3.5 | Enterprise, HIPAA | Global |
Model Capabilities
| Feature | Description | Providers |
|---|---|---|
streaming | Real-time response streaming | All |
function_calling | Structured function calls | OpenAI, Google |
vision | Image understanding | OpenAI, Anthropic, Google |
code_execution | Code interpretation | DeepSeek, OpenAI |
long_context | Extended context windows | Anthropic, Google |
Example Usage
cURL
bash
curl -X POST https://app.veriprompt.tech/api/gateway/routing-advisory \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <your-token>" \
-d '{
"requestId": "route_123",
"prompt": {
"text": "Write a Python function to calculate fibonacci numbers",
"type": "code",
"tokens": 150,
"expectedResponseTokens": 300
},
"requirements": {
"maxLatency": 2000,
"maxCost": 0.01,
"features": ["code_execution"]
},
"context": {
"userId": "user_456",
"companyId": "company_789",
"tier": "professional",
"priority": "normal"
}
}'JavaScript
javascript
const response = await fetch('/api/gateway/routing-advisory', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${token}`
},
body: JSON.stringify({
requestId: 'route_123',
prompt: {
text: 'Write a Python function to calculate fibonacci numbers',
type: 'code',
tokens: 150,
expectedResponseTokens: 300
},
requirements: {
maxLatency: 2000,
maxCost: 0.01,
features: ['code_execution']
},
context: {
userId: 'user_456',
companyId: 'company_789',
tier: 'professional',
priority: 'normal'
}
})
});
const routing = await response.json();
console.log('Recommended provider:', routing.recommendation.primary.provider);
console.log('Estimated cost:', routing.costAnalysis.estimatedCost);Python
python
import requests
import json
url = 'https://app.veriprompt.tech/api/gateway/routing-advisory'
headers = {
'Content-Type': 'application/json',
'Authorization': f'Bearer {token}'
}
data = {
'requestId': 'route_123',
'prompt': {
'text': 'Write a Python function to calculate fibonacci numbers',
'type': 'code',
'tokens': 150,
'expectedResponseTokens': 300
},
'requirements': {
'maxLatency': 2000,
'maxCost': 0.01,
'features': ['code_execution']
},
'context': {
'userId': 'user_456',
'companyId': 'company_789',
'tier': 'professional',
'priority': 'normal'
}
}
response = requests.post(url, headers=headers, data=json.dumps(data))
routing = response.json()
print(f"Recommended provider: {routing['recommendation']['primary']['provider']}")
print(f"Estimated cost: ${routing['costAnalysis']['estimatedCost']:.4f}")
print(f"Expected latency: {routing['performanceMetrics']['expectedLatency']['p50']}ms")Optimization Strategies
Cost Optimization
- Provider Selection: Routes to most cost-effective providers
- Model Matching: Selects appropriate model complexity for task
- Volume Discounts: Considers volume pricing tiers
- Regional Pricing: Factors in regional cost differences
Performance Optimization
- Latency Minimization: Prioritizes fastest available providers
- Geographic Routing: Routes to nearest available regions
- Load Balancing: Distributes load across multiple providers
- Caching: Leverages response caching when appropriate
Quality Assurance
- Model Benchmarking: Uses quality scores from standardized benchmarks
- Task Specialization: Matches models to specific task types
- Failure Detection: Monitors and routes around failing providers
- A/B Testing: Supports routing experiments for quality assessment
Advanced Features
Enterprise Load Balancing
Enterprise customers can access advanced load balancing features:
- Multi-provider distribution: Spread requests across multiple providers
- Failover cascading: Automatic failover through multiple fallback options
- Custom routing rules: Define custom routing logic based on business rules
- Traffic shaping: Control request distribution patterns
Real-time Adaptation
The routing engine continuously adapts recommendations based on:
- Live performance data: Real-time latency and error rate monitoring
- Capacity monitoring: Provider capacity and queue depth tracking
- Cost fluctuations: Dynamic pricing updates from providers
- Quality metrics: Ongoing quality assessment and scoring
Compliance Features
- Data residency: Ensure data stays within specified geographic boundaries
- Regulatory compliance: Route based on GDPR, HIPAA, SOX requirements
- Audit logging: Comprehensive logging of all routing decisions
- Encryption requirements: Factor in encryption and security requirements
Best Practices
Implementation
- Handle fallbacks: Always implement fallback logic for failed recommendations
- Monitor costs: Track actual vs estimated costs for budget management
- Cache decisions: Cache routing decisions for identical requests
- Update regularly: Refresh routing recommendations based on changing requirements
Performance
- Batch requests: Group multiple routing requests when possible
- Pre-fetch recommendations: Get routing advice ahead of actual requests
- Monitor metrics: Track latency, cost, and quality metrics
- Use webhooks: Implement async processing for high-volume scenarios
Cost Management
- Set budgets: Define clear cost constraints in requirements
- Monitor spending: Track actual costs against recommendations
- Optimize over time: Use historical data to refine cost models
- Volume planning: Consider volume discounts in routing decisions
Changelog
Version 1.0.0 (Current)
- Initial release
- Support for 5 major providers
- Cost and performance optimization
- Compliance-aware routing
- Enterprise load balancing features
