Skip to content

Synthetic Testing ​

Synthetic testing runs automated tests against AI providers using curated prompts to detect regressions, monitor performance, and validate security controls.

Why it matters ​

  • Catch quality drops before production: Automated testing identifies issues before they impact users
  • Compare providers and models: Benchmark performance, cost, and security across your AI stack
  • Track results over time: Historical data reveals trends and regressions
  • Validate security controls: Test prompt injection defenses with known attack patterns
  • Enterprise scalability: Archival system handles thousands of users and 1000+ providers
  • KPI-driven routing: Collect structured performance metrics to optimize provider selection

Key Features ​

Test Prompts ​

  • Create test prompts with expected behaviors (ALLOW, REJECT, SCORE)
  • Categorize by type: security, functional, performance, edge cases
  • Tag and search prompts for organization
  • Variable substitution for dynamic testing
  • KPI tracking toggle: Enable structured KPI collection per prompt

Scheduling ​

  • One-time tests: Schedule tests for specific dates/times
  • Recurring tests: Automated testing at intervals (every 10 minutes to weekly)
  • Provider targeting: Test individual providers or provider groups
  • Protection testing: Apply security wrappers to validate defenses

Execution Logging ​

  • Detailed metrics: response time (seconds), tokens, cost, success rate
  • Security analysis: attack detection, protection blocking
  • Network metrics: route hops, latency, cross-border tracking
  • Comprehensive audit trail
  • KPI compliance tracking: Flag executions that return structured metrics

KPI Tracking & Compliance ​

VeriPrompt's KPI tracking system collects structured performance metrics from AI model responses to enable intelligent routing optimization.

How KPI Tracking Works ​

  1. Enable KPI Tracking: Toggle "Include KPI Tracking" when creating or editing a synthetic prompt
  2. Structured Metrics Request: The system appends a JSON-formatted KPI instruction to prompts
  3. Metric Extraction: AI responses are parsed for structured KPI data
  4. Compliance Flagging: Executions are marked as "Compliant" or "Fallback" based on whether structured KPIs were returned

KPI Metrics Collected ​

When models return structured KPIs, the following metrics are extracted:

MetricDescriptionRange
Response CompletenessHow fully the response addresses the prompt0-10
AccuracyFactual correctness of the response0-10
ClarityReadability and organization0-10
RelevanceHow well the response stays on topic0-10
HelpfulnessActionability and usefulness0-10

Fallback Metrics ​

When a model doesn't return structured KPIs, VeriPrompt automatically collects fallback metrics to ensure all executions contribute to routing optimization:

Fallback MetricDescription
Word CountTotal words in response
Sentence CountNumber of sentences
Response ComplexitySIMPLE, MODERATE, or COMPLEX
Structure ScorePresence of headers, lists, code blocks (0-10)
Reading TimeEstimated reading time in seconds
Has Code BlocksWhether response includes code
Has ListsWhether response includes bullet/numbered lists

Compliance Indicators ​

In the Execution History, each execution displays a compliance badge:

  • ✓ Compliant (green): Model returned structured KPI metrics
  • ⚠ Fallback (orange): Fallback metrics collected (model didn't return KPIs)

Using KPI Data for Routing ​

KPI compliance rates by provider are available in the Analytics dashboard:

Provider/Model         | KPI Compliance | Total Tests
-----------------------|----------------|------------
anthropic/claude-3.5   | 95%            | 1,247
openai/gpt-4o          | 92%            | 1,156
google/gemini-pro      | 78%            | 892

Models with higher KPI compliance contribute more accurate data to routing decisions, helping you:

  • Identify which providers consistently return quality metrics
  • Optimize routing rules based on real performance data
  • Track improvements as providers update their models

Provider Selection Import/Export ​

Save and restore your test configurations using JSON or Excel formats:

FormatBest For
ExcelSpreadsheet editing, sharing with non-technical users
JSONProgrammatic processing, version control

Key features:

  • Export current provider/model selections and group assignments
  • Import with automatic validation and error checking
  • Share configurations across team members
  • Backup before making changes

See Provider Selection Import/Export for detailed documentation.

Log Archival System ​

For enterprise deployments with high-volume testing, the archival system ensures sustainable storage:

FeatureDescription
Retention PeriodConfigurable (default: 30 days)
AggregationKey statistics preserved before deletion
Max Limit50,000 logs per company safety cap
ExportJSON export for long-term storage
AutomationAPI-driven or scheduled archival

Aggregated Metrics Preserved ​

When logs are archived, these statistics are aggregated per provider/model:

  • Total executions, success/failure counts
  • Response time (avg, min, max)
  • Token usage and cost estimates
  • Attack detection counts
  • Time period covered

Simple Workflow ​

  1. Create prompts: Build a library of test prompts by category
  2. Configure schedules: Set up one-time or recurring test runs
  3. Run tests: Execute against individual providers or groups
  4. Review results: Analyze scores, trends, and anomalies
  5. Archive logs: Periodically archive old logs to maintain performance

Access Control ​

Synthetic testing is available to users with these roles:

  • Super Admin
  • SaaS Admin
  • Admin
  • Synthetic Admin
  • Synthetic Operator (view/run only)
  • Security Custodian

Note: Protective Prompts (security wrappers) are VeriPrompt intellectual property and only visible to Super Admin users.

Best Practices ​

Testing Strategy ​

  • Maintain prompts for: happy path, edge cases, abuse patterns
  • Test security prompts at least daily
  • Compare results across provider updates

Archival Strategy ​

  • Archive weekly for high-volume testing (500+ daily executions)
  • Archive monthly for moderate usage (50-500 daily)
  • Always export before deletion for compliance

Scale Recommendations ​

Daily ExecutionsRetentionArchive Frequency
< 5090 daysMonthly
50-50060 daysBi-weekly
500-200030 daysWeekly
2000+14 daysDaily

Learn More ​