Appearance
Synthetic Testing
Synthetic testing runs automated tests against AI providers using curated prompts to detect regressions, monitor performance, and validate security controls.
Why it matters
- Catch quality drops before production: Automated testing identifies issues before they impact users
- Compare providers and models: Benchmark performance, cost, and security across your AI stack
- Track results over time: Historical data reveals trends and regressions
- Validate security controls: Test prompt injection defenses with known attack patterns
- Enterprise scalability: Archival system handles thousands of users and 1000+ providers
- KPI-driven routing: Collect structured performance metrics to optimize provider selection
Key Features
Test Prompts
- Create test prompts with expected behaviors (ALLOW, REJECT, SCORE)
- Categorize by type: security, functional, performance, edge cases
- Tag and search prompts for organization
- Variable substitution for dynamic testing
- KPI tracking toggle: Enable structured KPI collection per prompt
Scheduling
- One-time tests: Schedule tests for specific dates/times
- Recurring tests: Automated testing at intervals (every 10 minutes to weekly)
- Provider targeting: Test individual providers or provider groups
- Protection testing: Apply security wrappers to validate defenses
Execution Logging
- Detailed metrics: response time (seconds), tokens, cost, success rate
- Security analysis: attack detection, protection blocking
- Network metrics: route hops, latency, cross-border tracking
- Comprehensive audit trail
- KPI compliance tracking: Flag executions that return structured metrics
KPI Tracking & Compliance
VeriPrompt's KPI tracking system collects structured performance metrics from AI model responses to enable intelligent routing optimization.
How KPI Tracking Works
- Enable KPI Tracking: Toggle "Include KPI Tracking" when creating or editing a synthetic prompt
- Structured Metrics Request: The system appends a JSON-formatted KPI instruction to prompts
- Metric Extraction: AI responses are parsed for structured KPI data
- Compliance Flagging: Executions are marked as "Compliant" or "Fallback" based on whether structured KPIs were returned
KPI Metrics Collected
When models return structured KPIs, the following metrics are extracted:
| Metric | Description | Range |
|---|---|---|
| Response Completeness | How fully the response addresses the prompt | 0-10 |
| Accuracy | Factual correctness of the response | 0-10 |
| Clarity | Readability and organization | 0-10 |
| Relevance | How well the response stays on topic | 0-10 |
| Helpfulness | Actionability and usefulness | 0-10 |
Fallback Metrics
When a model doesn't return structured KPIs, VeriPrompt automatically collects fallback metrics to ensure all executions contribute to routing optimization:
| Fallback Metric | Description |
|---|---|
| Word Count | Total words in response |
| Sentence Count | Number of sentences |
| Response Complexity | SIMPLE, MODERATE, or COMPLEX |
| Structure Score | Presence of headers, lists, code blocks (0-10) |
| Reading Time | Estimated reading time in seconds |
| Has Code Blocks | Whether response includes code |
| Has Lists | Whether response includes bullet/numbered lists |
Compliance Indicators
In the Execution History, each execution displays a compliance badge:
- ✓ Compliant (green): Model returned structured KPI metrics
- ⚠ Fallback (orange): Fallback metrics collected (model didn't return KPIs)
Using KPI Data for Routing
KPI compliance rates by provider are available in the Analytics dashboard:
Provider/Model | KPI Compliance | Total Tests
-----------------------|----------------|------------
anthropic/claude-3.5 | 95% | 1,247
openai/gpt-4o | 92% | 1,156
google/gemini-pro | 78% | 892Models with higher KPI compliance contribute more accurate data to routing decisions, helping you:
- Identify which providers consistently return quality metrics
- Optimize routing rules based on real performance data
- Track improvements as providers update their models
Provider Selection Import/Export
Save and restore your test configurations using JSON or Excel formats:
| Format | Best For |
|---|---|
| Excel | Spreadsheet editing, sharing with non-technical users |
| JSON | Programmatic processing, version control |
Key features:
- Export current provider/model selections and group assignments
- Import with automatic validation and error checking
- Share configurations across team members
- Backup before making changes
See Provider Selection Import/Export for detailed documentation.
Log Archival System
For enterprise deployments with high-volume testing, the archival system ensures sustainable storage:
| Feature | Description |
|---|---|
| Retention Period | Configurable (default: 30 days) |
| Aggregation | Key statistics preserved before deletion |
| Max Limit | 50,000 logs per company safety cap |
| Export | JSON export for long-term storage |
| Automation | API-driven or scheduled archival |
Aggregated Metrics Preserved
When logs are archived, these statistics are aggregated per provider/model:
- Total executions, success/failure counts
- Response time (avg, min, max)
- Token usage and cost estimates
- Attack detection counts
- Time period covered
Simple Workflow
- Create prompts: Build a library of test prompts by category
- Configure schedules: Set up one-time or recurring test runs
- Run tests: Execute against individual providers or groups
- Review results: Analyze scores, trends, and anomalies
- Archive logs: Periodically archive old logs to maintain performance
Access Control
Synthetic testing is available to users with these roles:
- Super Admin
- SaaS Admin
- Admin
- Synthetic Admin
- Synthetic Operator (view/run only)
- Security Custodian
Note: Protective Prompts (security wrappers) are VeriPrompt intellectual property and only visible to Super Admin users.
Best Practices
Testing Strategy
- Maintain prompts for: happy path, edge cases, abuse patterns
- Test security prompts at least daily
- Compare results across provider updates
Archival Strategy
- Archive weekly for high-volume testing (500+ daily executions)
- Archive monthly for moderate usage (50-500 daily)
- Always export before deletion for compliance
Scale Recommendations
| Daily Executions | Retention | Archive Frequency |
|---|---|---|
| < 50 | 90 days | Monthly |
| 50-500 | 60 days | Bi-weekly |
| 500-2000 | 30 days | Weekly |
| 2000+ | 14 days | Daily |
