Appearance
Provider Health Quality Metrics
Overview
/api/v1/providers/health now returns composite quality telemetry alongside the existing trust snapshot. The API blends execution data from scheduled benchmarks, synthetic prompts, and routing telemetry to expose a Veriprompt-specific quality score, gauge, and drift signal per provider/model/category.
Response Fields
The quality object is included for each snapshot:
| Field | Type | Description |
|---|---|---|
score | number | Weighted 0-100 composite quality score using Veriprompt axes (correctness, spec adherence, code quality, refusal handling, recovery, stability). |
gauge | number | Display-friendly value (0-100) used in dashboards. |
zScore | number | Z-score versus the trailing 30-day baseline for the same provider/model/category window. |
driftDetected | boolean | true when ` |
axes | object | Per-axis averages in the 0-1 range (keys: correctness, specAdherence, codeQuality, refusalHandling, recovery, stability). |
baselineMean | number | Historical mean quality score for the baseline sample. |
baselineStd | number | Standard deviation of the baseline sample. |
sampleSize | number | Number of historical snapshots contributing to the baseline figures. |
Axes and Weights
| Axis | Weight | Notes |
|---|---|---|
| Correctness | 0.30 | Derived from execution success rate and quality metrics per benchmark. |
| Spec Adherence | 0.10 | Aligns prompt outputs with specification and policy checks. |
| Code Quality | 0.20 | Aggregates structural and maintainability signals. |
| Refusal Handling | 0.15 | Rewards compliant handling of disallowed prompts while avoiding excessive refusals. |
| Recovery | 0.10 | Measures fallback effectiveness when primary routes fail. |
| Stability | 0.15 | Tracks response consistency and latency variation. |
Dashboard Integration
- Provider detail pages surface the
quality.gaugewith a five-level verification heatmap (UI update pending). - Drift flags (
quality.driftDetected) generate “needs attention” callouts under Monitoring. - Routing advisories will soon incorporate the z-score to modulate provider preferences.
Usage Example
json
{
"provider": "openai",
"model": "gpt-4.1",
"window": "ONE_HOUR",
"quality": {
"score": 82.4,
"gauge": 82,
"zScore": -0.8,
"driftDetected": false,
"axes": {
"correctness": 0.88,
"specAdherence": 0.84,
"codeQuality": 0.79,
"refusalHandling": 0.92,
"recovery": 0.61,
"stability": 0.74
},
"baselineMean": 78.3,
"baselineStd": 5.2,
"sampleSize": 26
}
}Notes
- Baseline calculations pull up to 60 prior snapshots for the same window (hourly, six-hour, or 24-hour).
- Axes fall back to success rate and latency variance when specialised metrics are unavailable.
- Additional human and prediction-based metrics will be added once annotation tooling is connected.
