Skip to content

Provider Health Quality Metrics ​

Overview ​

/api/v1/providers/health now returns composite quality telemetry alongside the existing trust snapshot. The API blends execution data from scheduled benchmarks, synthetic prompts, and routing telemetry to expose a Veriprompt-specific quality score, gauge, and drift signal per provider/model/category.

Response Fields ​

The quality object is included for each snapshot:

FieldTypeDescription
scorenumberWeighted 0-100 composite quality score using Veriprompt axes (correctness, spec adherence, code quality, refusal handling, recovery, stability).
gaugenumberDisplay-friendly value (0-100) used in dashboards.
zScorenumberZ-score versus the trailing 30-day baseline for the same provider/model/category window.
driftDetectedbooleantrue when `
axesobjectPer-axis averages in the 0-1 range (keys: correctness, specAdherence, codeQuality, refusalHandling, recovery, stability).
baselineMeannumber Historical mean quality score for the baseline sample.
baselineStdnumberStandard deviation of the baseline sample.
sampleSizenumberNumber of historical snapshots contributing to the baseline figures.

Axes and Weights ​

AxisWeightNotes
Correctness0.30Derived from execution success rate and quality metrics per benchmark.
Spec Adherence0.10Aligns prompt outputs with specification and policy checks.
Code Quality0.20Aggregates structural and maintainability signals.
Refusal Handling0.15Rewards compliant handling of disallowed prompts while avoiding excessive refusals.
Recovery0.10Measures fallback effectiveness when primary routes fail.
Stability0.15Tracks response consistency and latency variation.

Dashboard Integration ​

  • Provider detail pages surface the quality.gauge with a five-level verification heatmap (UI update pending).
  • Drift flags (quality.driftDetected) generate “needs attention” callouts under Monitoring.
  • Routing advisories will soon incorporate the z-score to modulate provider preferences.

Usage Example ​

json
{
  "provider": "openai",
  "model": "gpt-4.1",
  "window": "ONE_HOUR",
  "quality": {
    "score": 82.4,
    "gauge": 82,
    "zScore": -0.8,
    "driftDetected": false,
    "axes": {
      "correctness": 0.88,
      "specAdherence": 0.84,
      "codeQuality": 0.79,
      "refusalHandling": 0.92,
      "recovery": 0.61,
      "stability": 0.74
    },
    "baselineMean": 78.3,
    "baselineStd": 5.2,
    "sampleSize": 26
  }
}

Notes ​

  • Baseline calculations pull up to 60 prior snapshots for the same window (hourly, six-hour, or 24-hour).
  • Axes fall back to success rate and latency variance when specialised metrics are unavailable.
  • Additional human and prediction-based metrics will be added once annotation tooling is connected.