Appearance
Task-Aware Routing: Setup and Operation
How a platform administrator makes task-aware routing available, keeps its benchmark data healthy, and reads the evidence behind a ranking. The customer-facing explanation is in Task-Aware Routing.
Overview
Task-aware routing needs three things to be in place:
- The package feature, so the right companies can switch it on.
- Benchmark evidence, ingested automatically from public leaderboards into the capability matrix.
- Model-name mapping, so a leaderboard's name for a model lines up with the model in your catalog.
Once these are set, each company decides per routing policy whether to use it.
Access and permissions
| Task | Role |
|---|---|
| Assign the package feature to a plan | Platform super admin (Admin > Packages) |
| Capability Matrix and the unmapped-names queue | Platform super admin only. The role is re-checked on the server for every call; company admins do not see the page |
| Per-model evidence in the model configuration dialog | Whoever may edit the company's model catalog |
| Turn task fit on in a policy | Company roles that can edit routing policies |
1. Enable the package feature
Task-fit routing is the package capability Task-Fit Routing (allowTaskFitRouting). It is default-deny: a plan without the key does not have it.
- Open Admin > Packages, edit the plan, and tick Task-Fit Routing under capabilities. See Packages.
- Save. Companies on that plan see the effect within about a minute (entitlements are cached for 60 seconds).
Out of the box the feature is granted to Team Professional, Team Growth, Enterprise and the Sandbox Trial plan, and explicitly denied on the entry-level plans, so the package editor shows the decision rather than an empty box.
What the gate does, in both places it is enforced:
- Saving a policy with task fit enabled is refused (403,
FEATURE_NOT_ENTITLED) when the company's package does not include it. - At runtime, task fit and the classification spend apply only while the package grants it. A downgrade silently returns the company to standard routing. There is no grace period, because this is a value feature and not a compliance control.
Existing databases
On a new or reseeded database the plans already carry the key. On a database that was seeded earlier, existing plan records do not have it and the feature stays closed for everybody. Tick Task-Fit Routing on the plans that should have it in Admin > Packages, or ask your Veriprompt contact to apply the one-key update, which changes only this capability on the named plans and leaves companies, users, keys and usage data untouched. Do not reseed a shared environment for this.
2. Benchmark ingestion
A scheduled job refreshes the capability matrix from public leaderboards. You do not run it by hand.
| Source | Capabilities fed | License |
|---|---|---|
| LiveBench | Coding, Math, Reasoning, Instruction following, Writing, General | Apache-2.0 (project license) |
| Epoch AI | Reasoning, Math, Coding (agent-harness benchmark) | CC-BY-4.0 |
| LMArena | General, Coding, Math, Writing, Instruction following, German | CC-BY-4.0 |
| EuroEval | German | MIT |
| BFCL (Berkeley Function Calling) | Tool use | Apache-2.0 |
| Aider polyglot | Coding (agent-harness benchmark) | Apache-2.0 |
| Vectara hallucination leaderboard | Factuality | Apache-2.0 |
Only sources with permissive licenses are used. Sources with restrictive terms are deliberately not ingested.
How it behaves:
- Cadence. The job runs daily, but a source that was refreshed in the last six days is skipped without a network call. Healthy sources therefore refresh about weekly, and a source that failed is retried the next day.
- Sources fail independently. One broken leaderboard does not stop the others. A run counts as failed only when every source it attempted failed.
- Append-only. Each run writes a new snapshot; nothing is overwritten. Re-running on the same day changes nothing.
- Percentile storage. Scores are stored as the model's standing within that benchmark's leaderboard (0 is last, 1 is first), with the published raw value kept alongside. Routing maps this onto the same scale as the quality tiers so evidence and tiers are comparable.
3. The Capability Matrix
Open Admin > Capability Matrix (super admins only). The page has two tabs.
Matrix
A table of models by capability. The number in a cell is exactly the value routing uses for that model and capability, and the colour shows where it comes from.
| Kind | Meaning |
|---|---|
| External | Third-party benchmark evidence |
| Own data | Evidence from company ratings. This page is platform-wide and shows global rows only, so this kind appears for company-scoped evidence only in the company views below |
| Manual | The model carries a specialization tag for it |
| Tier | Only the static quality tier is available |
- The Latest snapshot per source panel lists each source, its snapshot, when it was fetched, how many rows it has and its age. A source marked Overdue is older than expected: look at it first when a ranking looks stale.
- Use the filters (model search, provider, capability, source, Only with evidence) to narrow the table.
- Select a cell to open the evidence drawer: every row behind the value, with the raw value, rank within the list, value on the tier scale, run date, license, snapshot and a link to the source. Rows older than the staleness limit are marked Stale and no longer count; if all are stale, the value falls back to the quality tier. Agent-harness benchmarks are marked Scaffold-dependent.
Unmapped names
Leaderboards name models loosely. A name that cannot be matched to a catalog model is not written to the matrix; it appears in the Unmapped names queue instead, with the benchmarks it appeared in, how many days it was seen, and when it was last seen.
For each name you can:
| Action | Effect |
|---|---|
| Map to model | Pick the catalog model the leaderboard means. The mapping applies from the next fetch; rows already stored do not change |
| Ignore | Keep the name out of the queue (for models you do not offer) |
| Reopen | Move a mapped or ignored name back to pending |
Every decision is recorded in the audit log. Work through the queue after you add models to the catalog, and whenever a model you care about shows only Tier evidence.
Interpreting evidence and staleness
- Evidence beats a tier above roughly the 44th percentile. A benchmark percentile is mapped linearly onto the tier range (0.5 to 0.95). The default tier (0.7) is reached at about the 44th percentile, so a model that is at or above that standing on a benchmark outranks an untagged model of default tier. Models in the quality, premium and enterprise tiers still beat a median evidenced model.
- Staleness. External rows older than Benchmarks expire after (days) in the policy (default 120) carry no weight. Some sources publish rarely; BFCL and Aider results are often older than 120 days and then simply do not count.
- Several benchmarks for one capability are combined with a weighted median; agent-harness benchmarks together can carry at most half of the external weight.
- Own evidence outranks external evidence once a model has enough samples (policy setting, default 30).
Internal (own) evidence and company scoping
Each company's own evidence is derived from prompt ratings (ratings of 4 or higher on the quality fields count as positive) of prompts with a task, over a trailing 90 days, recomputed daily.
- Every internal row is scoped to its company. A company's routing reads the global benchmark rows plus its own rows, never another company's. One customer's feedback cannot steer another customer's routing.
- Internal rows that are not restated within seven days expire, so evidence from a company that stopped rating does not linger.
- Chat thumbs and evaluation runs do not feed evidence yet (they carry no link to a task or a model).
- In the global Capability Matrix you see platform-wide benchmark evidence. Company evidence is visible to that company in the policy preview, the Model Radar and the model configuration dialog.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| A company cannot enable the option | Plan lacks Task-Fit Routing | Tick it in the package (older plan records do not carry the key, see "Existing databases") |
| Matrix is empty | No ingestion has completed yet | Wait for the daily job; check the snapshot panel |
| A source is Overdue | Repeated fetch failures (leaderboard changed format, rate limit) | Ask engineering to check the run log; other sources are unaffected |
| A model shows only Tier evidence | Its leaderboard name is unmapped | Map it in Unmapped names |
| Mapping saved but nothing changed | Mappings apply from the next fetch | Wait for the next run |
| Ranking differs between two companies | Their own evidence, policy thresholds or external-evidence setting differ | Compare the policy preview of each |
