Skip to content

Task-Aware Routing ​

Different models are good at different jobs. One model writes the cleanest code, another handles long German contracts better, a third is the cheapest sensible choice for a quick summary. Task-aware routing looks at what each request is about, checks which of your allowed models has the best track record for that kind of work, and ranks them accordingly.

You do not have to label requests. You do not have to maintain a model list. Your security rules do not change.

Overview ​

Without task-aware routing, a routing policy ranks models by price, speed, stability and one general quality number per model. With it, the quality number becomes specific to the task: a request that looks like coding is scored on coding evidence, a German contract question on legal and German-language evidence, and so on.

Three things make this safe to switch on:

  • It only reorders, never admits or excludes. Hard constraints, geofencing, preference tiers, provider groups and compliance rules decide which models are allowed. Task fit then ranks those models. A model your policy forbids can never be chosen because it is good at the task.
  • It is off by default. A policy has to opt in. Existing policies behave exactly as before until you turn it on.
  • It is explainable. Every ranking can show which task was detected, which benchmarks or own results were used, and from which date.

Who is it for: teams that send many kinds of work through the gateway (code, contracts, summaries, translations) and want each kind to land on a model that is measurably strong at it, without hand-writing a rule per task.

Access and permissions ​

WhatWhoNotes
Turn task-fit routing on in a policyAccount Owners, Account Admins and Admins, plus anyone whose access profile grants the routing-policy edit permissionThe company's package must include Task-Fit Routing
Pin a task on a promptAny company member who can see the promptStored even without the package feature; applied only with it
Preview a policy, read explanationsSame roles as policy simulation: the admin roles above, or the routing-policy view permissionThe policy must belong to your company (or be a global policy)
Capability Matrix, benchmark name queuePlatform super admins onlySee the admin guide

Package requirement. Task-fit routing is a package feature (allowTaskFitRouting). It is included in the Team Professional, Team Growth, Enterprise and Sandbox Trial plans and is not included in the entry-level plans. If your plan does not include it, the policy editor shows "Not included in your plan" with an upgrade hint, and saving a policy with the option on is refused.

How a request is routed ​

  1. Policy first. Your policy builds the admitted set: only providers and models that satisfy every hard constraint and preference tier.
  2. Detect the task. The task comes from the first of these that has an answer:
    1. Stated explicitly by the caller (metadata.taskCodes or specializations in the API).
    2. Prompt default: a task you pinned on the stored prompt, or the prompt's category.
    3. Profile: the execution profile, for example the coding lane.
    4. Detected from the text: pattern matching, then (if your policy allows) a short classification request.
  3. Score each admitted model on that task. For every capability the task needs (for example Coding, or Legal domain plus Long context plus German), Veriprompt takes the best available evidence for that model:
    1. your own results (once there are enough samples),
    2. third-party benchmarks (recent enough, if you allow them),
    3. the specialization tags on the model,
    4. the model's static quality tier.
  4. Rank. The task-specific quality replaces the general quality number in the usual weighted ranking.

If the task is unknown or detected with low confidence, the profile is recorded but not applied, and routing is identical to standard routing.

Turn it on (policy editor) ​

  1. Open Admin > Routing Policies and edit (or create) the policy that should use it.

  2. Find the Task-fit routing section. The badge reads Ranks, never excludes.

  3. Switch on Enable task-fit routing. If your package does not include the feature, you see Not included in your plan instead.

  4. Choose Task detection:

    OptionWhat happensCost
    OffNo content analysis. Only explicitly stated tasks, prompt defaults and profiles count.None
    Pattern matchingThe task is detected locally from patterns in the text.None
    Gateway classification (default)If pattern matching is not confident enough, one short classification request runs through your own gateway path. If it fails or takes longer than four seconds, pattern matching is used.One short request per uncertain prompt (not per version), billed to you
  5. Choose External benchmark data: Allow (third-party benchmarks may influence ranking) or Do not use (only your own measurements, model tags and the quality tier count).

  6. Optionally adjust the Thresholds:

    SettingDefaultMeaning
    Minimum own samples30From this many of your own ratings per model and capability, your data outranks external benchmarks.
    Benchmarks expire after (days)120Older benchmark rows carry no weight.
    Minimum task confidence0.5A detected task below this is recorded but not applied.
  7. Optionally open Capability weights to weight single capabilities up or down when a task touches several (neutral weight 1, range 0 to 10, 0 ignores a capability).

  8. Click Save.

Equivalent policy JSON:

json
{
  "version": "routingPolicyV1",
  "name": "Task fit, EU only",
  "scope": "company",
  "hardConstraints": {
    "geo": { "allow": { "memberships": ["EU"] } }
  },
  "preferences": {
    "weights": { "latency": 0.2, "cost": 0.3, "quality": 0.4, "stability": 0.1 },
    "taskFit": {
      "enabled": true,
      "classifier": "gateway",
      "externalEvidence": "allow",
      "minInternalSamples": 30,
      "staleAfterDays": 120,
      "minConfidence": 0.5,
      "dimensionWeights": { "domain_legal": 2 }
    }
  }
}

Note that the quality weight matters: with quality at zero, a better task fit has nothing to influence.

Start with the preview

You do not have to guess what changes. Use the preview below before you save and roll out.

Read the preview ​

In the same section, Preview shows how the saved policy ranks with and without task fit. It uses the saved version, not unsaved edits (save first, or you will see a note), only pattern matching, and sends nothing to any provider, so it costs nothing.

  1. Choose the tasks you expect under Or choose tasks directly (for example Contract and Legal), or paste a typical prompt into Sample prompt. A pasted prompt is detected by pattern matching only, so a short, single-topic text such as Review this contract for liability clauses shows Not applied (see the warning below); choosing the tasks shows the ranking the policy produces once the task is known.

  2. Click Run preview.

  3. Read the result:

    BlockWhat it tells you
    Detected taskThe tasks, the capabilities scored, the confidence, the source, and whether it is Applied or Not applied (with the reason).
    Ranking comparisonTwo columns, Without task fit and With task fit, per model: rank, quality score and where the evidence comes from. Movement is spelled out, for example "Moves up from rank 4 to 2".
    Evidence badgesExternal (third-party benchmark), Internal (your own measurements), Manual (specialization tag on the model), Tier (static quality tier).

If the policy has task fit off, the preview says so and shows what would happen if you turned it on.

Pattern matching is deliberately cautious

A prompt that matches only one pattern group is detected with confidence 0.4, below the default threshold of 0.5, so it shows Not applied. In live use, Gateway classification (or a pinned task) resolves these cases. If a preview says "Not applied" for a prompt you consider obvious, pick the tasks directly to see the ranking the policy would produce once the task is known.

The prompt editor: Detected task ​

A prompt is labelled with its task once. The label is detected on the first run and then kept for every later run and every version of the prompt, however the text changes. You therefore pay for at most one classification per prompt, not one per version.

You manage the label in the Detected task block. Open it in the prompt editor (Prompt settings), in the settings dialog of the prompt page, or from Studio (see below). The block shows what routing would use for this prompt right now, without sending anything to a model:

  • Fixed: categories you chose for this prompt.
  • From category: derived from the prompt's category.
  • Detected: the stored label from the first run (or the last re-detect). It is shown as stored; it does not expire.
  • Pattern match: detected from the saved text. Encrypted (zero-knowledge) prompts cannot be read by the server, so no detection runs for them; you can still choose categories.

Task category: Automatic or Fixed ​

The Task category control has two settings:

SettingWhat happens
Automatic (default)No category is chosen by you. The label is detected once on the first run and then kept. Until then the block says it has not been detected yet and will be on the first run.
FixedYou pick up to five categories from the drop-down (grouped by area) and click Save fixed categories. They apply to every run and every version, outrank any detection, and skip classification entirely.

Changes are saved immediately on an existing prompt. Switching back to Automatic removes the fixed categories. A detection is never turned into a fixed category on its own: only your explicit save does that. If your package does not include task-fit routing, fixed categories are stored but only applied once the package includes it.

Re-detect ​

If the purpose of a prompt has changed and you want the label renewed, click Re-detect on next run (Automatic only). This removes the stored label; it does not call a model. The next run classifies the prompt once and stores the new label. The button is disabled, with the reason shown next to it, when categories are fixed or when nothing has been detected yet.

Why the label does not follow edits

Keeping the label across versions is deliberate: a prompt with many versions is not classified again for each one. The trade-off is that a prompt whose purpose drifts keeps its old label until someone re-detects. If you do not want automatic labelling at all, choose Fixed.

In Studio ​

The task category belongs to the prompt, not to a single version, and applies to all its versions. In Studio > Versions, each version row has a Task category button (target icon) that opens the prompt's settings dialog with the block. From the prompt page, Settings opens the same dialog.

Model Radar: By task ​

Admin > AI Models shows the Model Radar. Besides the Overview axis set (price, latency, availability, reliability, policy fit) there is By task:

  • Select up to three models to compare.
  • Each axis is a capability: Coding, Tool use, Reasoning, Math, Instruction following, Factuality, Long context, German, Writing, Legal, Medical, Finance, Vision, General.
  • Values are resolved exactly as routing resolves them: your own data first, then external benchmarks, then tag and quality tier. Axes with no evidence at all are left out.
  • A filled dot is measured (benchmarks or your own traffic). A hollow dot is only a baseline from the tag or tier. A dashed outline means the model has no measured value on any axis. Evidence older than the expiry is shown as stale and not used.
  • Hovering a point shows the source, benchmarks, sample count and date.

The same view is useful when choosing which models to enable: a model that is hollow on the capability you care about is a claim, not a measurement. The model configuration dialog shows this per model under Evidence behind the specializations, including flags such as Tagged "coding", but there is no benchmark or own-data evidence for Coding. The tag is only a claim so far.

Worked examples ​

1. German contract review ​

A legal team reviews supplier contracts in German. Pin the task on the stored prompt "Vertragsprüfung":

  1. Open the prompt, go to Detected task, set Task category to Fixed, add Contract and Legal, and click Save fixed categories.
  2. The routing policy has task fit on.

Capabilities scored: Legal domain, Long context, and, because the text is German, German. Models are ranked by evidence on those three. A model that is strong at general chat but weak on long German documents drops; a model with solid long-context and German results rises, as long as your policy already admits it. Use the preview and choose Contract and Legal under Or choose tasks directly to see the movement. (Pasting only the text Prüfe diesen Vertrag auf Haftungsklauseln is detected with confidence 0.4 and shows Not applied; in live use the pin, or Gateway classification, supplies the task.)

Because a legal team usually cares more about legal knowledge than raw length handling, set dimensionWeights: { "domain_legal": 2 } in the policy (Capability weights in the editor).

2. Code review ​

Calls come from a CI job through the API. Say the task explicitly:

bash
curl https://app.veriprompt.tech/api/gateway/execute \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Review this diff and list bugs: ...",
    "policyId": "YOUR_POLICY_ID",
    "metadata": { "taskCodes": ["coding", "debugging"] }
  }'

An explicitly stated task has full confidence and no classification request is needed. Coding and Reasoning are scored; agent-harness benchmarks (for example SWE-bench-style results) count for at most half of the external evidence because their score depends on the scaffold, not only on the model.

3. Mixed team ​

Engineering, legal and marketing share one company policy.

  • Engineering sends tasks explicitly from its tools.
  • Legal pins Contract and Legal on its prompts.
  • Marketing pins nothing; Task detection is Gateway classification, so a short classification request decides when the text is ambiguous, and when it is still unsure the request simply gets standard routing.

One policy, three kinds of work, three different winners, all inside the same compliance boundary. Check the effect in Admin > Routing Policies with the preview, and per team in the explanation lines (see the FAQ below).

4. A policy that denies external evidence ​

Your counsel does not want third-party benchmark data to influence routing. Set External benchmark data to Do not use ("externalEvidence": "deny").

Now only three sources count: your own measurements (prompt ratings in your company), the specialization tags on your models, and the static quality tier. Until a model has at least Minimum own samples ratings for a capability, ranking is close to standard routing, and the preview will show Manual and Tier badges instead of External. As ratings accumulate, your data takes over. Your own evidence never leaves your company: it is only read for your routing.

API ​

Task-aware routing adds optional, additive fields to the gateway and routing APIs. Nothing existing changes when you omit them.

  • State the task on a gateway request: metadata.taskCodes or specializations (Gateway Execute).
  • Preview and explain: taskCodes, content and taskFit on policy simulation and ranking (Routing Policies) and on the subscriber routing advisory (Routing Advisory).

FAQ ​

Why was model X chosen? ​

Four places answer it:

  1. Policy preview (Admin > Routing Policies): the Ranking comparison shows both orderings and the evidence behind each score.
  2. Simulate API: taskFit.ranking[].explanations contains lines such as "Ranked first for coding: score 91/100 on livebench:coding (external, 2026-09-28)".
  3. Model Radar > By task: the evidence per capability.
  4. Capability Matrix (super admins): the exact value routing uses for a model and capability, with every row behind it.

If task fit was not applied, the explanation says why: policy has it off, the task was unknown, confidence was below the threshold, or the package does not include the feature.

Does classification cost money? ​

Pattern matching is free and runs locally. Gateway classification costs one short request (the first 1,500 characters of the prompt plus a task list of roughly 430 tokens as input, at most 80 output tokens) per prompt where the task could not be settled otherwise. It is billed to your company like any other request, runs through your own gateway path under your policy, and appears in usage as purpose task-classification. It is skipped whenever the task is stated explicitly, pinned, or set by a profile, and the result is cached per conversation and per prompt version, so repeated runs do not repeat the cost. The preview and the simulate API never call a model and cost nothing.

What data leaves my environment? ​

  • The benchmark data is public: Veriprompt fetches published leaderboards (see the admin guide for the sources). Nothing of yours is sent to them.
  • Your prompt is never sent anywhere your policy does not already allow. Pattern matching runs locally. A classification request goes through the same policy, the same admitted models, the same Shield sanitization and the same geographic and privacy rules as your own request would. It never uses a separate service.
  • Your own evidence (prompt ratings) is stored per company and only used for your company's routing.

What if the classifier is wrong? ​

Pin the correct task on the prompt, or state it in the API call. Both override detection. Below the minimum confidence, a detected task is not applied at all.

What happens if my package is downgraded? ​

The policy keeps its settings, but at runtime task fit is ignored and routing falls back to standard routing. The editor shows Task-fit routing is paused; to save the policy again, turn the option off (or upgrade). Pins stay stored.

Can task fit pick a model my policy forbids? ​

No. It runs after the policy has decided which models are allowed and only changes their order. This invariant is covered by automated checks that run against every policy shape.

Why does the ranking sometimes not change? ​

Common reasons: all admitted models have similar evidence, the task could not be determined, there is no benchmark evidence for the admitted models (they fall back to tags and tier), the policy's quality weight is very low, or only one model is admitted.

Troubleshooting ​

SymptomLikely causeFix
"Not included in your plan" in the editorPackage lacks Task-Fit RoutingAsk your account owner or Veriprompt to upgrade
Save refused with a feature errorThe policy has task fit on but the package does not include itTurn the option off or upgrade
Preview says Not applied with "Confidence is below the policy minimum."Only one pattern matched (confidence 0.4)Pin a task, state it in the API, or select tasks directly in the preview
Preview says Save the policy firstPreview runs on the saved versionSave, then run again
Same ranking with and without task fitSee "Why does the ranking sometimes not change?"Check the evidence badges; hollow dots in Model Radar mean no measurements
No Detected task for an encrypted promptZero-knowledge prompts cannot be read by the serverPin a task
Model shows only Tier evidenceNo benchmark maps to that model yetAsk a super admin to check the unmapped-names queue