Skip to content

Controlling AI spend ​

AI cost is unusual: it accrues per request, from many teams at once, and nobody sees the bill until the month closes. By then the money is spent. Controlling it means three separate things — knowing where it goes, capping it before it runs over, and making each call cheaper — and they need different tools.

What you wantWhereWhat it does
See where the money goes/sherlock/valueBreaks spend down by team, project, prompt and provider — far enough to see what is actually expensive.
Watch the trend/sherlockSpend over time for the whole account, so a change in direction is visible before month end.
Cap it/sherlock/budgetsHard budget caps per team, project or period, with alerts before the cap is reached.
Route to cheaper providers/ranger/routingPick the provider and model by cost as well as quality, per prompt.
Shorten the prompts/guru/optimizeFewer tokens per call, without losing output quality.
Report to finance/sherlock/reportsSpend data as a file, in the shape finance asks for.

If your AI costs are too high and you want to reduce them, work down this table in order rather than starting with whichever prompt is in front of you.

Where to start. Measure before you cut. /sherlock/value almost always shows that cost is concentrated in a few prompts or one provider choice, and cutting there is worth more than a broad efficiency push. Set budgets second, so the measurement has a floor under it.

Why the order is measure → cap → optimise ​

Optimising first is the intuitive move and the expensive one. Prompt shortening is real work, and it is usually applied to whichever prompt someone happens to be looking at rather than the one that costs the most. A breakdown at /sherlock/value typically reallocates that effort by an order of magnitude.

Budgets come second rather than last because they change the failure mode. Without a cap, an unexpected cost is discovered on the invoice. With one, it is discovered as an alert while there is still time to act — and a hard cap converts a surprise bill into a refused request, which is a much better problem.

Only then is prompt and routing optimisation worth the effort, because by that point you know which prompts and which routes to spend it on.

Budgets that are not just a number ​

A budget cap is only useful if it is scoped to something a person is responsible for. /sherlock/budgets supports caps per team, project or period, which lets a cap map onto a real owner. A single account-wide cap tends to produce the wrong outcome — the team that happens to run last in the month gets blocked for someone else's overspend.

Set the alert threshold meaningfully below the cap. An alert that fires at the same moment the cap does is a notification, not a warning.

Reducing the cost of a call ​

Two levers, and they are independent:

  • Route it differently. A routing policy can prefer a cheaper model where output quality allows, and reserve the expensive one for prompts that need it. See Intelligent routing.
  • Send fewer tokens. /guru/optimize rewrites a prompt to use fewer tokens while holding output quality. Combine it with A/B testing so "without losing quality" is measured rather than assumed.

Proving the saving ​

/sherlock/audit holds the audit console for spend, usage and findings — the record that shows a change in policy actually moved the number, rather than coinciding with a quiet month.