Inference cost
What each answer costs, and why the pilot was cheap and the rollout is not.
Training a model is a capital event; running it is an operating cost, charged per token in and per token out. This is why pilots mislead so reliably: fifty officers experimenting costs almost nothing, and the same system in the hands of ten thousand people, invoked inside a workflow they use forty times a day, costs a great deal.
Per-token prices have fallen steeply and keep falling, which makes budgeting awkward in the other direction — a business case built on today's prices may be conservative by the time it is approved. Meanwhile reasoning models, which generate long internal working before answering, can use ten or more times the tokens for a single response.
The honest unit is not the token. It is the cost per completed piece of work, including the human checking time.
Why it matters here
Agency business cases routinely model licence fees and omit consumption. The question to settle before a pilot, not after, is what a fully adopted year looks like at realistic usage — and who holds that budget when it lands.
The question to ask
What does this cost per finished task at full adoption, and who is paying for the checking time?
Reviewed 2026-09-20 · All decoders