Nobody is paid to reduce it
I have said it twice this year, in two different conversations, and I was wrong both times: that whoever burns the most tokens is the one who knows the least.
It is a good line. It is also only half true, and that half does more harm than good.
Six causes, one number
Consumption is not one thing. At least six different things drive it, and they have nothing to do with each other.
The setup. No persistent context, no caching, the same files read in again every turn. You pay for the same information over and over.
Not having kept up. The tools are different from nine months ago. Whoever works the way they used to works more expensively without noticing.
Volume in. The whole log file instead of the right minute. The whole codebase instead of the three files in question. This is the single largest lever and the only one that can differ by a factor of a thousand.
Model choice by prestige. The newest and most powerful because it feels safest, even on tasks where a simpler model returns the same answer for a fifth of the price.
Reruns without diagnosis. Trying again instead of understanding why it failed. Three attempts cost three times.
Produced code and verification. The actual work.
Only the last one is legitimate output.
And on top of the six sit account tier and model price, which can differ fivefold without anyone doing anything differently.
Which is why the number says nothing
High consumption can therefore mean someone is being careless. It can equally mean someone has moved a large part of their work into the machine: running agents in parallel, letting them work overnight, verifying with several models. That is advanced practice, and it costs more than doing everything by hand.
Two people with identical consumption can be each other's opposites. One has not learned the tool, the other has automated half their job.
What matters is therefore not consumption but consumption per delivered outcome. Four times the tokens and six times the verified delivery is efficient. Four times the tokens and the same delivery is not. Same number, opposite conclusion.
That is also why I was wrong. I read a figure as though it were a grade, when it is a map of where the work sits.
The uncomfortable follow-up
Five of the six causes are fixable, and they are cheap to fix. These are afternoons, not investments.
Which makes consumption a competence question rather than a budget question. Most conversations I hear about AI cost are about ceilings, licence tiers and volume discounts, when the largest lever is someone learning to send in the right material.
But here it gets interesting.
If consumption is billed onward as pass-through cost, from supplier to client, who in that chain is actually paid to reduce it?
The supplier gets it covered. The client cannot see what drives it. Neither party has a financial reason to spend those afternoons, and so they do not get spent.
This is not malice. The incentive is simply absent, and that is structural. Exactly the same thing happened with printing costs, with telephony and with cloud storage: as long as the cost was someone else's it never got optimised, and that only changed when it became visible to whoever was paying.
What actually follows from it
Asking a supplier why consumption is high rarely produces anything. They do not know, and they have not needed to.
Asking how it distributes produces more: how many objects pass through a model, what runs on which model tier, how much is reruns. An answer to that question reveals whether anyone has looked.
And making consumption visible to whoever pays changes behaviour faster than any training does. Not because anyone starts trying harder, but because there is suddenly someone in the chain who benefits from the number going down.
See also: An agent hour is not a unit (series 55) and An even number over uneven work (series 54).