For three years the pitch was capability, and this week it turned into a meter. Accenture launched a practice it calls Tokenomics and put a name to the thing every buyer was feeling: the agentic Jevons paradox, where the cheaper a unit of intelligence becomes, the more of it a system consumes, so the total bill climbs while the price of a token falls. On the same tape EY published the fifth wave of its AI Pulse survey, in which ninety-eight percent of the leaders investing in AI said token usage and its costs had forced them to reconsider their approach. KPMG advanced its own token-standardization work, and Cognizant tied its economics to the workforce it is certifying. Four firms, one week, one message: the cost of running AI is now the operating question.
It is worth being precise about what changed, because the firms were, for once, right about the diagnosis. Lan Guan, Accenture's chief AI and data officer, framed the trap plainly. Hand a powerful model to a whole workforce and people default to the most capable and most expensive option, ask it more, and let it call more tools, and that single behavior multiplied across a company is where the spend quietly balloons. Her prescription is a three-step discipline, see the usage, treat it, then manage it, with success measured not by cost per token but by value delivered. Read that last line twice. It is the argument this publication has made for a year, delivered by the largest consulting firm on earth.
So the honest question is not whether the diagnosis is correct. It is who ends up owning the cure. A tokenomics practice is a service that reads your meter, tunes your routing, and hands you a report every quarter on a bill you still do not control, because the meter belongs to the supplier and the discipline belongs to the firm you retained to run it. Manage the cost that way and you have rented a second layer of dependency on top of the first. The savings are real, and they are also a subscription. When the model price moves again, and it moved twice this week, the practice bills you to re-tune.
There is a different answer, and it starts one level down from the meter. Most of the work a company automates is deterministic, the same steps against the same inputs producing the same output, and deterministic work does not need a model at all once it is written as code the client owns. Encode the procedure, run the deterministic part without touching a token, and reserve a model for the narrow place where judgment is genuinely required, on whichever one is defensible that quarter. Now the meter runs on a sliver of the work instead of all of it, the bill is bounded because most of it was taken off the model, and the record of what the system did belongs to the company rather than to the firm reading its usage. The field spent this week learning to count the cost. The product is the thing that does not run it.
The market has stopped paying for AI spend and started asking for AI proof.Read the full story ↗




