Why the price fell
Three things happened at once. The chips got faster and, per unit of work, cheaper: each generation of Nvidia hardware roughly doubled the useful output per dollar. The models got more efficient, with techniques like distillation, where a big model teaches a small one to do most of what it does at a fraction of the size. And the labs started competing on price, because once two models are close in quality the only lever left is the invoice.
The result is that inference, which means running a trained model to answer a question, went from a scarce, expensive thing to something closer to electricity. You still pay for it. You just stop thinking about it.
Where the money went
Follow the dollar. A company paying for tokens sends money to a lab. The lab spends most of it on cloud compute, which means it sends most of it to Microsoft, Amazon, Google or a specialised GPU host. Those companies send a large share of it to Nvidia for chips, and Nvidia sends a large share of that to TSMC for manufacturing.
At every step in that chain, the margins are wildly different. Nvidia's gross margin has been above 70 percent. The cloud providers make healthy but normal margins. The labs, by most reporting, lose money on the frontier models and make it back, if at all, on scale. The company that sells the shovels made the money. The companies digging are still hoping.
Who pays now
Increasingly, not the user. The price of a chat with a model has fallen below what most people would notice, so the labs bundle it into subscriptions or give it away and charge businesses instead. The real paying customers in 2026 are companies running models inside their own products: customer service, coding tools, document processing, search. They buy tokens by the billion and negotiate prices that never appear on a public price list.
The other payer is the investor. Every large lab has raised money at valuations that only make sense if inference eventually becomes a very large, very profitable business. That money is subsidising today's prices. If the investors are right, the subsidy ends when scale arrives. If they are wrong, prices have to go up, and the products built on cheap tokens get more expensive.
What this means for a business built on AI
If your product is a thin layer over someone else's model, your costs will keep falling and so will your competitors'. That is good for users and bad for margins. The businesses that hold up are the ones where the model is one ingredient and the value is in the data, the workflow, or the relationship with the customer.
My read: token prices keep falling for another two or three years, then flatten as the labs need to show profits. The window where you can build something on nearly free intelligence is open now. It will not stay open forever, and the companies that treat cheap inference as permanent will be the ones surprised when the bill arrives.