The headlines are relentless: a top model drops its price by half, another cuts a variant by eighty percent, a cheaper open competitor undercuts them both. Every time it happens, usage does not just hold, it jumps several-fold, so the price cut ends up making the provider more money, not less. Strip away the noise and one clear trend remains: the cost of running a genuinely capable model is collapsing, and it is not slowing down.
Why the price keeps falling
Token pricing was always a little arbitrary. A provider’s real cost is fixed: the racks, the power, the people. They meter it out per token, but a data center at high utilization is far cheaper per token than one sitting idle. So cutting the price to pull in more usage is often the profitable move, and the hardware and efficiency gains keep pushing the floor down underneath everyone. The result for you is simple: whatever a given quality of AI costs today, it will cost noticeably less in a few months.
The trap of building around a price
The mistake is to treat today’s cost, or today’s cheapest model, as a fixed part of your plan. Prices move, and a business whose only edge is “we found the cheap model” has no edge, because so did everyone else by next quarter. Cheap, abundant intelligence is becoming a utility, like bandwidth, and nobody builds a durable company on having slightly cheaper bandwidth. We made the same point when the tools started going free: falling cost is a gift and a trap at once, and it moves the advantage somewhere the price cannot follow.
What to do with a falling curve
Two things. First, build now, and build a little ahead of the current price. Automations that look marginal at today’s token cost will be comfortably profitable at next quarter’s, so the backlog you shelved as “too expensive to run at scale” deserves a fresh look. Second, put your durable investment into the parts that do not deflate: the integration into your systems, the proprietary data you feed the model, the judgment around it. That is where value collects as the raw capability commoditizes, which is exactly why the paid work is in the integration. Watching which capabilities are getting cheap is also a decent map of where to build, and it lines up with the demand board.
Falling prices are the best news a business adopting AI could ask for, as long as you spend the savings on the things that stay scarce. That is the whole shape of our solutions: ride the cheap capability, own the parts of the stack a price cut can never reach.
Prompted by the video “OpenAI Ready To Kill China Prices” (2026). Credit to the creator; the read on what it means is ours.
Saraswati Stitch®contact@saraswatistitch.com