In December, Uber handed Anthropic’s Claude Code to its engineering organization and told them to build. By March, 84 percent of engineers were using it regularly, up from 32 percent just a month earlier. By April, the company had burned through its entire 2026 AI budget, four months into a twelve month plan. Uber’s president and chief operating officer later admitted the link between all that spending and anything customers actually noticed “is not there yet.” The company’s CTO said the team was going “back to the drawing board” on how it budgets for AI entirely.
Here is the part that should not add up. Over that same stretch, the price of the AI models Uber’s engineers were running got dramatically cheaper. Independent trackers put the drop in blended frontier model pricing at close to 98 percent since early 2024, roughly a thousandfold decline in cost per token across three years. By every normal rule of economics, a company using a resource that keeps getting cheaper should see its bill shrink. Enterprise AI spending is doing the opposite. Total corporate AI budgets grew an estimated 320 percent in 2025 alone, and one recent survey found 73 percent of enterprises now say their AI costs blew past what they had originally projected.
The Discount Nobody Actually Gets to Keep
The token math is real. Frontier labs really have cut prices that aggressively, driven by a mix of more efficient model architectures, brutal competition from Chinese labs shipping capable open models for a fraction of the price, and simple economies of scale as inference infrastructure matures. A task that cost a company a dollar in API calls two years ago might cost a fraction of a cent today.
What almost nobody predicted is that the discount would not translate into savings. Economists have a name for this pattern, borrowed from a nineteenth century observation about coal. Jevons Paradox holds that when something essential gets cheaper, people do not just buy the same amount for less. They find new ways to use more of it, and total consumption climbs faster than the price falls. Cheap coal did not shrink Britain’s coal bill in the 1800s. It expanded the range of things worth burning coal for. Cheap tokens are doing exactly the same thing to enterprise AI spending right now.
Agents, Not Chatbots, Are Driving the Bill
A year or two ago, most enterprise AI spending went toward something simple: a chatbot answering one question, generating one response, done. That kind of interaction still costs a fraction of a cent to run. It is not what is emptying corporate budgets anymore.
The shift is toward agentic systems, AI that plans a multi-step task, calls tools, checks its own work, and loops back to try again when something fails. Industry estimates suggest these agentic workflows burn somewhere between five and thirty times more tokens than a single chatbot exchange, and the gap keeps widening as the tasks get more ambitious. A simple linear workflow that cost about four cents to run in 2023 can cost well over a dollar today once it involves tool calls, reasoning steps, and iterative retries, even though the underlying per-token price fell the entire time.
Uber’s numbers make the pattern concrete. Engineers using coding agents were running up monthly bills of roughly 150 to 250 dollars each on average, with the heaviest users hitting 500 to 2,000 dollars a month. Multiply that across thousands of engineers adopting the tools within weeks of rollout, and a budget built for a full year disappears in a third of that time. None of that came from the price of any single token going up. It came from a workforce suddenly finding a thousand new reasons to spend, exactly the outcome coding assistants like Cursor have been riding to billions in annualized revenue as adoption spreads from individual developers to entire engineering organizations.
The Bill Behind the Bill
Even that understates the real cost, because token spend is only the visible slice of what companies are actually paying for agentic AI. Enterprise deployment audits consistently turn up hidden costs running an extra 40 to 60 percent on top of the raw inference bill most finance teams are watching. That gap comes from everything token pricing does not capture: the engineering time spent building evaluation pipelines to check whether an agent’s output is actually correct, the guardrails needed to stop an autonomous system from doing something costly or embarrassing, the data preparation work that has to happen before an agent can be trusted with a real workflow, and the compute infrastructure sitting underneath all of it.
Analysts increasingly split enterprise AI spending into three separate buckets that rarely get reconciled against each other: raw compute infrastructure, per-token API consumption, and per-seat productivity subscriptions. Each one is tracked by a different team, billed on a different schedule, and reviewed by a different budget owner. That fragmentation is a big part of why so many organizations cannot answer a simple question, what does AI actually cost us, until the number is already too big to ignore.
A Squeeze at the Top of the Market Too
The paradox cuts both ways. While corporate AI bills climb, the labs selling the underlying models are watching their own margins get thinner. As buyers get savvier about matching a task to the cheapest model that can handle it rather than defaulting to the most powerful option available, expensive frontier providers are losing routine work to cheaper competitors, many of them Chinese labs offering near-equivalent performance at a steep discount. One analyst summed up the shift by comparing premium frontier models to driving a Lamborghini to the grocery store for a carton of milk: technically impressive, and a waste of the car’s actual capability for the job at hand.
That dynamic is forcing a reckoning that shows up in how confident technology leaders feel about the whole endeavor. It is one reason confidence among CTOs in their ability to scale AI has now fallen for three straight years, with return on investment uncertainty cited as one of the top reasons scaling stalls. Spending more and feeling less certain about what that spending buys is not a comfortable place for any budget owner to sit.
What Companies Are Actually Doing About It
The organizations getting a handle on this are treating AI spend less like a software license and more like cloud infrastructure, something that needs the same discipline that finance and engineering teams eventually built around cloud computing bills a decade ago. In practice that means a handful of concrete moves. Model routing, sending routine, low-stakes requests to smaller and cheaper models while reserving frontier models for tasks that genuinely need them, is quickly becoming standard practice rather than an optimization for later. Per-employee and per-team spending caps are replacing open-ended access, especially for coding agents where usage can spike overnight. Tagging and observability tools that were built for tracking cloud spend are being repurposed to show finance teams exactly which team, project, or agent is generating which slice of the bill.
The stakes for getting this right are not small. Research from MIT’s NANDA initiative found that roughly 95 percent of generative AI pilots at large companies fail to reach production, a number that ought to worry anyone treating enterprise AI’s high failure rate as a footnote rather than the central risk. Spending more without a clear framework for measuring return is a fast way to end up as one more entry in that statistic.
None of this means the token price collapse was a mirage. It is real, and it is exactly what has made agentic AI economically viable at the scale companies are now attempting. What it did not do was make AI cheap. It changed what companies could afford to try, and companies, predictably, tried a lot more than anyone expected. The cost of a single AI request keeps falling. The cost of running an AI-driven company does not, and the businesses figuring that out first are the ones treating AI budgeting as seriously as they treat the technology itself.

