top of page

Your AI Bill Falls 10x a Year. Why Is It Rising?

Writer: The AI Daily
The AI Daily
Aug 12
6 min read

Updated: Aug 21

The price of a fixed unit of AI capability is collapsing about 10x every year. Your invoice, meanwhile, went up last quarter. Both of those things are true at once, and the gap between them is where most enterprise AI budgets get lost.


This is the question finance teams keep landing on, and the one that keeps surfacing across the tech desks we read, from Daily 24 Tech to the research shops publishing inference benchmarks. Per-token prices are in freefall. Bills are climbing anyway. Somebody is wrong about something.


Nobody is. Cost-per-token and cost-per-task are different animals, and confusing them is the single most expensive mistake in AI budgeting right now. Priced correctly, AI for business leaders is a forward question, not a snapshot of this month's invoice.


The AI Cost paradox

Why is the price of AI falling 10x a year?


The cost of hitting a fixed quality bar with an AI model falls roughly 10x annually, a trend a16z named LLM flation. Epoch AI's stricter measurement is even steeper, with a median around 50x per year across benchmarks. Performance that cost $60 per million tokens in late 2021 now costs pennies.


Here's the part most people get backwards: this is mostly not a hardware story.

Raw silicon price-performance improves only 33 to 38% a year. The rest of the collapse comes from software. Quantisation down to 8 and 4 bits. Distillation into small models that punch far above their size. Mixture-of-experts routing. Speculative decoding. Continuous batching. A 1-billion-parameter model today outperforms a 175-billion-parameter model from 2021, which tells you how much of the gain is engineering rather than transistors.


There's a catch, and it decides everything. You only capture the full 10x if you keep down-shifting to the cheapest model that clears your quality bar. Pin your product to the frontier flagship and you ride a much gentler slope: flagship output prices have fallen roughly 6x in two years, not per year.


The decline rate you get is a choice you make. It isn't a market condition you wait for.

So why is your bill going up?


Two forces push the other way, and together they routinely outrun the price collapse.


Jevons paradox. When something gets cheaper, people use dramatically more of it. Enterprise generative AI spend went from $1.7B in 2023 to $11.5B in 2024 to $37B in 2025, a 3.2x jump year over year while per-token prices were falling hard. Menlo Ventures treats this as the 2026 base case, not the exception.


The reasoning tax. This one is brutal and badly underestimated. Reasoning and agentic workloads burn far more tokens per task than a single prompt does. On the ARC-AGI benchmark, one hard task on OpenAI's o3 ranged from about $20 to about $4,560 depending on how long the model thought about it. Same problem, roughly 170x spread.

So the honest model has three dials, not one. Price per token falls. Tokens per task rises. Total volume rises. Multiply them and the same use case gets either dramatically cheaper or dramatically more expensive, and your architecture decides which.


Most teams track only the first dial. That's why the invoice keeps surprising them.


The memory problem nobody budgeted for

There's one place the "hardware always gets cheaper" assumption is currently wrong, and it's the one squeezing on-premise builds right now.


AI demand has pushed DRAM and high-bandwidth memory prices up as much as 90% since late 2025. Analysts describe the shortage as structural, tight through 2027, with no firm normalisation date. Memory now accounts for over 40% of an AI server's bill of materials.


Net effect: silicon deflation gets partly cancelled for the next year or two. Costs still fall, but the curve flattens and may briefly tick upward through 2026 and 2027 for memory-heavy configurations before resuming its decline.


The collapse is real. It just isn't a smooth line, and anyone modelling it as one will miss a budget cycle. We've tracked the memory squeeze week by week in our AI Weekly, and the supply picture has yet to loosen.


India's curve bends faster


If you're running workloads out of India, a second and steeper discount is available that dollar-priced buyers can't reach.


Servers and GPUs carry roughly 0% basic customs duty under the ITA-1 schedule, and the 18% IGST is fully creditable, so the cost base is set by power and utilisation rather than tariffs. The IndiaAI Mission has put more than 34,000 GPUs on tap at around ₹65 per GPU-hour after a 40% subsidy, roughly a third of global on-demand rates.


Then there's the sovereign model layer. Sarvam-105B is priced near $0.80 per million tokens, India-hosted, with voice at ₹3.5 a minute against an ₹8 to ₹12 industry norm. No FX exposure, no cross-border egress, lower latency. Models built with Indic-first tokenisers also spend fewer tokens per word of Hindi, Tamil or Bengali, which is a direct discount on any Indian-language workload.


At its Epoch showcase on 31 July 2026, Sarvam claimed its coding agent edged out Claude Code on Terminal Bench, scoring 72 of 89. Whether that benchmark holds up matters less than the pricing signal behind it. The sovereign track now reaches all the way up to frontier-class agents.


What the numbers actually look like


Take an enterprise support desk handling about 1.5 million queries a month. At today's frontier prices, that's roughly ₹2 crore a year.


Hold the capability fixed and project it forward on a base case: about ₹69 lakh in three years, about ₹34 lakh in five. A third of the cost, then a sixth. The ROI you can't justify today becomes obvious around year three. It just needed a calendar.


But run the same desk with agents bolted on and the model pinned to the frontier, and it costs more in five years than it does now. Two futures, same starting point, and the gap between them is entirely a set of decisions you control.


What to do about it


Three things worth taking into your next budget conversation:


  1. Stop extrapolating today's API bill in a straight line, in either direction. Aggressive model down-shifting can cut a fixed capability by 80%. Frontier lock-in plus agents can raise it.


  2. Budget for tokens per task, not just price per token. The price line falls on its own. The task line is yours to manage.


  3. Approve use cases that break even within two years. They look marginal today and compound into strong returns as costs fall underneath them. Design for optionality while you're at it, because nobody knows which providers win.


For Indian workloads, price the sovereign track explicitly. It's no longer only a cost story, it's a control story.


One practical note: token economics has quietly become one of the more useful data analysis topics a finance team can own. Pull your own logs, split spend by model tier and tokens per task, and you'll usually find one workload driving most of the growth.


The bottom line


Cheaper tokens don't produce cheaper AI. They produce more AI, and the bill follows whichever line you weren't watching.


The AI Daily covers this every morning, ranked by signal instead of volume, with a dedicated India lens. If you already read Daily 24 Tech for the industry news, treat this as the layer underneath it: fewer headlines, more of the reasoning that decides your budget. Subscribe free and get it by 7 am. 



FAQs


1. Why is AI getting cheaper per token but more expensive overall? 

Per-token prices fall about 10x a year, but two forces push the total bill up: usage explodes as prices drop (Jevons paradox), and reasoning or agentic workloads burn far more tokens per task. Cost per token and cost per task move in opposite directions.


2. What is the reasoning tax on AI costs? 

It's the extra token consumption from models that think through problems in multiple steps rather than answering directly. On the ARC-AGI benchmark, a single hard task on o3 cost anywhere from roughly $20 to $4,560 depending on reasoning depth, a 170x spread on identical work.


3. Will AI hardware keep getting cheaper? 

Silicon price-performance improves 33 to 38% annually, but memory is currently an exception. DRAM and HBM prices rose up to 90% since late 2025 and analysts expect tightness through 2027. With memory at over 40% of an AI server's bill of materials, the hardware curve flattens near-term.


4. How much cheaper is AI inference in India? 

Meaningfully. India AI Mission GPUs run around ₹65 per hour after subsidy, roughly a third of global on-demand rates. India-hosted models like Sarvam-105B price near $0.80 per million tokens, with no FX exposure and better token efficiency on Indian languages.


5. Should we wait for AI prices to fall before investing? 

No. Approve use cases that break even within about two years on a capex plus opex basis. They compound into strong returns as costs fall underneath them, and waiting costs you the learning curve while competitors build theirs.


 
 
 

Comments


bottom of page