Thirty cents.
That was the estimated cost of getting an AI model to a particular level of performance on a difficult benchmark in early 2025.
Less than 18 months later, reaching that same performance level could cost around $0.0004.
That’s a huge drop. And it happened in a field where better models have generally meant spending more money to run them.
The number comes from Epoch AI, which has been tracking how much it costs to reach the same level of performance across different AI benchmarks.
The cost of reaching the same level of AI performance is falling faster than it has for most technologies we’ve seen before.
So what exactly is getting cheaper? And if AI is getting this cheap, why can the bills still go up?
Table of Contents
The 47% Collapse in Token Costs
For most of computing history, adding more intelligence meant adding more people.
More code reviews required more engineers. More scientific research required more researchers. More analysis required more analysts.
AI changes that equation because the cost of running a capable model keeps falling.
Epoch AI tracks this by looking at how much it costs to reach a specific level of performance on AI benchmarks. Instead of comparing one model’s price over time, it measures the cheapest available way to achieve the same result.
Since 2023, that cost has been falling by roughly 47% every quarter.
The drop becomes easier to understand with one example.
When OpenAI released o3 in January 2025, Epoch AI estimated that reaching a 75% score on the GPQA Diamond benchmark cost around $0.30 per question. The benchmark tests advanced knowledge across fields like chemistry, physics and biology.
Later, GPT-5.6 Luna reached the same performance level at around $0.0004 per question.
The cost of the same level of AI capability had fallen hundreds of times in a relatively short period.
But the important part is what happens after intelligence becomes cheap.
Because cheaper AI does not automatically mean every AI task becomes cheap.
Cheap Intelligence Meets Expensive Problems
A model can solve a difficult math problem in a test environment, but a software engineer using AI has to deal with an existing codebase, unclear requirements, changing priorities, and decisions that are difficult to measure.
The gap becomes clearer when looking across different types of tasks.
Mathematics and pure logic have seen some of the fastest cost reductions, falling around 50% to 52% per quarter. Structured problems are easier to optimize because the path to a correct answer is clearer.
Hard science tasks have followed a similar downward trend.
But areas like combinatorial games and software engineering have moved more slowly. Chess and puzzle-style tasks fell around 39% to 43% per quarter, while SWE-bench data showed a slower decline of around 27.5% per quarter.
Software engineering is a good example of the difference between passing a test and doing useful work. Fixing a real bug often requires understanding why a system was built a certain way, not just producing a correct piece of code.
How Cheaper AI Changes Software Architecture
A coding agent can now write a change, run tests, inspect failures and try again without every step being treated as a major expense. A separate model can review the output. Another can compare approaches before the final result is returned.
The software is no longer built around a single AI response. It is built around a process that can generate, check and refine answers.
The focus shifts from How do we reduce the number of AI calls? to How do we design systems that use those calls effectively?
Also Read: OpenAI’s Agents Leaked 53 Images. What Else Did Its Agents Do?
The inference paradox: cheap calls, bigger bills
Lower AI prices create a strange situation, using AI more can still make companies spend more.
A simple chatbot interaction might use a few hundred tokens. An agentic workflow can use many times more because it has to read context, call tools, check results and repeat steps until it reaches an acceptable answer.
A coding agent working through a large repository is a good example. It may inspect files, make changes, run tests, review errors and try again multiple times. Each individual AI call may be cheaper, but the total number of calls grows quickly.
This is why cheaper models do not always translate into smaller AI budgets.
Gartner describes this as the Inference Paradox. As AI systems become capable of handling more complex tasks, companies often build workflows that require more inference, not less.
Gartner also expects the total inference cost of agentic workflows to increase more than fivefold through 2028 as these systems become more common.
The cost of each AI action is falling but the amount of AI work being assigned is growing even faster.
The New Challenge Is Orchestration
The cost of AI intelligence is falling quickly, but using that intelligence effectively is becoming a different challenge.
As software starts relying on more agents, more verification steps, and longer reasoning loops, the total cost of a workflow can still increase even when each individual AI call becomes cheaper.
That makes the design of the system just as important as the model behind it.
A well-built AI system knows when to use a cheaper model, when a more capable one is worth the cost, how much context to provide, and where additional reasoning actually improves the result.
The advantage will come from building systems that can get more value from every AI call without wasting the savings that cheaper intelligence provides.




