The cost of a given AI performance level has fallen about 47% per quarter
Epoch AI estimates a 13x annual drop in the price of hitting a fixed benchmark score, a pace it says outruns measured declines for DNA sequencing, compute, and lithium-ion batteries.
TL;DR
- Epoch AI estimates that the cost of a given level of AI performance has fallen about 47% per quarter, or 13x per year, over the past three years.
- On GPQA Diamond, the report estimates that o3 reached a 75% score at $0.30 per question on January 31, 2025, while GPT-5.6 Luna later matched that score at $0.0004.
- Epoch says the decline slows after a performance level debuts, from 66% per quarter to 32% per quarter two years later, and flags benchmaxxing plus incomplete data as caveats.
Epoch AI's September 22, 2026 report, "The plunging price of thought," estimates that the cost of reaching a given level of AI performance has fallen about 47% per quarter, or 13x per year, across five benchmarks covering mathematics, hard sciences, and games of skill. In log points, the authors say that pace is four times faster than DNA sequencing, six times faster than compute, and 18 times faster than lithium-ion batteries. [1]
The report's GPQA Diamond example estimates that OpenAI's o3, released January 31, 2025, could reach a 75% score at an average of $0.30 per question. Just under 18 months later, GPT-5.6 Luna matched that score at $0.0004 per question, a 725-fold drop. Epoch notes that its "75%" figure is 75% of the way from chance to a perfect score, which is 81.25% on this four-way exam. [1]
Epoch says the drop is not one curve. Game-based puzzles fell about 39–43% per quarter, while math problems fell 50–52% per quarter. Across the five benchmarks, cost falls 66% per quarter when a performance level is new and 32% per quarter two years later. The authors warn that benchmark-specific training, which they call benchmaxxing, can make measured gains look steeper than real-world work, and that buyers who do not switch to the cheapest capable model will not capture the full saving. [1]
TechSpot's October 1 report on the Epoch findings says cheaper access to a given capability could make more capable models affordable for everyday software use and make it harder for AI developers to hold onto pricing power. It also notes that the historical comparisons use uneven windows: the AI series is recent, the compute series ends in 2001, and the electricity series ends in 1973. [2]
Why it matters
A falling price for a fixed benchmark score changes what software buyers can afford, but Epoch's own caveat is that the figure is the cheapest model on a test, not the cost of a finished job. If labs keep raising the capability bar while inference prices fall, spending on chips, power, and frontier training can stay high even as a single question gets cheaper. That split sits on both the AI supply chain and the software market that buys the output.
Editor's note
Operator's view, recorded here as the editor's note: two forces are speeding the drop in AI cost — hardware improvement becoming general, and efficiency. Frontier labs such as OpenAI and Anthropic still face rising build costs, knowledge that once circulated as a public good is being absorbed into proprietary systems and sold, and the practical counterweight is models efficient enough to run on local machines. This view is not an Epoch finding. The report prices the cheapest route to a benchmark score. It does not establish hardware as the main cause, and it already costs some open-weight models on rented hardware. Smaller local models are already in use. The frontier scores in the report still depend on large compute. TODO-review: glossary has no entry for benchmaxxing; Korean renders it as 벤치마크 특화 학습.