What does it cost for AI agents to attempt a frontier mathematics problem?
OpenAI's Navier-Stokes announcement gives us a rare resource profile:
- Approximately 10,000 concurrent agents
- 2.7 million inter-agent messages
- 130 billion output tokens
- 88 hours to produce a proposed proof
- Another 17 hours for Lean formalization
At GPT-6 Astra's published prices, 130 billion output tokens would have a list-price equivalent of:
- $6.5 million at the standard rate
- $3.25 million with Batch or Flex
- $9.75 million at the standard long-context rate
That was not OpenAI's actual bill. The proof-search agents used an unreleased internal model. The calculation also excludes input tokens, failed paths, tool use, infrastructure, and agent-coordination overhead.
Across the broader mathematics campaign, the system generated approximately 300 billion output tokens — an output-only list-price equivalent of $15 million at the same standard rate.
But generation may be the cheaper layer.
Lean verifies that a formal proof follows from what was encoded. It does not establish that the encoding faithfully represents the written argument.
A recent preprint from researchers at the University of Cambridge and King's College London identifies mismatches between OpenAI's Lean formalization and its written Navier-Stokes proof. In one case, the Lean version proves a weaker estimate.
The authors are not claiming that the proposed result is wrong. They are showing that a passing Lean build does not automatically validate the paper it is intended to represent.
Then comes scientific acceptance.
Under the Clay Mathematics Institute's rules, a proposed solution must first be published in a qualifying refereed mathematics publication. It must then receive at least two years of scrutiny and achieve general acceptance within the mathematics community.
Machine generation and formalization: approximately 105 hours.
Independent scientific acceptance: potentially years.
OpenAI has also released more than 700 AI-generated mathematics manuscripts. Meanwhile, arXiv received a record 40,363 submissions in September and introduced a limit of two submissions per person each month.
These are different developments, but they point to the same constraint: AI can generate research faster than qualified humans can evaluate it.
As token prices fall, the cost of producing a plausible result may decline rapidly. The cost of producing an independently validated result may not.
The more useful FinOps metric is not cost per token or cost per generated proof. It is cost per independently accepted result.
Generation, failed attempts, formalization, and expert review — divided by the results the scientific community ultimately accepts.
References
- OpenAI — A Navier-Stokes existence and smoothness result
- OpenAI — Sharing progress in mathematics
- OpenAI — Mathematics research repository
- Bastounis, Circelli, and Hansen — Navier-Stokes lost in translation
- Clay Mathematics Institute — Rules for the Millennium Prize Problems
- arXiv — Content moderation and submission limits
- OpenAI — API pricing