AI News Feed
Market watch
Large Language Models

GPT-6 Astra Solves the Last Unsolved FrontierMath Tier 4 Problem, Epoch AI Declares the Benchmark Saturated

GPT-6 Astra has cracked the only FrontierMath Tier 4 problem no AI had solved, leading Epoch AI to declare the research-level benchmark saturated. OpenAI reports Astra's score at 97.6 percent, and the project is now moving to genuinely open problems.

Epoch AI's definition of saturation is cumulative rather than single-run: when attempts by different models at different times are pooled, every problem in Tier 4 has now been answered successfully at least once. The problem Astra solved was the only one still standing. The model did not answer all 43 problems correctly in one sitting, but the one it got right was the one no AI had reached before.

Tier 4 was introduced on July 11, 2025, when the best score on the leaderboard was about 5 percent. Fourteen months later the barrier has become a step. Jay Pantone, an associate professor of mathematics at Marquette University who wrote the problem Astra solved, said earlier AI systems looked for numerical shortcuts, while this time the model's solution came close to his own. Pantone added that after Astra's recent run of results in mathematics, he finds it hard to be surprised.

FrontierMath was first released on November 7, 2024, with the explicit aim of keeping mathematics benchmarks from being exhausted too quickly. Standard tests such as GSM8K and MATH were no longer separating the strongest models, so Epoch AI worked with more than 60 mathematicians, including Fields medalists Terence Tao, Timothy Gowers and Richard Borcherds, to design original problems that had never been published. After reviewing some of the research-level items, Tao called them extremely difficult and predicted that the hardest tier, Tier 3, might hold AI back for several years. In the first round of testing, leading models scored below 2 percent.

The original core set contained 300 problems split across three tiers. Tier 1 approximated hard undergraduate and olympiad material but allowed more advanced tools, Tier 2 reached senior graduate level, and Tier 3 resembled exploratory research questions from the early years of a doctorate. As reasoning models improved, Epoch added a fourth tier in 2025. Tier 4 problems were written mostly by mathematics professors and postdoctoral researchers, each spending several weeks on short-term research in their own area and compressing the result into a question that could be verified automatically. The initial set of 50 problems covered analysis, number theory, combinatorics, topology and algebraic geometry. At launch, only three had ever been solved across all model attempts, and those solutions rested on assumptions that were correct but not fully justified. Epoch's public sample page stated that some of the problems might not be solved by AI for decades.

The question bank itself then came under scrutiny. OpenAI found more errors in FrontierMath than expected during testing, and Epoch launched an independent audit, using GPT-5.5 and Claude Opus 4.7 to flag suspicious items before mathematicians reviewed them one by one. A v2 version released in June 2026 corrected 12 Tier 4 problems and removed seven, leaving 43. Scores kept climbing afterward: GPT-5.6 Sol reached 83.0 percent and Claude Fable 5 reached 90.2 percent, before Astra reached 97.6 percent and filled in the final gap.

Epoch is already moving past the tiers. The project now includes Open Problems, which draws on research questions that mathematicians have not yet resolved, and FrontierMath Erdos, which formalizes Erdos open problems in Lean and requires AI to produce complete proofs that pass formal verification. On the 68 Erdos problems, Astra solved two.