Google's Gemini Deep Think Solves Open Research Problems in Math
Google says Gemini Deep Think autonomously solved four open Erdős database questions and refuted a decade-old conjecture, in results spanning math, physics and CS.

Updated
Why it matters
- An advanced Gemini Deep Think version autonomously solved four open questions from Bloom's Erdős Conjectures database and contributed to five research papers.
- Gemini Deep Think scores up to 90% on IMO-ProofBench Advanced as inference-time compute scales, with scaling holding into PhD-level problems on internal benchmark FutureMath Basic.
- Google explicitly claims no 'Major Advance' or 'Landmark Breakthrough' results; Level 2 'publishable quality' work has been submitted to reputable journals, including an ICLR '26 acceptance.
Google says an advanced version of Gemini Deep Think, guided by expert mathematicians and scientists, has autonomously solved open research problems across mathematics, physics and computer science — including four open questions from a database of Erdős conjectures and a decade-old conjecture in online submodular optimization that it refuted with a counterexample.
The results appear in two papers published in the last week: "Towards Autonomous Mathematics Research" and "Accelerating Scientific Research with Gemini: Case Studies and Common Techniques." The effort was led by Thang Luong and Vahab Mirrokni, with deep technical expertise from Tony Feng and David Woodruff, and involved large author lists that include Demis Hassabis, Jeff Dean, Quoc V. Le, Koray Kavukcuoglu and James Manyika, alongside academic collaborators.
The stakes are straightforward. Since an advanced version of Gemini Deep Think reached Gold-medal standard at the International Mathematics Olympiad in summer 2025 — and a later updated version matched that level at the International Collegiate Programming Contest — the central question has been whether Olympiad-level performance translates into genuine research capability. Google now claims it does, at least in bounded collaboration with domain experts. That matters for how theoretical research gets done, how results get verified, and how the mathematical community documents AI-generated work.
A math research agent called Aletheia
Research-level mathematics differs from competition problems in one crucial way: it draws on advanced techniques scattered across a vast literature. Google notes that foundation models, despite their broad knowledge, often produce superficial understanding and hallucinations in advanced subjects because of data scarcity.
The company's answer is a math research agent, internally codenamed Aletheia, powered by Gemini Deep Think mode. It has two defining features. First, a natural language verifier identifies flaws in candidate solutions, enabling iterative cycles of generation and revision. Second — and Google calls this crucial — the agent can admit failure to solve a problem, which improved researchers' efficiency. The agent also uses Google Search and web browsing to navigate the literature, which Google says prevents spurious citations and computational inaccuracies when synthesizing published work.
On the benchmark front, Google reports that Gemini Deep Think has scored up to 90% on the IMO-ProofBench Advanced test as inference-time compute scales, and that the scaling law continues to hold beyond Olympiad level into PhD-level exercises, measured on an internal benchmark called FutureMath Basic. Notably, Aletheia achieved higher reasoning quality at lower inference-time compute.
Concrete results, carefully graded
Google breaks the mathematical outcomes down by degree of autonomy:
- Fully autonomous research. A paper generated by AI without any human intervention, cited as Feng26, calculates certain structure constants in arithmetic geometry called eigenweights.
- AI-guided collaboration. A paper cited as LeeSeo26 demonstrates human-AI collaboration in proving bounds on systems of interacting particles called independent sets.
- Semi-autonomous evaluation at scale. An evaluation of 700 open problems from Bloom's Erdős Conjectures database, including autonomous solutions to four open questions listed there. On Erdős-1051, the model autonomously solved the problem and helped lead to a generalization reported in a research paper cited as BKKKZ26.
The agent also contributed intermediate propositions to two further papers, FYZ26 and ACGKMP26.
Google has been deliberate about not overselling. After extensive discussions with the mathematical community, the team proposes a taxonomy that classifies AI-assisted mathematics research by significance and degree of AI contribution, feeding into a wider discussion on responsible documentation, evaluation and communication of AI-generated results. Level 2 work — "publishable quality" — has been submitted to reputable journals. The team explicitly states: "Currently, we do not claim any Level 3 ('Major Advance') and Level 4 ('Landmark Breakthrough') results."
Prompts and model outputs are publicly available, and the paper includes a "Human-AI Interaction card" documenting AI contributions.
Terence Tao appears among the external experts thanked for feedback and discussions, alongside a long list including Ravi Vakil, Ciprian Manolescu, Daniel Litt and Yufei Zhao.
Eighteen problems across CS, physics and economics
The second paper extends the same agentic reasoning approach to computer science and physics, collaborating with experts on 18 research problems. The team identifies effective collaboration "recipes," most notably an "Advisor" model in which humans guide the AI through iterative cycles the paper calls "Vibe-Proving" to validate intuition and refine proofs. Tactical techniques include "balanced prompting" — requesting simultaneous proof or refutation to prevent confirmation bias — and code-assisted verification.
The work builds on Google's earlier deployment of an advanced Gemini Deep Think version to assist in reviewing computer science theory papers for the STOC'26 conference. The headline results:
- Classic network problems unblocked. Progress on Max-Cut (efficiently splitting networks) and Steiner Tree (connecting high-dimensional points) had stalled. According to Google, "Gemini broke both deadlocks" by importing tools from continuous mathematics — the Kirszbraun Theorem, measure theory, and the Stone-Weierstrass theorem — into discrete algorithmic problems.
- A 2015 conjecture refuted. A theory paper proposed that in data streams, making a copy of an arriving item is always less valuable than moving the original. Experts spent a decade unable to prove it. Gemini constructed a specific three-item combinatorial counterexample, rigorously proving the intuition false.
- Machine learning optimization explained. Researchers had built a technique that automatically tunes the mathematical "penalty" used to train AI to filter noise, but couldn't mathematically explain why it worked. Gemini analyzed the equations and proved the method succeeds by implicitly generating its own adaptive penalty on the fly.
- Economic theory extended for AI markets. A recent "Revelation Principle" for auctioning AI generation tokens only held when bids were restricted to rational numbers; extending to continuous real numbers invalidated the original proof. Gemini used advanced topology and order theory to extend the theorem to continuous auction dynamics.
- Cosmic string physics. Calculating gravitational radiation from cosmic strings requires analytical solutions to integrals containing singularities. Gemini found a novel solution using Gegenbauer polynomials that absorbed the singularities, collapsing an infinite series into a closed-form, finite sum.
Because computer science has a fluid, conference-driven publication pipeline, Google describes these results by academic trajectory rather than a rigid taxonomy. About half target strong conferences — including an ICLR '26 acceptance — while most remaining findings will become future journal submissions.
What it means
Google frames the results as evidence that general foundation models, combined with agentic reasoning workflows, can act as a scientific collaborator rather than a novelty. The company's own language is measured: even when the model's contribution is identifying errors or refuting conjectures, Google writes, "these outcomes highlight AI's value as a high-level scientific collaborator."
The deeper claim is about workflow. Google describes Gemini as a "force multiplier" for human intellect — handling knowledge retrieval and rigorous verification so scientists can focus on conceptual depth and creative direction. The proposed taxonomy for grading AI contributions, and the decision to publish prompts and outputs, signal that the harder problem ahead is not model capability but trust: how research communities verify, credit and archive results that humans no longer produce alone.
With Level 3 and Level 4 results explicitly off the table for now, the benchmark for the next cycle is clear: a first AI-assisted result that the mathematical community itself would classify as a major advance.
Original: goo.gle
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
121 articles
Related articles
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
- Google DeepMind Launches AI for Math Initiative with Five Elite Institutes
- Google DeepMind Launches National AI Partnership With India
- Google Recaps 2025: Gemini 3, AlphaFold Milestones and a Physics Nobel
- Google DeepMind Joins DOE's Genesis Mission to Bring AI to 17 National Labs