Models

Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs

Google's upgraded Gemini 3 Deep Think claims 84.6% on ARC-AGI-2, a 3455 Elo on Codeforces and Olympiad gold-level results, and opens to enterprises via API.

Gemini 3 Deep Think: Advancing science, research and engineering
Gemini 3 Deep Think: Advancing science, research and engineeringAI-generated
By Sophie Lindqvist3 min read

Updated

Why it matters

  • Gemini 3 Deep Think scored 84.6% on ARC-AGI-2, a result Google says was verified by the ARC Prize Foundation.
  • The model set a 48.4% score on Humanity's Last Exam without tools and reached a 3455 Elo on Codeforces.
  • Deep Think is available today for Google AI Ultra subscribers and, for the first time, via the Gemini API through an early access program for researchers, engineers and enterprises.

Google has released a major upgrade to Gemini 3 Deep Think, its specialized reasoning mode, and for the first time is opening the model to external researchers, engineers and enterprises through the Gemini API via an early access program.

The updated Deep Think is available starting today to Google AI Ultra subscribers in the Gemini app. In its announcement, Google framed the release around a specific use case: research problems that "often lack clear guardrails or a single correct solution and data is often messy or incomplete."

The stakes are straightforward. Reasoning-focused models are the current front line of competition among frontier labs, and benchmark results on adversarial tests like ARC-AGI-2 have become the shorthand the industry uses to rank them. Google says it built the update "in close partnership with scientists and researchers."

The numbers

Google claims a set of results that, if verified independently, would set new highs on several widely watched evaluations:

  • 48.4% on Humanity's Last Exam, without tools, which Google describes as "a new standard" on a benchmark designed to test the limits of frontier models.
  • 84.6% on ARC-AGI-2, which Google says was verified by the ARC Prize Foundation.
  • An Elo rating of 3455 on Codeforces, the competitive programming platform.
  • Gold-medal level performance on the International Math Olympiad 2025.

The ARC-AGI-2 number is the most notable claim. The benchmark, run by the ARC Prize Foundation, was designed to resist pattern-matching shortcuts, and top frontier models spent much of 2025 scoring in the low single digits on it. Google did not publish per-task breakdowns in the announcement.

Beyond math and code

Google also reports results outside Deep Think's traditional strongholds of mathematics and competitive programming. The updated mode achieves gold medal-level results on the written sections of the 2025 International Physics Olympiad and Chemistry Olympiad, and scores 50.5% on CMT-Benchmark, a test of advanced theoretical physics.

The company also cites lineage for these capabilities: specialized Deep Think versions reached gold-medal standards at math and programming world championships last year, and more recently, according to Google, Deep Think has enabled agents to conduct research-level mathematical exploration.

The engineering pitch

Google is positioning Deep Think not just as a benchmark machine but as a working tool. The announcement highlights practical applications: helping researchers interpret complex data and letting engineers model physical systems through code.

One demonstrated workflow stands out for its concreteness. Deep Think can take a hand-drawn sketch, model the complex shape, and generate a file for 3D printing the physical object — a path from informal drawing to manufacturable output in a single step.

The distribution strategy signals where Google sees demand. Google AI Ultra, the subscription tier carrying Deep Think in the Gemini app, targets power users. The API early access program, by contrast, targets scientists, engineers and enterprises — the buyers Google needs if Deep Think is to move from leaderboard entries into research and engineering workflows.

"We can't wait to see what you discover," the announcement closes.

The open question is independent verification. Google says the ARC-AGI-2 score carries ARC Prize Foundation verification, but the other figures come from Google's own reporting. Expect scrutiny from benchmark operators and rival labs in the coming weeks, and expect the early access program's rollout to determine whether these results translate into research-group adoption.

Original: blog.google

Share this article:

More from Sophie Lindqvist

Sophie Lindqvist

Show full bio

Staff writer covering marketplaces and e-commerce at AI In Context.

114 articles

Related articles

  1. Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score
  2. Google Recaps 2025: Gemini 3, AlphaFold Milestones and a Physics Nobel
  3. Google Launches Gemini 3 Pro at $2/Million Input Tokens
  4. Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
  5. DeepMind's New Chief Prioritizes Gemini 4 Ship Date Over AGI Quest

« Previous articleNext article »