Google ships Gemini 3 Pro with a 1501 Elo LMArena score
Google launched Gemini 3 Pro with a 1501 Elo LMArena score, shipping it into AI Mode in Search on day one alongside the new Antigravity agentic coding platform.

Updated
Why it matters
- Gemini 3 Pro tops LMArena with a 1501 Elo score, the leaderboard's first-place result at launch
- Gemini 3 Pro posts 91.9% on GPQA Diamond, 37.5% on Humanity's Last Exam without tools, and 76.2% on SWE-bench Verified
- Gemini 3 ships in AI Mode in Search on day one, the first time a Google frontier model debuts in Search at launch
- The Gemini app surpassed 650 million monthly users and AI Overviews reached 2 billion monthly users, per Sundar Pichai
- Google launched Antigravity, an agent-first IDE built on Gemini 3 Pro, available the same day in AI Studio and Vertex AI
Google released Gemini 3 Pro in preview on Tuesday, claiming the top spot on the LMArena Leaderboard with a 1501 Elo score and a sweep of state-of-the-art results across the major reasoning, multimodal and coding benchmarks.
In a note published the same day, Alphabet CEO Sundar Pichai framed the release as a culmination of nearly two years of Gemini work. "AI Overviews now have 2 billion users every month," Pichai wrote. "The Gemini app surpasses 650 million users per month, more than 70% of our Cloud customers use our AI, 13 million developers have built with our generative models."
What does the launch change?
Gemini 3 lands inside Google products on day one for the first time. Pichai confirmed Gemini 3 will run inside AI Mode in Search from launch, alongside availability in the Gemini app, AI Studio, Vertex AI, Gemini CLI and a new agentic development platform called Google Antigravity.
Demis Hassabis, CEO of Google DeepMind, and Koray Kavukcuoglu, the unit's CTO and chief AI architect, called the release a step toward artificial general intelligence. "Today we're taking another big step on the path toward AGI and releasing Gemini 3," they wrote. "It's the best model in the world for multimodal understanding and our most powerful agentic and vibe coding model yet."
How does Gemini 3 Pro perform?
Google's flagship claims the strongest benchmark suite of any commercial model released this year. The headline numbers from Google's evaluation page:
- LMArena: 1501 Elo, first place
- GPQA Diamond: 91.9%
- Humanity's Last Exam: 37.5% without tools
- MathArena Apex: 23.4%, a new state-of-the-art
- MMMU-Pro: 81%
- Video-MMMU: 87.6%
- SimpleQA Verified: 72.1%
- WebDev Arena: 1487 Elo, first place
- Terminal-Bench 2.0: 54.2%
- SWE-bench Verified: 76.2%
Google positions the result as a generational jump from Gemini 2.5 Pro, which led LMArena for more than six months. Gemini 3 also sets a 1 million-token context window, matching the largest available from rival frontier labs.
What is Deep Think?
DeepMind is also shipping Gemini 3 Deep Think, an enhanced reasoning mode that pushes performance further on the hardest benchmarks. Deep Think scores 41.0% on Humanity's Last Exam, 93.8% on GPQA Diamond and a 45.1% on ARC-AGI-2 with code execution under ARC Prize Verified rules.
Google said Deep Think will reach Google AI Ultra subscribers "in the coming weeks" after additional safety evaluations.
How are agentic capabilities different?
Hassabis and Kavukcuoglu described Gemini 3 as their strongest agentic model yet. On Vending-Bench 2, a long-horizon planning test that simulates a year of operating a vending machine business, Gemini 3 Pro outscored rival frontier models by maintaining tool use without task drift.
Google AI Ultra subscribers can access the new Gemini Agent in the Gemini app starting today. The agent can organize Gmail inboxes and book local services through multi-step workflows, according to the company.
Pichai framed the agent push as the natural next step in Gemini's evolution. "Gemini 2 laid the foundation for agentic capabilities and pushed the frontiers on reasoning and thinking," he wrote. "Every generation of Gemini has built on the last, enabling you to do more."
What is Google Antigravity?
Google used the launch to introduce Google Antigravity, an agent-first IDE that runs on top of Gemini 3 Pro, the Gemini 2.5 Computer Use model for browser control and the Nano Banana image editing model. Agents in Antigravity receive direct access to the editor, terminal and browser so they can plan and execute end-to-end software tasks.
Third-party access arrives on day one through Cursor, GitHub, JetBrains, Manus and Replit.
How does reasoning behave in practice?
DeepMind's post described Gemini 3 Pro's output as more concise than 2.5 Pro, with what Hassabis and Kavukcuoglu called less "cliché and flattery" in responses. They framed the model as a "thought partner" that handles dense scientific concepts through code-generated visualizations, a capability Google illustrated with a tokamak plasma flow demo.
Pichai made a broader claim about the direction of the field. "It's amazing to think that in just two years, AI has evolved from simply reading text and images to reading the room," he wrote.
What about safety?
Google called Gemini 3 the company's "most secure model yet" and said it ran the largest safety evaluation suite of any Google AI model to date. The model card highlights reduced sycophancy, stronger resistance to prompt injection and improved protection against cyberattack misuse.
Independent assessments came from Apollo, Vaultis and Dreadnode. Google also granted early access to the UK AI Safety Institute and other bodies under its Frontier Safety Framework.
Why this matters
Gemini 3 is the first frontier model Google has shipped into Search on launch day. The move locks Google's flagship reasoning capability to the largest consumer AI surface in the industry, where AI Overviews already reach 2 billion monthly users.
For the broader market, the benchmark claims tighten the race with OpenAI, Anthropic and xAI. A 1501 Elo LMArena score, a 91.9% on GPQA Diamond and a 76.2% on SWE-bench Verified place Gemini 3 Pro at or near the top of every major leaderboard Google cited, while the simultaneous release of Antigravity signals that Google wants developer mindshare for agentic coding as well as raw reasoning.
What comes next
Gemini 3 Deep Think reaches Google AI Ultra subscribers in the coming weeks. Google said additional Gemini 3 family models will follow. "We plan to release additional models to the Gemini 3 series soon so you can do more with AI," Hassabis and Kavukcuoglu wrote.
The next test is whether Gemini 3 can hold its benchmark lead once independent labs and OpenAI's next flagship hit the same leaderboards.
Original: gemini.google.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
223 articles
Related articles
- Google Ships Gemini 3.1 Pro, More Than Doubling Reasoning Score
- Google Recaps 2025: Gemini 3, AlphaFold Milestones and a Physics Nobel
- Google Launches Gemini 3 Pro at $2/Million Input Tokens
- Google Ships Gemini 3.5 Flash, Promises Pro Model Next Month
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs