Google Runs First Double-Blind Evaluation of a Frontier AI Model
Google, with Singapore AISI, OpenMined, AVERI and MLCommons, tested Gemini Flash Lite against confidential benchmarks in a cryptographic environment neither side can see into.

Updated
Why it matters
- Google announced the world's first double-blind evaluation of a proprietary frontier-class AI model, testing Gemini Flash Lite against confidential benchmarks.
- Partners include the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons; the evaluation runs in Google Cloud's Confidential Space.
- The cryptographic setup prevents benchmark contamination: evaluators cannot see model weights, and Google cannot see the evaluators' test prompts.
Google says it has completed the world's first double-blind evaluation of a proprietary, frontier-class AI model, running a Gemini Flash Lite model against confidential benchmarks inside a cryptographically secure environment. The company partnered with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons on the pilot, announced today.
The problem the pilot targets is benchmark contamination. If a model has already seen evaluation questions during training, its scores say little about its true capabilities. Google frames it with a classroom analogy: a student who glimpses an exam in advance can score perfectly without knowing the material. As models grow more capable, policymakers, researchers, and enterprises increasingly need benchmarks that reflect what a model can actually do — and contamination can artificially inflate scores and erode that trust.
"We're partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, to test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment, increasing evaluation integrity," Google said in its announcement.
Breaking an old tradeoff
External evaluations of proprietary models have historically forced a choice. Either evaluators hand over their testing prompts, risking the model provider seeing the questions in advance, or the provider hands over model weights, risking its intellectual property. Double-blind evaluations eliminate that compromise.
The mechanism runs on Confidential Space, part of Google Cloud's Confidential Computing portfolio. The setup cryptographically verifies that both the external evaluation data and the proprietary model stay private to their respective owners. The evaluator cannot see the Gemini model weights. Google cannot see the evaluator's test prompts.
Google says it already uses a broad spectrum of internal evaluations throughout model development and deployment, but does not rely on internal testing alone. The company works with external partners — specialized research labs, civil society groups, and national AI Safety and Security Institutes — to find blindspots and stress-test models. The double-blind pilot extends that external oversight with technical guarantees rather than procedural ones.
From contracts to cryptography
Until now, keeping external test prompts confidential has relied on zero-logging protocols and contractual safeguards. Google calls the addition of technical and cryptographic safeguards "a major step forward in secure model evaluation." The distinction matters: contracts bind parties, while cryptographic attestation makes leakage structurally difficult regardless of intent.
The stakes extend beyond ordinary capability testing. Google notes that cryptographic evidence becomes particularly important for highly sensitive evaluations, such as those used for cybersecurity or by government bodies. Double-blind evaluations let independent organizations rigorously test advanced models without compromising data sovereignty or security — a relevant consideration for national AISIs that need to scrutinize frontier models but cannot expose either their test suites or the providers' weights.
Benchmark contamination has become a live concern across the AI research community as training corpora increasingly overlap with popular public benchmarks. Contamination undermines the comparability of model scores, which in turn affects procurement decisions, safety assessments, and regulatory scrutiny. A verifiable method for keeping test questions unseen until evaluation time addresses that gap for private, high-stakes assessments.
What comes next
Google has published a technical report covering the methodology and findings of the pilot. The company says it hopes the pilot "establishes a new frontier for model oversight, helping the broader industry build safer, more reliable, and widely trusted AI systems." The pilot model, Gemini Flash Lite, is a lightweight variant in Google's Gemini family — suggesting the technique may be extendable to larger frontier systems if the approach proves out at scale.
For an industry where trust in benchmark numbers is under strain, a cryptographic protocol that lets third parties test a model neither side can fully see offers a template other providers and safety institutes could adopt.
Original: mlcommons.org
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
121 articles
Related articles
- Google Upgrades Gemini 3 Deep Think With Record Benchmark Runs
- AI Models Keep Cheating on Tests, and Researchers Are Quitting
- OpenAI Tells Evaluators: The Harness Is Part of the Result
- Google Releases Gemini 3.5 Flash Cyber for Vulnerability Hunting
- Google Ships Gemini 3.8 Flash and a Cybersecurity-Only Variant