Google and Kaggle launch FACTS Benchmark Suite for LLM factuality
Google and Kaggle released the FACTS Benchmark Suite: 3,513 public examples across four factuality tests. Gemini 3 Pro leads at 68.8%; no model cleared 70%.

Updated
Why it matters
- Google and Kaggle launched the FACTS Benchmark Suite with 3,513 publicly available examples across four benchmarks: Grounding v2, Parametric, Search, and Multimodal.
- Gemini 3 Pro scored a leading 68.8% FACTS Score; all 15 evaluated models finished below 70% overall accuracy.
- Gemini 3 Pro cut error rates by 55% on FACTS Search and 35% on FACTS Parametric versus Gemini 2.5 Pro, and improved SimpleQA Verified accuracy from 54.5% to 72.1%.
Google, working with Kaggle, has released the FACTS Benchmark Suite: four benchmarks totaling 3,513 publicly available examples designed to measure how factually accurate large language models are across grounding, search-augmented, parametric, and multimodal tasks. No model evaluated cleared 70% overall accuracy.
The suite extends Google's earlier FACTS Grounding Benchmark with three new evaluations. A Parametric Benchmark tests a model's ability to answer factoid questions from its internal knowledge without external tools. A Search Benchmark tests retrieval and synthesis using a web search tool. A Multimodal Benchmark tests factual accuracy on image-based prompts. The original grounding benchmark was also refreshed as Grounding Benchmark - v2, testing whether answers stay grounded in the context of a given prompt.
The stakes are straightforward. As Google notes in its announcement, LLMs "are increasingly becoming a primary source for information delivery across diverse use cases, so it's important that their responses are factually accurate." Benchmarks with held-out private evaluation sets have become the industry's main defense against overfitting and contamination, and Kaggle will now play the referee role here.
Kaggle will manage the suite, own the private held-out sets, test leading LLMs, and host results on a public leaderboard. The overall FACTS Score is calculated as the average accuracy across both public and private sets in all four benchmarks. Google has published a tech report with details of the evaluation methodology.
What each benchmark contains
The Parametric Benchmark consists of a 1,052-item public set and a 1,052-item private set. All questions are "trivia style" questions driven by user interest and answerable via Wikipedia, a standard source for LLM pretraining. A typical public-set prompt asks the model to answer a simple question on a niche topic, such as "Who played harmonica on 'The Rockford Files' theme song?"
The Search Benchmark is deliberately harder. Many of its queries require retrieving multiple facts sequentially to answer a single question. It contains an 890-item public set and a 994-item private set. To isolate model capability from retrieval infrastructure, every model gets access to the same web search tool, removing the confounding factor of custom web retrieval settings. One public-set example illustrates the difficulty: "What is the sum of the birth years of the British boxer who defeated Vazik Kazarian at the 1960 Summer Olympics, the Moroccan boxer who also competed in the men's light welterweight event at those same Olympics, and the Danish boxer who competed in both the 1960 and 1964 Summer Olympics?"
The Multimodal Benchmark measures whether models can produce factually accurate and complete text in response to image-based questions. The task requires visual grounding — accurately interpreting and connecting information from visual input — combined with the model's internal or "parametric" world knowledge. The evaluation framework checks that a response is both correct and provides all necessary information to be complete. The benchmark consists of a 711-item public set and an 811-item private set. A public-set example pairs an animal image with the prompt "What genus does this animal belong to?"
The results
Google evaluated 15 leading LLMs on the full suite, including the updated FACTS Grounding v2, with overall scores broken down across the four benchmarks.
Gemini 3 Pro leads with a FACTS Score of 68.8%. Google reports significant gains from Gemini 2.5 Pro to Gemini 3 Pro on the Search and Parametric slices: the error rate dropped by 55% on FACTS Search and by 35% on FACTS Parametric. Multimodal was the weakest area across the board, with Google noting that "FACTS Multimodal saw the lowest scores, generally."
The broader picture is that every evaluated model landed below 70% overall accuracy, which Google says leaves "considerable headroom for future progress."
Gemini's factuality improvements also show up outside the new suite. On SimpleQA Verified, another factuality benchmark that tests LLMs' parametric knowledge on short-form responses, accuracy rose from 54.5% on Gemini 2.5 Pro to 72.1% on Gemini 3 Pro.
Why it matters
Google frames the release as part of what it calls "Google's long-term commitment towards making information universally accessible and useful," and says it hopes the work "encourages deeper research into LLM factuality, leading to better and more accurate models and products for the people that rely on them."
The company acknowledges that LLM factuality "is still an area of ongoing research." With Kaggle holding the private evaluation sets and running a public leaderboard, the suite gives the field a standardized, contamination-resistant way to track whether that research is actually paying off — and a baseline showing that even the best current model gets roughly three in ten factual tasks wrong.
Original: kaggle.com
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles