OpenAI Launches IndQA, a 2,278-Question Benchmark for Indian Languages
OpenAI built IndQA with 261 Indian experts: 2,278 questions across 12 languages and 10 cultural domains, adversarially filtered against GPT-4o, o3, GPT-4.5 and GPT-5.

Updated
Why it matters
- IndQA contains 2,278 questions across 12 languages (including Hinglish) and 10 cultural domains, authored with 261 Indian domain experts.
- Questions were adversarially filtered: only items where a majority of GPT-4o, OpenAI o3, GPT-4.5, and (partially) GPT-5 failed were kept.
- India is ChatGPT's second largest market, with about a billion people who don't use English as their primary language and 22 official languages.
OpenAI has released IndQA, a benchmark of 2,278 questions spanning 12 languages and 10 cultural domains, built to measure how well AI models understand and reason about questions that matter to people in India. The company built the benchmark with 261 domain experts from across the country and says it will use it to track model improvement over time.
The release addresses a gap OpenAI describes as structural. About 80 percent of people worldwide do not speak English as their primary language, yet most existing benchmarks that measure non-English language capabilities fall short, according to the company's announcement. Existing multilingual benchmarks like MMMLU are now saturated, with top models clustering near high scores, which makes them less useful for measuring real progress. Current benchmarks also mostly focus on translation or multiple-choice tasks.
"They don't adequately capture what really matters for evaluating an AI system's language capabilities—understanding context, culture, history, and the things that matter to people where they live," OpenAI writes.
India is the starting point for what OpenAI says it intends to become a broader effort. The country has about a billion people who don't use English as their primary language, 22 official languages—including at least seven with over 50 million speakers—and is ChatGPT's second largest market. OpenAI frames the work as part of its ongoing commitment to improve its products and tools for Indian users and to make its technology more accessible throughout the country. The company says it aims to create similar benchmarks for other languages and regions.
What IndQA covers
IndQA evaluates knowledge and reasoning about Indian culture and everyday life in Indian languages. It spans 10 cultural domains: Architecture & Design, Arts & Culture, Everyday Life, Food & Cuisine, History, Law & Ethics, Literature & Linguistics, Media & Entertainment, Religion & Spirituality, and Sports & Recreation.
Questions are written natively in Bengali, English, Hindi, Hinglish, Kannada, Marathi, Odia, Telugu, Gujarati, Malayalam, Punjabi, and Tamil. OpenAI specifically added Hinglish, the announcement notes, given the prevalence of code-switching in conversations.
Unlike existing benchmarks such as MMMLU and MGSM, IndQA is designed to probe culturally nuanced, reasoning-heavy tasks that existing evaluations struggle to capture. Each datapoint includes a culturally grounded prompt in an Indian language, an English translation for auditability, rubric criteria for grading, and an ideal answer reflecting expert expectations.
Rubric-based grading
Scoring works like an exam rubric for an essay question. Each response is graded against criteria written by domain experts for that specific question. The criteria spell out what an ideal answer should include or avoid, and each one carries a weighted point value based on its importance. A model-based grader checks whether each criterion is met. The final score is the sum of the points for criteria satisfied out of the total possible.
The construction process had four stages. First, OpenAI worked with partners to find experts in India across 10 domains. These experts—native-level speakers of the relevant language and English with deep subject expertise—drafted difficult, reasoning-focused prompts tied to their regions and specialties.
Second, each question went through adversarial filtering against OpenAI's strongest models at the time of creation: GPT-4o, OpenAI o3, GPT-4.5, and, partially after its public launch, GPT-5. OpenAI kept only questions where a majority of these models failed to produce acceptable answers, preserving headroom for progress.
Third, experts provided detailed grading criteria alongside every question. Finally, experts added ideal answers and English translations, followed by peer review and iterative fixes until sign-off.
What the benchmark shows so far
OpenAI used IndQA to evaluate recent frontier models and chart progress over the last couple of years. The company reports that its models have improved significantly over time on Indian languages, with caveats, but still have substantial room for improvement. OpenAI also stratifies performance by language and domain, comparing GPT-5 Thinking High against other frontier models.
The caveats are significant for anyone reading the results. Because questions are not identical across languages, IndQA is explicitly not a language leaderboard, and cross-language scores should not be interpreted as direct comparisons of language ability. OpenAI plans to use IndQA to measure improvement over time within a model family or configuration.
The filtering method introduces a second caveat. Because questions were selected specifically for being ones that GPT-4o, OpenAI o3, GPT-4.5, and GPT-5 could not answer sufficiently, question selection is adversarial against these models. OpenAI acknowledges this could confound the relative performance of GPT-5 and could disadvantage all OpenAI models compared to non-OpenAI models.
Who wrote the questions
The 261 experts include journalists, linguists, scholars, artists, and industry practitioners. OpenAI named several examples:
- A Nandi Award winning Telugu actor and screenwriter with over 750 films
- A Marathi journalist and editor at Tarun Bharat
- A scholar of Kannada linguistics and dictionary editor
- An International Chess Grandmaster who coaches top-100 chess players
- A Tamil writer, poet, and cultural activist advocating for social justice, caste equity, and literary freedom
- An award winning Punjabi music composer
- A Gujarati heritage curator and conservation specialist
- An award winning Malayalam poet and performance artist
- A professor of history, specializing in Bengal's cultural heritage
- A professor of architecture, focusing on Odishan temples
Why it matters
Benchmark saturation is a recurring problem in AI evaluation. When top models cluster near the ceiling of tests like MMMLU, researchers lose the ability to distinguish between them or measure real progress. IndQA attacks that problem in a market where the stakes are concrete: India is ChatGPT's second largest market, and roughly a billion people there don't use English as their primary language. If AI systems fail at culturally grounded reasoning in Bengali, Tamil, or Hinglish, they fail those users.
OpenAI says it hopes the release will inform and inspire new benchmark creation from the research community. IndQA-style questions are especially valuable, the company argues, in languages or cultural domains that existing AI benchmarks cover poorly. Similar benchmarks could help AI research labs learn more about where models struggle today—and, in OpenAI's words, "provide a north star for improvements in the future."
Original: huggingface.co
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI Launches 'OpenAI for India' With Tata Data Center Deal
- OpenAI Launches GDPval, a Benchmark Built From Real Jobs
- OpenAI Signals Data Shows ChatGPT Use Deepening Worldwide
- OpenAI Rewrites Its Model Spec Using Public Input From 1,000 People
- OpenAI Partners With Deutsche Telekom to Bring AI Across Europe