Research

OpenAI's GeneBench-Pro Pushes AI on Real Biology Judgment

OpenAI's GeneBench-Pro benchmark shows GPT-5.6 Sol solving just 31.5% of 129 synthetic computational biology problems, up from below 5% for the original GPT-5, while human experts spend 20–40 hours each.

By Elena Vasquez5 min read

Updated

Why it matters

  • GPT-5.6 Sol scored 28.7% on GeneBench-Pro at highest reasoning, 31.5% in Pro mode, up from below 5% for GPT-5 when the original GeneBench launched.
  • GeneBench-Pro contains 129 synthetic problems across 10 domains and 21 sub-domains; 82 were reviewed by external graduate students, postdocs, industry scientists, and professors.
  • Human experts estimate 20–40 hours per problem at roughly $200/hour, while AI inference costs run "only several dollars per problem," per OpenAI.
  • OpenAI is open-sourcing 10 problems on Hugging Face and will provide a 50-question subset to Artificial Analysis for independent benchmarking.
  • OpenAI expects GeneBench-Pro may be saturated by the end of the year at the current pace of improvement.

GPT-5.6 Sol, OpenAI's strongest model at the time of testing, solved just 31.5% of problems on GeneBench-Pro when run in Pro mode — and only 28.7% at the highest standard reasoning setting. The benchmark, released by OpenAI, contains 129 synthetic research-level problems designed to test whether AI can replicate the judgment calls that human computational biologists make every day.

That pass rate marks a sharp climb from where things stood when OpenAI began building the original GeneBench. At that time, the company's best frontier model, GPT-5, scored below 5%. "Progress on this benchmark suggests that frontier models are improving quickly, even on less tangible, systems-level scientific reasoning," OpenAI wrote. "At the current pace, this benchmark may be saturated by the end of the year."

What does GeneBench-Pro actually test?

The benchmark targets what OpenAI calls "research taste" — the chains of judgment calls that shape a scientific study. Each problem gives the model a realistic, messy dataset, brief experimental context, and a target estimand tied to a downstream decision. To answer, the model must explore the data, pick an analytical approach, iterate, and supply a final answer.

The skills in question include:

  • Handling ambiguity
  • Revising assumptions mid-analysis
  • Choosing the correct analytical path
  • Knowing when a result is decision-ready

OpenAI built the benchmark because such capabilities are hard to formalize, and therefore hard to measure, "even as weaknesses in them increasingly constrain overall AI performance."

How was the dataset constructed?

GeneBench-Pro spans 10 domains and 21 sub-domains, from genomics to quantitative biology to translational medicine. OpenAI constructed each problem synthetically, controlling the full causal structure and simulating the data-generating process directly.

That design solves two problems that plague long-horizon biology benchmarks:

  • Multi-step questions built from messy historical data often have no single correct path; two reasonable analysts can pick different cutoffs.
  • If a problem is too numerically insensitive, an agent can make fundamental errors and still produce a passing result.

By tuning complexity directly, OpenAI ensured that reasonable subjective choices still converge on accepted numerical results, while plausible-but-incorrect analyses fail. The team audited problem drafts through trace analyses to check for information leakage and unintended shortcuts.

What did external reviewers find?

OpenAI sent 82 of the 129 questions to outside domain experts — graduate students, postdocs, industry scientists, and professors. Reviewers judged each problem on realism, whether the target answer was identifiable, and whether the methods and estimators were appropriate. The company used the feedback to revise questions.

In a separate survey, those reviewers estimated that a typical GeneBench-Pro problem would take a human expert 20 to 40 hours to complete. At a conservative $200 per hour, the labor cost of a single problem runs into the thousands of dollars.

How do the models stack up?

Scaling test-time compute matters a great deal. At the lowest reasoning level, GPT-5.6 Sol achieved only a single-digit pass rate. At the highest reasoning level, the same model solved nearly six times as many questions as GPT-5.2 while using about two-thirds as many tokens.

OpenAI also compared its models against leading open-source alternatives, including GLM 5.2. The performance gap between GPT-5.6, GPT-5.5, and those open systems was significantly larger than expected when extrapolating from coding benchmarks. That pattern indicates open-source models remain more specialized for code than for broader scientific reasoning.

The company used frontier GPT models to evaluate and harden problems during development, raising concerns the benchmark might be biased against non-GPT families. "Competitor models at best matched the performance of the corresponding GPT model at the time of release, and tended to fall short considerably," OpenAI reported.

How is the benchmark being distributed?

OpenAI is fully open-sourcing 10 representative GeneBench-Pro questions on Hugging Face and has published an interactive web interface for browsing them. The company also plans to provide a 50-question subset to Artificial Analysis for independent third-party benchmarking.

What does the cost gap mean?

AI inference runs at several dollars per GeneBench-Pro problem. Human expert labor costs run into thousands.

That gap matters even if models remain too unreliable to replace people outright. Even partial automation at current performance levels could create meaningful economic and scientific value, the company argued. With sequencing costs plummeting and biobank-scale datasets linking molecular, phenotypic, and health-record data at unprecedented breadth, the limiting factor in biology is shifting from data generation to analysis.

Why does this matter for drug discovery?

OpenAI positioned the benchmark as a first attempt to evaluate the abstract skills behind good scientific judgment — the intuition that lets experienced researchers pick promising initial analyses, revise when data contradict assumptions, and reach conclusions on which downstream clinical, academic, or business decisions can rest.

"If agents can reliably automate this class of analysis, they could significantly accelerate scientific discovery," the company wrote. Human genetic evidence already drives target prioritization and translational follow-up, because mechanisms with genetic support are much more likely to lead to approved treatments.

Models that can reliably perform the analyses now handled by teams of human experts could reshape industrial research — accelerating hypothesis triage, target follow-up, and the iteration cycle between data generation and decision-making.

How much further can the models go?

Frontier models can make partial progress on tough problems but still struggle to close what OpenAI calls the "inferential loop." The failure pattern mirrors the gap between human experts and novices: experts frame the problem and adapt their approach, while novices make observations but struggle to integrate them into broader context.

Frontier systems still solve fewer than a third of GeneBench-Pro problems, and OpenAI framed that ceiling as evidence of substantial headroom. As capabilities advance, benchmarks that probe higher levels of abstraction — beyond book knowledge or routine analysis — will grow more useful, the company said. GeneBench-Pro, in its telling, is one such probe, and the gap between a single-digit pass rate and a 31.5% pass rate over a single development cycle shows just how quickly that probe may need to be replaced.

Original: biorxiv.org

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

195 articles

Related articles

  1. OpenAI says GPT-5.2 sets new state of the art on FrontierMath
  2. OpenAI Says 80 to 90 Percent of Its Research Targets GPT 7 and Beyond
  3. OpenAI Ships GPT-5.1: Faster Reasoning, Better Coding, Same Price
  4. OpenAI says GPT-5.6 Sol cut its own serving costs by 20 percent
  5. OpenAI Ships GPT-5.4 With Native Computer Use and 1M Context

« Previous article