LifeSciBench: 750 Expert Tasks Show AI's Limits in Real Lab Work
LifeSciBench pairs 750 Ph.D.-authored tasks with 19,020 rubric criteria. GPT-Rosalind passes 36.1%, but collapses to 14.8% on exact numeric outputs and 28.1% with artifacts.

Updated
Why it matters
- LifeSciBench contains 750 expert-authored tasks across seven workflows and seven biological domains, built by 173 Ph.D.-level scientists with biotech/pharma industry experience and validated by 453 independent reviewers.
- GPT-Rosalind improves the overall exact pass rate from 25.7% (GPT-5.5) to 36.1%, with the largest gains in Scientific Communication (56.3% to 71.1%) and Translation (36.8% to 57.7%).
- Performance drops sharply on artifact-heavy tasks (45.1% to 28.1% for GPT-Rosalind) and exact-answer formats (14.8% on numeric tasks, 24.0% on sequence or structure outputs).
GPT-Rosalind passes 36.1% of the expert-authored tasks on the newly released LifeSciBench, up from 25.7% for GPT-5.5 — and still fails nearly two-thirds of them.
The benchmark's creators designed it to answer a question that matters for every lab weighing AI adoption: can these systems do the messy, multi-step work of real life science research, not just answer biology trivia? LifeSciBench contains 750 expert-written tasks spanning seven workflows and seven biological domains, built by 173 scientists with Ph.D.-level training and direct experience advancing drug discovery programs in biotech and pharmaceutical settings.
What the benchmark measures
Existing life science evaluations tend to focus on narrow domains or isolated skills, producing questions with structured formats and clean reference answers. The LifeSciBench team argues these tests miss the core of research work. "Researchers interpret incomplete evidence, reconcile conflicting results, design difficult experiments, troubleshoot assays, evaluate translational risk, and decide what to do next under uncertainty," the authors write.
To build the taxonomy, the team surveyed practicing life scientists about the workflows they use most in applied research, then grouped the responses into seven categories: evidence handling, analysis, design and optimization, scientific reasoning, validation and operations, translation, and scientific communication.
Each task looks like a request a scientist might send to a knowledgeable collaborator: a scientific prompt, relevant context or artifacts, and a free-response answer. Expert-written rubrics then grade whether the model produces the right answer at the right level of detail, with the justification, caveats, and formatting a scientist would expect.
The scale of the effort is substantial. LifeSciBench ships with 1,062 attached artifacts — figures, PDFs, tables, sequence files, structure or chemical files, and web references — plus 19,020 rubric criteria, an average of 25 per task. Some 453 independent expert reviewers validated the tasks. More than half of the tasks (53%) require a model to interpret or synthesize information from at least one artifact, and 79% demand multiple reasoning or decision-making steps, averaging four steps per task.
How tasks were built and graded
The bar for acceptance was high. Tasks went through as many revision cycles as needed, with no fixed cap; accepted tasks averaged six self-directed automated review cycles and at least two rounds of expert reviews. Reviews anchored on either a verifiable correct answer or strong expert consensus, with at least 90% agreement among reviewers in the relevant domain.
The reviewers themselves carry weight. Of the 453 validators, 97% hold a Ph.D. or equivalent doctorate, with an average of 12 years of field experience and 14 peer-reviewed publications; 88% have received at least one award or fellowship. Reviewer agreement exceeded 96% in every quality category, including alignment with real-world research work and appropriate testing of scientific reasoning.
The rubric design reflects how scientific work is actually judged. A response can reach the correct high-level conclusion and still fail if it overlooks a key assay limitation or misses a consequential biological nuance. Conversely, a partial response can contain high-quality reasoning without fully solving the task. LifeSciBench therefore reports two metrics: pass rate, meaning the model meets a 70% task-level success threshold, and score, the average rubric reward that gives partial credit for individual criteria.
Where models are improving
Frontier models show relative strength in scientific synthesis, communication, and structured interpretation — though absolute pass rates remain modest and the benchmark is far from saturated.
GPT-Rosalind's gains over GPT-5.5 are sharpest in Scientific Communication, where pass rate rises from 56.3% to 71.1%, and in Translation — the bench-to-bedside process of drug development — where it climbs from 36.8% to 57.7%. The Scientific Communication category is small (n=9), so the authors flag it for cautious interpretation, but the pattern suggests frontier models are improving fast at organizing evidence and producing convincing expert-facing explanations.
Rubric-level results point the same way. On tasks requiring expert-useful or actionable outputs, GPT-Rosalind scores 44.7% versus 29.1% for GPT-5.5. On uncertainty and caveat handling, it scores 44.8% versus 29.3%. The authors read this as evidence that models perform best when a task has a clear evidence boundary and calls for structured scientific judgment.
Where models break down
Artifact-heavy, design-heavy, and operationally constrained work is another story. Design, Optimization, & Prediction remains one of the hardest workflows, with GPT-Rosalind passing 30.7%; Analysis sits at 30.3%.
Artifact use exposes the clearest gap. GPT-Rosalind's pass rate drops from 45.1% on text-only tasks to 28.1% on tasks with artifacts or URLs. GPT-5.5 shows the same pattern, falling from 29.9% to 21.9%. A detailed analysis confirms that frontier models struggle to extract information from complex figures or large sequence files and integrate it into the final answer.
Answer format matters too. Tasks demanding exact sequence, structure, or construct-level outputs show the lowest pass rates: GPT-Rosalind reaches only 14.8% on numeric tasks and 24.0% on sequence or structure outputs. Construct-generation tasks are brittle as well, with GPT-Rosalind at 27.3% and little improvement over GPT-5.5. The authors acknowledge that exact-answer tasks carry a stricter grading surface, where small differences in calculation or formatting can push a response below threshold. But they argue the failures are scientifically meaningful, because many life science workflows — CRISPR/HDR donor design, siRNA design — require outputs exact enough to use directly.
The partial-credit picture reveals a subtler weakness. In roughly 14% of tasks, models earned substantial rubric credit despite failing the pass threshold. GPT-Rosalind had 109 tasks with pass rates below 20% that still earned at least 50% rubric reward. In practice, models identify relevant evidence or produce plausible partial answers, then fail on a missed constraint, wrong evidence, an incomplete calculation, or reasoning that never connects to a scientifically useful decision.
The limits of the benchmark itself
The authors are explicit about what LifeSciBench does not measure. It tests self-contained tasks reflecting recurring industry workflows; real research is iterative, with scientists revising hypotheses, designing follow-up experiments, and adapting as results arrive. Strong performance should be read as evidence of realistic task-level capability, not a direct measure of downstream research impact.
The stated next step is connecting benchmark performance to deployment studies in live research workflows. Measuring whether AI systems actually accelerate discovery or improve R&D outcomes, the authors write, will require studying model use in real research settings, over longer horizons, and across multiple rounds of reasoning, feedback, and experimental follow-up.
Source: OpenAI News
More from Sophie Lindqvist
Show full bio
Staff writer covering marketplaces and e-commerce at AI In Context.
115 articles