LifeSciBench: 750 Expert Tasks Show AI's Limits in Real Lab Work
LifeSciBench pairs 750 Ph.D.-authored tasks with 19,020 rubric criteria. GPT-Rosalind passes 36.1%, but collapses to 14.8% on exact numeric outputs and 28.1% with artifacts.
Topic
Topic
LifeSciBench pairs 750 Ph.D.-authored tasks with 19,020 rubric criteria. GPT-Rosalind passes 36.1%, but collapses to 14.8% on exact numeric outputs and 28.1% with artifacts.