OpenAI's GeneBench-Pro Pushes AI on Real Biology Judgment
OpenAI's GeneBench-Pro benchmark shows GPT-5.6 Sol solving just 31.5% of 129 synthetic computational biology problems, up from below 5% for the original GPT-5, while human experts spend 20–40 hours each.