Safety & Security

OpenAI Partners With Los Alamos Lab to Test GPT-4o in Wet Labs

OpenAI and Los Alamos will test GPT-4o's vision and voice capabilities in a working wet lab, measuring how much the model upskills experts and novices at real biological benchwork.

OpenAI and Los Alamos National Laboratory announce research partnership
OpenAI and Los Alamos National Laboratory announce research partnershipPeter Blanchard / Openverse
By Elena Vasquez4 min read

Updated

Why it matters

  • OpenAI and Los Alamos National Laboratory's Bioscience Division will run the first evaluation of multimodal frontier models in a physical laboratory setting, testing GPT-4o and its unreleased real-time voice systems on standard wet lab tasks such as transformation, cell culture, and cell separation.
  • The study will measure 'uplift' — gains in task completion and accuracy — for both experts/PhDs and novices, using safe protocols as a proxy for dual-use tasks of biological concern.
  • The partnership follows the October 2023 White House AI Executive Order, which tasks DOE national labs with evaluating frontier AI models' biological capabilities, and is consistent with OpenAI's Frontier AI Safety commitments from the 2024 AI Seoul Summit.

OpenAI and Los Alamos National Laboratory have announced a research partnership to study how scientists can safely use frontier AI models in physical laboratory settings — what both parties describe as the first experiment of its kind to test multimodal frontier models in a working wet lab.

The study, run jointly with LANL's Bioscience Division, will evaluate how GPT-4o and its currently unreleased real-time voice systems can assist humans performing standard laboratory tasks using multimodal capabilities such as vision and voice. The announcement came from OpenAI's Chief Technology Officer Mira Murati.

"As a private company dedicated to serving the public interest, we're thrilled to announce a first-of-its-kind partnership with Los Alamos National Laboratory to study bioscience capabilities," Murati said. "This partnership marks a natural progression in our mission, advancing scientific research, while also understanding and mitigating risks."

Why it matters

The partnership lands at the intersection of two pressure points for the AI industry: mounting regulatory scrutiny of biological risks from frontier models, and growing commercial deployments of OpenAI's technology in life sciences.

The White House Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence, issued October 30, 2023, tasks the U.S. Department of Energy's national labs with helping evaluate the capabilities of frontier AI models, including biological capabilities. OpenAI's collaboration with LANL — one of the United States' leading national laboratories — positions the company directly inside that federal evaluation pipeline.

The commercial stakes are already concrete. OpenAI says Moderna is using its technology to augment clinical trial development by building a data-analysis assistant for large data sets, and Color Health built a copilot on GPT-4o to help healthcare providers make evidence-based decisions about cancer screening and treatment. Demonstrating that frontier models can be deployed safely in bioscience settings matters for both markets and policy.

What the evaluation will actually test

The experiment will assess the abilities of both experts and novices to perform and troubleshoot a safe protocol consisting of standard laboratory experimental tasks. OpenAI frames these tasks as a proxy for more complex procedures that pose dual-use concern — work that could be misused as well as beneficial.

The task list may include:

  • Transformation — introducing foreign genetic material into a host organism
  • Cell culture — maintaining and propagating cells in vitro
  • Cell separation — through techniques such as centrifugation

By measuring the uplift in task completion and accuracy enabled by GPT-4o, the researchers aim to quantify how frontier models can upskill both existing professionals and PhDs as well as complete novices in real-world biological tasks. That uplift question is central to biosecurity debates: the more a model increases what an untrained person can accomplish at the bench, the greater the dual-use concern.

"AI is a powerful tool that has the potential for great benefits in the field of science, but, as with any new technology, comes with risks," said Nick Generous, deputy group leader for Information Systems and Modeling at Los Alamos. "At Los Alamos this work will be led by the laboratory's new AI Risks Technical Assessment Group, which will help assess and better understand those risks."

Two departures from prior evaluations

OpenAI says the LANL study extends its previous biosecurity work along two dimensions.

First, it incorporates wet lab techniques. The company's earlier evaluations relied on written tasks and responses about synthesizing and disseminating compounds. Written answers, OpenAI argues, do not fully capture the skills required to conduct actual biological benchwork. Knowing that one must run mass spectrometry — or even detailing the steps in writing — is far easier than performing it correctly with real samples.

Second, it uses multiple modalities. Previous work focused on GPT-4 and text-only outputs. GPT-4o's ability to reason across voice and visual inputs could accelerate learning in a lab. A user unfamiliar with the components of a wet lab setup can point a camera at it, ask questions, and troubleshoot scenarios visually rather than describing the situation in writing.

Building on existing safety commitments

OpenAI positions the study as a continuation of established internal and international safety mechanisms. It builds on the company's existing work on an early warning system for LLM-aided biological threat creation and follows its Preparedness Framework, which outlines OpenAI's approach to tracking, evaluating, forecasting, and protecting against model risks.

The partnership is also consistent with OpenAI's Frontier AI Safety commitments agreed at the 2024 AI Seoul Summit, the follow-up to the 2023 Bletchley Park summit where major AI developers signed voluntary safety pledges.

The collaboration follows a long tradition of U.S. public sector institutions — particularly the national labs — working with private companies to translate innovation into advances in areas like health care and bioscience.

OpenAI describes Los Alamos as a pioneer in safety research and says it expects the joint evaluations to contribute to state-of-the-art research on AI biosecurity evaluations as model capabilities continue to improve. If the results quantify meaningful uplift for novices on real benchwork, the findings could shape how regulators and developers define — and test — biological risk in multimodal frontier models.

Original: whitehouse.gov

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. OpenAI Tells Energy Dept: Fast Permits and Federal Backing for AI Supercomputer Hubs
  2. OpenAI Signs Agreement to Deploy o-Series Models for U.S. National Labs
  3. OpenAI Signs Memorandum of Understanding With U.S. DOE
  4. OpenAI Backs Merge Labs' Seed Round to Build Brain-Computer Interfaces
  5. OpenAI Lays Out Vision for AGI That "Benefits Everyone"

« Previous articleNext article »