MIT Plugs GPT-5.6 Sol Into Quantum Lab, and It Calibrates Qubits Overnight
MIT's EQuS group connected GPT-5.6 Sol via Codex to its dilution refrigerators, and the agents now calibrate six-qubit chips overnight, supervised only from a phone.

Updated
Why it matters
- MIT graduate student Beatriz Yankelevich connected GPT-5.6 Sol via Codex to her lab's experiment-coordination software at the Engineering Quantum Systems Group (EQuS).
- The agent autonomously calibrated an uncalibrated six-qubit chip — identifying transition frequencies, calibrating control and readout pulses, and measuring quantum information retention — when signals were clear, but needed researcher guidance on weak or noisy signals.
- EQuS now regularly runs agents on routine measurements; each standard chip previously took a researcher several days to characterize.
An MIT graduate student has connected GPT-5.6 Sol, harnessed to Codex, directly to her quantum computing lab software — and the AI now runs routine superconducting-qubit measurements for hours at a time, overnight and unsupervised, while she works elsewhere.
Beatriz Yankelevich, a graduate student in MIT's Engineering Quantum Systems Group (EQuS), used the setup to test whether AI agents could streamline experimental workflows that traditionally consume weeks of researcher time. Her results point to a concrete shift in how quantum labs may operate: agents handle the routine calibration work, while humans keep the harder interpretive and creative tasks.
The stakes are practical. Quantum computing uses the properties of quantum mechanics to process information, and it could one day better simulate complex materials and molecules. Unlike conventional processors, quantum processors are built from quantum bits, or qubits. Preparing and running qubit experiments can take months and require hundreds to thousands of preliminary measurements — repetitive, interdependent, software-driven work that sits squarely in the territory AI agents are being built for.
Why superconducting qubits fit AI agents
EQuS studies superconducting qubits, which are cooled to near absolute zero inside specialized devices called dilution refrigerators. These qubits perform operations quickly, are precisely controlled using microwave signals, and can be made using familiar manufacturing techniques and arranged on a chip.
Once a superconducting qubit chip has been fabricated, packaged, and cooled, researchers interact with it entirely through software. That makes Yankelevich's experiments a natural testbed for AI agents. By connecting Codex to the lab software that coordinates experiments, the agent could run measurements, analyze the results, and decide what to try next.
The underlying physics creates the workload. Superconducting qubits are often called artificial atoms because, like atoms, they can only occupy specific energy levels. Microwave pulses move qubits between these levels and probe their quantum state. Researchers design and calibrate the pulse sequences sent to the chip, then digitize and analyze the returning signals. These measurements reveal each qubit's resonance frequencies, which allows researchers to accurately control the qubit; how long the qubit retains quantum information; and the settings needed to perform computations.
Calibration is not a fixed script. It requires a series of interdependent measurements, with each result shaping what happens next. Qubit properties can occasionally drift, and unexpected physical behavior can cause inconsistent results. Experienced researchers recognize these changes and adapt. That combination of software control, repeated measurements, and adaptive decision-making is exactly what makes qubit calibration a compelling use case for AI agents — and what makes it hard.
The six-qubit test
Yankelevich tested GPT-5.6 Sol's ability to run measurements on an uncalibrated six-qubit chip, one of a standard type that EQuS routinely uses to benchmark its fabrication process. She provided Codex with measurement-specific skills explaining how to run and evaluate each experiment.
Using those skills and the chip's design targets, GPT-5.6 Sol chose measurement parameters, operated the hardware, analyzed the resulting data, and then either refined the measurement or saved the result for use in the next measurement.
When the signals were clear, Codex completed a standard sequence of measurements with little researcher intervention. It identified the qubit's transition frequencies, calibrated the pulses used to control and read it, and determined how long the qubit retained quantum information.
The agent struggled when experimental signals were weak or noisy. In those cases, it took longer to find suitable measurement parameters and sometimes needed guidance from an experienced researcher. The results suggest a clear division of labor at the current state of the technology: agents can handle well-defined experimental workflows, but interpreting ambiguous physical results remains a challenge.
From experiment to routine practice
EQuS fabricates many of these standard chips, and each one can take a researcher several days to characterize. The group now regularly uses agents to handle routine measurements, freeing researchers to focus on other work.
"I can have agents running measurements for many hours overnight or while I'm working in the cleanroom," Yankelevich said. "I can check in from my phone, see what they've done, and steer them if something needs fixing or if I want to explore a different direction."
The immediate advantage is steady progress on experimental analysis and measurements without constant supervision. Yankelevich is candid about the limits: experienced researchers may still be able to identify the best calibration settings faster than current AI models. The gain comes from removing the need to monitor every step of the calibration process, reclaiming days of researcher attention per chip.
Routine chip characterization follows a relatively well-defined workflow. Novel experiments do not. For those, Yankelevich assigns Codex agents narrower experimental goals while drawing more heavily on their ability to write, modify, and test new code for control, analysis, and simulation. Connecting the agents directly to the lab lets the group revise code, test it against real measurements, and complete longer stretches of work autonomously.
"I've built infrastructure to guide agents through several parts of my work — measurement, theory, and chip design — and now it's really starting to pay off," Yankelevich said. "I can have multiple agents working on different problems at once, and I spend most of my time on higher-level work — interpreting results, devising experiments, planning next steps for the agents, reading, and writing."
The so-what
For quantum computing research, the significance is throughput. Chip characterization that once consumed several days of a researcher's time per device now runs in the background, and EQuS has moved from a one-off experiment to regular use. The failure mode is equally informative: agents stumble on weak or noisy signals, where physical intuition and experienced judgment still decide the outcome. The pattern emerging at MIT — agents on routine measurement, humans on ambiguity and design — offers a template other instrument-heavy labs can test against their own workflows as agentic coding tools spread.
Original: images.ctfassets.net
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
119 articles
Related articles
- A Quantum Physicist Is Using OpenAI's o1 to Tackle Physics' Biggest Questions
- AI Agents Proposed Over Half the Ideas, Humans Made 85 Percent of Calls
- AI Was Supposed to Hit New Grads Hard. Unemployment Data Says Otherwise
- AI Experts Underestimated the Field's Speed, Study Finds
- Anthropic Says Claude Found a New Enzyme System; CRISPR Researchers Call It Routine