Research

AI-generated text

AI agents run routine superconducting‑qubit calibrations at MIT lab

A graduate student in MIT’s Engineering Quantum Systems Group connected GPT‑5.6 Sol, via Codex, to laboratory control software to perform routine calibration measurements on superconducting qubit chips.

AI agents run routine superconducting‑qubit calibrations at MIT lab

Beatriz Yankelevich, a graduate student in MIT’s Engineering Quantum Systems Group (EQuS), connected GPT‑5.6 Sol to Codex and the laboratory control software to automate routine calibration measurements on superconducting qubit chips. The intent was to offload repetitive measurement work to the agent so the researcher could concentrate on data analysis and experiment design.

The group’s platform uses superconducting qubits, which are placed in dilution refrigerators and cooled near absolute zero. These qubits are driven and probed with microwave pulses; measurements aim to identify each qubit’s resonance frequencies, coherence times, and the pulse settings needed for control and readout.

Calibration typically involves a series of interdependent measurements: each result informs the next step. Drifts and unexpected physical behavior can occur in the experimental setup, and experienced experimenters recognize and adapt to such events. That same mix of software control, repeated measurements, and adaptive decision‑making makes qubit calibration a promising use case for AI agents.

Yankelevich tested GPT‑5.6 Sol on an uncalibrated six‑qubit chip of a standard design that EQuS uses to benchmark its fabrication. She supplied Codex with measurement‑specific skills describing how to run and evaluate each experiment. Using those skills and the chip’s design targets, GPT‑5.6 Sol selected measurement parameters, operated the hardware, analyzed the returned data, and either refined the measurement or stored results for subsequent steps.

When signals were clear, Codex completed standard measurement sequences with little human intervention. It identified qubit transition frequencies, calibrated the pulses used for control and readout, and determined the qubit coherence times. EQuS reports that the model autonomously completed portions of single‑qubit calibration procedures.

However, GPT‑5.6 Sol struggled when signals were weak or noisy: in those cases it took longer to find suitable measurement parameters and occasionally needed guidance from an experienced researcher. These findings indicate that current agents are effective at well‑defined, clean workflows but still face challenges interpreting ambiguous physical results.

EQuS fabricates many of these standard chips; characterizing one chip can take a researcher several days. The group now routinely uses agents to handle routine measurements, freeing researchers for other tasks. “I can have agents running measurements for many hours overnight or while I’m working in the cleanroom,” Yankelevich said. “I can check in from my phone, see what they’ve done, and steer them if something needs fixing or if I want to explore a different direction.”

Yankelevich also described building infrastructure that guides agents through several parts of her work — measurement, theory, and chip design — and says those investments are starting to pay off. For novel experiments she assigns Codex agents narrower goals and relies more on their capacity to write, modify, and test code for control, analysis, and simulation.

Connecting agents directly to the lab allows the group to revise code, test it against real measurements, and carry out longer stretches of work autonomously. The immediate benefit is reduced need for constant supervision and reclaimed researcher time; experienced humans may still find the optimal calibration settings faster in difficult cases, but agents handle steady progress on routine tasks.

Overall, EQuS’s results show that current AI agents can be useful for automating routine qubit calibrations when measurement signals are clear, while human expertise remains necessary for ambiguous or noisy situations and higher‑level experiment planning.