Researchers have demonstrated a reinforcement learning (RL) approach that continuously adjusts control parameters of a quantum processor during computation by using quantum error detection events as a learning signal. Described in the Nature paper “Reinforcement learning control of quantum error correction,” the method was tested on the Willow superconducting processor and in large-scale simulations. Under deliberately injected drift the RL controller improved logical stability by 3.5×, produced an additional 20% reduction in logical error rate after expert calibration, and achieved record-low logical error rates: fewer than one per 1,000 correction cycles in the surface code and one per 100 in the color code.
Why continuous calibration is needed
Quantum computers are fundamentally analog systems: qubits are driven by analog signals (frequencies, amplitudes, phases) that can drift in time. Today, keeping a processor calibrated typically requires stopping the computation to retune control parameters, which is a bottleneck if useful algorithms must run for days or months. Continuous calibration aims to correct control settings in situ so the device can maintain reliable operation without interruptions.
Error detection and quantum error correction
Direct measurement collapses quantum superpositions, so Quantum Error Correction (QEC) encodes logical qubits across many physical qubits and uses parity checks to convert analog noise into discrete error-detection events. These detection bits indicate that an error occurred somewhere within a limited space–time region of the circuit but do not reveal the exact location. Decoders such as the neural-network AlphaQubit (trained on real data) and the algorithmic Tesseract are used to infer likely error locations and determine corrections.
Limitations of traditional physics-based calibration
Conventional calibration relies on physical models refined over decades. Across technology domains, hand-crafted models eventually run into limits when complex, hard-to-model effects dominate. As quantum hardware improves, remaining errors increasingly arise from intricate phenomena and subtle hardware drift that are difficult to capture with simple models. Data-driven approaches have unlocked progress in many fields, and quantum control faces a similar transition.
Turning error detections into a learning signal
Reinforcement learning is well-suited to problems where behavior must be improved by experience rather than explicit programming. The innovation here is to repurpose the stream of QEC detection events: besides decoding and correcting errors, these events are fed to an RL agent as an active learning signal. While computation continues, the agent observes detection patterns and adjusts thousands of control parameters to counteract drift and prevent future errors.
Experiment on the Willow processor
The team implemented the RL control on the Willow superconducting processor. By injecting artificial drift into control parameters, they measured the RL agent’s effect on a running error-correcting code. The RL steering increased logical stability by a factor of 3.5, extending the time the device can serve as a reliable quantum memory. Even after extensive human-in-the-loop expert calibration, RL fine-tuning yielded an additional 20% suppression of the logical error rate.
Combining all technologies in the experiment produced record-low logical error rates: fewer than one logical error per 1,000 error-correction cycles in the surface code, and about one logical error per 100 cycles in the color code.
Scalability and simulations
To probe scalability, the researchers ran numerical simulations with systems of hundreds of qubits and tens of thousands of control parameters. The simulations showed that the number of RL training iterations (epochs) required does not grow with system size, thanks to the local sensitivity of QEC detection events. That said, fully realizing the approach’s potential will need tighter integration between the agent and the quantum processor, faster agent–processor communication cycles, and more advanced machine-learning techniques.
Implications for future quantum computers
This work enables a new paradigm: quantum computers that keep computing while learning from their own errors rather than pausing for calibration. Such continuous, autonomous calibration could be essential for sustaining long-running quantum algorithms on future large-scale processors.
Acknowledgements
The work was carried out by the Google Quantum AI team at Google Research in collaboration with teams from Google DeepMind. The authors thank co-authors for contributions to hardware, software, cryogenics and electronics infrastructure.



