Simulation Rankings Can Mislead Quantum Decoder Selection
A preprint using Google Willow records separates predicting error rates from choosing error-correction software.

Original conceptual diagram of the preprint’s simulation-to-hardware comparison. The shared comparison covers code distances 3, 5 and 7 over rounds 2–30; greater calibration did not consistently improve ranking agreement. Image creditOriginal QubitWire diagram. Source/method: Manor, Erhili and Jebbouri, arXiv:2609.04557v1 (2026). Use with the exact conceptual caption. Not a photograph, measured-results plot, claimed code release, peer-reviewed result or general advantage demonstration. · https://qubitwire.com/editorial-standards
Choosing software from simulated tests can be misleading, according to a September 3 preprint by Shay Manor, Leila Erhili and Yassine Jebbouri. Using public Google Willow records, they found that fitting simulated noise more closely to a device did not consistently improve the ordering of competing decoders.
Quantum error correction spreads information across and repeatedly measures checks for signs of faults. A decoder interprets those clues. Its job makes the comparison consequential: a software choice should be judged by how reliably it protects the encoded information, not merely by how realistic a simulator appears.
The authors compared four noise models with stored hardware measurements. Their shared ranking comparison covered code distances 3, 5 and 7 over 2–30 correction rounds. Device-calibrated models better predicted absolute error rates, but a simpler operation-specific model sometimes ranked decoders more faithfully.
This is not evidence that calibration is useless. The largest code distance had only one hardware patch, and the models omitted some correlated noise. The result also remains an unreviewed preprint, not an independently reproduced software recommendation. Reproduction is an immediate gap: when checked on September 8, the linked repository displayed only a license and a one-line README, despite the paper’s release claim. Inspectable implementations and additional hardware datasets would make the comparison more useful.