Science

Google’s Willow quantum chip completed a benchmark in under five minutes that Google estimates would take a leading supercomputer 10 septillion years — roughly 725 trillion times the age of the universe


Five minutes versus 10 septillion years is the kind of comparison that makes every ordinary unit of time seem useless. Ten septillion is 10 followed by 25 zeros. Divide that by the universe’s age of about 13.8 billion years and the result is roughly 725 trillion cosmic lifetimes.

Google’s 105-qubit Willow processor really did finish its test in less than five minutes. But the other side of the comparison was not measured with a stopwatch. It was Google’s estimate of how long the Frontier supercomputer would need to simulate the same specially designed quantum benchmark using known classical methods.

That distinction matters.

The result is evidence that a quantum processor can generate certain outputs that are extraordinarily difficult to reproduce classically. It does not mean Willow is 725 trillion times faster than a conventional computer at spreadsheets, weather forecasting, drug discovery or any other general task. Google itself says the benchmark has no known commercial application.

What Willow actually computed

The test is called random circuit sampling, or RCS. Engineers program the processor with layers of randomly selected quantum gates. Those operations make the qubits interfere and become entangled, creating a complicated probability distribution across all possible strings of zeros and ones. The machine is then measured repeatedly to draw samples from that distribution.

Willow was not asked to calculate a single answer that a scientist could use. It generated bit strings whose statistical pattern should match the output of the chosen circuit. The task is useful because a quantum processor can run the physical circuit directly, while a classical computer must represent or approximate a quantum state that grows exponentially as qubits are added.

In its December 2024 announcement, Google described RCS as an entry test for quantum hardware. If a processor cannot beat classical simulation on this artificial but demanding test, there is little reason to expect it to handle harder quantum workloads. The company also stated plainly that RCS has no known real-world applications.

That makes the benchmark closer to a test track than a delivery route. It can expose the processor’s speed, control errors and ability to coordinate many qubits without pretending that the sampled bit strings solve an industrial problem.

Where the 10 septillion years comes from

Google compared Willow with Frontier at Oak Ridge National Laboratory, which was the world’s highest-ranked supercomputer when the estimate was made. The comparison asks how much classical computation and memory would be needed to simulate Willow’s largest random circuits at a similar fidelity.

No one ran Frontier for 10 septillion years. The figure is an extrapolation from algorithms, hardware performance and memory requirements. Google said it evaluated several memory scenarios and even gave Frontier an unrealistic advantage: full use of secondary storage with no penalty for moving data to and from it. Under those assumptions, the estimate reached 1025 years.

The estimate is therefore meaningful within a model, but it is not an immutable property of the chip. Better classical algorithms can reduce simulation costs. After Google announced in 2019 that its Sycamore processor had done an RCS task that would take a supercomputer about 10,000 years, researchers developed much faster simulations. A team using China’s Sunway system later reported a comparable sampling task in 304 seconds.

That history does not erase Willow’s result. Willow ran larger, deeper and more demanding circuits, and its estimated gap is vastly wider than Sycamore’s original one. It does show why a benchmark comparison should be read as a statement about the best methods known at a particular time. Google also acknowledged that classical performance would continue to improve.

The less spectacular number may matter more

The five-minute run drew the headlines, but Willow’s error-correction data address a deeper obstacle. Physical qubits are fragile. Heat, stray electromagnetic fields, imperfect gates and measurement errors can destroy the quantum information needed for a long calculation. Simply adding more unreliable qubits normally creates more opportunities for failure.

Quantum error correction spreads one logical qubit across many physical qubits. If the physical error rate is below a threshold, enlarging that encoded unit should make its logical information more reliable rather than less reliable. Reaching that regime is necessary for a machine that can execute long algorithms.

In the peer-reviewed Nature paper on Willow’s surface-code memory, the Google Quantum AI team tested logical qubits of increasing size. Moving from a distance-3 code to distance 5 and then distance 7 cut the logical error rate by about a factor of two at each step. The distance-7 memory used 101 of Willow’s 105 physical qubits and recorded a logical error rate of 0.143 percent per correction cycle. Its encoded memory survived about 2.4 times longer than the best physical qubit in the device.

Those are not fault-tolerant computer specifications yet. The experiment protected a logical memory, not a complete useful algorithm built from reliable logical gates. The paper estimates that reaching a logical error rate of one in a million per cycle at the measured performance would require a distance-27 code using 1,457 physical qubits for a single logical qubit. The team also observed rare correlated error bursts that error correction could not handle.

An American Physical Society analysis likewise treated the below-threshold result as the more durable milestone while noting that the classical difficulty of RCS is inferred rather than directly timed. Error correction offers a scaling test that future devices must keep passing as they grow.

What the benchmark does not show

Willow is not a general-purpose replacement for Frontier. Classical supercomputers remain far better suited to almost every scientific and engineering workload in use today. Quantum processors are expected to help only on particular problem structures, and practical advantage will require algorithms whose useful result is worth the cost of preparing, controlling and checking the quantum computation.

The experiment also did not show that Willow can break modern encryption. That would require a fault-tolerant implementation of an algorithm such as Shor’s, with many logical qubits and an enormous number of reliable operations. Willow has 105 physical qubits and demonstrated one increasingly protected logical memory.

Nor does the speed comparison establish that quantum calculations occur in parallel universes. Google’s announcement mentioned that interpretation of quantum mechanics, but the RCS data test the processor’s output statistics, not competing philosophical accounts of what quantum theory means.

The next test needs to be useful

Google’s stated target after RCS is a calculation that is both beyond practical classical simulation and relevant to a real problem. Candidate areas include modeling molecules, materials and other quantum systems. In each case, researchers will need a credible way to validate the answer even when a classical machine cannot reproduce the full computation.

That combination is harder than winning a benchmark. A useful calculation needs reliable logical operations, enough encoded qubits, a suitable algorithm and evidence that the quantum result improves on the best classical approach. It must also survive the cycle in which classical researchers study the challenge and find better methods.

Willow’s five-minute run shows how sharply quantum and classical costs can separate on a task built to reveal that difference. Its error-correction result shows a path toward keeping quantum information alive as hardware grows. The next milestone will be an answer that matters outside the test itself.



Source link