Skip to content

A physicist dared AI to solve a nine-loop problem. Claude mostly did it unsupervised

Anthropic says Claude Fable 5.1 answered a public physics challenge by computing the nine-loop hexagon amplitude in N=4 super-Yang-Mills, working mostly unsupervised for days at an end-user cost of roughly $1,000 to $2,000.

By VibecodedThis 3 min read
Feynman diagrams for quantum chromodynamics scattering processes
Zephyr the west wind, CC BY-SA 4.0, via Wikimedia Commons

A theoretical physicist publicly dared AI companies to solve one of the hardest calculation problems in his field. Seven weeks later, Anthropic says Claude did it. On September 25, the company published a guest post reporting that Claude computed the six-particle scattering amplitude in planar N=4 super-Yang-Mills theory at nine loops, one loop past the previous record, with the model working largely unsupervised for days at a time.

The dare

The challenge came from Matt von Hippel, a former theoretical physicist who writes about the field at 4gravitons.com. On August 7, he asked AI companies to take academic-scale computing resources and crack one of two outstanding problems in scattering amplitudes: whether N=8 supergravity diverges at seven loops, or the nine-loop hexagon amplitude in N=4 super-Yang-Mills. His argument was that these problems are hard in a computational sense, scaling exponentially or worse, but solvable with known methods given enough compute.

At the end of August, two Anthropic physicists, Liam Fitzpatrick and Siddharth Mishra-Sharma, told von Hippel they had tackled the second problem. They gave Claude, running the Fable 5.1 model inside Anthropic's Claude Science research workbench, essentially a one-line prompt naming the task. Then, in the detail that will resonate with anyone who runs coding agents, they told it to keep working while they slept and to send updates every four to six hours.

What Claude actually did

The model carried the calculation through two independent routes: a direct bootstrap of the amplitude in hexagon-function space, and an indirect approach through a related form factor using a symmetry called antipodal duality. The direct route was written in Python with SymPy and ran the equivalent of 96 CPUs for a week, about $100 in classical compute. Anthropic estimates either route would have cost an end user around $1,000 to $2,000, mostly in model inference.

The result extends the eight-loop calculation that Lance Dixon and collaborators published in 2023. The two representations Claude produced agree across all 107,053 coefficients in the comparison, and the data release, published September 16, records checks on symmetry, vanishing conditions, and entry restrictions. Dixon, a physicist at SLAC and Stanford who built much of the bootstrap recipe Claude followed, validated the output and called it "quite a triumph" for a language model to execute the whole complicated procedure. Anthropic discloses that von Hippel was compensated for writing the post and Dixon received Claude usage credits.

The honest caveats

This is not a new law of physics. The agent worked inside methods, representations, and validation practices built by human researchers over years, and it introduced no new physical principle. The computation programs themselves are not distributed, only the machine-readable results, and no independent group has rerun the full calculation end to end. There is also a timing wrinkle: a group led by Song He at the Chinese Academy of Sciences reached most of the same nine-loop result around the same period with GPT-6 assistance, which suggests the field had made the problem computationally ripe.

Why developers should care

Strip away the physics and this is a story about what long-horizon agents can now do. The bootstrap calculation is exactly the kind of brittle expert workflow that defeats automation: one bad index, one wrong basis conversion, and every later step collapses. Claude reportedly wrote and debugged the implementation, organized days of classical compute, and kept going through failure states with minimal human input beyond "keep going." That is the operational milestone, and it is the same capability that matters when an agent is refactoring a 200,000-line codebase instead of a scattering amplitude. The next test, as Dixon put it, is whether AI systems start discovering new principles before humans do. This result is not that. It is proof that the persistence layer is real.