📑 Table of contents

Claude breaks the theoretical physics computation record: the Yang-Mills amplitude at 9 loops for $2,000 of compute

Deep Tech 🟢 Beginner ⏱️ 14 min read 📅 2026-09-28

Claude breaks the theoretical physics calculation record: the 9-loop Yang-Mills amplitude for $2,000 of compute

🔎 A physics record falls for the price of a good PC

On August 7, 2026, physicist Matt von Hippel issued a public challenge on his blog 4gravitons: solve a scattering amplitudes problem considered out of reach — the nine-loop amplitude in N=4 super Yang-Mills, or the seven-loop one in N=8 supergravity — with a compute budget accessible to an academic. One month later, the challenge was met. Not by an international collaboration of fifty researchers: by two physicists, one model, and a single prompt.

Liam Fitzpatrick and Siddharth Mishra-Sharma, both physicists at Anthropic, handed the target to Fable 5.1 in Claude Science, the company's paid research harness. The model worked for days almost entirely without supervision, wrote all of its Python code from scratch, and calculated the six-particle nine-loop amplitude — one loop beyond the human record set in 2023 at SLAC/Stanford.

The detail that separates anecdote from history: Lance Dixon, co-holder of the previous record, spent two weeks validating the result via an independent route before confirming it — "quite a triumph," according to the post published by Anthropic on September 25. Total bill: $1,000 to $2,000. This protocol — one prompt, days of autonomy, a result verifiable by a leading expert — that's the real news.


The essentials

  • The result: the nine-loop six-gluon MHV amplitude in planar N=4 super Yang-Mills — 107,053 nonzero coefficients, one loop beyond the 2023 human record (eight loops, Dixon and Liu).
  • The protocol: a one-sentence prompt, then days of autonomous execution with checkpoints every 4 to 6 hours and no scientific corrections along the way.
  • Internal redundancy: two independent derivations — direct bootstrap and the form factor route — converge on the same 107,053 coefficients.
  • External validation: two weeks of verification by Lance Dixon (SLAC/Stanford), via the computation of a related form factor.
  • The bill: ~$1,000–2,000 per approach, of which ~$100 in raw compute (Python/SymPy bootstrap on 96 CPUs for a week) and the rest in model inference time.
  • The counterpoint: Song He's team (Chinese Academy of Sciences) published on September 17 on Zenodo a nine-loop symbol obtained with GPT-6 assistance, via a human-led approach.

Tool Main use Price (September 2026) Ideal for
Claude Science Agentic research harness: long sessions, code execution, checkpoints Paid plan (pricing to be verified on anthropic.com) Labs wanting to delegate multi-day derivations
SymPy Exact symbolic algebra in Python Free, open source Exact computations without numerical drift — the backbone of the bootstrap
Zenodo Public repository for preprints and data Free Publishing community-verifiable results, fast

The August 7 Challenge: Nine Loops on an Academic's Budget

Von Hippel's challenge was not about whether the calculation was feasible, but about its cost: can you push a world record forward with the means of a lone researcher, without a supercomputer or a twenty-person team?

To grasp the audacity, you have to understand what a "loop" means. In quantum field theory, a scattering amplitude — the quantity that predicts the probabilities of particle collisions — is computed as a perturbative series. Each order of the series, each "loop," adds a quantum correction. And the complexity explodes at every step: going from eight loops — the record set in 2023 by Dixon and his colleague Liu — to nine loops is not 12% more effort. It's a combinatorial leap that had resisted every human team.

Von Hippel, a physicist turned science communicator, had calibrated the rules accordingly: each approach had to stay under roughly $1,000 to $2,000 in compute, essentially Claude runtime, as reported by Unite.AI. Two targets were permitted: N=4 super Yang-Mills at nine loops, or N=8 supergravity at seven loops. One month later, the first one fell.


What Fable 5.1 actually computed

The target was the six-gluon, nine-loop MHV (maximally helicity-violating) amplitude in planar N=4 super Yang-Mills — an object that kingy.ai breaks down well. Behind the jargon, three useful clarifications.

First, the theory. N=4 super Yang-Mills is not our universe: it's the drosophila of amplitude physics, a maximally supersymmetric theory where symmetries constrain the answers so strongly that they become almost hand-calculable. "Planar" refers to the limit where the theory simplifies further, and where the most powerful methods apply. "MHV" is the simplest non-trivial helicity configuration.

Next, the method. Summing Feynman diagrams at nine loops is beyond the reach of any machine on Earth. The bootstrap does the opposite: one writes down a general form of the answer, parameterized by unknowns, then imposes consistency constraints — symmetries, analyticity — until the solution is unique. The final product is a "symbol": an exact algebraic structure, not an approximate numerical simulation.

Finally, the execution. Fable 5.1 wrote all of its Python code from scratch, relying on SymPy for exact algebra — hence no rounding drift across 107,053 coefficients, a crucial detail reported by GPTS24. And most importantly: two independent derivations — the direct bootstrap on one side, the form factor route on the other — converge to the same coefficients. For an object of this size, that internal agreement is already a verification in itself.


The protocol: one prompt sentence, days of autonomy

Fitzpatrick and Mishra-Sharma didn't chain together dozens of engineering prompts. A single sentence, an instruction to keep working through the night, then checkpoints every 4 to 6 hours — and no scientific corrections along the way, according to Dataconomy.

For days, the model planned, coded, executed, failed, corrected. Human supervision boiled down to periodic continuation prompts, AlphaSignal summarizes. Nobody touched the physics between launch and arrival.

There is nothing automatic about this level of autonomy. Keeping an agent on the ramp for days requires rigorous context management: persistent instructions, context files, machine-readable resume points. This is precisely the work we detail in Context files: CLAUDE.md, AGENTS.md and beyond — and this use case validates it in hindsight. The difference between an agent that drops off after twenty minutes and one that lasts for days often comes down to context, not the model.


The validation: two weeks of Dixon, "quite a triumph"

The result was neither published nor celebrated before it had survived independent scrutiny — and not from just anyone. Lance Dixon, a professor at SLAC/Stanford and co-holder of the eight-loop record, spent two weeks verifying the amplitude by computing a related form factor, before confirming.

"Quite a triumph": the phrase carries weight coming from the man whose own record was just beaten by one loop. And the validation should be read for what it is: a check via a third route, on a result produced by a machine in a few days of autonomous work.

A word on transparency, because it matters. Anthropic's post is a guest post written by von Hippel, who was paid by the company, and Dixon received Claude usage credits in exchange for his verification. These conflicts of interest are disclosed in black and white; the scientific verification itself does not depend on them. That is exactly the level of diligence one is entitled to demand before piling on the superlatives — and which write-ups like techbriefly or cellcog largely adhered to.


The bill: $100 of CPU, $1,900 of thinking

The dominant line item in this feat isn't hardware: it's model inference time.

Cost item Amount Detail
Bootstrap compute ~$100 96 CPU cores for a week, Python/SymPy
Model inference ~$900–1,900 Days of Fable 5.1 reasoning, retries every 4–6 h
Total per approach ~$1,000–2,000 Estimate for an end user

The reading from AINvest is the right one: when most of the bill is model time, the economic signal is crystal clear. The scarce resource for this kind of computation is no longer the FLOP, it's sustained reasoning over long periods. For the industry, it's one more argument in favor of sustained inference demand rather than training.

Put the figure in perspective: two years of a PhD student, or $2,000 and a month. The ratio is absurd — with one caveat. You need a Dixon to check behind it. The real bottleneck of the protocol isn't generation, which costs almost nothing; it's validation, which costs two weeks of a world-class expert.


The Chinese response: Song He, GPT-6, and the "humans first" approach

Anthropic doesn't have a monopoly on nine loops. A few days after Claude's solve, Song He's team (Chinese Academy of Sciences) published on Zenodo, on September 17, a nine-loop symbol obtained with the assistance of GPT-6 — but through a human-led approach.

The two results mirror each other. On one side, an autonomous agent given a target and a budget, which organizes its work on its own. On the other, physicists who keep control and use the model as a computational instrument. That both paths succeeded within days of each other suggests the problem had become within reach — and above all that the human-AI division of labor remains a methodological choice, not a foregone conclusion.

Claude Science (Fable 5.1) Song He's team (GPT-6)
Steering Agent-led: one prompt, days of autonomy Human-led: the physicists direct
Deliverable 6-particle MHV amplitude, 9 loops, 107,053 coefficients Nine-loop symbol
Validation Dixon, 2 weeks, independent route Zenodo publication (September 17)
Stated cost ~$1,000-2,000 Not publicly detailed

This is no accident of timing: our September 17 briefing already showed the figures outlining this shift from AI toward physics. And on OpenAI's side, the result fits into a series of scientific performances, which we covered in GPT-6 Astra: 64.6% on Terminal-Bench science, ARC-AGI-3 nearly complete. The competition on long autonomous scientific tasks is now open — and it is no longer being played out on five-minute reasoning quizzes.


What it changes: a protocol, not just a record

The real novelty isn't the loop count. It's the triad: one prompt, several days of autonomy, a result verifiable by a leading expert. This triad now has a precedent, a record, and a price tag — in other words, the status of a method.

It's also the second occurrence of the pattern: Anthropic had announced shortly before that Claude had discovered an unknown enzymatic system in 21 hours (our article). With the nine loops, we now have a second data point. A one-off feat is called an anecdote; two, a few weeks apart, start to be called a method.

Three concrete consequences.

The bottleneck shifts to verification. Producing 107,053 correct coefficients: a few days and ~$2,000. Confirming them: two weeks of Lance Dixon's time. Labs that want to adopt this protocol will have to invest in verification — human and automated — as much as in generation.

Access is democratizing, for trained teams. $2,000 is a lab budget line item, not an allocation on a national supercomputer. But let's keep a cool head: the two authors of the prompt are physicists, and the solve relied on known bootstrap methods, with no new physics claimed. The model executed at scale a strategy the community has forged over twenty years.

The compute economy tips toward inference. ~$100 of CPU versus ~$900–1,900 of model time: for this class of workloads, the dominant cost is reasoning, not raw computation. Providers who optimize long-running inference — not just training — are on the right side of the shift.

My take, to be perfectly clear: the next decisive step won't be one more loop. It will be the day a model contributes a new method, not just one more result with existing methods. von Hippel's challenge, moreover, had a second target — seven-loop N=8 supergravity — where the bootstrap toolbox is considerably thinner. There's no indication that one has fallen. That's where the next chapter will be judged.


❌ Common Mistakes

Mistake 1: "AI made a physics discovery"

No. The solve used known bootstrap methods and claims no new physics. The feat lies in execution and autonomy, not conceptualization. The right way to read it: an advance in research automation, not a theoretical revolution.

Mistake 2: "$100 of compute, so any lab can replicate it"

The ~$100 only covers the bootstrap's CPU. The dominant cost item is inference: $900 to $1,900. And above all, you need an expert capable of validating — Dixon spent two weeks on it. Budget for verification, not just generation.

Mistake 3: "Theorists are being replaced"

The two initiators are physicists at Anthropic; the validator is one of the world's leading specialists in the field. The tool compresses derivation time, not scientific judgment. Treat these agents as force multipliers for trained teams, not as replacements.

Mistake 4: "Fable 5.1 beat GPT-6"

The two results are not comparable: one is agent-led with validation by Dixon, the other is human-led and published on Zenodo. Reducing the episode to a score duel is a category error. To position the models on a sound footing, see our Claude, GPT, Gemini, Llama 2026 comparison and our honest Claude 4 vs GPT-5 vs Gemini 3 comparison.


❓ Frequently Asked Questions

What is a "nine-loop" amplitude?

In quantum field theory, it's the order of the quantum correction computed in the perturbative expansion. Each loop adds precision — and a combinatorial explosion of complexity. At nine loops and six particles, the calculation exceeded anything a human team had ever produced: the record stood at eight loops, set in 2023.

Is the result scientifically validated?

On two levels, yes. Internally, two independent derivations (direct bootstrap and form factor) agree on all 107,053 coefficients. Externally, Lance Dixon verified it over two weeks via a third route. Peer-reviewed publication is still to come, but this level of cross-checking is rare for an announcement of this kind.

How much would a reproduction cost?

Between $1,000 and $2,000 per approach, most of it in model inference time; the raw bootstrap compute accounts for only ~$100 (96 CPUs for a week). Add to that the real hidden cost: access to an expert capable of validating the result.

Why N=4 super Yang-Mills rather than the real world?

Because it's the ideal theoretical laboratory: its maximal supersymmetry provides exact constraints that make the bootstrap possible. The results don't directly describe our universe, but the tools — now automated — permeate amplitude physics at large.

Fable 5.1 or GPT-6: which to choose for science?

Both reached nine loops, in opposite roles: Fable 5.1 as an autonomous agent within Claude Science, GPT-6 as an assistant to a human team. The right choice depends on your protocol more than on any ranking — our 2026 comparison details the strengths of each family.

Will theoretical physicists disappear?

No, and this episode proves it: the instigators are physicists, validation required two weeks from a leading expert, and the bootstrap method itself is a human legacy. What disappears is the monopoly of very large teams on the heaviest calculations.


✅ Conclusion

For $2,000, a single prompt sentence, and a few days of autonomy, the frontier of amplitude physics has been pushed back by one loop — and it's the protocol, more than the record, that laboratories will take away from this. If you need to choose a model for this kind of long-haul work, start with our selection of the best LLMs for coding: writing exact code, like the SymPy bootstrap from this episode, remains the first filter to pass.