In this visitant post, physicist and discipline author Matt von Hippel shares what happened whenever he issued a difficulty to AI companies concerning a issue in his erstwhile subfield of theoretical physics.

It’s not frequently that you matter a challenge, lone to see it beaten a duration later. But we’re living in different times.
Let me current myself: I’m Matt von Hippel. I used to be a theoretical physicist; these days I’m a discipline writer. Throughout, I’ve been a blogger, penning weekly at 4gravitons.com concerning discipline and the group who do it.
More and more, blogging concerning discipline has meant blogging concerning AI. That’s a problem, since I’m certainly not an AI expert. I’ve dabbled in it, sure. I likely cognize additional than your grandma. But I mostly have to stage rear and rely the experts. And frustratingly, the experts disagree! I’ve heard from smart, well-informed group who are assured that AI is a few years distant from superintelligence, and that superintelligence volition be capable of really terrifying things. And I’ve heard from smart, well-informed group who are likewise assured that LLM-based AI is near to a ceiling, that models akin Claude won’t equal be capable to do notable activity in physics, let solitary conquer the world.
I’ve been reluctant to create my own predictions. Before forming an opinion, I wanted to see an LLM create advancement on item familiar, item I knew was difficult to do since I’d tried to do item akin myself.
In supplement to that, I wanted to see an LLM do item that I expected to be computationally hard. LLMs have made notable strides in math, certainly, and this duration alone has apt changed many peoples’ minds. But advancement in math comes from new ideas, and ideas are mysterious things: one never fairly knows how difficult they are to discover until they’re found. Computation felt additional solid. I wanted to see an LLM tackle a difficulty that seemed out of attain not since researchers didn’t cognize how to do it in principle, but since doing it seemed akin the benevolent of item that would obtain additional computers and period than the researchers fairly had admission to. I wanted to see if those researchers were wrong: if a smarter, synthetic investigator could use the identical computers, and resolve the issue anyway.
So, I issued a challenge:
“If AI companies desire to power group akin me (or scare us, for that matter), afterward they need to tackle my old field. Show that an AI can obtain the kinds of device resources an scholarly has admission to, and resolve among the scattering amplitudes field’s big outstanding problems. Show that a computational bounds everyone expected to be a issue doesn’t really matter. Give us N=8 supergravity to seven loops, or N=4 excellent Yang-Mills to nine loops.”
In short: can AI resolve a frontier issue in my erstwhile subfield of theoretical speck physics? And can it do it on a budget?
The challenge
My old site is a branch of theoretical speck discipline called amplitudeology. When another speck physicists foretell new particles, they create certain they can do the calculations to test those predictions. They compute formulas called scattering amplitudes, which let physicists use the momenta and energies of subatomic particles to compute how apt they are to react in particular ways. If physicists can create additional exact predictions for these reactions, they can inspect whether results from experiments akin the Large Hadron Collider equivalent those predictions. A mismatch could be evidence for a new theory, one that could explain several of physics’ big lingering mysteries, akin the nature of dreary matter, or the balance between matter and antimatter in the universe.
These scattering amplitude formulas are difficult to compute, so difficult that physicists nearly continually use approximations. They do partial calculations, cut off at a particular figure of “loops,” a measure of how complex interactions between particles are allowed to get. The additional “loops” they contain in their calculations, the nearer they get to the genuine answer, and the harder, computationally, the calculation is to do.
In practice, most scattering amplitude formulas have lone been calculated to two loops. A few have three. The most exact prediction in speck discipline you power have heard of used five.
Amplitudeologists desire to do better. They create experimental new techniques, and test them on particular “toy model” theories. By trying the method alongside a toy example anywhere the calculation is easier, fairly than the additional challenging particles of the genuine world, amplitudeologists can stress-test the new methods and see how far they can go.
I posted challenges for two of those toy models. The one the folks at Anthropic chose to tackle was to go up to nine loops alongside a particular toy example theory, called N=4 excellent Yang-Mills.
“Yang-Mills” is a specialized name for a category of theory that explains most of the earth about us. Three of the four essential forces of nature: electromagnetism, the powerful nuclear power that holds the nuclei of atoms together, and the feeble nuclear power that causes radioactive decay in things akin bananas, are all Yang-Mills theories.
The “N=4 super” comes from supersymmetry. Physicists have speculated that all speck has a “supersymmetric partner,” a speck alongside the identical charge, but of a distinct type, matching matter particles akin electrons to power particles akin photons. At one period they were optimistic these particles could explain dreary matter, via undiscovered partners of additional acquainted particles. Those speculations used “N=1” supersymmetry. In “N=4,” all speck has four supersymmetric partners, not fair one.
That surfeit of particles makes the theory extremely unrealistic. N=4 excellent Yang-Mills isn’t used as an clarification for dreary matter, or for item in the genuine world. Instead, amplitudeologists use it to hone their techniques, since N=4 is paradoxically easier to compute with. The susceptible balance between the distinct particles method lone certain combinations of variables are needed, streamlining calculations.
These calculations were done alongside an experimental method called a bootstrap, which ended up bizarrely well-suited for use of AI. To bootstrap an amplitude, you don’t have to obtain into document all imaginable speck interaction. You fair need to cognize approximately what the answer ought to appearance like, keeping track of all possible in device records in a specialized alphabet. Then you commencement checking everything you know: predictions from another calculation techniques, rules the answer has to obey, links to connected problems anywhere the answer was easier to find. It’s a bit akin Sudoku, anywhere you commencement alongside a grid alongside all imaginable numbers, afterward cross them out as you go. In the end, you’re expecting to discover that lone one possible satisfies all the checks, during having adequate checks remaining complete to create certain you didn’t create a mistake.
That meant that Lance was already fine set up to inspect if person had handed him the next amplitude formula, alongside nine loops. It would be an engaging answer, not fair as a validation of the bootstrap technique, but as a rare example of an amplitude alongside that many loops of complexity, an answer that could be value studying in its own right.
But he hadn’t computed it, and neither had anyone alternatively in the field. The way he established the eight-loop answer was already a bit indirect, via a surprising link to a distinct but connected equation called a form-factor, a benevolent of partial amplitude involving distinct particles that turns out to be a bit easier to calculate. He was expecting to discover the next iteration equal additional indirectly, possibly by a distinct benevolent of AI method. If group idea it was imaginable to fair run the customary bootstrap method for one additional loop, person would have done it.
Then group did it
Apparently, there are folks at Anthropic who peruse my blog.
At the end of August, Liam Fitzpatrick and Siddharth Mishra-Sharma, two physicists at Anthropic, reached out to me to say they had tackled among the challenges in my post. After verifying the outcome alongside Lance, they talked me through how they got it.
True to the soul of the challenge, they didn’t use millions of dollars in device power. They used Fable 5.1, operating inside Claude Science, a phase scientists can pay to use. Claude Science is what folks in the biz call a “harness,” a program that uses the Claude LLM alongside organized rules and prompts in command to get additional sturdy and scientifically helpful behavior.
Apparently, following asking Claude which issue it was most apt to be capable to tackle, they gave it a uncomplicated prompt:
“The issue is to compute the Six-particle (hexagon) amplitude in planar N=4 SYM at nine loops.”
From there, they fair kept telling it to keep going, alongside comments like:
“I'm going to sleep and won't be accessible for another multiple hours. Keep operating on this until I inform you to stop. Give me updates all 4-6 hours.”
Claude ended up doing the calculation two distinct ways: the first bootstrap, and the indirect form-factor approach. Either method would have disbursal an end-user about one or two thousand dollars, mostly because of the disbursal of operating Claude for so long. The bootstrap calculation, done alongside the Python programming tongue alongside bundle SymPy, took about $100 of the budget, corresponding to operating 96 CPUs for a week.
Running 96 CPUs for a week power have felt akin a lot whenever I was doing this benevolent of activity ten years ago, but it’s beautiful affordable now if you have a fine reason.
As it turned out, the outcome wasn’t all that far distant for humans either. A few days following I heard from Anthropic, we heard from Song He, an amplitudeologist at the Chinese Academy of Sciences in Beijing. Song’s collection had already gotten the bulk of the result. They’d used several AI assistance, according to GPT-6, but not the benevolent of one-shot nearly human-less method Anthropic used.
Everyone has been affable here, which is a bit of a relief. The humans, Lance and Song and their collaborators, volition get to publish the results, taking period to explain them and analyze them for the advantage of forthcoming researchers. Claude’s function is done, for now.
So, issue solved?
I set my difficulty since I wanted a improved awareness of what current AI can do, and anywhere it could go from here. So what have I learned?
I’d idea this could be a chance to see AI conquer a computational obstacle in a amazing way. Instead, it did item it turned out humans were additionally capable to do. Claude used known methods, alongside a bit additional compute than group had tried to use before. It may have gotten a boost from using Python, and not Maple (Lance’s favorite program for math) or Mathematica (mine), and it may have used much improved application engineering practices than we would have, but not super-intelligently so.
My biggest takeaway is that there is additional low-hanging create out there than you’d expect. Even whenever a goal is uncomplicated and well-defined, sometimes it’s going to appearance much small achievable to experts than it really is. There are group alongside a device discipline backdrop who’ve been telling me for years that amplitudeologists could create a lot additional advancement fair by hiring a few programmers. They should awareness vindicated.
It’s additionally noteworthy that Claude Science accomplished this in one shot, without any specialized oversight additional advanced than “keep going.” These are finicky, messy calculations. If I’d used a week of period on 96 CPUs to do this benevolent of calculation, afterward I’d nearly certainly end up using two weeks: it’s practically guaranteed I’d screw up item on the archetypal try. I don’t cognize how many mistakes Claude made internally on the way, but the harness got it to the end without an exterior collaborator’s input. I’m not certain that surprises me, at this point. But if you didn’t cognize it could do that since you’re motionless thinking of AI as so error-prone that it’s unusable, afterward this have to be your takeaway: It can do this benevolent of item reliably now.
Things certainly appear to be moving fast. In March, AI was accomplishing discipline projects akin a student: smaller-scale tasks alongside a lot of hand-holding and mistakes. In contrast, this is a genuine frontier calculation, the benevolent of item normally tackled by the top experts in amplitudes. While it’s imaginable that this is fair a much additional AI-friendly problem, I don’t think it’s fair that: I think the innovation has genuinely gotten better.
How far can I generalize this? That I’m not certain of.
These toy example theories lean to be the concentration of small sub-communities. The real-world amplitudes calculations are a wider field, alongside many groups trying to attack all another to the frontier. It’s imaginable there’s small low-hanging create there. But I wouldn’t figure on it. I cognize group who activity on those calculations have been increasingly using AI for coding. If group aren’t already checking whether AI discipline harnesses can one-shot frontier calculations there, they ought to (and they ought to have a scheme for how to inspect the results). I wouldn’t be all that amazed if it was imaginable to compression another iteration out on a sensible budget.
Then it becomes a inquiry for the community to discuss: anywhere is the new frontier, and what needs to be figured out next? Unlike many problems in mathematics, amplitudes aren’t fair a training dirt for new methods. There’s a goal, to create predictions exact adequate to difference alongside upcoming experiments. How much nearer is the site to that goal?
More broadly than that, though, I didn’t really get an answer.
I went into this inquisitive not fair concerning what AI can do in investigation today, but concerning the future. When you peruse predictions concerning superintelligence from the days before LLMs, they frequently propose fantastical-seeming risks. People imagined AI that could simulate group to foretell their reactions and manipulate them, or fig out how to build a species-ending microorganism or world-devouring nanotech from archetypal principles. And the customary exception to these risks is that they conflated intelligence, the vague and mysterious origin of new ideas, alongside computational power. Critics asserted that equal a fleet of new datacenters wouldn’t have the computational power to do any of those tasks, that they were nightmares of a sci-fi forthcoming that wasn’t coming any period soon.
I don’t awareness akin I have a improved answer for those critics. I learned a bit concerning what AI can do now, that it can do activity that matters in my old site on a sensible budget, and do it beautiful much autonomously to boot. But I’d hoped to see item stranger, new methods for the calculation itself alongside unexpected power. I’d hoped to get a glance of the future, item that would provision me an informed opinion in debates concerning superintelligence. I wanted to cognize how far AI could shove computational limits… and I awareness akin what I learned current is fair that I was too naïve concerning anywhere the bounds was.
An addendum: How does it awareness to be scooped by a machine?
By Lance Dixon, Professor of Particle Physics and Astrophysics at SLAC National Accelerator Laboratory and Stanford University, who checked Claude's nine-loop result.
Most theoretical physicists I cognize acknowledge that the current era of ample tongue models is going to entirely change the way we think concerning physics. The inquiry was just: whenever was it going to really hit home? For me, it happened on September 1, whenever Liam Fitzpatrick and Siddharth Mishra-Sharma at Anthropic told me that Claude had computed the nine-loop MHV six-particle amplitude in planar N=4 excellent Yang-Mills, and asked me to validate its result.
I'm not going to explain all the specialized conditions in that final sentence; Matt has covered the backdrop above. I do need to citation that there are really two connected objects, the "amplitude" and item we call the “form factor.” Each has an connected figure of loops: one, two, three, and so on. Every iteration command is harder than the former one, computationally, equal following finding many of tricks to create things easier. Also, the form aspect is easier than the amplitude at the identical iteration order. In 2023 Andy Liu and I showed how to use the form aspect and a weird symmetry we call antipodal duality to get the amplitude at eight loops.
Since 2023, my collaborators and I have eyed getting to nine loops, archetypal for the form aspect and afterward for the amplitude, using our 2023 idea. I idea it would be too difficult to do the amplitude directly. So I was really fairly impressed that Claude could do it directly. Not so much since it was a big computational task, but since the entire setup is extremely fragile: if you create any error at all in the computational recipe, it all crashes downward akin a unsuccessful soufflé, and you are remaining to amazement why (and debug). Also, there are so many particulars of the building that are too boring to document completely in a publication. So Claude had to create all that code from scratch.
From the nine-loop amplitude it is comparatively uncomplicated to go rear to the form factor, and it was easier for me to validate the outcome mostly that way. That meant that for the final two weeks I've been validating a result, the nine-loop form factor, that our squad had been operating toward for a brace of years. And a device had solved a issue that I idea was too difficult to do directly. Does that annoy me personally? Is it soul-crushing?
No, for two reasons. One is that our squad already had a run to use tradition transformer models to foretell higher loops, and part of our motto was: “We have all the tools to validate any applicant resolution a device would provision us.” Claude is a distinct benevolent of transformer model, likely complete a myriad times bigger than our tradition one. But sure, we stated we could validate any outcome an AI example would provision us, so we can and should do it. The second logic is that, if you appearance at how Claude solved the problem, it used all the methods my collaborators and I developed complete the years, and it presented the resolution (maybe as a favor to us) in the identical format we had already set up. So during I'm validating Claude's result, Claude is validating all of our former work. In fact, I would province that Claude understands our 2019 and 2023 document improved than any human, apart from my co-authors.
After I wrote this, Song He told me that his collection had additionally computed the part of the nine-loop amplitude called the symbol. (People fair appear to akin to inform me concerning their nine-loop successes, for any reason.) Song's collection used AI (GPT-6) to assistance them compute several of the constraints, but not for the general framework. So now I've been scooped by the two a device and by humans affirmative a machine, inside two weeks.
Going rear to the Claude computation: it's fairly a triumph, in my opinion, for a ample tongue example to execute all of the steps in the complex formula we laid out, and to arrange the computational horsepower. But the additional soul-searching moments volition arrive whenever ample tongue models commencement to arrive up alongside new bodily principles and insights before humans.
Additional material
- The full nine-loop result, in the format used for the before iteration orders;
- The concurrent nine-loop result by Song He, Jirong Jing, and Xiang Li.
Disclosure
Anthropic welcomed Matt von Hippel to compose this article and compensated him for his time. Anthropic personnel gave feedback on drafts; the satisfied and opinions are his own. Lance Dixon validated the outcome independently and received Claude use credits.