Summary: In this visitant post, Prof. Matthew Schwartz returns to depict a new method to AI-accelerated science. In Vibe Physics, Schwartz discussed similarities in capability between Claude and a discipline alumnus student. Here, he describes what happened whenever he stopped fighting Claude and allowed Claude to discover “Claude-shaped” problems: ones finest suited to the capabilities of the current generation of LLM tools. This led him to build BootLoops, a toolkit for exact calculations in quantitative science. Because akin calculations frequently rotate up throughout extremely disparate areas of science, Claude established connections to ecology, community genetics, and a dozen another fields. These connections were frequently technically accurate but scientifically unremarkable at first, so Schwartz worked alongside domain experts to steer BootLoops toward questions those sectors attention about. Below, we portion additional concerning these projects and how BootLoops came about.
Agentic AI is improving rapidly. Everyone notices the models appear smarter: they cognize more, create small mistakes, and have improved ideas. If you prosecute the trend lines, it is uncomplicated to conversation confidently concerning the possible of AI to revolutionize science. However, scholarly scientists trying to use the models today in their own activity frequently awareness a disconnect. The models may be solving challenging and longstanding problems, but so far these have mostly been well-scoped applications of existing techniques. Many of the headlines appear to be in mathematics, the one part of discipline anywhere a issue can be stated entirely and an answer checked absolutely. But most of discipline is not akin that. And for many of us researchers and students, the extend between those headlines and what happens whenever we use these models can depart us disappointed and anxious.
The center conflict, as I see it, is that although these models are brilliant, operating akin a individual researcher is not what current LLMs do best. Claude and GPT are fine at science, but they are not scientists: yes, they are smart, but it can obtain a lot of hand-holding to get them to create item of specialized value. Physicists would call this an “impedance mismatch”: two systems that all activity fine but are poorly matched, so most of what one puts in never gets through to the other. Here, the mismatch is between what scientists desire and what AI does well.
So how can we fix it?
I started looking for examples anywhere the impedance mismatches are small acute. I began by having Claude build an accessible suite of tools for mathematical physics. Before long, the tools established uses for another problems. This iterative procedure generated a set of application and specialized protocols, which I call BootLoops. BootLoops functions as a benevolent of harness for the LLM, much akin Claude Code or Claude Science is a harness for Claude, or Codex is a harness for GPT. I’ve established BootLoops particularly well-suited for a category of quantitative problems in science. It is additionally open-source, so it can be used alongside any example you like.

Once Claude had BootLoops, it kept noticing the identical pattern: many sectors have problems that a method from mathematics, physics, or device discipline would resolve outright if anyone knew it existed. I started calling these “Claude-shaped” problems, and followed them exterior high-energy theoretical physics—my residence turf—into geology, biology, economics, and linguistics. In these another areas, I could not depend on my own ability to cognize whether what Claude established was interesting. So I established several experts and asked. With their guidance, BootLoops was capable to create substantive advances in many investigation areas.
Below, I portion how BootLoops came about, depict several first findings in areas anywhere I have been applying it, and portion several of my thinking on how to determine the impedance mismatch between individual and AI scientists today.
Claude, obtain the wheel!
Last December, I tried using Claude as a investigation assistant, and established that Claude Opus 4.5 performed akin a powerful alumnus pupil at 20 times the speed. Despite Claude producing a high-quality document at the end of the experiment, it was a slog to get there. I had to accurate all declaration it wrote, steer it distant from irrelevant threads, and drag it rear from deceased ends.
This summer, I tried doing item different: alternatively of treating Claude akin the collaborator I wanted it to be, I started to treat it akin the collaborator it really is. This required looking for problems suited to its strengths. Right now, Claude is fair not capable to assistance me alongside profound conceptual questions—but it does have a practically unlimited breadth of cognition throughout all domains, amazing coding skills, leading-edge cognition of math and statistics, and the capability to parse papers, appendices, and data at device speed.
A natural location to commencement looking for Claude-shaped problems was in areas anywhere coding could help. When Anthropic released Claude Fable 5 in Summer 2026, I wanted to see whether its cyber capabilities would translate to specialized computing. So I sought to test it by having it port, code up, and enhance assorted methods from a fistful of my document and the neighboring writings on scattering amplitudes.
Scattering amplitudes are how we construe data from the Large Hadron Collider: smash two protons at 13 trillion speck volts, and the amplitude is the theoretical extend between the debris and any new particle, a Higgs boson or item unknown, the collision produced. At their center are Feynman diagrams, multidimensional integrals of a particular form. The ones we are struggling alongside now can all be a PhD thesis, or inhabit a collection for years.
Over the former 20 years an substitute has grown up: the S-matrix bootstrap. In the bootstrap approach, alternatively of grinding out the integral, you enforce bodily constraints until lone one answer is possible. Knowing anywhere the amplitude is infinite (its “singularities”) power narrow it to 20,000 options; a symmetry cuts that to 500; and so on downward to one. The traditional bootstrap is purely analytic and has gone furthest in the most symmetric theories, anywhere the constraints attain all the way to a sole choice (the nine-loop amplitude in N=4 super-Yang–Mills is an example). Closer to the genuine earth you frequently run out of constraints before the end. A newer pivot, the semi-numerical bootstrap, closes the gap whenever you can additionally compute the amplitude at a fistful of points to ridiculous precision (sometimes 1,000 digits): if few adequate options remain, those numbers pin downward the remaining coefficients exactly.
The semi-numerical bootstrap seemed ideal for agentic AI. It draws on mathematics, physics, and device discipline that no one individual has mastered; it needs a awesome agreement of coding and algorithm development; and it is checkable, since the identical numerics let anyone, expert or not, verify the final answer against the integral to as many digits as they akin by operating two scripts. The community's ability is additionally unevenly distributed: fine ideas sit in Wolfram Language, C++, Python, or Julia, and many additional sit in document alongside no code at all. So my archetypal project for Fable 5 was to harbor all of it to a average framework, and to compose the code the document never provided.
Claude did this effortlessly. I was amazed whenever it reproduced the results from my document in about 20 minutes, during the code I wrote to do it took me weeks. However, I was not amazed whenever it informed me that I was doing item extremely inefficiently and that there was a improved algorithm I was unaware of.
Then, I asked Claude to hunt for unsolved amplitudes it could compute. It turns out that problems that are uncomplicated adequate for the S-matrix bootstrap are additionally uncomplicated adequate for humans to do—and, indeed, most have been done. It nevertheless established a few unsolved problems. After several conversation alongside the model, it became apparent that Claude was limiting itself to amplitudes alongside the simplest family of functions: logarithms. So I asked, could it do the identical item for the next simplest family, elliptic functions?
Elliptic integrals are really hard, equal for group akin me who expend a lot of their period computing integrals. Only a fistful of elliptic Feynman integrals have always been computed, and none entirely by the bootstrap, at smallest to my or Claude’s knowledge. The matter is not that the methods wouldn’t work, but fairly that nobody had tried, since the ability needed to do so is distributed among many humans. Claude, by contrast, effortlessly generalized all of the machinery it had ported and built for the logarithmic case to these another integral classes. This period it wrote most of the application itself, or borrowed it from math fairly than physics. As the toolkit grew, it started to district one integral following another. Soon we had 30 integrals BootLooped from end to end, comprising 15 reproductions of known results by this new method and 15 that had never before been computed.
That was all following lone a few weeks. Initially, I idea I would be satisfied fair to do a write-up on that, but I was too tempted to see what alternatively we (that is, me and Claude alongside BootLoops) could do.
“I cognize Kung Fu”
A serendipitous characteristic of discipline is that the identical equations frequently appear complete and complete again in distinct contexts. In physics, for example, the diffusion equation, Fokker-Planck equation, and Schrödinger equation all have the identical mathematical form, so if you create a method to resolve one problem, you can frequently use it to many others. I knew that computations BootLoops was fine at were applicable elsewhere: in cosmology and cord theory, for example. What I didn’t know, but Claude was blessed to inform me, was that these integrals could additionally map onto Bayesian evidence integrals in community genetics, or that the finite-field methods used for Feynman integral decrease could additionally use to problems in evolutionary biology.
This kicked off a particularly productive duration of searching for and solving Claude-shaped (and, additional narrowly, BootLoops-shaped) problems. Some ideas were immediately apparent as awesome applications. For example, the discipline of phylogenetics studies how to arrange category into family trees using DNA; the applicable computation is a Bayesian evidence integral, whose output is a sole figure saying how fine a applicant tree accounts for the observed DNA, an integral of the identical benevolent BootLoops was initially built to do. As the ideas came, I insisted that Claude the two use old tools and build new ones, so that the harness would grow. Each tool, equal from a unsuccessful project, opened up new doors. Claude was getting additional and additional capable. It was akin Neo in The Matrix following waking up from the Kung Fu download.
As the projects drifted from my expert comfort zone, however, I started to worry. When Claude claims item it did in my site is fantastic, I can fairness whether that’s true or not (it frequently isn’t). But whenever it claims item it did in another site is fantastic, I discover myself agreeing. My suspicion heightened, I knew I needed to bring in several experts to be sure. Indeed, I established that in nearly all cases, Claude was technically correct, but the outcome was not all that engaging until the expert helped steer us.
One specified project was on neutral biodiversity theory in ecology. In any ecosystem, several category thrive during others die off. Neutral theory asks how much of that life history is because of random chance. Building on before work, in 2001, the ecologist Stephen Hubbell suggested, provocatively, that perchance it was all random. In 2005, the ecologist Rampal Etienne put an equation to Hubbell’s theory allowing it (at smallest in principle) to be tested in a exact way. Unfortunately, for 20 years nobody could resolve the equation at scale. Claude recognized Etienne’s equation as BootLoops-shaped, and solved it. Applying the computation to data, we established that in ecology’s most studied timber on earth, Barro Colorado Island in the Panama Canal, the mix of tree category changes 4.5 times faster than neutral theory allows.
Excited by this result, I brought it to James O’Dwyer, a prof in factory existence discipline and an expert in neutral theory. James was tolerant alongside me. Though he was impressed by the specialized feat, he stated the results would apt “be met alongside a gesture by many ecologists.” Ecologists had already observed, additional qualitatively, that neutral theory can’t keep up alongside genuine forests. But James had a improved idea: subtract the neutral prediction and study the remainder. This would provision us a improved awareness of what changes were attributable to natural selection, competition, and the differences between species.
James and I proceeded to activity intensively alongside Claude to accomplishment a successor model. James has a discipline background, so the disciplinary tongue obstacle was additional effortlessly surmountable, but, akin many scientists, he had not yet really appreciated the power of agentic AI. As the collaboration progressed, I became a kind of Claude handler, translating Claude-speak to James and keeping the example on track, during James pushed the example to create item ecologists power value. The final outcome is item all three of us are arrogant of: a minimal predictive example of existence histories in outstanding accord alongside data. We are currently extending the example from Panama to another earth timber plots, using datasets that Claude has helped curate.

Another project engaged looking at exact calculations in community genetics, a site that studies how genes change inside a community and why. Claude archetypal used methods imported from mathematical discipline to resolve a 30-year-old integral expression for how natural choice shapes rare mutations. We applied it to gnomAD, the largest community catalog of individual familial variation. Claude was extremely enthusiastic concerning this result, but I wasn’t so sure. I had to compose to three distinct biologists for validation before one responded, but at final I was capable to enlist my coworker Michael Desai, who plant in the field, although not on this exact problem. As alongside James, Michael was impressed by the specialized outcome but not compelled by the science.
Michael noted, however, that additional impactful results power arrive from studying correlations between pairs of mutations on a sole chromosome, using the identical or akin methodology. Within this scope, Claude established and built new tools, akin the inequality certificates used for computer-assisted proofs in mathematics, and added them to the BootLoops kit. We afterward analyzed 5.7 milliard pairs of nearby mutations in genomes from the 1000 Genomes Project, and established evidence for a scheme called cistron conversion. This is an crucial finding since nearly all inspection that uses connected familial variation, from community former to sickness mapping, ignores cistron conversion. The bigger instruction is that AI can assistance discover existence discipline hidden in enormous genomic datasets, of which there are thousands sitting in community archives.
Other projects had a akin arc: Claude comes up alongside an first finding I discover compelling. I bring it to an expert, who is unmoved but sees the potential, and together we sculpt it into meaningful science. Interestingly, several of these drifted from BootLoops-shaped quantitative calculations to additional broad Claude-shaped work. For example:
- Economics. Economics journals now ask authors for a replication package, and an publishing company is frequently tasked alongside checking it. That is careful, manual work. Working alongside two economists, we built an AI data editor: it ported the packages of 4,452 document in five foremost journals from MATLAB, Stata, and another business tools to open-source code (some 30,000 routines) and checked basically all validatable figure against the published tables. The results appear in an NBER operating paper.
- Linguistics. Word-stress catalogs are typically assembled by hand, alongside a few hundred languages, choice effects, and individual preferences. Claude scoured all openly accessible origin and, operating alongside three linguists, we produced AccStack: a database of term stress covering 6,072 languages, quoting the deciding passage for nearly all entry, affirmative a bibliography of 160,000 phonology works.
Some additional highlights, all of which was done in collaboration alongside experts, and all of which is undergoing additional exploration and verification, include:
- Phylogenetics. We made the Bayesian evidence for evolutionary trees accelerated to compute and exactly checkable, and established that sole genes are frequently tied between competing trees by small than the error of norm sampling programs.
- Earth science. We blended equations and data from atmospheric science, geochemistry, and climate modeling to build a quantitative, predictive example of how the Great Oxidation Event unfolded through four glaciations.
- Genomics. We studied the data of single-cell RNA counts at measure to characterize bursting rates and the departure from the textbook telegraph example of cistron expression.
- Sunspots. We deduced the lifecycle of sunspots using the methods of individual community demographics, afterward extended the methodology to characterize starspots, a confounding aspect in many exoplanet searches.
- Mathematical physics. We solved Watson’s “final problem”: the exact come back probability of a 3D random stroll alongside three unequal hopping rates, the final of the lattice integrals George Watson began in 1939.
- Cosmology. We built a complete theoretical toolkit to fit cosmological parameters to data on the universe’s large-scale structure, including the complete two-loop power spectrum and the one-loop trispectrum.
- Statistics. Mixture models are notoriously ill-behaved; we derived exact approximations to their Bayesian evidence that create example choice practical, alongside applications from biostatistics to textual analysis.
Overall, this method of iterative betterment of BootLoops, searching for Claude-shaped applications, and pushing to the specialized frontier alongside experts has been astonishingly productive. More data concerning these projects and others can be established on bootloops.ai.
The specialized details
Running many projects at formerly (36 manuscripts in 18 sectors alongside 19 coauthors complete three months, out of several 400 applicant problems) takes a lot of coordination, the two for me and for Claude. The setup that makes this manageable is fairly simple. The Claude Code sessions all run in terminals on Google Cloud virtual machines, connected to assorted GitHub and Overleaf repos depending on my collaborators’ preferences. I have distinct sessions for all project, affirmative a expert meeting that coordinates the others, allocates compute, and validates results. Each meeting involves agents that run the computations in the background, alongside intermediate results stored in markdown records in their corresponding folders. These subagents are helpful since I regularly run into Fable 5’s classifiers; alternatively of these blocks corrupting my entire session, they fair close downward a sole agent. I additionally have distinct sessions for penning results; creating the GitHub repos, tool manuals, and the BootLoops website; penning and validating all the distinct tool codes; and checking and rechecking results as an adversarial referee.
Still, the sessions necessitate a lot of guidance. For example, Claude has no awareness of time, and loves to be dramatic. At the end of one project, it told me “Matt — the equation is solved. Two years of campaign, four profound marches, …”. It had lone been operating for three days. When I ask for ETAs, the estimates are continually either way too lengthy or way too short. I tried telling Claude not to commit to an evaluation until it ran a test to confirm, but that didn’t activity fine either. Over time, I started to get a awareness for how lengthy things would obtain myself, since I couldn’t depend on Claude to inform me.
A bigger issue was that Claude’s default method seemed to be to try to grind through a long, multiday calculation fairly than build a new tool that would create that identical calculation obtain uncomplicated minutes. Time and again, I had to inform it to think smarter, not harder. Over the way of these lengthy projects, compaction would frequently boot in, causing Claude to endure crucial context. To defender against this, I had Claude periodically arrange its records and consolidate them, so it continually had admission to the latest type of the plan. I’ve written protocol skills and another tools into the BootLoops harness to location this. Some of these problems have additionally been solved in fine business harnesses, akin Claude Science. But many are fair the sorts of irritations inherent in the current generation of agentic AI that I anticipate the models volition eventually outgrow.
Here are a few additional nonaccomplishment modes I encountered, and tips for dealing alongside them:
- Claude loves to province victory. “Done, alongside one asterisk” is frequently “not done at all.” “Exactly that, alongside one refinement” normally method “no.” Giving Claude apparent and rigid standards for what achievement looks akin can assistance alongside this. On one project, it was arrogant of its proof, up to “one unproven lemma.” That lemma was the entire proof! I stated no unproven lemmas. Then “done” again, but now alongside a new axiom.
- Look at everything yourself. I continually ask to see plots. Even alongside all the monitors I set up, I’ve established the automated checks motionless can’t be trusted. Beware of qualitative claims akin “good agreement.”
- Question the conclusions. Claude is fine at performing calculations, but the conclusions it draws can be wrong. Have it explain what it established until you accept it.
- Supply the taste. Claude can discover Claude-shaped problems, but there are thousands, perchance millions, of them. Claude seems to favor old debates, extremely cited but lengthy forgotten. Asking it to concentration on what is newly possible, not what is merely faster or involves additional digits, helps, but you motionless can’t rely its judgement of whether item is interesting.
- Follow up. Once Claude has a awesome result, ask for an equal greater one. It volition continually arrive up alongside something, and occasionally it’s brilliant.
- Watch out for the grind. Honestly, I never succeeded in getting Claude to evaluation period well. Instead, I developed a awareness of how lengthy item engaging should take, and I learned to awareness whether Claude was making what I considered advancement on a sensible period scale. The example volition grind everlastingly if you let it.
It additionally bears mentioning that these projects were compute- and token-intensive. Part of the expense, for me, came from trying to build tools that would be helpful beyond any one project: connecting fields, and assembling a bundle that others could choice up. I anticipation the BootLoops harness can be used broadly throughout quantitative discipline without all person having to rebuild everything from scratch. Agentic AI makes it easier to nexus ideas and portion codebases, and if we study to use that strength, the community can move faster than any idiosyncratic could alone.
Outlook
No matter how astute the models become, most specialized advancement comes from real-world data that have to be acquired, understood, and checked, alongside all circular notifying the next question. AI can sit inner that loop, and that is genuinely valuable, but it does not downfall the iteration to a point. Understanding is built up in increments, all one resting on the last; a cure for malignancy or a fusion reactor volition attain that way too, not from a sole prompt.
The concentration on Millennium Prize Problems and Big Science may be, as Claude likes to say, a “footgun”: a characteristic that makes it uncomplicated to fire yourself in the foot. Here, the hazard is that unrealistic expectations may discourage productive uses that are already possible. Real advancement volition likely arrive the way it continually has: by strengthening the foundations of specialized investigation and constructing increasingly advanced applications upon them. It volition be sped up dramatically—as illustrated by the activity I’ve described—but I don’t see any evidence or need to revisit the specialized method.
A term I akin that is sometimes used to conversation AI discipline is the convex hull. Take a form and nexus all brace of its points alongside a direct line; the new form you get is the smallest convex form that contains the original: its convex hull. Science is now extremely disjointed: there are jagged frontiers everywhere. Biology pokes out one way, math another. One lab power expend 20 years learning everything concerning a sole set of genes alongside a sole method they have mastered, during neighboring genes and substitute methods lie fallow. A harness akin BootLoops can nexus these points, filling in the hull.

Things are changing fast, and it’s practically unattainable to scheme ahead. Why use for a aid for three years of backing to compute a issue an AI example power end up solving overnight? I motionless don’t cognize how to train alumnus students. Since I published my part on vibe physics, that emotion has lone rotate into additional acute. In several fields, specified as device science, the disruption is downright scary. Two years ago, I would have recommended a way on “Python for engineers” as essential. Today, it’s unnecessary. I have spent a lot of period on device learning, construction models to study assorted bodily phenomena. Now, all of those models can be built by AI, so equal learning the ins and outs of neural networks is of constricted utility. Just inform Claude to discover the state-of-the-art ML method and it volition do it for you.
I think BootLoops shows that if you discover the correct problem, much of the specialized activity of solving it can now be automated. That is liberating, in a way. But humans are motionless needed for the conceptual part. Because of that, I anticipation we can fig out a way to provision humans the credit they deserve equal for Claude-shaped science. Curiously, AI runs the hazard of inverting the customary endeavor model, anywhere the individual at the top gets the credit and the group operating on the issue hands-on get none. I don’t cognize how this resolves, but it won’t be resolved by pretending the individual contribution was the typing.
So anywhere does this depart Joe Scientist? If you ask me, in a awesome place, actually. We now have tools at our fingertips that can create advancement on apparently intractable problems alongside lightning speed. I can foresee huge growth on the horizon in data-rich-but-theory-starved sectors akin systems biology. I can additionally effortlessly ideate individual minds being freed from the tedium of calculation and inspection to do the deeper intelligent activity of direction and guidance. And most of all, I anticipate scientists branching out from their residence turf to discover synergy and collaboration, as I have alongside BootLoops. The most refreshing outcome of focusing on Claude-shaped discipline is that it helps us value what human-shaped discipline looks like.
Details of the BootLoops harness and the discipline it has enabled can be established at www.bootloops.ai. The BootLoops harness is on GitHub for anyone to use and contribute to.
Acknowledgements
The discipline I produced alongside Claude and BootLoops would not have been anyplace near as engaging without the direction and collaboration from my chap humans: Isaiah Andrews, Nima Arkani-Hamed, Michael Desai, Scott Edwards, Noam Elkies, Cecilia Garraffo, Matthew Gordon, Thomas Grimm, Martin Hemberg, Mikhail Ivanov, David Johnston, Gary King, Paul Lewis, Brendan Meade, Amara McCune, Joe Pater, Subhabrata Sen, Siddharth Mishra-Sharma, James O’Dwyer, Kevin Ryan, Jesse Shapiro, and Xiaoyuan Zhang.
Disclosure
During this project, Schwartz has been operating as a visiting investigator at Anthropic. BootLoops is not an Anthropic project; it is owned and maintained by Matthew Schwartz.