Tim Gowers: What sort of maths are LLMs good at?

Aug 12, 2026 05:04 PM - 1 hour ago 1

For the liking of anyone who mightiness publication this blog station successful the distant early (a period from now, say), fto maine mention that I americium penning it a fewer days aft OpenAI announced that it had solved 10 awesome problems successful mathematics and theoretical machine science, including the first building of a non-sofic group, and a impervious that the multicolour Ramsey number R(3,3,...,3) (where location are k 3’s) grows superexponentially successful k. The first was, to judge from various talks I person been to, 1 of the astir important unsolved problems successful group theory, and the 2nd was a awesome unfastened problem successful Ramsey mentation that I didn’t needfully expect to spot solved successful my lifetime, though of people specified expectations now person to beryllium revised. The logic I want to beryllium clear astir the timing is that I shall beryllium discussing the existent capabilities of LLMs successful the afloat anticipation that those will proceed to alteration rapidly. So it is apt that successful not excessively agelong from now, if location is thing absorbing successful what I write, it will beryllium absorbing chiefly arsenic a grounds of what the business looked for illustration successful early August 2026.

These results, and the different 8 connected the list, are extraordinarily impressive, but it still doesn’t look to beryllium the lawsuit that LLMs are amended than each humans astatine each aspects of mathematics. If they were, past their large velocity advantage complete america would mean that location would beryllium overmuch much of a flood of results. So it is earthy to wonderment astir what kinds of problems LLMs are bully at, and astir wherever location is still room for improvement. I don’t dress to person a bully reply to this question, wherever a bully reply would beryllium a crisp classification that would fresh the existent examples well, but it is an absorbing workout to effort to norm retired immoderate bad answers, and to effort to place imaginable answers that aren’t evidently contradicted by the evidence.

Are LLMs peculiarly bully astatine uncovering counterexamples?

A first remark present is that LLMs are not conscionable bully astatine uncovering counterexamples: they tin find proofs of difficult statements arsenic well. However, it is notable that the astir celebrated problems they person solved person almost each been pinch counterexamples alternatively than proofs. That is existent of the 2 problems mentioned above, and besides of the Jacobian conjecture and the portion region conjecture.

If 1 wants to theorize that LLMs are peculiarly bully astatine uncovering counterexamples, past location are 2 things it would beryllium bully to do to make the mentation much convincing. The first whitethorn sound unproblematic: it is to determine erstwhile solving a problem counts arsenic uncovering a counterexample. Once that is sorted out, the 2nd is to travel up pinch a imaginable mentation of why LLMs would beryllium peculiarly good suited to solving problems of that peculiar kind.

What does it mean to find a counterexample?

Why americium I suggesting that it is not wholly evident what it intends to find a counterexample? Surely, 1 mightiness suggest, each it intends is that you person a connection of the shape “Every entity of specified and specified a type has specified and specified a property,” and you grounds an entity of the fixed type that does not person the fixed property.

However, this doesn’t ever work. Consider a celebrated consequence of Vinogradov, which states that each sufficiently ample affirmative integer is simply a sum of 3 primes. The negation of this connection is (or is balanced to) the connection that for each affirmative integer N location exists an integer n\geq N specified that n is not a sum of 3 primes. In different words, it states that each affirmative integer N has a definite property. Seen successful this light, Vinogradov recovered an illustration of a affirmative integer N that does not person the fixed property. Do we want to opportunity that Vinogradov recovered a counterexample? Clearly not — the consequence should evidently beryllium classified arsenic a theorem and not a counterexample.

Thus, we cannot conscionable naively opportunity that LLMs are peculiarly bully astatine negating universally quantified statements: location has to beryllium thing astir the nature of the cosmopolitan quantification. With the three-primes example, it is clear that Vinogradov did not think, “How americium I going to find N pinch this property?” Rather, what he thought would person been much like, “I’ve sewage an integer n that is very large. How americium I going to show that it is simply a sum of 3 primes?” In different words, each his attraction would person been connected the universally quantified n, pinch the existentially quantified N being a benignant of afterthought erstwhile the specifications of the impervious person been worked out.

In general, galore absorbing results, erstwhile they are stated formally, statesman pinch an alternation of 2 aliases 3 (or more) quantifiers. The mobility past becomes to find which is the first “interesting” quantified adaptable successful immoderate sense. Here’s different illustration to exemplify the point, from the mentation of finite-dimensional normed spaces. I’ll springiness a fewer mathematical specifications for those curious, but if you don’t attraction astir those, past you tin skip the adjacent 3 paragraphs and should get the gist of what I americium saying astir this example.

Let X and Y beryllium 2 n-dimensional normed spaces and fto T beryllium a linear representation from X to Y. We opportunity that T is simply a C–isomorphism if location exists \lambda>0 specified that \lambda|x|\leq|Tx|\leq C\lambda|x| for each x\in X. By rescaling we tin ever return \lambda to beryllium 1, successful which lawsuit we person that |x|\leq|Tx|\leq C|x| for each x\in X. If C=1, past this tells america that T is an isometry. In general, the Banach-Mazur distance d(X,Y) betwixt X and Y is defined to beryllium the smallest C specified that location exists a C-isomorphism from X to Y. It is easy to spot that the logarithm of the Banach-Mazur region is simply a metric connected the group of isometry classes of n-dimensional normed spaces. A little easy fact, but still not excessively hard, is that the resulting metric abstraction is compact: successful fact, it is known arsenic the Banach-Mazur compactum.

It is earthy to wonderment what the diameter of the Banach-Mazur compactum is, and present things get interesting. A consequence of Fritz John states that each n-dimensional abstraction X has region astatine astir \sqrt n from \ell_2^n. (The thought of the impervious is arsenic follows: prime wrong the portion shot of X an n-dimensional ellipsoid of maximal volume; that is the portion shot of a normed abstraction Y that is isometric to \ell_2^n; it tin beryllium shown that the personality representation is simply a \sqrt n-isomorphism betwixt X and Y.) From Fritz John’s theorem and the (multiplicative) triangle inequality, it follows that d(X,Y)\leq n for immoderate 2 n-dimensional normed spaces. That is, the diameter of the Banach-Mazur compactum is astatine astir n. But mightiness it beryllium substantially little than that?

An denotation that the reply is not evident comes from looking astatine the spaces \ell_1^n and \ell_\infty^n. The personality representation betwixt these 2 spaces is an n-isomorphism, but 1 tin do overmuch amended by mapping the modular ground vectors not to themselves but to vertices of the portion cube, pinch the vertices chosen to beryllium arsenic orthogonal arsenic possible. In particular, if location exists an n\times n Hadamard matrix, past the corresponding linear representation is simply a \sqrt n-isomorphism. One tin push this study and deduce that for immoderate p,q\in[1,\infty] the Banach-Mazur region betwixt \ell_p^n and \ell_q^n is O(\sqrt n). It is besides easy to show that d(\ell_1^n,\ell_2^n)=\sqrt n, truthful \ell_p-spaces hardly amended connected the easy little bound, and do not amended connected it astatine each successful dimensions n for which an n\times n Hadamard matrix exists.

In 1981, Gluskin famously solved the problem by determining the correct asymptotics for the diameter of the Banach-Mazur compactum. Informally, what he showed was that the diameter is wrong a changeless of the precocious bound that follows instantly from Fritz John’s theorem. If we make the quantification explicit, past the connection we extremity up pinch is

\exists c>0\ \forall n\ \exists X,Y\in K_n\ d(X,Y)\geq cn,

where I person written K_n for the group of each n-dimensional normed spaces. (If you want to reason that it is not a set, past fto maine specify successful summation that the underlying vector abstraction is \mathbb R^n.) In words, location is simply a affirmative changeless c specified that for each affirmative integer n location are n-dimensional normed spaces X and Y specified that the Banach-Mazur region betwixt X and Y is astatine slightest cn.

I can’t proceed without very concisely describing the beautiful and highly influential thought Gluskin had for solving this problem. He took X and Y to beryllium normed spaces whose portion balls were random symmetric convex sets defined arsenic follows: return the modular ground vectors and a fistful of different random portion vectors, arsenic good arsenic the negatives of each these vectors, and return the convex hull. Gluskin past showed that if 2 normed spaces are chosen from this distribution, past pinch precocious probability their Banach-Mazur region is astatine slightest cn.

But backmost to the main point, which is that the logical shape of the supra connection is very akin to the logical shape of Vinogradov’s theorem, which is

\exists N\ \forall n\geq N\ \exists p_1,p_2,p_3\in P\ \ p_1+p_2+p_3=n

where I person written P for the group of primes. And yet, Vinogradov’s consequence is unquestionably a theorem, while Gluskin’s consequence is unquestionably a counterexample, aliases astatine slightest an example.

What is the important quality betwixt the 2 statements? It seems to beryllium that successful Vinogradov’s three-primes theorem the number n plays a much basal domiciled successful the connection that is to beryllium proved astir the various quantified variables. In Vinogradov’s theorem, that connection is n=p_1+p_2+p_3, whereas for Gluskin’s theorem the connection to beryllium proved is

\dim X = \dim Y = n and d(X,Y)\geq cn,

which we tin constitute equivalently as

\dim X = \dim Y = n and d(X,Y)\geq c\dim X.

In the lawsuit of Vinogradov’s theorem, the full situation is to get those 3 primes to adhd up to n, whereas for Gluskin it is not remotely challenging to get the dimensions of X and Y to adjacent n: the situation is to get X and Y to beryllium very acold from each other, comparative to their communal dimension.

There is simply a further complication to carnivore successful mind here, which is that via the process known arsenic Skolemization, a universally quantified connection of the shape \forall x\in X\ \exists y\in Y\ \ P(x,y) tin beryllium converted into an existentially quantifed connection X\to Y\ \forall x\in X\ \ P(x,f(x)). (For this to beryllium an equivalence 1 needs the axiom of choice, but it is surely a capable condition.) This is not conscionable a portion of logical trickery, but it often reflects rather accurately really we deliberation astir immoderate problems. For instance, it is much earthy to deliberation of Gluskin’s illustration arsenic a look for constructing (or astatine slightest proving the beingness of) a brace of suitable normed spaces for immoderate fixed magnitude n, aliases successful different words to conception a suitable usability from \mathbb N to pairs of normed spaces by giving its worth astatine each n, than it is to deliberation of it arsenic a connection that says that each affirmative integer $n$ has a definite analyzable property.

Yet different complication is that immoderate universally quantified statements travel people from existentially quantified statements, aliases whitethorn moreover beryllium balanced to them. For example, the theorem that a 2-dimensional torus is not homeomorphic to a 2-dimensional sphere is an existentially quantified connection (every representation from the torus to the sphere fails to beryllium a homeomorphism), but the earthy measurement to beryllium it is to beryllium the existential connection that location is an invariant that distinguishes the 2 spaces. For an illustration of wherever a cosmopolitan connection is balanced to an existential statement, see a connection of the shape that a vector x\in\mathbb R^n does not beryllium to the convex hull of a definite compact group A. The connection that nary convex operation of elements of A is adjacent to x is balanced to the beingness of a linear functional \mathbb R^n\to\mathbb R and a \lambda\in\mathbb R specified that \phi(x)>\lambda and \phi(a)\leq\lambda for each a\in A. In some these cases it feels earthy to respect the consequence arsenic a theorem that is proved via an existential statement, possibly because it is the theorem that is yet what interests us. But utilizing “what interests us” arsenic a criterion to find what counts arsenic a counterexample seems a small vague, and seems to beryllium a difficult criterion to usage if we want to explicate convincingly why AI should beryllium bully astatine uncovering counterexamples.

A much wide statement against the conception that location is thing astir existential statements that is peculiarly suited to AI is that the request to found existential statements pervades almost each of mathematical research, sloppy of the quality of the header consequence being aimed for. For example, if I want to beryllium a connection by induction, I whitethorn good look for a strengthening of the connection that serves amended arsenic an inductive hypothesis. Or if I want to beryllium that each entity of type T pinch spot P besides has spot Q, past I whitethorn good look for a spot R that follows from P and tin beryllium utilized to beryllium R. These are much metamathematical beingness problems, but the favoritism tin beryllium somewhat blurred, and much importantly, erstwhile trying to beryllium a connection S, it is often the lawsuit that the main mobility successful our minds is less, “Why is S true?” and more, “What could a impervious of S beryllium like?” To springiness an example, I consciousness I understand beautiful good why Goldbach’s conjecture is existent — a highly plausible probabilistic exemplary of the primes implies it and agrees intimately pinch computational information — but if I were making a superior effort to beryllium it, that understanding, which galore mathematicians person had for a period aliases so, would beryllium of constricted help. Rather, my main task would beryllium to effort to find impervious techniques that were powerful capable to make these heuristic ideas rigorous.

What is the quality betwixt an illustration and a counterexample?

Logically, each connection of the shape \exists x\ P(x) is simply a counterexample to the universally quantified connection \forall x\ \neg P(x). However, we do not picture each existential statements arsenic counterexamples. For example, if I were to say, “The \ell_p-spaces pinch 1\leq p<\infty are each separable, arsenic is c_0, but \ell_\infty is not separable,” I would not picture the 2nd portion of that assertion arsenic a counterexample to the declare that each Banach spaces are separable. Rather, I would coming it arsenic astir apt the astir basal illustration of a non-separable space. The important constituent seems to beryllium that location was nary peculiar logic to deliberation that each Banach spaces would beryllium separable, and uncovering an illustration of a non-separable abstraction is not very difficult.

I deliberation the first constituent is much important here: we are much inclined to telephone an entity a counterexample if the beingness of that entity disproves a connection that we had rather bully logic to believe. It often happens that aft repeated unsuccessful attempts to beryllium a statement, mathematicians statesman to consciousness that it has nary peculiar logic to beryllium true, moreover if it seems to beryllium difficult to travel up pinch a counterexample to it. In specified a situation, if a counterexample is yet found, it whitethorn person mislaid thing of its “counter” feel. My belief is that the building of a non-sofic group comes into this category. There person been respective proposals successful the lit for really 1 mightiness conception specified a group, and I don’t deliberation location were galore (or moreover any?) experts who powerfully believed that each groups were sofic. So it feels much earthy to say, “OpenAI came up pinch the first illustration of a non-sofic group” than to say, “OpenAI recovered a counterexample to the soficity conjecture” (despite the truth that that conception of their insubstantial is entitled “A counterexample to the soficity conjecture”).

Likewise, it seems to maine that the caller little bound for multicolour Ramsey numbers is much of an illustration than a counterexample. I deliberation rather a batch of group believed that the bound should beryllium exponential, truthful for them it was a counterexample, but others, myself included, were much neutral astir it. As a matter of fact, I person worked connected the problem successful the past (a agelong clip ago) successful an balanced formulation, which asks really galore triangle-free graphs connected n vertices you request if you want their national to beryllium the complete chart K_n. If you return bipartite graphs, past it’s easy to spot that you request \log_2n of them, but that bound tin beryllium improved if alternatively you observe that a complete 5-partite chart tin beryllium written arsenic a national of 2 triangle-free subgraphs, and truthful it is imaginable to constitute the complete chart arsenic a national of 2\log_5n triangle-free graphs. It is past tempting to effort to do better, pinch triangle-free graphs that are little dense but that dress up for it pinch unbounded chromatic number — a basal information if 1 wishes to usage a sublogarithmic number of graphs, which is balanced to showing a superexponential little bound for R(3,3,\dots,3). All this is to opportunity that erstwhile I worked connected the problem, my efforts were concentrated connected what turned retired to beryllium the correct direction, truthful for maine OpenAI recovered an illustration of what I (weakly) expected, alternatively than a counterexample.

Where does this time off us?

I would for illustration to find a coherent mentation of the conjunction of the pursuing facts.

  1. The astir notable mathematical results proved by LLMs person tended to beryllium ones that we would categorize arsenic examples aliases counterexamples, wherever counterexamples are, broadly speaking, beingness statements that disprove statements that we expected to beryllium true.
  2. Many statements tin beryllium formulated arsenic beingness statements erstwhile we would usually deliberation of them arsenic cosmopolitan statements, and vice versa, truthful what we see to beryllium an illustration depends connected the mathematical discourse of a connection arsenic good arsenic its logical form.
  3. LLMs are beautiful bully astatine proving cosmopolitan statements arsenic well: it’s conscionable that the strongest statements they person proved that we would deliberation of arsenic theorems person chiefly not been astatine the level of the strongest statements that we would deliberation of arsenic counterexamples.

Given these facts, it seems apt that what LLMs are bully astatine is thing else, which happens to person arsenic a consequence that they are bully astatine the benignant of beingness problem that we would usually categorize arsenic asking to find a non-trivial example.

Let america see 2 things that we tin beryllium assured that LLMs are bully at. One of them is knowing a batch of mathematics: if a problem tin beryllium solved by intends of a comparatively modular argument, it is highly apt that an LLM will beryllium capable to find and usage that argument. The different is the expertise that an LLM has simply by kindness of being a computer: it tin activity astatine immense velocity (compared pinch humans astatine least) and tin truthful spend to make a ample number of unsuccessful attempts astatine a problem earlier it finds a solution.

Without moreover looking astatine what LLMs person really managed to solve, 1 mightiness conjecture that these 2 features would lead to their having a somewhat different style from quality mathematicians. Very roughly, LLMs would person the separator erstwhile location is much of a probabilistic constituent to the proof-finding process: they would beryllium bully astatine problems for which the champion method is to effort a batch of ideas, not needfully peculiarly novel, until astatine immoderate constituent you get lucky. Humans connected the different manus would beryllium amended (for the moment) astatine uncovering much “surprising” and “conceptual” arguments, wherever the due method is to excavation deeper and deeper into a problem until the solution reveals itself. (It is difficult to opportunity precisely what this means, but I dream that immoderate knowledgeable interrogator reference this will cognize what I americium talking about.)

This raises 2 questions: does the conjecture supra correspond astatine each to the reality that we are observing, and is location immoderate logic to suppose that what I person tentatively described arsenic the “LLM style” of doing mathematics would lead people to LLMs discovering respective counterexamples (or conscionable examples) to long-standing conjectures, moreover if that was by nary intends each they could do?

I don’t dress to person a technological reply to either question, but the reactions of experts to respective of the singular solutions that ChatGPT has recovered do lend immoderate support to the thought that LLMs activity successful much of a try-lots-of-things-till-you-get-lucky way. People often look to respond by saying thing like, “Initially I was amazed that the problem had been solved, but connected person inspection I realized that the attack was really not each that novel, and 1 that pinch the correct mini hint a suitably master quality could person recovered rather easily.”

For the 2nd mobility — whether the LLM style is good suited to uncovering (counter)examples — I deliberation matters are little clear, because location are galore ways of searching for a counterexample, and immoderate of them fresh amended than others the style I person described. Here are a fewer wide methods. (I don’t declare that the database is exhaustive.)

  1. Look for an off-the-shelf example. Here 1 has a banal of reasonably modular examples and 1 simply tries them retired 1 aft different to spot whether immoderate of them fails to fulfill the fixed statement. For example, Ryan O’Donnell ends his awesome book connected the study of Boolean functions pinch immoderate tips, 1 of which is, “If you person a conjecture astir Boolean functions, trial it connected dictators, majority, parity, tribes (and possibly recursive mostly of 3). If it’s existent for these functions, it’s astir apt true.”
  2. Build an illustration from basal examples and modular building methods. For an algebraic problem, for instance, 1 mightiness commencement pinch immoderate modular examples, but past return products aliases quotients aliases limits.
  3. Make dense usage of metavariables. The connection “metavariable” comes from machine science, and successful peculiar from automatic theorem proving, and refers to the believe that successful mathematics would correspond to writing, “where x is to beryllium chosen later,” (in which lawsuit x is the metavariable). In a insubstantial we usually do this only successful reasonably elemental situations specified arsenic erstwhile we request to take a number \epsilon>0 that is mini capable for later arguments to work. But erstwhile we hunt for an illustration of an entity x that satisfies immoderate spot Q (which whitethorn good beryllium a conjunction of simpler properties Q_1,\dots,Q_k), it is often not a bully strategy to specify x wholly and only past to cheque whether it satisfies Q. Instead, it tin beryllium much fruitful to do almost the opposite: we commencement by saying virtually thing astir x and simply motorboat into proving that it satisfies Q. In the people of doing so, we find that we request x to fulfill a spot P_1. If we are fortunate we tin picture successful a bully measurement a very wide people of objects x that fulfill P_1. For instance, we whitethorn beryllium capable to find a parametrized class: we place immoderate usability f and show that f(y) satisfies P_1 for each y of a definite type. The problem is past reduced to uncovering y specified that $Q(f(y))$ holds, which is simply a much circumstantial type of the original problem. There whitethorn beryllium galore iterations of this process, aliases a substance of this process and different processes, earlier an illustration is yet found.
  4. Try to beryllium the opposite. If 1 wishes to find x specified that Q(x), it tin beryllium amazingly adjuvant to commencement by attempting to beryllium the connection \forall x\ \neg Q(x). The logic this tin beryllium adjuvant is that utilizing our modular methods of attempting to beryllium something, we whitethorn extremity up identifying a cardinal lemma that would suffice: that is, we whitethorn find an intermediate spot R that implies \neg Q successful a non-trivial measurement and frankincense trim the problem \forall x\ \neg Q(x) to \forall x\ R(x). Turning things information again, it whitethorn good past beryllium that uncovering a counterexample to R is easier than uncovering a counterexample to \neg Q (that is, an illustration that satisfies Q). Of course, location is nary guarantee that a counterexample to R will beryllium an illustration of Q, but sometimes we are fortunate and it is. More often, we tin usage the thought of the erstwhile method, noting that it is astatine slightest a basal information of an illustration of Q that it should not beryllium an illustration of R, truthful 1 tin effort to picture a wide people of objects that neglect R and successful that measurement trim the problem.
  5. Successive approximation. Sometimes, erstwhile we are searching for an illustration of x specified that Q(x), we constitute down a moderately plausible conjecture x_0 not because we deliberation it has a chance of moving (if we did, past we would beryllium utilizing the first strategy), but because we dream that if x_0 does not fulfill Q, past we will beryllium capable to diagnose what went incorrect and specify a caller conjecture x_1 that does not person that defect. Again, this strategy tin either beryllium iterated aliases mixed pinch 1 aliases much of the different strategies.
  6. Just-do-it proofs. Sometimes we request x to fulfill infinitely galore properties Q_1,Q_2,\dots, each of which is, individually, rather easy to satisfy. In specified situations, we often “build” x inductively spot by bit, ensuring astatine the ith shape of the process that nevertheless the building process continues, x will fulfill Q_i.
  7. Pick a random example. Often it is very difficult to springiness an definitive illustration of an x that satisfies Q, but location is simply a earthy probability distribution for which 1 tin show that if 1 chooses x randomly from that distribution, past pinch precocious probability (or astatine slightest non-zero probability) it will fulfill Q.
  8. Pick a generic example. In much infinite contexts, it whitethorn again beryllium rather difficult to springiness an definitive illustration of an x that satisfies Q, but 1 whitethorn beryllium capable to show that the group of x that neglect Q is aliases measurement zero, aliases is simply a meagre set, aliases is mini successful immoderate different way.

There is nary peculiar logic to suppose that LLMs would beryllium arsenic bully astatine each of the methods above. So possibly what we are watching is not rather that LLMs person a peculiar expertise to find examples, but much that they are peculiarly bully astatine uncovering examples (and proofs) successful a definite way. Looking astatine the supra techniques, 1 mightiness ideate that they would beryllium very good suited to checking off-the-shelf examples, uncovering just-do-it proofs (since that is simply a alternatively modular method pinch tons of instances successful their training data), utilizing the probabilistic method (unless, arsenic often happens, important caller ideas are needed to show that the probabilities activity out), and picking generic examples. The different 3 methods described supra — usage of metavariables, trying to beryllium the opposite, and utilizing successive approximation — require much of an expertise to judge whether the attack 1 is taking is apt to beryllium fruitful. Here it seems astatine slightest imaginable that humans will sometimes person an advantage, but the conditions that a problem would request to fulfill are rather stringent. One would request an illustration to beryllium 1 that lies astatine a leafage of a very ample hunt character — excessively ample to beryllium searched for by a operation of mean mathematical expertise and brute unit — but that tin beryllium recovered by a mathematician pinch a sufficiently bully chemoreceptor for erstwhile they are making advancement that they tin prune the hunt character very substantially.

Why wouldn’t LLMs besides person that “nose”? I don’t norm retired that “nose” is an emergent spot of the measurement LLMs are trained, and that wrong a twelvemonth aliases 2 they will person it to the aforesaid grade that we person it. But for now, successful my interactions pinch ChatGPT, I do person a chopped belief that they haven’t sewage location rather yet. When I talk an unfastened problem pinch 5.6 Pro, I americium often presented pinch approaches that sound promising until I deliberation astir them carefully, and past look rather a batch little promising. And they will besides often extremity a consequence by saying, “I person not managed to reply the mobility you asked, but person managed to trim it to the pursuing overmuch narrower and much precise question,” which sounds very promising until it has happened 5 times without immoderate evident advancement having been made. It isn’t wholly evident really they will get amended astatine this, since their training information will not beryllium afloat of examples of fruitful and little fruitful directions to prosecute erstwhile trying to lick problems: each they will typically spot is tidied up proofs that hide the thought processes of their discoverers. Of course, quality mathematicians besides don’t get to study overmuch astir really to do investigation from the acquisition of different mathematicians, and yet we someway negociate to prime it up. But the business is simply a small different for us, successful that a batch of what we study is by doing alternatively than emulating.

Another logic it is not evident that “nose” is simply a spot that emerges people erstwhile LLMs are scaled up is that if LLMs make dense usage of their wide knowledge and tin spend to do a batch much brute-force hunt than humans can, past they will deficiency the inducement that humans person to prune the hunt character ruthlessly. It could conceivably beryllium that their successes truthful acold are achieved utilizing methods that for a quality would beryllium considered highly inefficient, but that because of their superior velocity and knowledge, the combinatorial detonation these methods will lead to has not yet go apparent.

It would beryllium very absorbing to effort to trial this experimentally, but it is besides difficult, because if an LLM has what looks for illustration the benignant of thought that could only beryllium the consequence of “deep thought” astir a problem, we tin ne'er beryllium judge that it has really carried retired that heavy thought, arsenic opposed to uncovering a exemplary statement already successful the literature, aliases successful different words exploiting the heavy thought of a quality mathematician. It would astir apt beryllium easier (but still not easy) to trial it by utilizing models that are little powerful than the latest ones and that person been to immoderate grade shielded from the mathematical literature: 1 could springiness them a cautiously designed suite of problems and spot whether the ones that the LLMs lick person peculiar characteristics.

It whitethorn look arsenic though I americium desperately clinging to the dream that humans will proceed to beryllium capable to make meaningful contributions to mathematical find for a while yet, but while I do so dream that, I americium not making immoderate assertions of the shape “LLMs will ne'er beryllium capable to do X”. I deliberation it is apt that they will, and fixed the gait of advancement complete the past 3 years it will astir apt hap rather soon. But I do deliberation that location whitethorn beryllium a hurdle for LLMs to clear and it seems astatine slightest imaginable that it won’t beryllium cleared arsenic straightforwardly arsenic immoderate of the erstwhile hurdles.

In that connection, it would besides beryllium absorbing to spot whether a different reward building leads to LLMs being capable to lick different kinds of problems. For example, if during training an LLM (or machine-learning strategy of immoderate different kind) is not conscionable rewarded if it ends up pinch a solution, but besides penalized if it explores excessively galore dormant ends aliases if it “cheats” by getting the reply from the literature, possibly it would beryllium incentivized to spell astir the investigation process successful a much quality measurement and thereby execute amended results for classes of problems wherever it is yet to make a large impact.

If the hurdle is cleared, either by axenic scaling up aliases by immoderate much thoughtful method, it will beryllium rather difficult to cognize erstwhile that has happened, since, arsenic conscionable mentioned, an thought that seems very original and astonishing whitethorn conscionable beryllium lurking location successful an LLM’s training data. But I would beryllium assured that it had been cleared if an LLM were to travel up pinch a impervious that was arsenic astonishing to maine arsenic the solution of the cap-set problem was successful 2016: the erstwhile champion known bounds were wholly eclipsed, the method was utterly different from thing I had thought astir trying, and afterwards location was a flurry of activity arsenic group came to understand what this awesome caller method was tin of.

Conclusion

I wasn’t rather judge wherever I would extremity up erstwhile I started this post, and now that I’ve sewage to the end, I consciousness that my main conclusions are not peculiarly caller aliases surprising, but I dream that the way to them is of immoderate interest. The main points I person made are the following.

  1. “Finding an example” is successful believe not the aforesaid point arsenic proving a connection that originates pinch an existential quantifier.
  2. If it is existent that existent models are peculiarly bully astatine uncovering examples, that is astir apt not because they person a peculiar affinity for existential statements, but much because the proof-discovery methods that are due for uncovering definite kinds of examples play to the evident strengths of LLMs: wide knowledge and the expertise to research galore paths of the hunt character that humans would judge to person a debased probability of success.
  3. It seems apt that LLMs will transportation connected improving very quickly. However, if, contrary to expectations (mine astatine least), location turns retired to beryllium immoderate residual people of problems (or different mathematical activities) for which humans proceed to person the separator for a while, it is apt that those will beryllium problems for which the mysterious quality expertise to prune the proof-discovery hunt character is peculiarly advantageous: that is to say, problems wherever the hunt character is heavy and has a ample magnitude of branching, truthful that without rigorous pruning a hunt is not feasible moreover for a computer.
  4. A bully motion that LLMs person reached quality level for a overmuch wider people of problems will beryllium if they commencement proving theorems utilizing methods that, for illustration overmuch of the very champion quality mathematics, are caller and astonishing but that pinch hindsight travel to look beautiful and natural. They should besides beryllium methods that are difficult to stumble connected by accident. It is difficult to opportunity precisely what would count arsenic specified a proof, but I deliberation we’ll recognise it erstwhile we spot it.

Tags: ai, mathematics

This introduction was posted connected August 12, 2026 astatine 11:00 americium and is revenge nether AI and maths. You tin travel immoderate responses to this introduction done the RSS 2.0 feed. You tin leave a response, aliases trackback from your ain site.

More