Lest someone gets the wrong impression, I think that any mention of the a priori strange-looking resolution to the Jacobian conjecture should be accompanied with a read of pages 160–161 of https://arxiv.org/abs/2609.05746. The tl;dr being that the very same construction featuring in the resolution appeared earlier in a draft paper that was accidentally made publically available for a travel award application. The same story has accompanied several other of the major announcements made by now. We haven't have a move 37 for maths yet, despite what OpenAI's marketing department might want you to believe.

Similarly, if all you ever read were OpenAI blog posts, you would get a very wrong impression of the usefulness of large language models of today in maths. For a working researcher, it's not a magic wand that you point at any given proposition and it tells you whether that proposition is true or not. It does appear to help if, while pointing your wand and utter the magical incantation “do it up bro”, you also make it convert $15 million into heat, but for most people, this kind of inverted Midas touch isn't quite accessible yet.

Instead, the reality seems to be closer to this, projecting a fair bit: a given mathematician will have a collection of propositions that they care about, and that they'll use as their own internal benchmark as new models come out. Very rarely will anything come out of it, but sometimes, in particular if you make sure to provide the wand with all relevant context, papers that could be relevant, proof strategies and lemma structures that you suspect are useful, something (which may or may not be plagiarism) will pop out, and that's really nifty. Moreover, it is not unimportant what the proposition and the relevant proof is like. And what does come out tends to be quite bizarre; proofs that use terminology that doesn't exist, seem overly pretentious, based on nonsense analogies where it's surprising that it even works at all, and the only comfort is that you can join it with an equally unreadable Lean blob. And where you would be _crazy_ to just publish those artifacts and think that you have contributed much of anything to maths.

But sometimes it works. It's still very unclear what kind of maths the models are good at, but it seems to certainly be an advantage if what you're looking for is a counterexample hidden in a pile of otherwise similar-looking non-counterexamples, if your proof is one that requires considering 36 different cases, each of which are so tedious that no researcher would have the patience to go through them by hand, or if the proof is an amalgamation of several existing structures, some of which are only documented in Georgian.

The gold rush, more than anything else, seems to be populating the convex hull of existing maths.

This can all change. The $15 million wand requirement today will be less tomorrow. Whether we ever get a move 37 is less clear, or whether we will eventually reach stagnation as all low-hanging fruit is picked, and the convex hull is populated; call this cope if you like. But maybe we do get move 37s all over the place, and it's fine that people think about what that future will look like.

Until then, and while we're still picking friut, let us rather have a think about what we can do to fix the incentive mismatch, to ensure that we increase the prestige of digestion over being the first to convince the LLM to do it up. Since that's the one thing everyone seems to agree, chances are it'll probably converge to something that doesn't have to be written in commandment form, but out of the guest posts hosted by Tao so far, the one by Antieau has some useful suggestions for standards (that aren't entirely unlike those from Leiden): https://terrytao.wordpress.com/2026/09/15/fast-math-slow-mat...

[deleted]

> It does appear to help if, while pointing your wand and utter the magical incantation “do it up bro”, you also make it convert $15 million into heat, but for most people, this kind of inverted Midas touch isn't quite accessible yet.

That's like where chess was when Deep Blue was built by IBM. Productivity improved. There was someone complaining on here recently that the seat-back entertainment system on some airline had a chess program set to "trounce all humans".