As a mathematician maybe I am a little more optimistic than this declaration.
I am thinking of Mochizuki's abc conjecture: He worked in relative isolation, and dumped a huge incomprehensible proof on the community (to oversimplify a bit). That's not totally unlike what might happen if AI generates a huge, incomprehensible proof of let's say RH.
Well, what is the result? In the Mochizuki case, it was a lot of skepticism, but it also generated conferences, papers, talks in the hallway, discussions with students, and so on--a flurry of exactly that kind of community process that the declaration says is the main driver of mathematics.
Ultimately we think a fatal flaw was found in Mochizuki's proof, so it didn't lead anywhere in particular. But in our hypothetical "AI lean-verified proof of RH" situation, it would presumably generate substantially more of that community activity we saw in the Mochizuki situation. And if it's correct, that community activity would be productive (expository talks, students given problems to flesh out or generalize, etc).
Maybe mathematics just becomes a little more like other fields--relying on labs with lots of money for compute, digging through a corpus of AI-generated proofs, etc.
It’s frustrating that this comment is at the top because it, along with lots of the replies it inspired, absolutely misrepresents the actual declaration. The declaration is not making any statements about not using any AI in mathematics. The entire point is to push the use of the technology in a direction which is compatible with positive pre-existing features of the math community, and to make it better known what some of the current problems are.
It includes Terrance Tao who has made the front page several dozen times at this point for his usage of AI such as to help the community write proofs for Erdos problems.
But to many commentors he's a now gatekeeping AI-hating Luddite clinging to a dying profession out of bitterness and envy because his position is more nuanced than "throw AI at everything and turn off your brain".
I certainly didn't intend to imply that the declaration is against using AI.
My point is just that I feel a little optimistic that the human culture around math ("ideas we disseminate in talks, private discussions and careful writeups, connecting them to the previous ideas of others" and so on, to quote the declaration) is robust even against relatively irresponsible use of AI (i.e., an onslaught of proof-slop).
And I guess we'd better try to be optimistic, because even if the declaration results in some realignment between the community and big AI companies, the models capable of this work are not always going to be exclusively under the control of those aligned parties.
That said, I think the declaration is great and I support it--let's see what comes out of it.
> the models capable of this work are not always going to be exclusively under the control of those aligned parties.
Yes, the trailing edge will catch up quickly.
I wonder how long until P vs NP falls.
The maths community is now in the antithesis phase, synthesis will take a while ;)
Lee Sedol said in an interview that "losing to AI, in a sense, meant my entire world was collapsing. ... I could no longer enjoy the game. So I retired", and I think there will be folks in the mathematical community who would feel the same when the solutions pages to hard problems are suddenly available.
But on the other hand, people learned a lot from chess engines. After decades of chess computers beating humans, there was still a renewed interest in watching Leela beat Stockfish, with many people trying to understand the strategy Leela used.
If your happiness comes from grinding on a problem and making progress, the prospect of having to dig through a corpus of AI-generated proofs might be hard to swallow. But if you're willing to do that, you will still find beautiful things that only so many people can truly appreciate.
I really hate this overly condescending takes. First of all, what do you know about the internals of math research that allows you to speak with so much confidence. Second, you're not even addressing the issues raised by the letter! This is not about "oh they made a bunch of problems easier". There are huge economical interest behind: who owns and has access to models? are these companies interested in developing research or they just grind PR stunts without worrying about externalities in how research is actually conducted? Etc etc.
If it's worth anything: I have a PhD in (theoretical) mathematics and I entirely stand by stabbles' comment.
There is a real, undeniable possibility of AI becoming better at mathematics in the same way that it became better at chess and Go, and in such a scenario, one may expect the community's response to be comparable.
It makes no sense to compare mathematics with chess. Chess is a sport. No one is interested in watching two machines compete. Chess doesn't have a practical impact. Etc. What you seem to suggest is that AI will be able to completely (or at least in a great part) replace mathematicians. It could be the case in the future, but no one knows right now, and more importantly: tech companies don't even think about it! they don't think on the externalities.
Does mathematics still have a practical impact without humans in the loop? I don't think there's one single answer to that question, but I think it's worth considering exactly what that impact may be.
Tech companies are as much the topic of this post as AI, I think that's the immediacy.
To a significant extent, the pursuit of mathematics research is a pursuit of human understanding of mathematics, without knowing where it might lead, or whether it might lead anywhere at all. I don't see how the motivation for that goes away on its own, but the institution supporting it is certainly threatened by the potential loss of grant money and graduate student applications.
> Does mathematics still have a practical impact without humans in the loop?
There's a single answer to that question: yes, math very much has an impact without humans in the loop.
Math has a lot of applications, and those applications don't care whether eg the new faster matrix multiplication algorithm was found and proven correct by a machine or a meatbag.
Well, what you are saying that the situation is even better for math than for chess?
Chess is only valuable as an entertainment. So no one really gains anything from computers becoming really good at chess.
But with math, everyone in the world would gain from computers becoming really, really good at it.
But solving math problems is not the same as increasing understanding of math concepts.
> No one is interested in watching two machines compete
I’ve watched quite a lot of YouTube videos where two machines compete, so you may not be completely right here
The Square One commentary on the AlphaZero v Stockfish game from 2017 is pretty entertaining: https://www.youtube.com/watch?v=LnVDUQksIDk
In other videos he's called out the influence that this and similar games have had on human players in recent years, particularly around square denial and thorn pawn strategies.
https://tcec-chess.com/
Top chess engine championship is pretty fun to watch.
> What you seem to suggest is that AI will be able to completely (or at least in a great part) replace mathematicians.
There is no bound on the amibitions of AI. AI is set to replace anything done by people, and there won't be any room left for people. There isn't any task done by humans that AI won't be better at.
This is not a tenable outcome.
We should never have built machines with agency, rather than optimization processes that operate as subroutines of humans.
You say it like it's a bad thing.
Replacing people leaves no room for people.
> We should never have built machines with agency...
We haven't quite crossed that bridge, but we do appear to be standing on it.
Stockfish isn't owned by a club of three trillionaires. It does not cost $15 million to achieve a result in Stockfish.
Stockfish does not steal research or scoop researchers.
The concentration of computing resources and capital should be examined by the math community.
The trailing edge of AI is catching up fast. There's plenty of open source (and even more open weight) AI models and they are getting better and better.
If that's your only objection: in a few years you can prove Rieman's hypothesis on your smartphone, no need for any trillionaires to give you permission. Does that make any change to your argument, or did it not actually matter?
This is the real problem. We're looking at a future where those who control AI have an insurmountable advantage in everything. They can control the amount of intelligence the masses have access to -- for their own safety, of course -- and they will never, ever be able to close the gap.
There's lots of competition in AI. Where do you see the 'insurmountable advantage in everything'?
Say they do. What then? Will people buy from them? What if we decide not to?
At a certain point, you don't have a choice. Before China got into the game, the only way to avoid giving Luxotica money if you wanted a pair of glasses was to essentially not buy glasses. This is the same for many industries -- consolidation behind the scenes.
Huh? I've been able to buy reasonably priced glasses for all my life. (However, I've never lived in the US nor China.)
Who cares about the giant labs? The trailing edge will catch up quickly, and in a year or three you can prove Rieman's hypothesis on your smartphone.
>> "are these companies interested in developing research" judging from the money, resources spent and the value they derive from this the answer is very definitively yes.
What makes you think these companies (and I'm not a fan of all their motives) are not interested in developing research? The motives may be self-serving, but it is undoubtedly and objectively accelerating research.
Chess is kept afloat by a couple of billionaires like Sinquefield, MBS and the guy who sponsors freestyle (Fisher random) chess.
Carlsen is bored by studying engine lines.
The popularity is boosted by YouTubers because chess is very suitable for somewhat higher class content.
I'm not sure we'd want that world for math. Positions will be cut just like archaeologist positions are cut now.
Chess is kept afloat by chess players, not by billionaires. If all the billionaire backers stopped sponsoring tournaments, people like me would still play, still pay for chess club memberships, still pay entry fees for tournaments, and still buy chess books, and so on.
I think the parent comment meant professional, high-level chess. The kind people get played to play, not just do for a hobby. That's absolutely on life support.
I'm not sure what the equivalent would look like in the math field, but it probably involves a lot of mathematicians losing their jobs and the quality of human-produced math decreasing overall.
The quality of the math in general would be fine, since in this scenario cpus will keep producing it. The quality of cpu-cpu chess games is quite high, beyond human understanding in many cases.
Chess is a weird example because it doesn't really have any utility beyond itself. Even pure math sometimes ends up having use in the strangest places. Although if no one understands the frontier math (because no one is getting paid to), I'm not sure it even matters what the quality of the cpu math is?
It's a bit like a tree falling in a forest. If an LLM proves a theorem but no one understands it, did it make a sound?
"Although if no one understands the frontier math (because no one is getting paid to), I'm not sure it even matters what the quality of the cpu math is?"
Presumably AI will be connect the dots to the applications. As the declaration says, this isn't just about math. Human understanding is losing economic value. You can understand stuff on your own time, I guess.
The standard justification for pure math to holders of purse-strings is something like "it might lead to a useful application down the road, like crypto, who knows". That looks pretty inefficient now. We have to entertain the possibility that AI can develop the math needed for any application we put to it. Eg if number theory didn't exist, we could have asked AI for a way to transit messages securely and it would maybe come up with fermats little theorem as part of its solution or maybe come up with an approach we can't conceive of right now seeing as most of us are constrained to available number theory. Like how in the last year when I give an LLM a programming project I see it doesnt even bother with of the many software libraries I and others have written and just codes up the calls it needs on the fly or finds some other ad hoc solution.
When there's a billion people playing something, money will never be an issue for those at the top. Even things like chess.com was able to sponsor a tournament with a million dollar prize pool.
Also I'd argue that chess's utility is ultimately the same as pure math, particularly in esoteric fields. These things are highly unlikely to ever lead to any sort of real world breakthrough or application. The main benefit is an outlet for human logic, creativity, and exploration - which significant self improvement possible along the journey for players.
Though I think even that's probably too socially utilitarian. I think ultimately the 'real' drive is the same in both fields - it's fun and personally rewarding.
> Though I think even that's probably too socially utilitarian. I think ultimately the 'real' drive is the same in both fields - it's fun and personally rewarding.
That's fine, and no amount of machine excellence will keep you from enjoying recreational chess or recreational math.
I would argue people getting paid to play chess was a short lived phenomenon anyway if you put it in context. The transition there is less related to the introduction of chess engines and more related to the shift in the media landscape.
> It's a bit like a tree falling in a forest. If an LLM proves a theorem but no one understands it, did it make a sound?
But in future most proofs will be for consumption by other AI models in the pursuit of yet other proofs.
It's kind of surprising so many mathematicians act surprised by this given this was clearly where automated proof assistants would lead. I guess they assumed they'd always be the ones guiding them.
> But in future most proofs will be for consumption by other AI models in the pursuit of yet other proofs.
What is the purpose of that?
Its like art being produced for AI to consume. What is gained from that?
It's going to be hard to compete with something that has access to all of math at once and can find connections between elements that appear unrelated to humans.
And at some point AI will start suggesting - or doing - physical experiments.
Presumably some of the proofs will have applications beneficial to humans beyond impressing other mathematicians, and AI will surface them, or use them directly.
Well if it proves useless presumably they'd stop doing it.
But if AI is to recursively self improve understanding and evolving its own foundations, which are clearly mathematical, is essential. There is no need for humans to grasp what is going on in that loop.
Math also has applications, and they don't rely on human mathematicians doing the math.
I mean, was the point of math ever just because some humans enjoyed doing it? Even though a lot of it is theoretical, there's been all sorts of useful things that have come out of it as well due to an improved understanding of the universe through new ways of thinking about it. If it got to the point where no human could understand it and there were no ways to actually use it, I don't think anyone would bother having their computers doing it at all.
> was the point of math ever just because some humans enjoyed doing it?
Yes.
Friends and I often work on Putnam problems and this series:
The (Almost) Impossible Integrals, Sums, and Series by Cornel Ioan Vălean
wouldn't AI solving problems in science, engineering, economics, etc be able to apply the new AI math?
[dead]
Mathematics is more than establishing arbitrary facts (although some look like curiosities), it's also defining what interesting research directions are and establishing common language/notation. I think that will stay relevant?
In theory you can automate finding interesting research directions by identifying conjectures with many dependencies. And notation has never been mathematicians' forte, with them trying to cram the entirety of universe into single letters.
Why? How do you define interesting research directions? That used to be defined by testing the limits of human understanding i.e. some people can't figure something out. AI might have very different ideas about what is interesting and I am not sure what humans would get out of putting years into understanding AI proofs for what? What are we doing at that point? Like if you spend years understanding some AI proof of theorem 123456, why is that meaningful? I am actually asking why you think defining interesting research directions will stay relevant. In my opinion, people spend years acquiring knowledge so they can work on problems which is separate.
There's two points about this I am assuming 1) Mathematics actually has a significant subjectivity to it and is community oriented and not just climbing a never ending list of theorems that exists in the universe 2) A lot of mathematical research work is inside of a subfield and isn't directly motivated by applications. Sometimes it is but e.g. people don't work on obscure theorems about elliptic curves because of a dire need for that but more because the community found it interesting.
People become interested in things when they become invested in it personally, because they've contributed to it. So I don't think it will stay relevant...
>but no one understands it, did it make a sound?
Does your "one" only contain humans or does it also contain other AI systems. AI math is not a single monolithic thing, but a distributed one. I see value in sharing proofs even among just AI.
[dead]
Also, you learn to be a better chess player by... playing better players. The widespread availability of chess engines has made flawless opponents available to every player.
If your goals are understanding the game, self improvement, building thinking skills-- this is the best chess has ever been. It's only if your goal is to beat every opponent you can find that chess is in a bad place.
It's not that easy. Playing stockfish is like playing tennis against the wall (for untitled players at least).
Even if you don't blunder anything, you'll still find yourself in a worse position without any clue as of what went wrong and why.
Whereas when playing humans, they can usually explain their approach and when they noticed errors in your play.
I mean, I'm not a great player but I've learned a lot from working through games with stockfish using a git repo and a small script that lets me rewind to different moves and try different approaches. It may not explain its moves but if you're thinking through what's happened you can usually debug your game anyway.
> Playing stockfish is like playing tennis against the wall (for untitled players at least).
It's the same for Magnus Carlsen. Even with Queen odds, Stockfish is literally unbeatable for the best players in the world. It's just too strong at evaluating all kinds of random tangent moves (and ensuing positional advantage) which no human player can possibly pay attention due to the time required.
Stockfish vs any human is like Carlsen vs other players by about 3-5 orders of magnitude[0]. It's that stark.
[0] A wild pun appears.
EDIT: To avoid having to respond to each responder, fair comments about Queen odds. Maybe I was thinking Rook odds? Also, I kinda lumped Stockfish in with all the other engines, but I realize there are other engines with different properties ofc.
Well, that's actually not true at all. Stockfish is not a very good odds player and at queen odds is easily beatable even by bad players like me. It will just trade down into more trivial and easier to win positions that it perceives as "less bad", since everything is super-losing anyway when you start down a queen.
Leela odds networks, on the other hand, are an entirely different beast. I cannot beat Leela queen odds, much less rook or minor piece odds, and even GMs struggle against Leela knight odds.
Without odds though, yeah, Stockfish is just incomprehensibly strong by human standards. All top chess engines are, but Stockfish moreso.
No worries, I know stockfish is unbeatable by humans.
But sometimes, these GMs can flag it, which counts as a win (especially when it's proxied by a cheater). Sometimes they can also explain the idea that cost them the game, so they've learned something maybe.
Whereas us scrubs literally cannot do anything at all for reasons completely beyond our understanding.
> Whereas us scrubs literally cannot do anything at all for reasons completely beyond our understanding
Computer moves are typically much more concrete than human moves: a human will play based on pattern matching ("intuition") and can only make explicit calculation of a small fraction of possibilities, after which decisions are guided by guesswork. The computers are unbeatable in practice because they can calculate concretely in seconds what might take an expert human long intensive study to notice, and they don't make the same kinds of oversights humans can make.
But if you stop and explore a particular position for an extended time, and if you have an intermediate level of chess skill, you too can probably often (usually?) figure out why it's doing something. Sometimes understanding the computer's reasons takes searching multiple branches of a tree several unlikely looking moves deep, but the collection of threats the computer was preemptively thwarting, traps it was setting, etc. are comprehensible to humans with enough effort, especially in games between the computer and a human.
The frustrating thing about playing against the computer is that it notices and thwarts every plan you might come up with, before you make up the plan yourself, and it doesn't make (human-apparent) mistakes, so the game ends up feeling hopeless. Nothing you try works on it, and if your idea is even slightly inaccurate it will be exploited.
You put it better than I could. The lesson of "in this exact position you can kick the pieces for 7 moves to get a fork, so instead you should play a4" is not something that I can implement into my games
I mean, this is the problem with analogies and trying to use them to prove things, right? People working through problems from an analysis book with their friend (or an LLM) is not the same as research mathematics. People playing in a chess club is not the same as what makes for a good chess tournament. Lumping everything together is just making this branch of the conversation less relevant.
Chess is a fun game. That's why it's been around for 1000+ years.
There was a renaissance during Covid and due to 'The Queen's Gambit' where it gained much more mainstream popularity, but... Chess AI was already far far (like 1000+ Elo) ahead of human players at that point.
The thing is... chess is humans playing (communicating) with humans and that's what keeps it interesting. Check out the view counts of chess AI tourneys vs. human tourneys.
[flagged]
Is this the scenario described in Ted Chiang's short story https://en.wikipedia.org/wiki/The_Evolution_of_Human_Science where scientists are "catching crumbs from the table" trying to decipher the results generated by superhuman intelligence?
It's still an optimistic scenario. Artificial superintelligence may develop hypermathematics of a kind that never will be accesible to human mind, enhanced or not. One can't teach geometry to ants even if you put them on a Moebius strip.
It would be more like Lem's novel where it completely disappears from the human horizon: https://en.wikipedia.org/wiki/Golem_XIV
Which is why if humanity had empathy, it would be working on how to make smarter ants, so that they can learn more advanced geometry.
I love that as a goal!
Perhaps indeed a better understanding of what intelligence really is would allow this sort of Uplift (as in Brin's books).
A fellow Brin aficionado!
Sounds like a great use of AI.
It's quite possible ants understand a geometry more advanced than our own.
What do you think we're doing?!
Ant Intelligence will be next big thing in machine learning after this bubble bursts
What’s optimistic or non-optimistic specifically about the machine having a system of mathematics beyond our comprehension within it? Why should we care about that in itself?
In the first case, we'd still have a chance to take a glance at the frontier of discovery (even if ordinary human mathematicians had to spent years translating what metahumans achieved).
In the second case all human-level maths would be solved and what lies beyond would be always out of our scope.
Why would this be optimistic? This sounds extremely negative to me…
But the work the Mochizuki case generated can also be done by AI. AI could generate a landmark proof and then people could use it to solve or simplify intermediate problems and you could use a different AI prompt to try to disprove it if you were really skeptical. From my memory I think they said it took 88 hours to solve a Millenium Problem versus the decades of time humans have put into it.
I don't like nuance here. I think progress is really measured by what humans are able to do and understand, not machines. It is significant if we find problems we struggle to solve. That tells us something. What does it take for humans to solve these problems is related.
The best analogy I can give is if you wanted to climb Mt. Everest you might ask someone for guidance. Would it be better to ask someone who has climbed Mt. Everest or someone who took a helicopter ride up near the top and then went to the peak? This is like the AI versus human gap to me. The helicopter is like using AI to generate a proof. The person who actually climbed Mt. Everest has firsthand knowledge of the experience. Same thing for a difficult proof. The struggle people have is actually valuable here. Likewise, we know people are actually capable of climbing Mt. Everest but if they had only ever rode a helicopter to the top, the knowledge of climbing it would not exist, and surely that is meaningful knowledge given the risks.
So if we rely on AI for proofs I think we lose a sense of what is difficult and why. We lose a sense of what human achievement is. Surely climbing Mt. Everest means more than taking a helicopter up? For students, why bother grinding through all the material of climbing Mt. Everest and then attempting it if the helicopter ride is how things are done now? This would have the affect of destroying knowledge.
(please do not nitpick the analogy because it's the best but perhaps a clumsy way to describe my thoughts)
Tangent:
> From my memory I think they said it took 88 hours to solve a Millenium Problem versus the decades of time humans have put into it.
Keep in mind those ~88 hours were spread across ~10,000 simultaneous agent instances.
So, roughly 880,000 hours of compute.
Assuming a fifty-year career, and forty-hour workweeks, a human mathematician's career is about 100,000 hours of "compute".
I suspect that with six good mathematicians spending their whole careers primarily focused on it, and working together closely, Navier-Stokes might well have fallen already.
The perverse incentives of academia mean this has never occurred.
The perverse incentives of industry mean OpenAI intentionally scooped researchers who were getting close (granted, with AI help).
I'm not trying to dismiss the achievement - if the proof turns out to be solid, it's quite impressive (though much less so if the training data included the recent human breakthrough, which seems pretty plausible).
I'm just pointing out that "88 hours" is a very misleading way of framing this.
> I suspect that with six good mathematicians spending their whole careers primarily focused on it, and working together closely, Navier-Stokes might well have fallen already.
> The perverse incentives of academia mean this has never occurred.
This. Mathematicians in their most energetic years are trying to get tenure or land a tenure-track job. They are disincentivized to go all-in on ultra high risk, high-reward problems. The potential downside is just too forbidding. It's much safer to develop a research program in a mainstream field that affords many opportunities for partial progress that can translate to a robust publication record.
ok I realize this is a tangent but you're saying my post is very misleading and then also saying that a human mathematician's career is about 100,000 hours of compute and that Navier-Stokes could've had a solution by now if not for perverse incentives. You may be right but I don't think this is a great argument because in a year I would bet that those numbers change since computing power tends to increase or get cheaper over time. So I am taking the stance AI can outdo people if not now, perhaps soon.
> Would it be better to ask someone who has climbed Mt. Everest or someone who took a helicopter ride up near the top and then went to the peak?
Depends on if I want to go by helicopter myself.
I agree entirely with what you're saying, right up until your final question:
> why bother grinding through all the material of climbing Mt. Everest and then attempting it if the helicopter ride is how things are done now?
I think you answered this yourself earlier:
> I think progress is really measured by what humans are able to do and understand
People want to make this progress. Therefore people will "grind Everest" as a mathematical community, and that is maybe not so hugely different from a lot of previous mathematical work.
There's still ample room for creativity: simplifying, generalizing, asking new questions humans are interested in, ...
That grind is emotional. And it's something AI will never have. The desire to solve a problem.
Citation needed. Who are you to say large enough clusters of neurons can't develo emotions?
I think progress is really measured by what humans are able to do and understand, not machines.
Building a machine that solves Millennium problems is pretty cool too. You wouldn't know it from reading these stories, though.
>Maybe mathematics just becomes a little more like other fields--relying on labs with lots of money for compute, digging through a corpus of AI-generated proofs, etc.
Dr. Tao said the same thing. Somehow, this letter came through. He wants to conduct Math competitions where participants who don’t have formal credentials can contribute to mathematical research through AI.
Title: Terence Tao - SAIR Competitions and the Future of Experimental Mathematics
https://www.youtube.com/watch?v=rB9YOi3lb7w
and this:
Daniel Litt - Working with LLMs to do high quality math
https://www.youtube.com/watch?v=0wL8NlhxXcU
So he got exuberant because he is funded by SAIR and the "AI for math" fund.
And embarrassingly they used him for a "coal miners should learn math" moment that just benefits the AI industry.
He has severely reversed course in the past week. Without concrete propositions it remains to be seen how much of the new resistance is for show.
I have zero formal math training beyond my Grade 12 Pre-Calculus class. Yet with an LLM I have recently devised an architecture with incredible math potential. Math is a language like any other, and without LLM's I never would have developed the techniques that I have.
AI is a tool. It speaks languages I don't (Math, Science, Code). I would love to participate in a Math competition without a hint of any formal advanced math training because my experience so far tells me I will do well.
‘Incredible math potential’ .. Yeah right, your AI said so ?!
How could you possibly know if it has potential or not? Sounds like AI psychosis setting in.
Because every test I run with it is telling me so? The great thing about math is it can be externally validated.
> Dr. Tao said the same thing.
Apparently he has since changed his mind.
Did he say so somewhere? I don't think these ideas are contradictory. It's just an pro AI tooling but anti-slop stance.
Where is the difference?
He sees value in mathematicians using AI to carefully study mathematics, develop an understanding of both old and new things, and help others understand the new things.
He doesn't see value in scrolling through unsolved problems asking an AI to please solve them. In his view, this is a fundamental confusion about what mathematical research is for. Knocking down unsolved problems without developing the community's understanding of them is like prompting Claude to go through a Jira board, write code for all the open tickets, and then close them without merging or deploying the code.
> He doesn't see value in scrolling through unsolved problems asking an AI to please solve them.
Yet that's exactly how the field works. A new grad student is tasked with finding a suitably difficult problem from a list of unsolved problems. The sweet spot is obscure, so that fewer people are working on it, but not too obscure that no one knows about it. It works the same way in theoretical physics and theoretical Comp Sci, and I speak from insider knowledge. The rosy view of mathematicians in the media is largely a product of marketing.
The authors of the declaration agree with you that this is how the field works today. They think that fact causes AI use to produce bad results, and they want to reformulate how the field works so that AI use will produce good results instead.
The letter addresses this specifically.
Isn’t it more like it merges the code without a dev reviewing or understanding it?
Pretty close, but IMO not quite. A math proof in and of itself is useless unless either:
If you merge and deploy code, you have released a tool that can be used. If you ship a gibberish math proof, it's not useful unless someone else can understand and deploy it to some other means. Now, it's possible AI could understand and make use of the math proofs, even if we can't, which refutes some of my hair splitting :)Not necessarily. That's the best case scenario, but proofs can be intrinsically useful in and of themselves. It's just that for problems of that nature, speculative work is often done ahead of time, e.g. the body of work that already exists assuming the Riemann hypothesis is true.
No. Merged code can perform actions with effects on the world, even if a human being never saw it. Constructing a giant Lean formalization that nobody understands simply doesn't do anything.
taste
"I believe I did, Bob" lives rent free, every time someone tells another person to fuck themselves using technology.
Thank You!
lmao, thank you, i think
Please consider making more comics. You have it.
Yes. I find it really interesting to consider what the machines do and will think of as intrinsically interesting to them. Will they develop their own theories of beauty, mathematical and otherwise?
>Dr. Tao
Professor Tao.
> Ultimately we think a fatal flaw was found in Mochizuki's proof, so it didn't lead anywhere in particular. But in our hypothetical "AI lean-verified proof of RH" situation, it would presumably generate substantially more of that community activity we saw in the Mochizuki situation. And if it's correct, that community activity would be productive (expository talks, students given problems to flesh out or generalize, etc).
This also sounds like a vector for trolling the community with complex putative proofs hiding a known flaw.
Not if it's lean-verified.
"Lean-verified" is not some magical incantation that makes a supposed proof irrefutable. Even disregarding potential bugs in the kernel as others have said.
Say that AI gives you a Lean proof and says it proves Theorem X. It could just as easily give you the same proof but claim that it proves (not X). How would you know the difference?
Nothing can really be considered proven unless a human expert can read the Lean proof and determine that (X as defined in the Lean proof) corresponds to X. The proof (at least the statement of the theorem) must be intelligible to humans to have value.
It's possible people will just start taking AI at its word. Maybe AI says "Here is a Lean proof of X" and we all just shrug and go "Okay, X is proven." But that's not how it works right now for human mathematicians. Why would we apply that standard for AI?
I'm not totally sure what you mean. You can validate the proof using lean, which is what it's for--the whole magic of it is you don't have to just trust what AI says. That said, to your larger point, there are loopholes, and we'd certainly better be able to read the statement in lean, etc.
Didn't some of the recent proofs exploited a couple of blind spots of lean, and they were invalidated?
Edit: Yup. A bug report to Lean was disguised as a "Collatz" proof in a humorous way. Links below.
- https://x.com/gro_tsen/status/2082483878480977959
- https://infosec.exchange/@0xabad1dea/117002106099986943
There was a hash collision bug in the main Lean kernel that was patched, but AFAIK nothing relied on it. You'd have to know what you were doing to accidentally get there...
The incident in the links I posted exploited several bugs AFAICS, so it's a different story than a single hash collision bug, it seems.
I don't know ... do you have a reference?
Yup, found it:
https://news.ycombinator.com/item?id=49101465
Cool, thanks!
> Maybe mathematics just becomes a little more like other fields--relying on labs with lots of money for compute, digging through a corpus of AI-generated proofs, etc.
I think a better comparison is: mathematics just becomes like mining bitcoins.
I think you might have to explain that comparison a bit more to be honest. How are math proofs like bitcoins? A bitcoin has a pre-defined value, a math conjecture / proof is a bit more complicated.
A Bitcoin does not have any pre-defined value. The value of Bitcoin keeps fluctuating, and historically has risen dramatically from its initial value of 0. If this weren't the case, there would be no investment/speculation in Bitcoin because there would be no potential for any ROI.
In reality, the value of Bitcoin is determined by humans (even if indirectly, not by planning), and I think the OP’s point may have been that maths proofs can be regarded similarly. No intrinsic value, just what humans find in it.
If there's a machine that automatically turns electricity into proofs then mathematics becomes a different thing altogether.
Mochizuki's claimed proof of the abc conjecture was extremely unusual for the reason that nobody was able to extract a single useful idea from the argument. I was starting grad school when it came out, and my immediate visceral response was "if this is what number theory is going to look like in the future, then I will leave mathematics."
The current wave of AI slop mathematics might end up driving the next generation of mathematicians away from the subject for the same reason that Mochizuki would have convinced me to quit if his proof had been accepted by the community. Luckily, my professors had the taste to immediately recognize that it was garbage.
The difference is the scale. A few incomprehensible long papers per year, sure, we will study it. A flood of AI results closing research directions left and right, that will be a problem.
> closing research directions left and right
Why would research be closed in one direction? Even if AI or human says "Tried that, didn't work" or whatever, someone (or something I suppose) might very well retry it in the future, if nothing else to reproduce it didn't work, in theory at least.
AI tends to take nearly finished research directions and push it to the conclusion in one step. If deployed massively, it will pluck all the low hanging fruits causing a drought of near term promising research project. Because people who start promising research directions do not get to see it finish, over the long term fewer people will start new directions, causing the field to slowly whither.
It's not just the isolated dumping, it's the fast, isolated, possibly untraceable dumping, without long term support.
It'll basically become slop fatigue if OpenAI starts dumping out proofs faster than the community can keep up, and some turn out to be wrong, never formalize it, don't stay to support it, etc.
I wonder if they will continue to dump proofs, though? Their point has been made, the novelty will wear off, and it maybe won't be a priority use of their resources to spend however many millions on another big proof--they will move on to the next thing to show off I'm sure. At that point, the ones generating proofs will be, I hope, mathematicians (professional and otherwise) that are more interested in the results and community discussion.
(Well that's my hopeful, optimistic take, anyway.)
They aren't going to stop at one, that's for sure. They already claimed they have "made substantial progress" on another millenium problem. Let's say they bag another one (Hodge and/or BSD according to the rumors), if it looks like their internal model could solve P/NP or Riemann Hypothesis, you think they wouldn't take that chance ?
They already told the NYT that they’ve made “substantial progress” on another one of the MP Problems (most likely the Hodge Conjecture).
That said, that’s probably just because of the drama miring their most recent one. After 2 I don’t see why they’d bother anymore.
>That said, that’s probably just because of the drama miring their most recent one. After 2 I don’t see why they’d bother anymore.
P/NP and the Riemann Hypothesis are part of the milllenium problems. They will 100% keep trying to crack those regardless.
I think AI companies making a point is not the only thing at play. Discovering new maths ultimately leads to new technologies and applications. It may start theoretically but end up being of practical use in the future. Even if humans do not understand it (lose interest, too complex, or just way too many new proofs to go though) AI can use this AI derived math corpus which will help it in other fields.
So it wasted everyone's time, thousands of hours of research trying to disprove something said very loudly. What OpenAI is doing is a DoS of the scientific community: wasting your time trying to check if they're not wrong, and claiming glory in the mean time.
That's true, but the story would have unfolded differently if Mochizuki had a lean-verified proof and was correct. I guess baked into my premise is that AI is producing reliable proofs (in the long term at least).
Agreed! Although, if done by a mathematician, it's not a ~complete waste. I think the community learns something along the way.
Is there an established term for the idea of "DoS"? I've taken to calling it slop fatigue.
Denial of service is the established term. Hammering their API (reviewer committees) would be an informal one
OpenAI avoids this by formally verifying the proof.
https://github.com/openai/NavierStokesAndEuler
Even before AI we used to say if you write code that you only barely understand, then it will be to complicated to debug. (and/or maintain)
Mochizuki was still one human and it required legions of other humans to unpack and untangle to confirm that it didn't lead to anywhere in particular.
AI is now capable of constructions so complex that no human or human team can unpack. And its ability to increase that complexity is growing while our human ability is stagnant.
meta-AI analysis cannot help. We (software professionals who use AI regularly) already know that if you run into a situation where a Fable/Astra-generated analysis reaches the limits of our comprehension/complexity due to their subjectivity, throwing more AI at the problem doesn't always converge.
There are many reasons to feel optimistic about AI, and ultimately its general ability to help science and mathematics.
I see no reason to feel optimistic about the future of mathematics and AI based on the current path of frontier labs, unless the misalignment Tao is writing about can be reconciled.
> AI is now capable of constructions so complex that no human or human team can unpack.
How can we possibly know this when we haven't even seriously started on the endeavor of actively reverse engineering these AI-generated proofs? That's a proper job for human mathematicians, because the AIs themselves are demonstrably clueless about what steps in a proof are genuinely interesting and load-bearing from a human POV. This is evidence of a limitation in AIs' capabilities, not of any kind of misaligned behavior. The fact that Tao actually uses that term in his complaint is deeply disappointing.
Not to mention, there are already (pre AI) machine-generated proofs that we've pretty much agreed not to try to explain fully, like the four-color theorem which ends up with brute-force verification of 600+ cases (down from close to 2,000 when first demonstrated)
> there are already (pre AI) machine-generated proofs that we've pretty much agreed not to try to explain fully, like the four-color theorem
Algorithmic verification is a very unsatisfying answer to the problem (e.g., surely it's not just dumb luck that every single case happen to have this exact property), but that's an entirely different issue than saying that no one follows logic of the proof method itself.
For the interested; the saying I believe you are referencing in regards to writing code / debugging is from Brian Kernighan, specifically:
(from, 'The Elements of Programming Style')It's prescient.
It’s a cute statement, but it doesn’t really match reality. Programs written by humans can generally be debugged by humans.
Could AI write programs that humans can’t understand or debug? Probably, but that’s not what Kernighan was describing.
> AI is now capable of constructions so complex that no human or human team can unpack
Can you give an example of this?
The first major computer-assisted proof, of the four-color map theorem in 1976, was an example of this. It created a lot of controversy at the time. It used proof by exhaustion, i.e. essentially analyzing every possible relevant case, something that no human could do without the assistance of, at the time, a supercomputer.
I really like this take, and while I hate math I value it. Your position sounds extremely plausible and it fits with the pattern we see in the community here. Regardless if it's ai slop or not we still debate the value and attempt to understand. In the process generating new insights and ideas. Life will go on.
I get excited at the idea of a world in which advanced mathematical problems (and their solutions) become much more accessible to a much greater number of people. As a result, making mathematics much more loved at a societal level.
Imagine a world where these most complex mathematical problems are not accessible to a few hundred people, but a few hundred thousands people. ...Those original few hundred gifted mathematicians would have an even more prominent role, and their names and achievements would be known by orders of magnitude more people that they are now.
This is the hope, but I suspect the reality is that we see an ever widening gap between the fortunate and the unfortunate. We're looking at the automation and commodification of all knowledge, and the best models will be kept locked behind closed doors so that they can't be stolen. And, of course, "for our own protection".
As a laymen, I wish the same. But I also hope it doesn’t disincentivize those that dedicated themselves to the study
Math department administrators fire Terence Tao.
Based on current reward models, the frontier AI labs will burn down mathematics as an impressive display of capabilities and in doing so, will make it impossible for people that get paid to do mathematics to stay employed.
If your job is literally to publish papers, and OpenAI and Anthropic decide that making an infinite-paper-printing machine is the best thing to show how effective their tech is, then as a demo, they destroy that industry.
I wouldn't expect them to destroy that industry, imagine global squadrons of academics and mathematicians focusing their attention on LLM's, training algorithms, scaling laws, ... they're gonna try and beat the incumbent frontier AI labs, eye for an eye, tooth for a tooth
The apocalypse being triggered by frontier labs picking a fight with mathematicians was unexpected.
I've met a few Ph.D Mathematicians in Academia socially. My unfortunate experience was that they were insufferable,borderline hostile people. I tried to genuinely engage with them too. I've met one Ph.D Mathematician that left the industry whom was very enjoyable to talk to. I have a feeling that my experience was not unique and the Math world is mostly a bunch of too good for everyone on their high horse a-holes that are now being knocked down a peg. They don't like it obviously.
I'm not a fan of knocking down things that work, however I also find it hard to be against death of the gatekeeping old guard of any industry.
I think math is just gonna have to suck it up like every other industry now. Math productivity is longer out of reach of the average grad student. Like every other industry they are no longer untouchable and are gonna have to adjust to the new way of things or market forces will do what they always do which is refuse to fund ineffectiveness.
I've had to accept that tech/IT will never be the same. Just how it is. You can thrash against it all you want.