Their claim is only indirectly related to the motivations of the people using their models. What they're saying is that doing math in this way does not produce the same value as traditional mathematical research, and the people using these AI models aren't concerned about that because their marketing objectives don't depend on whether their results produce mathematical value. If people doing valuable work are made irrelevant by people doing a larger volume of non-valuable work, that's not a positive outcome.
But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.
I actually found this to be the case with some basic linear algebra notes I was recently doing in Lean (without using mathlib). The model could generate working proofs, but they obscure the basic ideas (actually I wonder somewhat if this is because the Lean code that's out there to train on doesn't make a huge effort to read like textbook proofs, which was my motivation in the first place). I give it a skeleton of a couple lines of `calc`, letting it fill in the reasoning for each line, and it does much better. Then ask it about making some macros to simplify "trivial" or "obvious" things, and it does even better. etc.
I suspect there's a good workflow where a big SOTA model makes an impenetrable proof (or code) and then a human works with a FIM model to simplify it (with the larger gnarly proof right there in context for FIM), but unfortunately everyone seems to only care about agents right now.
>> Humans can then work on clarifying why it's true.
Presumably you're a human. Are you going to do that?
There are vanishingly few research mathematician positions and it's one of the most competitive fields, so no. But I'm not sure how that's relevant. As the OP says, usually the value of a proof is not the knowledge that something is true per se, but the reasoning techniques to understand why. How can it be anything other than helpful then to have a truth oracle as you try to figure out why things are true?
> But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.
To me this analogy points in the complete opposite direction. Imagine somebody takes a half-completed project design you're trying to figure out, vibecodes a rough prototype of it, emails your manager to announce that the project just launched in alpha, and then dumps it back on your lap for approvals and testing and productionization. Would you say that they've added value to this process? Or did they just strip away all the hard parts of the problem so they could claim credit for the easy part?
One of the awesome things about LLMs is they make it quick and easy to make PoCs, so yes. Proving that an approach will work before spending a bunch of deep design effort is absolutely valuable. Your exact scenario is something I've literally done: give a half-completed design to a team member and asked them to vibecode a PoC to prove the approach will work and figure out some of the details, explore scaling and failure characteristics, etc. Or I do the PoC vibecoding myself too. LLMs have been a gamechanger here.
It's valuable for you, the person who's going to spend a bunch of deep design effort, to make POCs. Is it valuable for someone else to drive by, dump some POCs on your lap, and then leave you to do the deep design effort while they run away to study AI?
If that person then runs around telling people that they're the real author of your project, because they generated the original POC, would you consider that an accurate assessment?
Your analogy is far enough away from the way that the real world works that I'm not sure that I can really even strain my experiences to fit within it. Sure, I guess that would be annoying?
But mathematicians define their field. They're smart people. They're capable of recognizing when someone just did a vibecoded throwaway PoC and when someone has a well structured proof. Actually even before LLMs they'd publish new, clearer or more elegant proofs of old results. They can say that inscrutable proofs are exactly as valuable as they are, and that the first explanation people can actually understand carries its own prestige.
Yes, mathematicians will clearly need to rewrite the qualifying criteria for prizes to better align with the actual goals and value they were hoping to get from a solved problem. The field as a whole assumed good faith actors and collaboration, not expecting a few trillion dollar companies to walk in and start turning in piles of Lean no human understands to be able to claim "first".
This letter includes someone like Terrance Tao who publicly expressed a lot of optimism about AI for solving novel math like with the Erdos problems. It's not sour grapes but the first steps to define those new expectations for the future to reduce the perverse incentives.
And yet, predictably, people are accusing him of "gatekeeping" and ignoring the arguments he has made here and elsewhere about the benefits vs. harm in different ways of using AI.
It's a real example that's happened to me twice in the past year, so I'm not sure what to make of the idea that it's far away from how the real world works.
I'm also not sure I understand what you're objecting to if we agree that mathematicians define their field. The source link is a declaration from 25 Fields Medallists with precisely that goal. They believe/define/declare that the type of AI-generated proofs we've seen are vibecoded throwaway PoCs; they feel that a well-structured proof must include factors such as "a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and the success criterion is not a true/false conclusion but rather "development and integration into the mathematical canon".