To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that understanding.

I do see how this is a problem in terms of assigning credit, but I think the cat is already out of the bag in terms of these models being capable. Even without AI labs spending millions of dollars to solve millennium prize problems, there are plenty of other people who will use them to pick low hanging fruit. I don't think any social solution is going to make things go back to the way they were, where you could share your progress towards a famous open problem without risking someone "scooping" you within a couple of days.

I think that the most likely outcomes are either mathematics becomes more secretive, or there is a more deliberative approach to assigning credit than who was "first" to solve some problem. In the former case, this may slow down progress, and in the latter case, this could mean that credit would become more subjective, and be a continual source of controversy.

The statement is not about AI but about the behaviour of AI companies. OpenAI have put vast resource into solving open maths problems: many millions of dollars of compute just on the Navier-Stokes result, plus whatever they spent on the broader Millenium Prize problems initiative and the other results they have published. Anthropic are doing the same. The statement is asking them to stop doing this.

AI companies are investing these resources primarily as a marketing exercise. There is no near term commercial value to a 100 page Lean proof of blow up in an extreme special case of Navier Stokes, besides the bragging rights. As the statement says any commercial value in this stuff comes a very long time later after new insights and techniques have been digested, integrated into the mathematical canon, expressed in ways that don't take a lifetime of study to understand, etc. (things that AI is not yet capable of doing itself). The bragging rights, on the other hand, are massively valuable. There is a mystique to maths that makes "our AI solved a Millenium Prize problem" an irresistable headline for a company like OpenAI.

What the mathematicians are saying is stop pouring resources that most mathematicians can only dream of accessing into projects that are actively damaging to their field. They face a massive challenge of figuring out how maths can evolve in the face of this new technology, and this is not helping.

What do we do about the problems that don't require many millions of dollars in resources?

Last weekend I spun up a small agent swarm and pointed it at a field of math I have some affinity towards. Within four hours I had settled three conjectures, one of which is rather famous (for the field, not in general). It cost me about four hundred dollars.

I am at a loss about what to do with these results. On one hand I feel like the mathematicians working on these should know about them, but on the other I feel a bit like a barbarian who suddenly finds themselves sacking Rome.

I'm not a mathematician but it seems to me that if all it took to solve the problem was an enthusiast level understanding of the domain and a few hundred dollars of tokens then the result probably isn't that valuable. Even assuming you are the first person in the world to solve it, these kinds of LLM-friendly problems that are now easy to solve and easy to verify will almost certainly be picked off by one person or another in the near future.

It's also possible that the result is already known and you just weren't aware of it. It's easy for someone outside of a field, or even one steeped in it, to not be aware of certain solutions.

sounds a bit like cope. Crouzeix's conjecture was solved exactly under these circumstances and I wouldn't describe it as "not that valuable". In any case, AI capabilities will increase dramatically over the next few years while human math capability will not. That means that there's a fixed target regarding whatever is currently considered a "serious math problem" and its difficulty level. Soon the average problem solved with a few hundred dollars of compute will be at that bar.

> while human math capability will not.

I actually partially disagree with this. What happened to all the excitement about Intelligence Augmentation (IA)? Now it's AI instead of IA. I think there's so much untapped potential for augmenting our intellect with the likes of https://dynamicland.org and https://folk.computer, as well as the work that's been going on in college math education, things like Lean, etc. I think the only reason human math capabilities haven't expanded that much is a failure of our imagination, not our potential.

> I am at a loss about what to do with these results.

I would recommend publishing them to Palomar (https://palomar-registry.org/) - I have no affiliation, this is an online registry of Lean-verified proofs created by Terrence Tao.

I have submitted a proof there that's also minorly important in an extremely niche field.

Anyway, I feel like it's a good place to dump AI slop lean proofs because the main point of the registry is that it verifies that: 1) your Lean challenge statement is the same as what you informally state you're trying to prove; 2) your Lean proof actually compiles.

This could be useful to future AI slop researchers who want to know if a given result has already been formalized, and they may be able to mine some lemmas from your work. Also, it's good to know for the field in general what has been proven.

I'm fairly certain you can set your publishing name to be whatever you want, so you could set it to be just the word "Anonymous", or the name of the model you used.

[deleted]

It is interesting isn't it.

You asked for a painting. A robot made the painting. You looked at it and said, "well, I guess it's good. Should I put it online or something? Dunno. Hey Fred, what do you think of this?"

Meanwhile, your next door neighbor spends their entire life developing their understanding of life through art. They "understand" (maybe not in a way they can articulate) art. You go next door, you look at their painting and say, "well I guess it's good." But you also understand that your neighbor is just like you, and maybe you are a painter in another way.

I find it strange that, people can't see that, we don't need to solve hunger and poverty and work balance, and etc, by a round-about make-super-intelligent-AI. We could just solve it. It's pretty obvious how to, as well.

We can all be painters, if we put restrictions on the psychopaths.

what about cancer?

How do you know the proofs are correct?

The numerical results are trivial to check. I wrote the analytical results in Lean by hand before asking a former professor to confirm after asking him to keep this private.

They're valid.

It's really tough to come to terms with it but making academic contributions in general from the outside (with or without AI tbh) is not often welcome and the whole process feels very gate-kept.

(academic.) Unfortunately the overwhelming majority of outsiders are missing core knowledge (or are cranks), so the optimal prior from a time management perspective is to ignore them. AI just makes engaging more costly because there’s more volume and it’s harder to get signal on whether they know what they are talking about.

If that's the case, that makes me much less sympathetic, even though I can understand how it's very disturbing to see the field suddenly changing like this.

I mean, let's say you spun up a swarm of agents to rewrite a large component of a well used open source library to be memory safe. You could dump it in a big PR and walk away (we all know how that would go), or you could try engaging, see if they're interested, write something up and see where it goes.

The biggest problem is, IMO, drivebys uninterested in actual results, just getting a check mark, and the equivalent of dropping a 200k line PR on people and expecting them to be interested and do the work for you. These are things many on HN are familiar with and know how to do better :)

> and the equivalent of dropping a 200k line PR on people

I can understand why the community is pissed. So now, lean proofs can be churned out at scale, and the community is left to decipher all of that slop into human understanding. There are bad actors with misaligned incentives coming in with drive-by proofs upending what the community holds dear which is to practice and propagate the art. I applaud them for this declaration.

To re-align incentives the following could happen. AI slop lean proofs are dumped unceremoniously into a lean dumpster, and what gets rewarded are results that could digested into human understanding - via the already followed human review process. Prizes are not given to lean proofs since anyone with sufficient compute can churn them out.

It depends.

If you are trying to understand better the field, then do a good write up of the proofs so that people can learn from it.

If you want to earn the respect of people because you found interesting proofs. Then do a good write up of the proofs so thst people can learn from it.

If you want to plant flags and pollute peoples minds. Then please publish it anonimously, no one wants to correct LLM slop for you.

Probably we should build a repository of AI slop proofs that are only allowed to be publish anonimously. That way people may be more inclined to work on it because they would feel like they are cleaning your house for free.

Maybe I'll end up doing the write ups pseudonymously. I have taken care to make sure the results can meaningfully contribute to field but I don't want to plant flags or really receive credit of any kind. I just think they are interesting.

I like my current life and don't want to get dragged into the current fracas surrounding the use of AI in math.

If you have taken care, then share it with the community. People will be thankful and happy.

The problem is with people that may do it without contributing to the community.

If this was limited to just three results, I would agree with you. But those three are just the ones I've managed to verify myself. The list of unverified results is quite a bit larger.

The field I've been investigating is not large. Even if I were to take the time and care to beat the interesting results into something meaningful, I'm afraid the pace at which I'm able to produce these results would not be well received.

That solution (stop pouring resources in to proofs, stay in your lane) works today. How does it work 5, 10, 20 years from now? The software and hardware advances will continue.

> primarily as a marketing exercise

Like the mathematicians working on famous problems in private until they could claim full credit for something interesting wasn't also a marketing exercise for their own careers. The commercial value (or lack thereof) of a proof doesn't depend on whether it was done by a human or a machine.

These mathematicians dedicated their life to math and were working for a long time to achieve the pinnacle of their careers.

OpenAI just burned millions of dollars over a weekend after hearing that someone else was close to solving the problems. Their interest was in their AI system more than the actual math problems.

Don’t you see how that’s different?

I see how it can be devastating to their ego, but no, I don't see a particular difference in a company spending money for clout vs. a person spending time for clout. The underlying motivation is the same.

Do you really think people dedicate their life to mathematics for clout?

If clout was the goal I don’t think becoming a lifelong mathematics academic would be the first step

for many fairly strong mathematicians, the career calculus is fame and status (mild though it may be) through mathematics or anonymity but financial reward in tech or finance. Yes, clout by becoming a lifelong academic is in fact rational for some and part of their motivation. Of course they really like what they do as well, but earning the respect of the peers they know are also respected by a large swathe of society is very important.

Yeah that’s totally fair!

I guess that type of “clout” feels different to me.

Wanting to be validated by peers for your talents in a niche field vs. using millions to try to solve a math problem that you don’t really care about with AI to market the gigantic company you work for.

I think you’re right that some of the angst here is due to a fairly insular community having its culture disturbed by commercial forces.

Yeah bro, same for probably the majority of this site that write software for a living and/or hobby.

Sorry I’m not entirely sure what you’re trying to say

How good were you at those “A is to B as C is to D” SAT questions?

No need to be rude. I didn’t want to misinterpret when responding.

I can see how what is happening to mathematicians is similar to what is happening to coding.

I’m not sure what your broader point is? Mathematicians shouldn’t be upset? Coders should? Something else?

The point is that nearly every white collar worker is in the same boat right now, and the majority of them probably have already had much more profound impacts to their fields. So yeah, we can certainly imagine what it’s like for mathematicians.

Okay. I agree with all that you said in your latest comment.

From your first comment it seemed like you were disagreeing with me but I’m not sure how.

That’s why I asked for clarification.

Edit: BTW I’m a web developer and designer, Not a mathematician

> There is no near term commercial value

How much do you think other AI companies would offer to get access to the transcripts of the generation that led to the proof? No doubt OpenAI will include it in their training data somehow and use it to build the next generation.

There is already economic value.

I suspect that beyond just marketing, these pursuits yield plenty of useful information about model design that will likely lead to model improvements and optimizations for both mathematics and general reasoning going forward.

Are they claiming that the only value in solving these problems was for their field's personal development process? I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology. It would be insane to demand that people avoid making progress on technology that can save lives or improve general quality of life, just to protect the sanctity of your karate belt system. Perhaps in lieu of open problems left to solve, mathematicians should be welcome to take up chess or sudoku to keep their minds spry.

No, there's most likely zero practical benefit of having found a pathological edge case in which the Navier-Stokes equations do not work. We're most assuredly not talking about "saving lives" here. Unless advanced aliens show up and tell us they'll destroy Earth unless a counterexample to the N-S equations is provided within 24 hours.

> I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology.

This is the core misunderstanding that the open letter is attempting to correct.

Developing a better understanding of the Navier-Stokes equations could have a number of implications for useful technology. They're fundamental to fluid dynamics, and turbulence in particular is something that many people feel we could work with more effectively if we better understood how and why it's generated. The Navier-Stokes smoothness problem is an interesting and long-standing benchmark for this understanding; we don't know why it should be so hard to answer, so we hoped that the process of developing a proof to the problem would produce more understanding. (We may still be able to extract this understanding after the fact, if OpenAI's proof is fully human-comprehensible.)

Simply knowing that there exists a finite-time blowup is not practically useful. We know that fluids in the real world don't produce random singularities, so the result can't really have much physical meaning. What it illustrates is that the Navier-Stokes equations fail to model physical fluids in some yet to be characterized way.

Humanity is better off for knowing these proofs. This strikes me as academic NIMBYism.

Do you "know", in any meaningful sense, any of OpenAI's recently publicized proofs? Do you suppose that there is any large community of non-academics that does?

One of the points the parent makes, along with the TFA, is that academia -- or more specifically, the "mathematical community"-- is a setting primarily for creating and ingesting mathematical knowledge, and disseminating it to the next generation and to other fields. Humans absorb this material slowly, through lots of discussion and collaboration -- it is necessarily a slow process. Facilitating this is one of the important functions of academia. Your usage of academic as a slur here is a bit silly for this exact reason.

I don't claim it is perfect, and we can argue about pedagogy in elementary courses till the cows come home. That's not really material. But this is one of the only settings in which such knowledge is broadly valued for its own sake, and in which there is a semblance of incentive to help others "know" this stuff as well, be they future generations of mathematicians, science and math educators and communicators, practitioners in other fields, or genuinely curious amateurs.

Your point is largely addressed in the article, did you try reading it? "In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align."

The point is that these proofs are largely useless without the insights. The value of a proof is largely in the travel, not so much in the destination.

Different person here, I read the article and they are all wrong. Hope that helps.

Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing. And instead of them - and nobody - spending millions of dollars to solve the problem, successfully, they want every problem of their academic industry to persist because even though they never solve the problem, they synthesize and solve lots of other problems nobody asked for. And get to boost their egos?

Yeah, stop that. Actual alignment is on the humans themselves, if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades and don't worry about the narcissistic elements that slow their industry down.

Mathematical breakthroughs with commercial relevance are few and far between, and often depend on dusting off old results which were, at the time of discovery, "solutions nobody asked for."

The NS counterxample is actually, by any market measure, a "problem nobody asked for" in the sense that its existence doesn't have any commercial relevance (beyond juicing OpenAI's IPO). So the only long-term value solving it could have is by virtue of whatever reusable theory/insights were generated along the way to the counterexample itself. The letter is absolutely right on that point.

It's not actually clear that those insights will come faster from reverse engineering this LLM proof vs. humans building theory to solve the problem themselves. So what you're saying may or may not even be an efficient way of operating. Also, it implicitly depends on mathematicians to do the hard work of creating problems and then deciphering LLM hieroglyphics for essentially free while the only immediately profitable component gets outsourced to a frontier lab. In what world is that model going to work?

Reading between the lines, it seems like maybe you have a personal grudge for some reason and simply think the technology will advance enough to where we won't need academics at all. But you should say that in the first place.

What irks me is the ego

My stance is that solving the problem is aligned with humankind

the rest is just hypothesizing a way that academics fit in this world at all

Their claim is only indirectly related to the motivations of the people using their models. What they're saying is that doing math in this way does not produce the same value as traditional mathematical research, and the people using these AI models aren't concerned about that because their marketing objectives don't depend on whether their results produce mathematical value. If people doing valuable work are made irrelevant by people doing a larger volume of non-valuable work, that's not a positive outcome.

But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.

I actually found this to be the case with some basic linear algebra notes I was recently doing in Lean (without using mathlib). The model could generate working proofs, but they obscure the basic ideas (actually I wonder somewhat if this is because the Lean code that's out there to train on doesn't make a huge effort to read like textbook proofs, which was my motivation in the first place). I give it a skeleton of a couple lines of `calc`, letting it fill in the reasoning for each line, and it does much better. Then ask it about making some macros to simplify "trivial" or "obvious" things, and it does even better. etc.

I suspect there's a good workflow where a big SOTA model makes an impenetrable proof (or code) and then a human works with a FIM model to simplify it (with the larger gnarly proof right there in context for FIM), but unfortunately everyone seems to only care about agents right now.

>> Humans can then work on clarifying why it's true.

Presumably you're a human. Are you going to do that?

There are vanishingly few research mathematician positions and it's one of the most competitive fields, so no. But I'm not sure how that's relevant. As the OP says, usually the value of a proof is not the knowledge that something is true per se, but the reasoning techniques to understand why. How can it be anything other than helpful then to have a truth oracle as you try to figure out why things are true?

> But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.

To me this analogy points in the complete opposite direction. Imagine somebody takes a half-completed project design you're trying to figure out, vibecodes a rough prototype of it, emails your manager to announce that the project just launched in alpha, and then dumps it back on your lap for approvals and testing and productionization. Would you say that they've added value to this process? Or did they just strip away all the hard parts of the problem so they could claim credit for the easy part?

One of the awesome things about LLMs is they make it quick and easy to make PoCs, so yes. Proving that an approach will work before spending a bunch of deep design effort is absolutely valuable. Your exact scenario is something I've literally done: give a half-completed design to a team member and asked them to vibecode a PoC to prove the approach will work and figure out some of the details, explore scaling and failure characteristics, etc. Or I do the PoC vibecoding myself too. LLMs have been a gamechanger here.

It's valuable for you, the person who's going to spend a bunch of deep design effort, to make POCs. Is it valuable for someone else to drive by, dump some POCs on your lap, and then leave you to do the deep design effort while they run away to study AI?

If that person then runs around telling people that they're the real author of your project, because they generated the original POC, would you consider that an accurate assessment?

Your analogy is far enough away from the way that the real world works that I'm not sure that I can really even strain my experiences to fit within it. Sure, I guess that would be annoying?

But mathematicians define their field. They're smart people. They're capable of recognizing when someone just did a vibecoded throwaway PoC and when someone has a well structured proof. Actually even before LLMs they'd publish new, clearer or more elegant proofs of old results. They can say that inscrutable proofs are exactly as valuable as they are, and that the first explanation people can actually understand carries its own prestige.

Yes, mathematicians will clearly need to rewrite the qualifying criteria for prizes to better align with the actual goals and value they were hoping to get from a solved problem. The field as a whole assumed good faith actors and collaboration, not expecting a few trillion dollar companies to walk in and start turning in piles of Lean no human understands to be able to claim "first".

This letter includes someone like Terrance Tao who publicly expressed a lot of optimism about AI for solving novel math like with the Erdos problems. It's not sour grapes but the first steps to define those new expectations for the future to reduce the perverse incentives.

And yet, predictably, people are accusing him of "gatekeeping" and ignoring the arguments he has made here and elsewhere about the benefits vs. harm in different ways of using AI.

It's a real example that's happened to me twice in the past year, so I'm not sure what to make of the idea that it's far away from how the real world works.

I'm also not sure I understand what you're objecting to if we agree that mathematicians define their field. The source link is a declaration from 25 Fields Medallists with precisely that goal. They believe/define/declare that the type of AI-generated proofs we've seen are vibecoded throwaway PoCs; they feel that a well-structured proof must include factors such as "a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and the success criterion is not a true/false conclusion but rather "development and integration into the mathematical canon".

[deleted]

> What the mathematicians are saying is stop pouring resources that most mathematicians can only dream of accessing into projects that are actively damaging to their field. They face a massive challenge of figuring out how maths can evolve in the face of this new technology, and this is not helping.

Is it reasonable for any field to make such demands? If this were doctors objecting to AI becoming good at medical practice would you have the same concerns?

While any idea of OpenAI spying on people to pursue their goals is disgusting, the rest of this is par for the course, as Kasparov experienced with IBM in the 90s. Humans still play chess after all.

It's not a simple matter of "becoming good at", and yes, I could very well have similar concerns, depending on how it impacts the field. The statement itself mentions that such concerns exist in many other fields.

I doubt OpenAI will take such a combative stance and accuse these mathematicians of "demanding" things, as you do. As I said, the purpose of this is marketing and the statement simultaneously undermines the value of that marketing (showing these projects as irresponsible) and gives these companies an even better piece of marketing in its place: "our AI got so good at maths the mathematicians begged us to stop". It's entirely possible they will stop pouring millions into these projects.

So let’s say OpenAI cure cancer and put every cancer researcher out of work depriving them of intellectual satisfaction, this would also be a problem? It would certainly impact the field.

The fact is these fields are supported by society because of the benefits to everyone else. Once the same results can be achieved in a cheaper and faster way that is what will be done. We should mourn this in the same way we do buggy whip manufacturers. Again people still ride horses.

> So let’s say OpenAI cure cancer and put every cancer researcher out of work depriving them of intellectual satisfaction, this would also be a problem? It would certainly impact the field.

Would you expect Fields medalists to cure cancer if you moved them from the math department to a medical research lab? This is precisely the fallacy that the frontier labs are counting on to inflate their valuation as their IPO approaches. They want to use headline-grabbing problems in pure maths to make their models look "smart" in the public eye. But what does "smartness" in mathematics really mean in terms of economic value? It is not at all obvious whether success in abstract mathematics should translate to successes and, more importantly, profitability, in more grounded endeavors.

Look at OpenAI's job postings (https://openai.com/careers/search/). Those roles involve far more pedestrian yet profitable duties than research mathematics. So why isn't OpenAI automating them with their vaunted models? Success in one field, no matter how "difficult", does not predict results in another field.

> The fact is these fields are supported by society because of the benefits to everyone else. Once the same results can be achieved in a cheaper and faster way that is what will be done.

Maybe it's worth double checking that you know how these fields benefit everyone else? Proving the blowup of the Navier Stokes equations in 3D isn't going to make your gas cheaper or make harvesting food easier or make drones easier to protect against. Maybe consider the deeper effects at work?

Math people absolutely are important to things like the SpaceX landing control systems, hypersonic gliders, 5G networks and so on.

If you can make breakthroughs on such areas as fluid dynamics, control theory or information theory with AI then that absolutely is a big deal with real technological implications.

A long complex proof to something we have been able to simulate in fidelity forever doesn't do any of that though.

Why don't we first assume OpenAI can create a cold fusion reactor and then give them a trillion dollars to get it done?

Chess is even more mainstream and accessible now!

[deleted]

yeah beacuse no one trust any of their benchmark results now they are scrambling to find a signal thats undeniable

I expected this. They prove a millenium result, but it doesn't count because they are bad people.

It's this sort of thing that motivates people to burn down the institution you might be trying to defend.

> It's this sort of thing that motivates people to burn down the institution you might be trying to defend.

lol, yes, this sort of thing is what many people who voted for Trump were saying, and things are going great for them.

> They prove a millenium result, but it doesn't count because they are bad people.

OpenAI has only themselves to blame for this, and they know it. They could have handled this so much better. I'd bet there's more meeting time right now going into how to unveil future math results than on meeting about the actual math research.

The issue of credit is a relatively minor point in the declaration.

It's more about bypassing the culture and processes mathematicians have developed that lead to human understanding, generating new ideas, and bringing up new generations of mathematicians. (See also his article about "non-renewable mining" of good problems.)

Reducing mathematics to "let's just generate results through an isolated and automated system" is a misalignment since it bypasses those processes.

> The issue of credit is a relatively minor point in the declaration.

What a load of croc. This entire debate is fueled by a perceived lack of attribution. The AI learnt from researchers and did not give them a sporting chance of being first before scooping them. They were expecting some sort of fair play, instead they got a ruthless machine. Every other tangent to this debate is irrelevant, the culture, the community, the shared symbolic growth. Every mathematician I know is secretly trying to one-up their peers.

I wonder if this is a root of the complaints across fields, how AI is ruining the greater picture and process in writing, acting, drawing, filming, coding, and more.

That it makes life more ends and less means.

It's also known as "commodification of labour", and AI is just the latest and greatest tool to do it.

Luddites complained that the trajectory of technology was to allow less skilled workers to mass produce goods via machines owned by factory owners, as opposed to helping skilled workers build up and use their skills while passing them on.

Now we have a lot of money and time focused on LLMs owned by a few companies, making it easier for them to monetize low skill labour(prompting versus art/research/artisanry)

Let's take your argument one step further.

Suppose that tomorrow we learn that AI just exploited a bug in Lean and the proof is, in fact, bullshit. Or suppose it is the case, but we never learn that.

Where are "ends" and where are "means" here?

Well, given a proof of something then a system can be built that relies upon the inviolability of that thing.

Should the proof turn out to be bullshit, then that system will be revealed to be unreliable. Maybe.

I don't think it really attacks human understanding though. You can still read and understand an AI written proof. If another person comes up with a solution to a problem, you can read their methods and understand it. It doesn't matter if a human came up with that or not. It's really only attacking the "generating new ideas" part.

That's precisely the problem though. You cannot still read and understand an AI written proof at the current skill level of the AI being applied, because they're orders of magnitude longer than human written proofs even when they don't need to be, and spend most of that length on the parts that aren't important. This has been really thoroughly documented by expert mathematicians who are engaging with AI in public like Terence Tao and showing in detail how much work it takes working alongside AI to figure out how to understand AI generated proofs. With human generated proofs that process is forced to happen before publishing the proof because the new style of AI generated proofs validated only by formal verification is supplanting the old human peer review process that forced the burden of understanding onto the publisher and not the reader.

> You cannot still read and understand an AI written proof at the current skill level of the AI being applied, because they're orders of magnitude longer than human written proofs even when they don't need to be, and spend most of that length on the parts that aren't important.

Just like how they write software, then :-)

That doesn't seem to be true. The OpenAI NS paper was 166 pages. Wiles-Taylor proof of Fermat's last theorem is 129 pages. The length is not unprecedented for a difficult unsolved problem.

To be honest, I feel like the difficulty of reading AI proofs is due to the fact that we are on the verge of being beyond human comprehension. This is a demonstrable fact as no human has figured this out despite the problem being open for almost 100 years.

> To be honest, I feel like the difficulty of reading AI proofs is due to the fact that we are on the verge of being beyond human comprehension.

I can see where that's coming from, but I really don't think it's the case. Even with Astra, the proofs you get are just off in a way that doesn't signal superhuman comprehension. As 9question1 says, a common theme is that they dwell on insignificant steps. Another one is that they'll often be full of terminology that either doesn't exist, or has this weird quality where it looks like it is trying to make some minor insight seem much greater than it is. At first glance, that'll often make it look like it knows more than you, but when it's really just doing the same thing but in a more complicated and worse fashion, that to me isn't a signal of comprehension at all. The bizarre thing is that despite all the "stochastic parrot" style nonsense you'll get in individual proof steps, they still often combine to something valid.

In either case, what all of this means is that the working mathematician still needs to go through, and generally completely rewrite, any proof output by an LLM. Otherwise you are passing the burden of unreadability onto the reader.

Yeah, that mirrors what I've seen throwing some of the leading models at a set-theory problem that's stumped me (https://mathoverflow.net/q/511601): in this case, the problem does not easily yield to the standard tools, but the LLMs do not recognize it as a major open problem they should give up on. So they seriously try it, but typically end up in a loop of inventing certain classes of simple solution or counterexample attempts, defeating them, and trumpeting each one as a major result, each time inventing some new terminology.

It's definitely quite curious that the AI labs are able to push these results through seemingly with pure brute force. Perhaps it's largely a function of how many monkeys you have attempting various constructions on top of the known results and strategies the models have memorized.

That matches my experience with AI writing in software engineering so I’m not surprised

> This is a demonstrable fact as no human has figured this out despite the problem being open for almost 100 years.

That's not true. Alpoge and Buckmaster's related LLM-assisted blowup result (https://news.ycombinator.com/item?id=49605915) utilized a strategy developed recently by Cordoba and Martinez-Zoroa.

This is a very token-brained take. The length of a work has no bearing whatsoever on its comprehensibility.

Not "a" human's understanding; Humanity's understanding. Understanding the research problem, and the solution especially, is a lot more involved than simply "read their methods". That's the whole point being made.

It matters if a human came up with it because of everything mentioned in the article... A mathematician's solution is necessarily built on other's ideas that have been disseminated, internalized, pressure tested etc. Methodologies differ too. AI can abuse its compute resources and generate a true/false or counterexample statements, without laying the foundation that a decade of globalized research would have.

>You can still read and understand an AI written proof.

No you can't lol, they're multi million lines of Lean, which is already an obscure language to understand. It's an assault on your senses.

It's not only Lean code, there are English writeups too. To my understanding the pipeline for these problems is 1) solve in english 2) formalize systematically to check. No one is tackling problems purely in Lean, to my understanding.

https://cdn.openai.com/pdf/32d9f210-8b73-45e0-91bc-82a30aef8...

[dead]

I agree that the declaration doesn't focus on credit, but I think it's still at the root of the problem. Because ask yourself: if the AI generated proofs are not creating any new ideas or insight, just brute forcing a boolean true/false result, then why can't mathematicians simply ignore their results? Why does it matter if OpenAI or even amateurs with AI are "solving" these problems, without contributing to any deeper understanding?

I don't think "intellectual poisoning" is really the mechanism that harms the mathematics community.

The harm is if you have a community of mathematicians who are focused on expanding human understanding, then having instant access to a bunch of AI proved results muddies the water about who has contributed what. If someone could scoop any significant theorem at any time by pointing an AI at it, how do you really demonstrate that you have created new understanding? Or that your new understanding is about something important? How do you prove that the AI needed your new concepts to be able to solve it?

Open problems are not that yardstick. Fermat's Last Theorem is the result of Wiles and Wiles-Taylor, but without key results from Serre, Ribet, Ihara, Langlands-Tunnels, and Frey's program none of what Wiles did would work. But Wiles did get the prize. Nowadays I think the inputs have shrunk a bit by doing more in the R=T theorem so less other cleverness needed.

Doing the last step of solving an open problem is the yard stick. But not for much longer: https://davidbessis.substack.com/p/the-fall-of-the-theorem-e...

>or there is a more deliberative approach to assigning credit than who was "first" to solve some problem

it feels like an unintended consequence of the millennium prize is that people view the [last contributor to the solution] as the only one to make progress on the problem. I've never viewed Poincaré as solved by one person and the objective of the prize was to encourage more people to make attempts and contribute towards progress.

this issue is independent, but in these circumstances perhaps interweaved, with the 'ai is taking over math' concerns

> destroyed the yardstick … that has traditionally been used to measure how much they have contributed

This goes much broader than mathematics or academia. This is the entire basis via which society distributes its wealth: based on a labour market derived valuation of ‘contribution’.

> This is the entire basis via which society distributes its wealth: based on a labour market derived valuation of ‘contribution’.

Correction: that's not how society distributes its wealth, it's how it throws some bones to the masses. I wouldn't be surprised if over half the wealth goes to people who don't sell their labor at all.

> but I think the cat is already out of the bag in terms of these models being capable.

When there's a discussion about doing something against the damage of the AI industry: "whoopsy, sorry, another cat escape, nothing can be done".

When there's a concrete mention of an actual solution to avoid more cats escaping: "that won't happen, and even if it did, the damage is already done, and in fact it’s not that bad you all just have to go with the future we decided for you."

So the bag is wide open, more cats will escape, and nothing can be done about any of it. not about the ones that got out, and not about the ones still inside. Sounds more like a preemptive excuse for inaction, cosplayed as pragmatism

I feel a similar fate will befall engineers too. Your predictions anre quite interesting from that perspective.

Markets defined entirely by law have distorted our collective understanding of what can actually be built with the knowledge our species has accumulated thus far. How will traditional shields that have protected capital accumulation in tech to survive in a world where governments now realize control of technology is a national priority? Especially as we see its impact on modern warfare, and that such conflict looks like it’s only escalating over time.

Mathematicians appear to me (as an outsider) to exist in a field without such distortions, and I think offer engineers a preview of what’s to come. I certainly have completely ceased sharing original ideas online at this point.

Engineers are needed to build real things. AI isn't there yet.

Agreed. The problem is not with AI per se, but with the reward mechanism in academia in general.

For those outside academia, the ”reward mechanism” is a choice between A) being a genius and working hard to become a leader in your field, B) becoming very good at writing grant applications, C) capitulating to corporations and living with the moral burden of their exploits

Right. It seems like reading an AI proof (although it may not be well written) will provide the same insights as reading a proof from another mathematician, assuming it's been reviewed and edited, just like any human-authored publication. If the work is inherently valuable on it's own, I feel like that's mostly what matters.

At issue is the fact that it doesn't typically work like: mathematician produces a proof in isolation, generates a PDF, and shares it with a bunch of people. There's a whole culture and community going on behind the scenes with conferences, seminars, lectures, chats in the hallway, advising students, etc. that AI-generated proofs bypass.

I don't think that AI would disrupt any of that. Even if AI solves a problem, you can still discuss the methods at conferences, seminars, lectures, chats in the hallway, advising students, etc.

Hopefully that's true. Instead of bypassing the community maybe it stimulates the community. I think it does disrupt things a bit, but I'm optimistic at the moment!

Tao’s key point is that the way people are using AI today does disrupt that. Because people and companies are valuing the results over understanding. So we’re getting slop results rushed out that are automatically verified.

Moreover in the past, discussion and idea sharing would happen naturally to overcome the friction of the process. But now when OpenAI is stuck on a particular part of NS for example, they can just throw more capital & tokens at the problem.

I won’t speak for Tao. But it is not “people” who have valued the result over understanding. “People” are a bit impressed that the machines are as good or better than the priests who have been praying at the inscrutable altar of pure mathematics. But it is mathematicians—I am speaking as a mathematician—who have adopted a culture of valuing results over exposition and transparency. The field has been rife with extremely opaque papers for many years and the character of research participation has been one of exclusion and competition over proof priority at the expense of understanding and transparency. It is unbelievably ironic to listen to mathematicians complain about “AI slop” when they have built careers upon human “slop” if slop means papers crammed with technical density that prevents all but specialists from reading the work.

I don't think that's entirely fair.

I think it's OK that there are some materials designed for a specialist audience, and some materials designed for a wider audience. Technical density serves a real purpose in the former (I'd like to just say "stack" without including an explanation 10 times longer than the rest of my paper about what a stack is and why we are talking about them), and some people expend a lot of effort on the latter (think lectures, lecture notes, textbooks, seminars, blogs, etc--totally appropriate to dig into the motivations here).

It's just the fact that it's a really vertical field, not some cultural failing, that results in pretty opaque stuff sometimes.

Scale the problem down to an email. You're saying AI won't disrupt writing an email, when the text is vetted and authored by a person. True!

But now let chatGPT write lengthy emails unrestricted and now no one wants to read your slop anymore. That's what is being advocated against.

Is it not an option to make academia more like a normal job, where people focus on collectively achieved outcomes rather than credit, priority etc.?

Almost every "normal" job has regular performance reviews where individual contributions, not collective outcomes, are reviewed and used as a sole input for raises, promotions, and firings. If you can't sufficiently document what you personally did, you might've as well not done anything at all.

Yes, that's true, but it is still generally accepted that a completed project is a result of some collective effort where people of varying degree of seniority and ability contribute. There is also not a singular event of a project being completed with a list of heroes/geniuses making it happen, but rather a whole lifecycle of gradual development, maintenance and going out of relevance with contributors coming and going.

I can imagine mathematics of the future being more like that rather than history of discoveries with dates and names

How is that like a normal job?

Nothing will be lost if the credit system disappears. History of Science has many examples where wipe outs happen. The chimp brain cant survive without creating elaborate stories about how important it is, more as a cope to its own limitations and what it cant predict or control. Humility is good for health. 3 inch chimp brains didnt create the universe.

[dead]