I think this sort of small scale research on this problem is inherently pointless and will be a lagging indicator of diffusion, not a leading indicator of capability.
The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.
If the AI could do this task we would see this happening in places where the economic incentives let them spend millions of dollars on this problem, not on an eval like this.
This specific form of eval where you just ask the agent to solve it with no specific scaffolding besides GPU access (e.g. nothing like AlphaEvolve, ArchPilot, etc that try to work around model shortcomings) is also going to further trail what is possible at small scale. It's good that we at least give them execution environments now, but this feels like the experiments that were worked on figuring out how to get LLMs to do native arithmetic rather than just giving them a calculator/python env.
> The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.
Not to mention a century of theoretical foundations and innovations by humans, written for human understanding.
I've wondered how; "The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result." can possibly not also include many dollars worth of duplicated effort? I find it hard to believe every agent was doing something novel. What this means beyond wasting dollars I'm not sure of, but just flinging around the numbers I don't think should be considered impressive or even required to get the results found.
Regarding the Navier Stokes theorem, one of the best posts I've seen about this (and the other OpenAI math proofs) recently is Nestor Guillen's post where he introduces "the convex hull of ideas". I love this because I've had some vague ideas about the limitations of LLMs that aligned with this, but Guillen really clarified the idea and explained it with a great analogy: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...
Basically Guillen is arguing that LLMs are great at finding results within the "convex hull" of existing literature (i.e. their training data). They can make connections across different parts of the literature where it would be impossible for a human to be an expert in all these areas. But when it comes to truly novel, original ideas, there is no proof yet that LLMs are able to go there. That's honestly a limitation I'm really rooting for because otherwise I think the future of humanity is generally fucked.
Love that link. Thank you. It's helpful to find others thinking this way :)
I work in collective intelligence, and been thinking about what surprise is, and why the ability to sense it is baked into each and every one of us. It's a feeling that moves our attention
I've been riffing on the idea of "mapping" collective surprise by gathering information on which statements give us the feeling. The feeling of something being dropped just beyond the "boundary of our knowing", which I think of as a shell we are each building through our models, like something pushed out into semantic space. I'm saying it funny, but it's really just about collecting data on what surprises people.
And in aggregate, maybe you could have a collective map of this shifting thing (really a map of the education system in the sweeping life-long sense), of the whole human cultural swarm. Surprise is the what we each feel as we cross a threshold that we can't easily return over. We forget today what surprised us yesterday. But maybe it's possible to learn something about the structure we build together (in knowledge), if we gather little radar blips on when we are passing over such moments.
Like if a collective over the course of a week, all has a moment of surprise as they learn a specific thing. If you could see that happening to your group, we could all know that we've all integrated (or at least encountered) something that can be built on.
This framework feels part of talking about it: https://openresearchinstitute.org/onboarding/A_B_U.html
Anyhow, just sharing in case it lines up with anything you think about :)
This is just cope dressed up in math formalism.
Before the convex hull, we had stochastic parrots.
Search is an incredibly powerful method in both human and machine endeavors for generating new ideas and we will see Move 37s in math soon enough, even if we have not already.
I have heard that before, it's essentially the LLM cannot be creative packaged differently. But it always falls flat because there is just no way to measure it let alone define it. In the end it always give me a vibe of them wanting that to be the case rather than it being it case.