Even under your interpretation, OAI pushed a button and solved NS. Yes, that is very impressive. Are you kidding me? Imagine building an automated system that can solve NS.
Even under your interpretation, OAI pushed a button and solved NS. Yes, that is very impressive. Are you kidding me? Imagine building an automated system that can solve NS.
> OAI pushed a button and solved NS
Not exactly...
From here:
https://cims.nyu.edu/~tristanb/statement.pdf
"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem".
The evidence we have (not much) leaves the facts severely under-constrained. All the below are consistent with what we observe:
- OpenAI lying their face off for marketing purposes
- Buckmaster sore about getting beaten to NS, misrepresenting what was told
- Buckmaster not being an expert in ML, not understanding what he was told and hearing what he wanted to hear as it vindicated him
- Others...
OAI claimed that none of the people on their team were domain experts. So it's more akin to Deep Blue than Stockfish (actually slightly better than Deep Blue which did hire domain experts) but nevertheless represents a highly sophisticated automatic theorem prover.
That's not my interpretation - that is literally what OpenAI say in that press release.
Even if true, I don't see why this is an issue. Are they not allowed to work on problems others are working on? Did Anthropic get first dibs on this problem? Competition is good. And I don't exactly have tons of sympathy when the other side is just a leading AI lab. It's not like it's some scholar who dedicated his life to this problem.
The "steal their thunder" is interpretation. What I'm saying is that you believe they solved NS on a lark to bully some other researchers, and that this is not impressive?
What's impressive for a human and for an AI are two different things.
Magnus Carlson had a peak ELO rating of almost 2900.
Would you be impressed with someone with an ELO of 3700?
Would you still be impressed if I told you it was Stockfish?
OpenAI didn't go looking for a tough-for-an-AI problem to solve - they went looking for one that looked like it was easy since it they had heard it had already been solved.
Do you find this impressive?
Yes, Stockfish is genuinely impressive. I would be proud to author Stockfish. You don't think so?
As a developer yes, especially given that it runs on a PC, and DeepBlue in it's day was really more impressive since is used custom ASICs.
But, I assume the Stockfish developers aren't comparing themselves to Magnus.
Let's see if OpenAI, or someone else, can get these sort of physics/math results out of a desktop PC - that would also be an impressive piece of engineering!
Possibly after being given the significant part of the solution from actual human researchers. Which they then bullied. And they beat them to the finish line only because they heard rumor and threw everything at the problem. It doesn't look good for openAI in any way. I see more reasons to avoid using them rather than use them from this story.