> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
See:
> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)
Sure and if I make a half court shot after an hour of trying, the result only took 1 second.
Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.
They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.
[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...
That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.