That second person stated that for many years they tried that particular graph problem on various AI models, starting with o1 and o3.

Its quite likely they now found the counterexample with a more serious prompt, and then for virality re-tried a few times with meme-prompts like "you should do a breakthrough", knowing that the model is capable of solving this particular one. Worst case the meme-prompts don't work and they share the real one they initially used.

>Its quite likely they now found the counterexample with a more serious prompt, and then for virality re-tried a few times with meme-prompts

i am not sure why this is "quite likely". it'd be pretty silly to get a mathematical breakthrough and then hide it for an undisclosed amount of time to get a few more likes on a tweet, when the impressive part is the breakthrough.

not saying your theory is impossible, but i think the simple answer is that the model is just smarter than o1 and o3.

and, in any case, the model ended up getting the result with the meme prompt and "keep going", which was what i find fun. just like how the crypto results were from prompts of, more or less, "keep going", and that's pretty damn cool.

You can open 10 tabs and prompt 10 times. We are talking about a few hours delay.

I agree with you that obviously no prompt engineering was needed just "solve this problem", but imagine it was you doing this problem with every model, wouldn't you have tested a new model with the best prompt you had from previous iterations, maybe with some partial previous results in it, exactly to maximize your probability for a mathematical breakthrough?

Not everything is a conspiracy