I don't know, people. We still really don't know how OpenAI or others are producing these results. It's all very hand wavy and trust-me-bro. How much money/time/compute have they really thrown at these problems? How much human involvement was there? What LLM did they even use? How much regular software was involved? They have given answers to some of those questions but no proof that that's actually what they did. I don't know if it's worth giving them this much credit (which is what we are doing by writing these essays and spending so much time debating). Anthropic wrote a C compiler that turned out to not really be a ready made replacement for GCC. Did they ever do any more work on it? Has anyone else produced a C compiler? It seems like that and these proofs are just demoware that are not (yet? Who knows?) production ready to turn the world upside down. Impressive one-off demos, yes, but companies have been pulling those off for centuries without ever going anywhere afterwards.

A very important point! In Tristan (NYU prof)’s write up he noted evidence of the OpenAI mathematician team doing a lot of correction and guidance along the way. We are never told about this with openness and clarity. At a minimum, complete disclosure and honesty is needed by the companies and about the precise role of their staff members.