Can we audit the CoT and work the AI did to generate such a remarkable cancellation?

I doubt Anthropic will share the details (or at least the full true details). The mystery of the magic makes for much better marketing.

I think a reasonable assumption is that there is an interaction between an LLM, a https://en.wikipedia.org/wiki/Computer_algebra_system tool, a human prompting with deep math expertise, and lots of compute that explains hitting upon the remarkable cancellation.

I think you can reasonably assume that frontier models are using SymPy or something like it any time interesting math gets into the picture, and the person driving Fable here is an accomplished mathematician, but I don't think we can reasonably assume either extensive prompting or brute-force compute in any sense other than what it normally takes Fable to, say, whip up a calculator app.

I have no idea what actually happened behind the scenes, but the human prompter, Levent Alpöge, indeed has deep math expertise. Princeton PhD, Harvard postdoc, and some excellent research (prior to this) to his name.

https://alpo.ge/

Not just an assumption, I saw the LLM saying it used sympy.

[deleted]

> a human prompting with deep math expertise

The original tweet implied that the whole thing was done while the author was watching the World Cup final.

I know it’s tempting to hope that a human did the “real” work here, but if some special insight was put into prompting, the author kept it to himself, and there is no reason why they would hide this since it would elevate their own status.

I don't think it is as much about 'real' work or a special insight as it is being willing to push back multiple times, or simply asking in a way that steers it towards actually 'giving enough of a fuck' to even bother. We tend to be ~blind to how differently we would ask about something we know compared to a novice, this is what makes some better teachers than others.

Have encountered a similar flavor in programming, wrote it off until I saw someone point out how garbage in garbage out they tend to be. If you hand any frontier model dogshit and ask it to do something simply, the result is often not great.

But! If you spend 20 minutes having it comb through and clean up with something like jscpd, then tell it to step through with a debugger, gather profiling traces, etc... very likely it will yield meaningful improvements or catch some corner cases. If it doesn't, anyone with experience is going to tell it to try something else, or that it isn't good enough, as opposed to accepting the first result.

You can recreate this by disabling web search and asking a model about the conjecture and then giving it his post. I've tried a few and their initial responses range from "this is a meme I'm not even going to verify it" to vaguely insulting chains of thought, concerns about the need to be careful because you're clearly nuts or stupid, then falling back on remedial explanations. After a few nudges they all eventually work through it, accept it, and apologize.

IMO its reasonable to imagine a situation where someone is having a beer or two watching The Big Game, asking an LLM to do something stupid for fun and landing somewhere like this on the magic jump to conclusions mat.