This looks cool and the mechanism looks plausible. I found the experience of trying to understand whether the claims here are legit to be aggravating.

First of all, the whole readme section about benchmarks appears to be Claude/Codex written. What a slog to read this.

Second, they claim success on SWE-Bench Verified, but it's only on 50 tasks, not making clear how these tasks are chosen. I know from experience that you can keep selecting sets of tasks until you get a result. Also, this was run 1x, and the lift they show actually only has a p val of 0.22.

There's a famous book "how to lie with statistics". I don't think the continual posts about benchmark results here are purposeful lies, but I think it's just so easy to fool yourself (and I've been burned, most recently building http://pellmell.ai ). Also I don't think this post is the most egregious example.

we are running them on a recurring basis, so will keep on updating that 50 number. and hence as it's currently running that section gets updated by claude only.

Wow, this is really weird. I looked at the github and see you published benchmark updates in these sizes of N:

9 → 20 → 36 → 50

These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:

> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.

> You chose different sized task groups (almost certainly not the case).

> You excluded timed out trials, or something (also would be dishonest).

Hopefully there's a better explanation here!

that's just when our claude limit's were about to exhaust :/ as we have to spin up a whole claude session to test it.

but you could 100% be Sherlock Holmes!

HN commenters shouldn't have to be Sherlock Holmes!

Extraordinary claims should be supported by evidence. You could have easily posted the seed, the 50 questions, and the results for claude and graft on each, but didn't.

I also see Graft got 16/16 on one update and 2/14 on another. Further evidence that each group of questions was selected somehow, not random.