If anyone wants this kind of benchmark for their own codebase, happy to set you up with superconductor.com/benchmark
1. Import your own PRs 2. We find the original spec or infer the spec 3. Agents you select (eg claude code opus 5, codex gpt 6 astra high, pi kimi k3, etc) implement the spec (starting from the parent commit of the PR, with the git history is pruned so they can’t look up the implementation) 4. Three judge LLMs grade agent solution given spec, PR implementation, and a rubric 5. You see the scores