This is great. People always talk about how important the harness is and yet we have so few harness benchmarks. Agree with sibling commenter that ‘harness x model’ is needed.

People always talk about how Cursor harness has some secret sauce; would like to see how that one stacks up.

Thanks for making this and filling a real gap!

Thanks! V1.1 is set to add more harnesses and evaluate a broader range of models. We're moving to a full harness × model matrix to uncover interaction effects and release a complete harness-model compatibility map.

I am looking forward to your findings on home-field advantage. It's natural that Claude Code will work better with Anthropic models. If we know by how much, it will inform many of us.