The breakdown with which model, per-task cost, and methodology is in the "Full results and methodology" link in the post, not omitted. Definitely check it out if you haven't.

On the saturation point, we agree that a 95.8% result on a mature benchmark isn't the main proof, which is why we're currently running against harder, less saturated benchmarks like Terminal-Bench, CursorBench, and SlopCodeBench (going to publish results on these hard benchmarks shortly). Apart from current user experiences, that will show the value of our harness.

I did read "Full results and methodology". It doesn't seem to show which agents you are routing to. Am I missing something? And how do those agents score on SWE-Bench Verified?

The benchmark wasn't evaluated with the router, we wanted an apples-to-apples comparison of our harness against other harnesses like mini-swe-agent on the same models to see if we added speed + cost value beyond routing.

Okay, that makes more sense, but then this post is very confusing. Why mention the router at all? Or does your live system have a router and the eval doesn't? In which case why not publish an eval with the router?

Fair point, yes, the router is only in the live system (we'll try to make that more clear). We thought that comparing the router to other harnesses on the benchmark wouldn't be a fair evaluation because the speed/cost gains from our harness would mostly be from using cheaper and simpler models rather than actually having a better harness.