Happened to be testing a "review patches on a mailing list" harness I was developing; here are a sample of the latest results, testing 12 patches containing a total of 14 issues:

Opus 5.5: Found 8/14 issues. Total cost: $15.40

Fable 5.1: Found 7/14 issues. Total cost: $66.34

Opus 5: Found 6/14 issues. Total cost: $15.19

Sonnet 5: Found 2/14 issues. Total cost: $19.15

This is a relatively small sample size, but it was both the best and the cheapest.

ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.

Heh I’ve been doing almost the same - back testing against PR comments - and opus 5.5 matched fable 5.1. About $24 for 10 PRs.

I was surprised how much worse Astra did on correctness; I stopped testing with it. Gonna try sol and Luna but low confidence

I doubt your test if you cant even notice that it is not the cheapest with just 4 numbers to compare.