This reminds me of many years ago when Mozilla/Firefox (I think it was) said that they stopped focusing on mainstream benchmarks because they didn’t really translate to real world browser performance gains.

I view these AI benchmarks the same. No I do not care that GPT got 1200 on FartAGIMaX-4.0-Extreme and Claude got 1350. I care about how much it costs and how correctly it does the tasks that I give it. Unfortunately the only way to know is to use them all myself and measure it myself.

At the end of the day these things are all so damn close in how they behave in whatever harness so it realy just does boil down to whatever is actually cheapest.

This is why Deepseek is great: it’s so much cheaper it doesn’t matter if I burn way more tokens because it’s still orders of magnitude cheaper than the US SotA models. If it doesn’t get it quite right immediately I just do a few more turns and then it’s fine. Barely an inconvenience.