I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.

I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)

I would be surprised, considering that OpenRouter requires disclosing the quantization and shows automatic benchmarks to compare between providers for the same model.

The latter is a joke

I saw that it runs GPQA Diamond and TAU-Bench Airline and shows the results over a 32 day rolling average.

Other than that they track Tool call error rate and Structured output error rate.

I only discovered this today, and it seems like a good idea. What are the problems in practice?