> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem.

The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.

The tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price.

It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.