I find that a lot of the recent allegedly great open models are cranking their reasoning way further than I find reasonable for interactive use. I’m writing this while waiting for the new Deepseek V4 Flash to finish its task, which is taking way longer than the older version.
What gets reported is always the benchmark result, but rarely the real-world trade-off made to achieve it. That’s an obvious incentive for the labs, so I think Simon is correctly zeroing in on it. Please continue doing so for models that don’t go too far as much as this release.
Don’t get me wrong, I think it’s amazing what we can get out of smaller models with more reasoning, but we should be super aware how very much not-free it is.
This is a good opportunity to call out models that reason quickly: Meta’s Glimmer seems to be pretty token efficient so far, as do the GPT 5.6s.
The model itself is excellent, the defaults are bad. As Simon and other pointed out, medium is great.
Reminds me of Gemma4 and the official (or at least popularly used around launch) Jinja templates being wrong and broken for tool calling.
Seems less like "bad"/broken defaults and more defaults tuned to the max for benchmarks.
All the positive PR from "Opus 4.6 level" online buzz is well worth the minor annoyance from taking a half hour to solve a simple problem since a user just needs to turn down the reasoning knob if it bothers them.
It’s both: The default is bad, but not by accident. Since they certainly chose this default intentionally to be evaluated by it, they entirely brought it upon themselves for it to be evaluated as slow, overthinking and overcomplicating things.
It’s a similar level of dishonesty as trying to conflate “starts at” vs. “tested configuration” car prices.
We change the incentive to do this by evaluating it exactly as advertised.
[flagged]