Complaining about overthinking in xhigh then pointing out output had bugs with thinking turned off seems like it’s missing the obvious compromise?
Complaining about overthinking in xhigh then pointing out output had bugs with thinking turned off seems like it’s missing the obvious compromise?
Agreed. Test driving the bad default is what they deserve (they brought this upon themselves as a benchmaxx attempt), but comparing it to running with reasoning completely disabled is also weird.
I think the point is that a larger model with far less reasoning time solves the problem just fine. Which of course is the tradeoff: The smaller the model, the more reasoning you need to get decent answers to tough questions.