I always find it confusing that a meaningful volume of the comments are saying "this reached parity with SOTA models. Best $/task."
And a meaningful chunk of the comments are saying "this piece of garbage isn’t even at the level of gpt-oss 20B".
I always find it confusing that a meaningful volume of the comments are saying "this reached parity with SOTA models. Best $/task."
And a meaningful chunk of the comments are saying "this piece of garbage isn’t even at the level of gpt-oss 20B".
I was in the first group up until last week.. now, the second one.
For anything even moderately complex.. like, even low end of complexity, this model behaves maximum like gpt-5.6-luna-high .. nothing more.
Yesterday itself I gave it a coding task in some existing moderately complex small project, and i was using xhigh thinking effort, it was unable to cover all edge cases... and i had already got it to review, and then fix, 3 more times, after the first initial one.
Still it left 2 edge cases.
Then, reverted full code, gave sol-high the same task, it took well over 20 minutes, and completed it in one go with zero edge cases remaining.
I am not using it for anything serious anymore.
I guess both are true and for everyone at some point. All models, even SOTA, fail. When they fail, it is quite frustrating. Additionally, some models are very cheap to run and use. When Deepseek fails the cost was minutes and pennies.
Yours is the only mention of GPT-OSS in the whole thread. I count about 2-3 comments saying the model is meh and many more saying it’s a big step up.
I am among those with real life experience with the model that used the previous as well and will attest that the new model is a big improvement