> Beats Opus 4.7 Max
I'm a huge open model fan, and have used them since forever, even have daily drivers for on-prem dev, but no. They do not beat opus on real-world usage.
Qwen models are impressively good for what they are, are "good enough" for plenty tasks, can be ran locally on decently priced hardware, and so on. They certainly have their uses, and the field in general has advanced faster than my early expectations. But to compare a 27B model to SotA behemoths from a few months ago is doing everyone a disservice, especially people who pick it up, try to use them just like API models, and leave disappointed and confused. Number goes up on a benchmark isn't it.
> They do not beat opus on real-world usage
We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios (mostly Rust, some C for microcontroller stuff). Qwen3.6-27B scores only 4% lower for pass@1, n=250 compared to Opus-4.8.
For the labeled dataset, the average PR size they're being measured against is around 1.5k SLOC.
This is very much "real-world usage" for us. The sort of change sets that come in daily/weekly and are solving non-trivial issues in the respective codebases.
As is usually the case, the most broad claims from both the labs and from the consequent pushback are talking past each other.
Let us know when you have Qwen vs Qwen comparison stats. As long as there's not a regression, that'd be awesome.
4% is within the margin of error anyways for pass@1, so I think pass@k > 1 is gonna be the better indicator of any movement (still need to calibrate the optimal k to re-test). 10 seems too tolerant even though that tends to be the next tranche I reach for.
Depends on where you sit on the binomial curve. At p=0.04 for n=250 4% points would not be within margin of error.
[dead]
[dead]
>> We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios
Okay but the parent said real-world usage, presumably meaning coding tasks.
We have a whole bunch of complex evals that Haiku 4.5 passes. That doesn't mean it is a good model for coding.
Yes, these are coding tasks in the embedded systems domain (I mentioned Rust and C).
How much does it score though? 0% would be 4% less if Opus was at 4%. Unless you mean relative fraction not percentage points - but people usually mean percentage points in such situations.
0% is not 4% less than 4%, that would be 3.84%.
0% is 4 percentage points (pp) less than 4%.
[dead]
> ...but no. They do not beat opus on real-world usage.
I agree, but then we just need meaningful benchmarks that clearly show that! Otherwise it's hand waving about something that should be put on paper in quantifiable terms.
If you are working in a company and using language models, it is a very good idea to hold a bunch of evals you can trust and use to validate new models. Calibrate every once in a while with prod data. We have our own and the only numbers on quality and cost I trust come from this setup.
A wise man once said, "not everything that counts can be counted, and not everything that can be counted counts".
In the end, the only benchmark that matters is your own.
Only useful benchmarks are those you (in particular) don't have access to.
The only useful benchmarks are those you've created for your specific workflow. Only then can you assess whether a given model is better or worse for what you are using it for.
There are tools like promptfoo designed for this.
> but then we just need meaningful benchmarks that clearly show that!
That's the rub. AI benchmarks are IMO, by and large totally unreliable. We think of them as similar to traditional benchmarks of deterministic processes where the number of variables is low. But they're anything but that. Non-deterministic processes with an astounding number of variables and fuzzy acceptance criteria.
It leads to results like these, where if you take it at face value, the only conclusion you can draw is "wow Anthropic must be stupid if Opus takes 1T parameters to do what Qwen can do in 27B."
How can you say this when you haven't even tried it yet? Is it just hypothetical vibes?
Yep. These small models are actually worse than GPT 3.5 at some tasks (like recalling facts). You can definitely make models smarter at specific tasks (like tool calling, coding) but you can't compress the entire human knowledge into a 30GB file. It's just not enough bits.
But why would you use a model to store factual knowledge, that is stupid. We want intelligence, not a database.
Models cant make any decisions if they have 0 idea that the feature exist in this language or in some general fact.
For example you would tell a model hey, become an expert in this language for me, search it online, it would still need to learn it and download the data to it's context and then increasing the memory usage, there's no way around it.
that is why we enable web search for the agent. the memory can come from the internet.
deepseek-v4-flash needs web search to return true facts.
Yeah but results are worse, see https://news.ycombinator.com/item?id=49301574
There is 0 shot you can make that claim about this model you have not used or downloaded yet
"Benchmark is stupid" and "model beats model on benchmark" are two different things, though. The second one is objectively true regardless of your views on the first one, right? To expect everyone to share your opinion that benchmarks are stupid is pretty weird, and just saying "no" to an objective truth is the definition of delusion.
If a benchmark is a measure of nothing useful, then model beats model is an objectively useless fact