It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
I keep saying “I’d be so
Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.
The most visible leap between 4.6 and 5.5 seems to be that the latter has gotten much more computer-use training, so there's a clear progression in the ability for the model to use Blender. But catching up on that is just a matter of training on the same thing.
Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)
Seconding this. Flash-next (and let's not forget, it's a PREVIEW of the 4 architecture - with the "real" 4 rumored coming later this month) is the first model I can run on reasonable hardware (2 thoroughly obsolete P100s off ebay at ~$100 each plus the RAM I could scavenge from other PCs at home) at a reasonable speed (22, with GPUs in layer-split due to llama-cpp's limitation on qwen4-exp arch, and no MTP. Strata could double these numbers).
It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.
(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)
Running an nvidia card at full load, would cost me ~100 euro of electricty each month (europe). Of course one wouldn't have usage caps.
Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.
Even with a country. Here in Norway we have multiple price zones, and at times there can be 100x difference between them, often 10x. All due to lack of transmission capacity between northern and southern zones.
We are already there. Prior to Strata the best I could run was Qwen3.8-27B at Q6, which itself is already at like Opus 4.5/4.6 level, and now with Strata on an R9700 and 64GB of RAM I can run Qwen3.8-Flash-Next IQ3_XXS at 60 t/s. It's even better. You can run it on even more modest hardware with Strata too.
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
That is right now. This model is easily as good as Sonnet5 / Opus 4.7 on DeepSWE. I have benched over half of DeepSWE now on a 3 bit Flash Next quant and it is at parity with Sonnet and Opus 4.6/4.7. It finishes most of the tasks they do. Overall it is within 1 point.
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
Secure systems are possible, but now we have a compelling reason to actually write them. Everything can be trivially hacked because the industry is pathologically adverse to security being part of the design process.
I'm hopeful that this security situation is just short term turbulence and we come out the other end with companies taking security seriously, updating things on time, and writing code in more secure languages and frameworks.
Nothing is stopping it. This model is woefully bad at accurate creative red teaming however. GLM finetunes on the other hand are pretty good. And I'd bet they are already deployed and doing all sorts of deeds.
if the alternative is all human intelligence is cucked by 2-3 amoral American labs then we've had a good run, don't care.
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.
But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).
Opus 5.5 was a step function change. Really crushes on multi-hour coding compared with prior models.
Highly agree.
I keep saying “I’d be so Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.
The most visible leap between 4.6 and 5.5 seems to be that the latter has gotten much more computer-use training, so there's a clear progression in the ability for the model to use Blender. But catching up on that is just a matter of training on the same thing.
I would be forever happy with Opus 4.8.
Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)
Seconding this. Flash-next (and let's not forget, it's a PREVIEW of the 4 architecture - with the "real" 4 rumored coming later this month) is the first model I can run on reasonable hardware (2 thoroughly obsolete P100s off ebay at ~$100 each plus the RAM I could scavenge from other PCs at home) at a reasonable speed (22, with GPUs in layer-split due to llama-cpp's limitation on qwen4-exp arch, and no MTP. Strata could double these numbers).
It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.
(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)
Running an nvidia card at full load, would cost me ~100 euro of electricty each month (europe). Of course one wouldn't have usage caps.
Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.
> (europe)
Given the massive difference in electricity price between different european countries, adding "Europe" doesn't bring much context.
Even with a country. Here in Norway we have multiple price zones, and at times there can be 100x difference between them, often 10x. All due to lack of transmission capacity between northern and southern zones.
These GPUs use about 400W at full tilt (200W each), for reference.
We are already there. Prior to Strata the best I could run was Qwen3.8-27B at Q6, which itself is already at like Opus 4.5/4.6 level, and now with Strata on an R9700 and 64GB of RAM I can run Qwen3.8-Flash-Next IQ3_XXS at 60 t/s. It's even better. You can run it on even more modest hardware with Strata too.
Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.
Opus at home is a thing now.
Another 2-5 years from now, it may even become accessible to most people (current barriers being cost and technical expertise, bottleneck is cost).
The mentioned “Qwen3.8-27B at Q6“ can be run on a <1500$ Mac mini and with the barest technical expertise.
That is right now. This model is easily as good as Sonnet5 / Opus 4.7 on DeepSWE. I have benched over half of DeepSWE now on a 3 bit Flash Next quant and it is at parity with Sonnet and Opus 4.6/4.7. It finishes most of the tasks they do. Overall it is within 1 point.
FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)
Dream or nightmare?
In face of the recent Hugging Face incident we should really be concerned about the security implications.
What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.
We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..
https://huggingface.co/blog/security-incident-july-2026
- OpenAI hacked Hugging Face
- OpenAI models refused to help Hugging Face during incident response
- Hugging Face turned to GLM, who helped in the defense
That pattern repeats over and over. https://www.felonybench.com/
You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.
Secure systems are possible, but now we have a compelling reason to actually write them. Everything can be trivially hacked because the industry is pathologically adverse to security being part of the design process.
I'm hopeful that this security situation is just short term turbulence and we come out the other end with companies taking security seriously, updating things on time, and writing code in more secure languages and frameworks.
Nothing is stopping it. This model is woefully bad at accurate creative red teaming however. GLM finetunes on the other hand are pretty good. And I'd bet they are already deployed and doing all sorts of deeds.
if the alternative is all human intelligence is cucked by 2-3 amoral American labs then we've had a good run, don't care.
my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.
[flagged]