China has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you.
What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.
Something I don't think many have internalized is that China has been as good or better for quite a while now (long before anyone was pointing distillation fingers) and enough people have finally tried it for themselves that the understanding has reached critical mass and the careful narrative of american companies is collapsing.
When I finally put $15 into Deepseek and it beat the brakes off Codex 5.5 on multiple rather complex projects without any of the obnoxious mistakes, I was sick to my stomach with buyers remorse. I couldn't believe I ever felt like I was getting my moneys worth at $200/mo. I wouldn't even use OAI's models if they were free and unlimited at this point, I'll happily pay for what I already know works. No reset bingo, no cache errors, no annoying shitposters as a primary source of info. Oh, and I still had $10 of tokens left
And yes, 3.6 is excellent locally. The rest of this year is gonna be awesome
Another potential takeaway is that the models all gathering around the same point supports the idea that there is a ceiling to LLM capability.
They always all gather around the same spot then that spot moves every 6-9 months. I think the clustering is more likely evidence of distillation. I don't personally think distillation is a bad thing. If the LLM providers can distill all of human output into their models for 'free'. I don't think distilling a model from the output of those models is morally wrong.
I gathered that the most recent advances haven’t been in capabilities of the model but more the way that it’s able to be employed (most recently agents).
What an interesting take. One question, do you think stealing from a thief is morally okay? I'm just asking no judgement on my side.
I'd say it's more "Downloading LimeWire Pro from LimeWire" than actual theft.
If the legal system declares the first thief’s theft not theft then all bets are off.
> If the legal system declares the first thief’s theft not theft
But they didn't find it. The Big LLM provider accepted guilt and paid a fine.
You can argue whether it was a fair amount they paid, but there is no legal precedent that was set. It's still considered theft.
As i understand it, they accepted guilt for downloading stuff illegally. They didn’t accept guilt for incorporating all of human output into their model without consent.
Is it theft if another thief steal's the first thief's theft?
Don't say that too loud, you may burst the bubble prematurely.
the ceiling is to eval's quality
Qwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.
I find 35B A3B viable as well, but your harness and runtime really matters to get tool calling and such dialed in. In fact, I would encourage you to experiment with it some as I find I get more reliable output from 35B A3B, though 27B is still generally smarter. A3B with a review cycle or two from 27B is great for me.
One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B model, but it is close enough it can usually figure it out with the right tools.
Works great with room to spare on my lenovo pgx too
I'm still skeptical of the smaller models after the talent exodus a few months ago.
At this point I feel like the only factor differentiating SOTA models now is who they’re propagandizing you on behalf of (not considering agentic tooling/state management, etc).