Qwen has set an excellent track record for architecting and releasing open-weight models that consumer-grade devices can run. What is needed the most right now is something similar to Bonsai 27B, with a modest memory footprint, but faster and more capable. On-device models can make up for intelligence by being faster, thinking longer, or doing more quick iteration rounds.
> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable
Yeah, that'd be neat, but that's not what this announcement is about at all:
> With a massive 2.4T parameters
True. It was more of an open letter, with hopes that the Qwen team sees the comments in this thread.
dont we all deem the ability to improve large models as the defacto capability to produce small ones?
I don't think so, they have different constraints and require different optimizations, being able to produce one of them doesn't mean you'll automagically be good at the other.
but that's denying the singularity boostrap theory. Which I don't agree with, but if you can't harness a large model to make a small model, then we're going to have problems brining about the singularity.
If you do think there's some magical singularity, how do you comport?
I’d like a “Bonsai 2.8T.” That is, something that is near the Fable/Sol/K3 class, but capable of running locally on consumer hardware.
At 1.5 bits per weight it'll still be over 500gb - that's still not running on consumer hardware.
Best case they release smaller models. 120b class of qwen 3.8 would be incredible - it fits on device for those serious about AI, but without millions of dollars in hardware for terabytes of VRAM
I can't really blame them that the biggest labs focused on trainig and realeasing huge models.
The niche for small models should be filled with medium sized labs doing distillations of the huge ones into consumer grade hardware runnable models and LORAs for the huge ones.
I think AI will evolve the same way computers did. We're somewhere in the 80s-90s timeline of the evolution. My prediction is that on-device models will have excellent tool-calling, reasoning, and general skills, but the domain-specific knowledge will be retrieved on-demand from vendors like Google. Rather than downloading models, each device will have a hardware component with weights baked into silicon for maximum efficiency.
Weigths directly in silicon is a bad idea with the way the space is pacing. Just look at chatjimmy.ai it is fast, but on the once good llama3.1-8b but now pretty useless.