> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable
Yeah, that'd be neat, but that's not what this announcement is about at all:
> With a massive 2.4T parameters
> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable
Yeah, that'd be neat, but that's not what this announcement is about at all:
> With a massive 2.4T parameters
True. It was more of an open letter, with hopes that the Qwen team sees the comments in this thread.
dont we all deem the ability to improve large models as the defacto capability to produce small ones?
I don't think so, they have different constraints and require different optimizations, being able to produce one of them doesn't mean you'll automagically be good at the other.
but that's denying the singularity boostrap theory. Which I don't agree with, but if you can't harness a large model to make a small model, then we're going to have problems brining about the singularity.
If you do think there's some magical singularity, how do you comport?