I understand people just get off posting stuff like this. But creating LLMs from the entire corpus of human text was a huge achievement. Distilling those models is much less of an achievement. It means China is further behind than we thought.
I understand people just get off posting stuff like this. But creating LLMs from the entire corpus of human text was a huge achievement. Distilling those models is much less of an achievement. It means China is further behind than we thought.
They're using competitor output as additional training input. Shrug.. It's not like they have access to the weights.
First movers rarely take the prize, though, do they? Les Paul and Mary Ford pioneered overdubbing voices back in the 1940s and 50s. The Beatles stole it from Buddy Holly. And Elton John from the Beatles.
Distilling is a massive achievement. I can run Qwen. I can't run GPT (TM). It's not a matter of X is better than Y. It's a matter of Y exists, X does not.
How is China further behind if distillation cannot stop? I think it's a reasonable strategy to follow, even if they could train from scratch.
distillation results in a worse product than the actual teacher model iirc
Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.
Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.