What would the benchmarks be if GLM5.2 and Kimi K3 be if they didn't use distillation of frontier models?
I don't think any of the Chinese AI companies have stolen IP from the US AI companies (at least, I've seen no evidence of it), but the evidence of distillation is pretty apparent.
My issue with this is that if distillation occurs frequently enough, it is going to zero out most SOTA research into AI. Having cheap AI is great, but that alone won't advance the state of the art. There needs to be groups that are pushing the boundaries, and unless some kind of protection is put into place, there will be zero financial incentive to do so if anyone can come along and effectively steal your model and get financially rewarded for serving it far cheaper than the original group can, because much less R&D cost is needed. I say this as someone who sees distillation as "legal", since if you can train anything you can look at, and you can look at the output of those frontier models, then you can train on them.