So, when you said “proven” I thought you might have meant like, a mathematical proof. It appears that this is not what you meant. (Right?) (If it is what you meant, I was asking for like, the particular theorem.)
Also, ok, when I said “intelligence”, it was because you were already talking about how “smart” the model could be. So, I thought you were already on board with using the word “intelligence” to refer to the phenomenon where these kinds of models produce outputs that satisfy the kinds of tasks they are pointed at.
None of those links give an argument that the current 1GB models are the best they can be.
My understanding is that so far when training a model by distillation (using the logits of the teacher model), one can achieve better outcomes than one could if training the 1GB model from scratch on the same training data as the large model, and that so far, better models as the teacher model have yielded better results for the student model.