Fwiw people have told me that GPT2 doesn’t qualify as an LLM at 550MB despite being one of the first LLMs.

So the practical answer to your question is: not much.

I don't think some people are aware that "large" has always referred to the training inputs not resulting the size of the model.

LLM is not "we made a language model and it is big" -- it's "we trained this model on a lot of language"

Not aware of that etymology