If there's any hope for AI sovereignty and equality, we would have to either make expensive models cheap to run or make cheaper models do less work.

Making the latter happen involves either reformulating work in ways less intelligent LLMs can work better with. Or condensing intelligence into smaller models.

I think condensing intelligence into smaller models is the way to go.

Also, “smaller” can mean many different things. The cost is not in storing the weights on disk.

It is in the power required to do the inference with the “active parameters”.

There is increasingly more evidence that those two can be decoupled and more power to those who are pushing on that lever!