We’re already seeing incredible knowledge compression out of the models the GP mentioned like Qwen 35B-A3B, which feels well within the realm of “runs on a phone” in the next handful of years.

And by then we’ll probably have been further surprised by just how much information and capacity for reasoning can be crammed into a few gigs of weights. Models just keep getting better for a given size, it’ll be interesting to see where the limit of that is.