I disagree. I own over 1TB of vram at home. I can tell you that it's not a bubble. From my builds, I would rather have cloud, cloud is easier. From running small models like Qwen3.8-27B to large models like Qwen3.8-2.4T. I can tell you that small models will never be enough or match up. Everyone will want the smartest model, not just a good enough model.

> Everyone will want the smartest model, not just a good enough model.

Not so sure about this. There’s always a potential threshold. After all, we don’t all use the most powerful computers, the latest phones, the highest resolution cameras, the fastest or best cars.

I am already not interested in cloud LLMs and I don’t even use the best (on paper) model that I can run locally. I prefer a model that people insisted (here) was “dead on arrival” but appears to work better for me.

I think the difference is that AI as an edge. That edge will turn into more money, better quality of life, etc. Of course, with serious skills, you might be able to use use a not so smart model to keep up with folks with smart models. People are lazy tho, and will prefer for AI to do all the work if it means they do none.

Can't be an edge if everyone has access to it.

The edge is somewhere else.

... and everyone won't have access to it, look at Fable. How many people in the world can afford Fable or are using it?