what??? not true!
for inference the compute is the last thing we need more of.
memory bandwidth is the numebr one blocker, after that the inefficiencies that where introduced with MoE models (and all new large models are made that way)
Here is a quick read: https://news.ycombinator.com/item?id=49324600