Am I correct in understanding that this is just 730B-ish parameters as an MOE? That sounds like incredible performance per parameter. The new Deepseek was also very impressive with its 280B or so. Plus the most recent 30B-ish Qwen and Muse.
I find the performance to size ratio of these models to be way more interesting, selfishly because it makes me bullish on what I'll be able to run on a machine I own over the next few years. The progress is just incredible.