I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?

6B activated weights per token vs 27B. Something like DGX Spark is way better suited for Flash Next.

It is on paper, but crazy enough, both models at NVFP4 run similar speeds for decode! The reason is that much more sophisticated speculative drafting is available for 27B. I’m hoping this will come to Flash Next, but I know MoEs pose challenges with that.