6B activated weights per token vs 27B. Something like DGX Spark is way better suited for Flash Next.

It is on paper, but crazy enough, both models at NVFP4 run similar speeds for decode! The reason is that much more sophisticated speculative drafting is available for 27B. I’m hoping this will come to Flash Next, but I know MoEs pose challenges with that.