Flash[1]: 309B total / 15B activated parameters

Pro [2]:, 1.02T total / 42B activated parameters

[1]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL

[2]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL

There's also a Qwen 3.5 9B distill

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B

Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?

It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data

[flagged]

curious why the HF pill (on the right) always has inaccurate values

I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.

I noticed the same, and I wonder as well.

I suspect they are calculating something in the weights or config, I see it pretty consistently with quants

more like 500B in FP8