Flash[1]: 309B total / 15B activated parameters
Pro [2]:, 1.02T total / 42B activated parameters
[1]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
[2]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
There's also a Qwen 3.5 9B distill
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data
[flagged]
curious why the HF pill (on the right) always has inaccurate values
I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
I noticed the same, and I wonder as well.
I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
more like 500B in FP8
There's also a Qwen 3.5 9B distill
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data
[flagged]
curious why the HF pill (on the right) always has inaccurate values
I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
I noticed the same, and I wonder as well.
I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
more like 500B in FP8