It needs 3-4 Sparks to run well (at an acceptable quantization and sufficient KV cache):

https://github.com/christopherowen/spark-ds41f

Ah that’s a shame. GLM 5.3 Flash is honestly as good IMO and can run on two pretty successfully from what I understand.

I’m quite spoiled with how good Qwen 3.8 Flash Next is on a single spark though: shocking how good local models are getting on attainable-ish hardware

DeepSeek V4 Flash runs well on two Sparks, I documented that here:

https://blog.jonathanpage.com/

GLM 5.3 Flash runs fine on two Sparks and Qwen 3.8 Flash Next on one is indeed incredible! I made this 3D game with it in two days using Qwen Code as agent:

https://games.jonathanpage.com/