I've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time.
More or less, yeah. I've found mild success with deepseek-v4-flash though, and also Qwen3.5-122B-A10B-NVFP4 running locally, especially in terms of "doesn't overthink every single prompt" and somewhat reasonable quality. Really wishing for a 3.8 update of the 122B variant, that'd be really competitive (for local usage) :)
I've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time.
Same problem with every chinese model currently, they overthink way too much and take too much tokens and time.
More or less, yeah. I've found mild success with deepseek-v4-flash though, and also Qwen3.5-122B-A10B-NVFP4 running locally, especially in terms of "doesn't overthink every single prompt" and somewhat reasonable quality. Really wishing for a 3.8 update of the 122B variant, that'd be really competitive (for local usage) :)
A consequence of aggressive distillation?
[dead]
Why? Price? If the reason is performance, I've been using non-frontier models for cheap, and they run great for my needs (GLM 5.2, DeepSeek v4 Pro).
We'll probably use it, but for images. Qwen is still the best for images
What's the price difference?