Absolutely! DeepSeek-V4-Flash-0731 has become my daily driver. It's pretty amazing what it can do for what it costs at deepinfra.com (I don't use deepseek as a provider since they train on your data [at least their honest about it]). GLM-5.1 was my daily driver before that and Kimi K2.5 before that.

Are you finding DS better then kimi k3 and glm-5.3? Do you mind sharing your primary use case?

My primary use is AI coding agent. Its vastly cheaper than Kimi K3 and I haven't found a scenario where I really need Kimi K3 versus smaller models. GLM-5.3 Flash is good but there is series of bugs in the vllm middleware that prevent GLM models from getting all of their reasoning content returned to them that impairs inference quality. A lot of inference providers use vllm which makes it hard to find a good provider for GLM. I've been using friendli.ai but using GLM-5.3 Flash from them is more expensive then using DS V4 Flash from deepinfra.com simply because deepinfra.com is so cheap. The DS V4 Flash cost at together.ai is similar to the GLM-5.3 Flash from friendli.ai or at least that's what I found in my benchmarks a week ago: https://www.linkedin.com/posts/joshheitzman_i-ran-a-fuller-r...

I've used Kimi K3 for a few months as my main model and DeepSeek 4.1 is as fast and about 10x cheaper.

I just had like four big sessions going today, paid about $8 in tokens. I see no reason to pay more, this is more than I need for intelligence.

4.1 consistently surprises me in capability for the price. And I don't think I'm the only one. It's been dominating the leaderboard at OpenRouter, and I just got an email today from Fireworks saying they were _raising_ the price by about 30%. I'll probably switch, because their infra doesn't support being the highest-cost, but it's still telling.

How does it compare to 4.1 flash? Curious why folks don’t use the more “modern” one.

4.1 flash is very fast and capable. Token efficiency is not great so it fill up context window much faster compared to similarly capable models.

glm 5.3 flash is a tad slower but a bit more capable and way more token efficient.

Source: self hosted tested on rented GB200 node at 8bit.

Wow, I'm surprised you are saying GLM 5.3 Flash is more capable. Isn't is like half the price of 4.1 Flash?

I haven't tried 4.1 flash as I'm assuming its a preview. I did not get good results from the preview version of 4.0 flash (i.e. the one that did not include the month and date of release in its name).

4.1 Flash is a horse of a very different color. It cooks. IMHO it's probably a preview of DS5, rather than a true DS4-series model.