Yes. V4.1 flash performs really well. I don't know what to say, maybe you don't believe, but my team has been using it mainly for almost a month now for programming. It is as bad and annoying as any of the US sota, but costs pennies. And if you look close enough you find providers that can push it 300-500 tokens per second...

Maybe it is due to us being all very experienced devs. And can steer the model. But my daily routine is just to have 8-9 Zellij tabs open, DeepSeek in omp in each, and grind research and code day and night. Really nice model...

Hold up, 300-500tps at p50? At p99, even something like Opus 5.5 on fast reliably hits 330tps+, so that'd be expected but if truly p50, wow. And no regressions in tool call and structured output vs the official endpoint? Would love to test that with my evals, please share.