Is GPT Luna 6 dethroning Deepseek V4.1 Flash? It's price seem to be undercutting flash at a relatively similar capability.

It is not super good in long horizon tasks and worse than 5.6 in our evals. It really failed the agentic evals where DeepSeek, Kimi and Opus are the winners.

It is great on creating summaries and content.

Yeah, very similar benchmarks at 1/4 the price. I've been pretty happy with Luna 6 though I still think the gap between small and frontier models is larger than many people want to admit.