Even in this very thread the feedback on Kimi's actual efficacy is debated. I personally feel its worse than both Fable and 5.6 Sol, but I feel like the conversation isn't really about whether its good or not, but a backlash against the U.S governments foray into regulation. So I think people _want_ it to be superior out of anger/frustration with the current situation.

It's really good. I'd put it between Sol and Fable. I'm not super impressed by Sol's UI design skills, something K3 is strong at. Fable is still overall the fastest, most consistently well-performing model, though.

This does depend heavily on the kind of work you do and how you use these models, but the idea that K3 isn't right up there with US SOTA models doesn't match my experience.

[deleted]

This seems like a replay of what happened with DeepSeek. They put out v3, or whichever one it was, and everyone said it was over for US companies... then everything continued on.

The markets can be irrational for a while, but if the Chinese models are about 90% of the performance of OpenAI and Anthropic models, and the Chinese companies are < 10% of OpenAI and Anthropic's proposed IPO valuations, something has to give eventually.

This isn't just the AI race, but the end to perceived American exceptionalism (where USA wins by default). It's going to take a while for people to recognize that. Before that the markets will still go crazy, but that's not evidence things will continue on as "normal".

Continued how? I have switched most of my personal LLM coding to DeepSeek V4 Flash since it was launched.

And now 100% to a mix of K3 / DeepSeek V4 / MiMo 2.5.

It's nice not being called a terrorist just because I told it to reverse engineer something.

At work they are still hemorrhaging money to Western providers due to enterprise contracts but I foresee they won't renew for much longer. Specially of the upcoming final version of DeepSeek V4 proves to be Opus+ level.

Anecdotes. Where's the data?

https://openrouter.ai/rankings#top-models

Chinese models dwarf USA models usage. And now there's a Fable/5.6 alternative. The gap widens.

Now go and ask for your GP poster for their data as well. Unless you're only interested in data that supports your bias ofc.

I've seen that, but that's clearly a skewed dataset, right? People using openrouter will be those seeking other models. They are not the same kind of people who mostly just use gpt or claude.

I would agree if openrouter was the only provider of those.

But those LLMs are also offered outside openrouter.

I can't count the number of times I've heard people here say the frontier models are 6 months or more ahead of the open-weights models. That's not true anymore. So the goalposts are shifting.

I feel like that was true up until 6 months ago.

That would make sense, what the US government has done this year with regards to AI is unacceptable

Agree completely.

When you net out across benchmarks and firsthand reviews it seems like it's maybe a little behind. There seems to be a consensus it's token hungry and a little slower. So maybe it's a point release behind.

That's weeks maybe months behind, not months maybe a year behind. It's "would my life really change if Claude was gone, not really" behind.

I actually haven't used it much, because Claude started kicking ass again the last few days. Like, way too much of a difference to be normal load-based variance. I got more done in the last 48 hours than week before that.

So, fuck yeah competition.

Just for the sake of argument and using some admittedly insane numbers, give me Opus 4.5 at a tenth the cost and running ten times as fast and I'd take that for almost any coding task over any current frontier model. There was a real phase transition somewhere in that range and improvements since then, while impressive and useful and by the benchmarks quite large, have in practice not been anywhere near as big a phase change. Honestly until we get to the point where the models don't need any checking at all, incremental improvements on how much checking they need don't do all that much for me. In practice "they get 90%" doesn't differ much from "they get 94%".

> give me Opus 4.5 at a tenth the cost and running ten times as fast and I'd take that for almost any coding task over any current frontier model

This is pretty much where I'm at

I think my main problem with the current 'aligned' version of models is that they're aligned with the very worst of the California social justice warriors.