except besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience

I have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it just answers and does what I want.

This is from somebody who put thousands of dollars every month to Opus. Now it's 40% of that and I get as good or better results without having to turn the caps lock on before lunch...

Edit: yes company money. We don't get subscriptions we pay per token.

Opus 5 is the least reliable frontier-class model in the market

In what way? It has worked well in my experience. It holds up with long context windows, unlike many, too.

Eh, I use Opus professionally and DS v4 Flash for personal work. I honestly don't notice the difference too often other than Flash being twice as quick and an order of magnitude cheaper.

The reality is most work people do doesn't need the very cutting edge and these open weight chinese models more than cut it most of the time.

[deleted]

Honestly I really like GLM 5.2 a lot for coding. There’s some weird failure modes in Anthropic’s models where it just does absolutely idiotic things.

You'll find that hard to prove objectively and conclusively.