Maybe others have found otherwise, but I find the benchmarks drastically different to real world "feel" of a model, even within the same harness. I'm not sure if this just reflects personal interaction styles, or if it is indicative of benchmaxxing or unrealistic automated benchmarking methodology.
Opus 5 consistently comes at or near the top, but outputs constant unreadable jibberish. Meanwhile GPT 5.6 Luna medium tends to be rated pretty poor on agentic tasks compared to the Chinese lab open models, but I find the latter much more likely to lose track of their own behaviour during a long-horizon task or get stuck in a doom loop.
(This isn't a comment on GLM-5.3 Flash as I've not used it!)
Opus is a yap god. I've found it much, much better with Claude Code's output style set to `Concise` and this:
https://news.ycombinator.com/item?id=49413456
We shouldn't have to resort to this, but it can be mitigated enough that it stays as my daily worker agent. Although, I mostly use Fable to farm out to Opus agents so I don't have as much exposure to what kind of blathering is going on in there.
It's crazy that sometimes I ask Opus 5 to explain what it just wrote to me, and it declares "that was word salad" (its words, not mine, without any hint from me other than "explain it").
I've been using 5.3 since they initially announced it and my gut feel is that it's not as good as 5.2 for agentic tasks. I'm still using it - I don't think it's bad. I'm just not convinced it's better.
idk i think that i spend significant tokens with both to be able to tell 5.3 is way better overall.
https://kommodo.ai/i/IFSUUQYT522uZXWvePZ1
it still fells stupid sometimes and it is benchmaxxed for sure. but its good enough that im building all the hobby projects with it.
I have had the opposite where I felt 5.3 as a stronger model than 5.2. Its feels way more in tune with my code, and does more nuanced edits. Though I do handhold my models a lot, so might fall outside the agentic term.