People will claim it's not comparable to Opus despite it beating the score. I'm not sure I disagree, but I'm also unsure whether I care. Most new models nowadays are "good enough". I cannot complain because I'd rather spend that time improving my prompts and docs. Opus might be a _slight bit better_ at picking up vague hints, but it's also extremely expensive, and I hit the 5 hour limit way too quick.
I care a lot about speed and efficiency right now. For my setup I would like to have 2-3 different model families. I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting, Deepseek V4 Pro 0813 for developing, and Gemini flash lite (any recent cheap model) for repo scouting. I'll add another one in the mix for reviewing (in this case Gemini 3.7) and that's all I need.
I've tried most models except Grok.
Qwen is too expensive IMO (Alibaba Cloud subscriptions are hard to come by and I'm not spending 50 euros a month for a tool, so 18 euros it is). If it ever becomes efficient enough to run locally I will definitely look back.
Claude is slow and expensive (the cache hit prices are absurd).
OAI is pretty good, I might add it to my arsenal seeing how cheap it is.
These opinions change every day. Last week I would've never picked Deepseek until I read about the pricing. even post aug 16 it's worth it (although it's getting close to gemini pricing).
Right now my costs are 12 euros a month (z.ai) + whatever deepseek consumes. This typically isn't more than 8 euros a week. 44 euros a month and I have a setup that is doing pretty well.
> I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting
Dude, GLM-5.3 released _today_.
The phrasing "I've settled on" is incorrect for this context.
hence the "former deepseek v4 pro". I tried it out this morning and have had no complaints. I already liked glm 5.2
> Deepseek v4 pro 0813
Which itself released yesterday? You're writing, reading, and evaluating enough software in a ~36 hour period to form, reject, and form another opinion about which model makes better _architectural_ choices?
The sentence still doesn't make sense, because "settled on" implies a long testing phase with a verdict eventually emerging out of that.
What you're currently doing is "testing out"
Honest question, how do you assess models this quickly? What metrics are you using? Would love to get my suite from multiple days and hundreds of prompts down to minutes. Got a few first pass tasks I run upon release for an initial experience, but those only work because even Fable and Sol fail despite objectively correct solutions existing, so it works because most models fail, but then, those are consciously not enough for coding, tool use, adherence or task specific inference and assessment…
What are you working on? That can dictate which models are best.
You should check out Grok, it's quite a good deal from the Cursor subscription side but it's cheap even by API prices.
[dead]
No serious person or sane person uses the LLM that's constantly being tweaked by an anti-woke white-genocide-supporting weird little man. Don't feed the totalitarian wannabe's (or the totalitarians in general, for that matter).
Yeah I'm sure everyone on r/cursor or in previous HN threads about Grok 4.5 or 4.6 are all unserious and insane.
No one actually cares about the politics as long as the model codes well.
Edit, quite interesting to see the reception to this comment compared to essentially the same type of comment I made on a Grok 4.6 benchmark HN post: https://news.ycombinator.com/item?id=49275385#49275571
It's true that Cursor gives a lot of usage with Grok, most users of Cursor don't care about Musk.
I would say it is sad that there are people who use Grok when there are so many other choices available which don't come with the issues of supporting Musk.
It is not all just 'politics'. Take a stand on some issues. It doesn't cost much not to use Grok.
There are a lot of people who are apathetic to what musk is, people that don't care are not people who should inspire you. What the hell is so inspiring about apathy anyway?!
And yeah, people that don't care DO make the world worse through their apathy.
This is the “Mussolini made the trains run on time” of ai hot takes.
(Btw, mussolini didn’t make the trains run on time)
"Sure I'm indirectly funding the erosion of basic human rights in the States, but at least I made my Hello World app cheaper!"
Also: "What do you mean everybody at this party is a Nazi? They have made me feel so welcomed!"
They are all unserious and insane, yes.
i care about not financially supporting a person that is actively trying to disenfranchise me, why is that a difficult concept for some people? that not everyone is motivated exclusively by financial profit? is moral bankruptcy so pervasive that some people assume it is unanimous?
i think you will like luna if you haven't tried it yet
Luna is twice the price of Deepseek V4 Flash 0731, and less capable :/
I'm convinced a lot of the anti-open-weight model comments at this point are inorganic traffic - there's trillions in investor money riding on a world where these models aren't cheap commodities. Having actually used things like the recent GLM, Kimi, and Qwen I think any edge the labs have is marginal at most and actually prefer the open weight models in most day to day usage.
Anthropic's recent releases are wordy to the point of exhaustion. Every time I use opus recently I find myself wanting to yell "GET TO THE POINT" at a terminal, which is exacerbated by it being slow.