I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do.
Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
When comparing OpenAI and Claude thats pretty much true, but not Gemini... And have you tried Antigravity? Yikes
The CLI version of agy is great. Have you tried it?
Do you dangerously allow permissions? I absolutely cannot use it until they ship an auto approver. As it is now I have it write one bash/python script to do everything it wants to, then I review that. Otherwise it is COMPLETELY unusable and it shocks me when I hear people are using it.
Sounds like they shipped some changes today that might reduce approvals: https://x.com/antigravity/status/2100001904969297980
is anti gravity open sourced just like codex or grok code?
Compared to gemini-cli that they took out behind the woodshed, I hate it.
I've used Antigravity as my main coding agent on one of my biggest projects for about a year. It's been great for me. (and I use Claude, Codex, Grok and Muse for all the other projects)