I don't see how it makes sense to use models outside of Grok or Composer in cursor unless you like the tooling to such degree that you're willing to pay out the nose for access to OpenAI and Anthropic to use with that tooling. And Codex is regularly scoring towards the top among harnesses, so you might as well use the best harness available for the given model. I'm happy with Grok and Composer in Cursor. If Cursor wants to add more models, they should host more open models.

The reason it makes sense is because different models are better at different things. Allowing yourself to be locked in to a particular model by these companies is a bad idea and will lead to worse outcomes for yourself and the market as a whole.

Because I pay $60 per month and not using the included API tokens on other models is wasteful.

Is there a good place for harness comparison/scoring?

https://www.harness-bench.ai/leaderboard.html

This provides a method, but the data looks stale and perhaps a bit thin compared to say, Cursor, or even AntiGravity data.

never heard of nanobot. how reliable is this benchmark?

I can’t speak to the veracity of the benchmarks but it appears their methodology is sound. Nanobot has 47k stars, fwiw https://github.com/HKUDS/nanobot

It has been more of an OpenClaw or Hermes alternative than a coding agent like OpenCode or Pi, so it’s likely to do well given less context bloat.

IMO it's not. It's benchmarking GPT 5.4 and Opus 4.6. It's also missing Claude Code... one of the most popular harnesses (the most?)

No Pi, no Aider either.