Why are models better than agents, isn't it supposed to be the opposite? I don't understand the difference and what you are measuring.

Some agents have specific tools that the models have been trained to use. E.g. diff formats for editing that aren't the same as the standard unified diff format. Access to specific thread / subagent / etc. tooling or the base prompt can perhaps also impact how the tasks are completed.

[flagged]