I checked out the development benchmarks for agents, and was suprized to find that they're not really representative of many of my own day-to-day development workflows.

It appears that they're mostly testing the ability to make business tasks autonomous, with ~20% associated to development tasks (ssh here, install this, etc.), but not actual programming.

That's probably because the actual coding benchmarks were saturated several years ago.