GP kinda explains why: "You have to lead it by the nose".
If you're happy to "lead it by the nose" you can do well with a lot of very low end models.
If you want to kick off a "/goal run until [complex verification passes]" and let it run for a week with minimal intervention, then not so much.
You can make do with cheaper models for long running agentic runs too, but it tends to require a lot of extra scaffolding and additional review steps.