Without a human to benchmark against it's really tough to gauge how good these models are vs how good the task definitions and existing codebases are.

My intuition from the example full instructions are that the tasks are poorly specified which results in ~60% failures due to bad assumptions and missing requirements.