> current frontier models need laborious oversight and guardrails on even the simplest task

As models advance, we shift the goalpost for what "simplest task" means. Before, "simplest task " meant "write a coherent English sentence." Now, "simplest task" means autonomously fix, review, and merge a bugfix.

Eliza wrote coherent English sentences.

And you know compare Eliza with what an LLM can do today?

Or do i miss the point you are trying to do?