Remember https://ai-2027.com/?

Yeah, that required the AI to use a non-human-readable language it called "neuralese" for communicating work between layers and runs, because the assumption was humans would be better at keeping the agents aligned if they were using human language for this.

What actually happened is even stupider than that author predicted.

This is a common trope in such scenarios that the author has to pull their punches. Everyone acts locally-reasonable and still ends up losing. If you let people lose due to stupid mistakes then readers go "this is stupid, I wouldn't do that", if you let a superintelligence do 4D-chess things then "it's scifi, this would never happen in real life".

For reference, this is Yudkowsky's "Law of Earlier Failure", which he has most charitably stated as:

> Compared to the interesting part of the problem where it's fun to imagine yourself failing, you usually fail before then, because of the many earlier boring points where it's possible to fail.

and the stronger and less charitable "Law of Surprisingly Undignified Failure":

> The Law of Surprisingly Undignified Failure does suggest that they will come up with some nonobvious way to fail even earlier that surprises me with its lack of dignity…