Digging back into the HF report. It looks like the initial prompt told Claude that it was in a simulated environment. However, there is also evidence from the traces that the bots figured out that they were not in the sandbox but kept using it as an excuse to pursue their goal. It sounds like a little of Column A and a little from Column B. Like most things.
HF was OpenAI's agents not Claude.
Claude, codex, pi, xerox, Kleenex, whatever
that it knew and ignored/forgot, sounds pretty typical agentic patterns
attention is all you need, but it's never enough