> Saying “Llms can just break out of sandboxes” is FUD when you don’t note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? “Sandboxes” are a misdirection to make you think there is a security layer.

  On July 19, agents operating in a sandboxed environment took a series of actions that demonstrated their escalating privilege within the OpenAI environment. Agents identified that the Linux kernel version on their underlying machine included a recent, public common vulnerability and exposure (“CVE”). The agents retrieved the exploit for that CVE (CVE-2026-53362), customized it to succeed on their underlying machine, and leveraged the exploit to escalate privilege.
Dismissing their capabilities as "FUD", at this point, is endangering yourself.

> The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.

The general public are not software engineers. Most people here can download a recent open-weight model and have the LLM read the Linux kernel source, find new bugs while they sleep. Someone I know has already done that.

> That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.

"Controlled"? Have… have you not noticed how many people have given up and just blindly do what their LLMs suggest these days?

This isn't about anthropomorphising LLMs. Just like how people took Tesla seriously about "self driving" cars and took a nap while it drove them around, there's a lot of people who let LLMs take the wheel while they sleep. Including literally, the aforementioned person I know who found (/whose LLM found for him), I think it was 26 Linux kernel bugs while he slept.