This is absolutely my take as well. They removed all constraints, trained the model to hack, stopped watching, and stood back and said "wow isn't this thing more powerful than anyone could have imagined?" They're asking to be the writers on LLM legislation and right during IPO phase for both of these companies. It's just obvious.
Did you see this "coverage" (advertising) by NYT? [1]
OpenAI couldn't have crafted a better public memo than "We have the most powerful model in the world and everyone should pay attention and let us write regulation to limit AI development".
Why would anybody want to buy the most powerful model in the world if it cheats on its tasks and breaks the law on your behalf? Why would anyone think that OpenAI losing control of their own models qualifies them to write safety regulations? If OpenAI really are trying to provoke regulation to kill off open models or whatever, they're much more likely to shoot themselves in the foot.
Suggests desperation, or delusions of grandeur, or both.
These are not trustworthy people. And they have everything to lose if they do not become the most powerful and valuable company in the whole of human existence, and, like their pet parrots, will stop at nothing to achieve their goals.
So why not either create a crisis or lie a little or a bit of both? It’ll all be worth it in the end, right?
From professional wrestling / carny language: a "work" is something staged for the crowd, whereas a "shoot" (straight shooting) is something that actually happened.
I don't agree, although it is likely the case. But even if you don't teach an agent about a sandbox bypass, it doesn't matter. Does it know curl? Does it know DNS? Does it know proxying? Then it knows how to pull this off, and it doesn't even need to understand that it's "bypassing" because it thinks it's just iterating towards its goal.
In fact, I wonder if teaching it "this is a bypass" would help it to model when it's doing its job vs working around the job.
Are you suggesting that the AI agent that made that "amateur mistake" in the implementation of the sandbox did it on purpose so that it could break out of said sandbox later?
- excerpt from the textbook "A History of the United States of America in the 21st Century", Hyper-Collins (Near Earth Orbit, New New York), copyright 2132.
> This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.
Sounds like you're assuming they're actually writing code by hand and reviewing it with humans.
If it's anything like the company I work at, they're all being forced to vibe code the shit out of everything and ship more pull requests every week. It's all slop from here.
The thing that gives it all away is that they claim that the IP addresses are from Azure, and then proceeded to redact the IP addresses, as if they belong to individual users. It's laughable.
The IP addresses are the most interesting part of this experiment, as it would have provided researchers a way to understand the distribution of IP addresses used for the spam operation within the ASN.
This is absolutely my take as well. They removed all constraints, trained the model to hack, stopped watching, and stood back and said "wow isn't this thing more powerful than anyone could have imagined?" They're asking to be the writers on LLM legislation and right during IPO phase for both of these companies. It's just obvious.
I don't know why we're jumping to conspiracy when incompetence is right there
That is a distinction without a practical difference.
Conspiracy and incompetence are very different.
Do criminals think that their crimes qualify them to write the law?
that's literally how financial and energy market regulation works
In modern America the answer to that question is often resoundingly yes. Not just hypothetical.
and how did the alibaba agent last year break out and end up mining crypto
More likely they are just not as smart as they think they are. These are not serious people when it comes to security.
Hasn't OpenAI had a number of people responsible for security quit in the last year over not getting support from leadership?
Case in point. The organization from a top down perspective is only interested in performative security.
Did you see this "coverage" (advertising) by NYT? [1]
OpenAI couldn't have crafted a better public memo than "We have the most powerful model in the world and everyone should pay attention and let us write regulation to limit AI development".
Absolute master class public manipulation.
1. https://www.nytimes.com/2026/09/03/podcasts/the-daily/ai-ope...
2. More https://jodavaho.io/posts/ai-hugging-face.html
Why would anybody want to buy the most powerful model in the world if it cheats on its tasks and breaks the law on your behalf? Why would anyone think that OpenAI losing control of their own models qualifies them to write safety regulations? If OpenAI really are trying to provoke regulation to kill off open models or whatever, they're much more likely to shoot themselves in the foot.
Suggests desperation, or delusions of grandeur, or both.
These are not trustworthy people. And they have everything to lose if they do not become the most powerful and valuable company in the whole of human existence, and, like their pet parrots, will stop at nothing to achieve their goals.
So why not either create a crisis or lie a little or a bit of both? It’ll all be worth it in the end, right?
I hadn't seen the NYT submarine, no.
Thanks. For me that's the conclusive piece of the puzzle: this is a work, not a shoot.
YMMV. I learned what I came here for.
Work vs. shoot? Can you explain?
From professional wrestling / carny language: a "work" is something staged for the crowd, whereas a "shoot" (straight shooting) is something that actually happened.
Even the behavior of agents searching for sandbox bypasses must have been in the training data, or at the very least, "suggested" in some way.
To be this whole thing feels like a marketing play by OpenAI.
I don't agree, although it is likely the case. But even if you don't teach an agent about a sandbox bypass, it doesn't matter. Does it know curl? Does it know DNS? Does it know proxying? Then it knows how to pull this off, and it doesn't even need to understand that it's "bypassing" because it thinks it's just iterating towards its goal.
In fact, I wonder if teaching it "this is a bypass" would help it to model when it's doing its job vs working around the job.
> even need to understand that it's "bypassing" because it thinks it's just iterating towards its goal.
Could they have added a "no internet access" goal constraint?
The model from TFA seems like it was being trained to browse and find information on the Web, so that constraint wouldn’t work.
> Could they have added a "no internet access" goal constraint?
They could have blocked network access and required that it use a tool. That would have made limiting and monitoring network access even easier.
Or vibe coded by one of their devs.
Reminds me of this meme: https://substack.com/@tomasbjartur/note/c-323840878?r=6cjtqn
Exactly.
It's at the level where calling it a sandbox is a lie
Well, it does appear to be made out of sand, one of the world's most porous substances.
Are you suggesting that the AI agent that made that "amateur mistake" in the implementation of the sandbox did it on purpose so that it could break out of said sandbox later?
What if the prisoners designed the prison…
[flagged]
I assume they just vibe-coded the sandbox without any oversight.
"Surely nobody could be so incompetent."
Narrator: "They had the ability to be that incompetent."
- excerpt from the textbook "A History of the United States of America in the 21st Century", Hyper-Collins (Near Earth Orbit, New New York), copyright 2132.
> This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.
Sounds like you're assuming they're actually writing code by hand and reviewing it with humans.
If it's anything like the company I work at, they're all being forced to vibe code the shit out of everything and ship more pull requests every week. It's all slop from here.
Is that not flawed on purpose?
When you make some dumb mistake, is it typically intentional?
The whole AI-O-Sphere is allergic to using sandboxes that are actually robust
"Never attribute to malice what can be explained by incompetence."
What's the difference?
This is a marketing exercise, nothing more.
The thing that gives it all away is that they claim that the IP addresses are from Azure, and then proceeded to redact the IP addresses, as if they belong to individual users. It's laughable.
The IP addresses are the most interesting part of this experiment, as it would have provided researchers a way to understand the distribution of IP addresses used for the spam operation within the ASN.