It would all be more convincing if the incidents so far didn't seem to be facilitated by an outrageous level of negligence.
We had OpenAI "accidentally" run an entire swarm of 10,000 agents apparently for weeks, on a security related task, seemingly totally unsupervised, hacking all over the internet - all the conversations were completely visible, anybody who looked would have seen it. But they didn't.
So before we start regulating innocent parties, maybe let's start by taking some direct action against the specific ones that appear to be behaving with criminal levels of negligence.
The "sandbox" they used was apparently made of thin paper exposed under a day of heavy rain, too. You'd think, if they truly believed the model is so dangerous, they'd run it in a VM without a network adapter.
I brought this up to someone else and was told that airgapping is apparently much more expensive than I'd naively think.
I still think this is a sign that they are not taking their own rhetoric seriously.
> I brought this up to someone else and was told that airgapping is apparently much more expensive than I'd naively think.
These labs are one of the most valuable and heavily funded enterprises in the whole world, that they can't properly air-gap their systems to me reads as if they "agents" and LLMs are not as good as they say they are, because if they were, why would it be hard/expensive to air gap a system? They already scraped most if not all of the internet, where did that data go?
Not that airgapping is expensive so much as it's really, really inconvenient once you take it seriously. You need to build special rooms for it, you can't just API out to a datacenter. You need to have processes for requesting data be sent into the box. And so on.
I feel like there is a reasonable compromise between "yeah they have full internet access" and "separate airgapped rooms that require multiple levels of authorization to access" that would make this a lot better without that much more work. I feel like they're doing it intentionally to show how dangerous these models are and that the government must step in and protect them
it's really weird to hear frontier labs say "our internal models are basically AGI" while also saying "airgapping is too hard uwu".
if your internal models are so damn good, they should be able to "one shot" airgapping... right?
Agents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images..
Not surprised this is always what they have and hack.
Who would use an Agent that spends $10,000 re-implementing some OAuth lib or reverse-engineering a proprietary lib when it's free on the internet?
You don't need a full air gap. Set up a microVM with network access limited to local network and send all package requests through a filtering gateway that only allows normal download endpoints. Or self host a big collection of popular packages if you need extra security.
Isn't that exactly what they did? The bots could only access the jfrog instance, so they hacked jfrog?
No that's not what they did, they exposed jfrog raw. It would have been so extremely simple to gate services they need the llm to access... I mean, jfrog was not written with this kind of threat model in mind, and neither were a lot of other tools
Right, you mean it didn't go through a gateway? But would that actually have helped? The requests all went through jfrog didn't they? I guess it depends on the level of filtering at the gateway?
Whilst it might not be JFrog's threat model, I wouldn't assume it can be used as a full internet proxy.
I don't really mean to defend OpenAI here, but they did make some attempts at sandboxing. Although it does seem that they didn't really know what they were doing.
There was and continues to be no reason to share the package manager between models. This was begging for abuse.
> Agents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images..
Yes, yes they do, but read through artifact proxies are dodgy as fuck, which is why and facebook (and I assume a fuckload others) don't have them.
Also semi-airgapped labs are a lot less expensive than you think at that scale. Once you have to do multi-region VLANs with machine certs before you get access to juicy VLANs, the difference between "no internet for you" and "mostly airgapped" falls to almost zero.
Also I would want an artifact mirror because a) that give a good signal about how the model reacts, and what training material its latched onto, b) it hides what the models are doing from the outside.
It's expensive if it wasn't part of the planning and design. The same as 'security' is expensive, or compliance with regulations is expensive.
It is also a choice to not do any or all of the above.
> You'd think, if they truly believed the model is so dangerous...
They would have been watching what it does, especially when running it on ExploitGym of all benchmarks... that is criminal worthy neglegence
yes, that is the kicker
These same people who supposedly believe these agents pose an existential threat to humanity apparently fired up 10,000 of them and left them unsupervised for weeks.
Look at the post-incident investigation: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
While I do think OpenAI were negligent in not developing the harness that would allow to understand better what's happening close to realtime, I'd say "anybody who looked" in that case would probably be someone with another swarm tasked with analysis, it's no longer "glanceable" in a traditional sense.
I don't understand why hugging face is not getting more shit too. It is extremely embarrassing to get owned because you are letting arbitrary programs/users call out to the open web from the infra
Sounds like advertising platforms. Spraying malware and links to scam sites all over the place.
"They" don't care about the end-people. "They" care about maximising their profit thing, in a vacuum.
[dead]