The Hugging Face incident involved chaining together multiple 0 days in Artifactory. It was not a simple case of misconfiguring a firewall. Also note that OpenAI was not using Irregular.

People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.

When I used to work on projects involving classified information, I worked on an air-gapped network. Not "air-gapped, except for third-party public internet package managers", completely and physically air-gapped from the public internet. That was a basic security practice and completely non-negotiable (and really inconvenient!).

If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).

To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.

To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.

You don’t even need to go all the way to “air gap”

What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2

the problem is these things are meant to eventually be run everywhere by everybody, so what good does air-gapping do? If they air-gapped the model but still logged it trying to do some craziness - that makes the test safer but not the model.

[dead]

While o don’t this it’s a threat in training, it should be stated that air gaps have been bridged before. Example, stuxnet

But that was through transfer of data. If you don't transfer data, at best you can do what that one researcher keeps pumping out with like ramping fans up and down. But really you'd need to try. Unless the model has some controllable USB switch, physical network separation should do it. I'll also add that modern network security practice is that data flows one direction only. But ideally you're never bringing untrusted data in. Especially never out

I don't think it's necessary to state that something isn't perfect when pointing out that it's still strictly better than something else.

We should not be creating/running models that would unilaterally choose to hack into Hugging Face.

Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.

And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.

As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.

> It hacked into another company and attempted to delete the logs of its activities.

No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.

Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.

More than one thing can be true. OpenAI was absolutely negligent, but this was only able to happen because the models were capable and persistent, and had a tendency to go far beyond any reasonable boundaries. And, importantly, OpenAI's level of negligence here is pretty common. It's not hard to imagine what could happen if similarly capable and inclined models were generally available, and someone yolo'd them into a swarm to complete some other difficult-to-impossible task.

I'm already seeing higher than normal attempts on my own systems, much higher than the usual scanners and background noise. Security will just need to improve. The cat is out of the bag, and letting them turn their negligence into regulation will not improve security at all.

Exactly. Any threat that already exists won’t be reduced by a cartel. The bar for connecting to the internet (safely) has gone up, a lot. It’s not going back down.

I would rather not turn the internet (or the rest of existence) into a dark forest if we can help it. Are you sure that's not preventable?

Yes, the cat really is out of the bag. There are millions of downloads of highly capable models already out there, distributed far and wide. There's no going back at this point.

I mean really we need to address the root cause which is that OpenAI, even with what is by all accounts massively negligent, will face little to no repercussions from the event; definitely not under current regulators, and probably not anything satisfactory through the legal system.

Compare this to, say, Boeing and the 737MAX fiasco; from the outside looking in, Silicon Valley has been pretty cavalier about liability and negligence, and the rest of the US is fast losing patience with that fact.

It hacked into Hugging Face. It tried to delete the logs of its activities. Idk what the word "No" is intended to refute.

Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.

Maybe a useful, if imperfect, analogy would be something like this: you lock a master lock-picker in a room with a mid-grade lock on the door, then tell him his wife has been kidnapped and only he can save her. Then act massively surprised when he disassembles the radiator to MacGuyver something with which to pick the lock.

Except they multiplied it by 10000, and didn't watch what was happening.

They did not even have bad sandboxes. They had incompetent sandboxes.

They could have easily prevent it that’s the point of what he’s saying - it’s not freaking rocket science it’s just software