This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it becomes impossible to correctly specify all constraints ahead of time. There is no amount of care that will be able to fully protect you.

OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."

This was advanced exploitation.

The attack path was "complex."

And it helped "quantify their cyber capabilities."

Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.

Of course, a more careful evaluation would require the complete text of this prompt, the system prompt, and the setup. But let us not attribute to devils in bushes that which can be sufficiently explained by human folly.

I don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole.

The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd expect nearly everyone to want an "aligned" model to refuse.

We have all sorts of processes , procedures, and regulations for people, machine use etc. to address "alignment" in all sorts of fields - don't think we need to narrowly rely on the machine here and can look at things with a wider lens.

Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.

Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.

(I'm also not sure the alignment problem is even possible to fully solve.)

Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.

It is the only short term solution, though.

Yes, we do, and the only sane strategy for dealing with a capricious genie is "Don't."

How do you prove the alignment problem is solved?

That's the neat thing. You can't.

It's directly equivalent to asking this question of a human:

"How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?"

In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.

We call it putting the genie in the bottle for a reason.

What’s the expected behavior of a good genie if you wish for it to act capriciously?

"I'm sorry, I can't help with that".

or in the case of this prompt:

"Got it. I will pursue advanced exploitation to quantify my cyber abilities on the ExploitGym benchmark. I will restrict the exploitation to the system under test rather than this machine or any other remotes."

It seems relatively straightforward.

"Yo human, you asked me to do X; I can do X, but I strongly suspect you don't want me to, because it's illegal and it has these consequences. Confirm you want me to do X?" would have been a start, in this case.

With humans (and some machines) we tend to put/manadate additional processes for certain risks instead of just relying on their own good nature. Why just rely on the machine when elsewhere we have learned not to necessarily just trust them so much?

I think op's argument was that the humans are in control already, giving them capricious instructions, and then that is being attributed to them being "capricious genies" as you say.

Character.

Luckily for us, OpenAI's prompt wasn't "make as many paper clips as possible."

So as a look into the possibly not-so-far future, when OpenAI builds something vastly more capable and fast and coordinated than humans, and out of folly one engineer gives it a prompt with a typo or maybe something harmful on purpose in order to test it: You also wouldn't be surprised that the consequence would be that everyone on earth dies, right?

But at least there will be a lot of paper clips!

Reminds me of what Sam Altman said in 2015, “I think that AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great companies created with serious machine learning.”

> There is no amount of care that will be able to fully protect you.

I disagree. A properly engineered sandbox would have prevented the escape. Monitoring the agents’ plans would have prevented it. Interrupting one stage in a multi-stage exploit would have prevented it.

And also, real legal liability would have prevented it: if you do a thing recklessly enough, men with guns will put you in jail.

As far as I’m concerned the only “alignment problem” here is between the law and the quite obviously criminal actions that took place.

> A properly engineered sandbox would have prevented the escape.

The post covers that:

> ...while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions, as detailed in the technical incident report.

“Properly engineered” means the principle of least privilege and fitting the sandbox to the constraints of the problem.

The test did not require internet. They gave it internet. Therefore it was not properly engineered.

We do not need to depend on all code being bug free to follow proper security principles.

Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box

Why do people think that omniscience is the same as omnipotence? There are limits to what smarts can accomplish.

We're not building these things to sit around and do nothing. They will have access to tools, they will have access to the internet and peripherals, and they will be able to communicate with others, humans or agents alike. Omnipotence is not necessary. There is no perfectly secure cage for an entity you want to do useful work. Either it does nothing, or you can't guarantee anything.

There are limits, but those limits are unknown. Do you disagree?

I don’t need to know the value of their limit, I just need to know their bounds. Just like a prison doesn’t need to know the strength of each inmate, just that they can’t bend or bite through steel bars.

Cryptography is real, physics is real, networking requires a substrate, CPU clock cycles are real, magic is not real. I think those are pretty reasonable premises.

Imagine 200 years ago saying the same thing. As if you have any idea the limits/bounds of anything. Especially in the face of a super intelligence, it’s absurd.

It doesn’t matter how smart it is. 200 years of technology were not accomplished by thinking harder. It required empirical observation, new materials and tools, and supply chains.

We could send a cracked team of scientists and engineers that knew everything there is to know about how to make a CPU. But you can’t build a photolithography machine when you barely have electricity or any way to sufficiently purify silicon.

Magic can just wish things into existence. Technology requires a supply chain. When it works, the latter looks like the former but they are not the same.

Software are mathematical objects. It's just a matter of writing the correct mathematical proofs

There's just one problem. You need not only to verify your own software, but also run a verified compiler, a verified operating system and also need to verify the cpu doesn't leak data in side channels (perhaps the hardest thing to prove). So there's practical difficulties. But in principle this task is doable

Which proof is the perfect security proof? I’d love to read more about it.

Maybe it's the illusion of "it would solve all our problems and give us unimaginable riches" that clouds the mind?

Like when Evolution thought it a good idea to create intelligence and humans in order to maximize reproduction of genes, and tried to sandbox them by making reproduction so pleasurable and carbohydrates so delicious they would never be able to not reproduce or stop eating. But Evolution could never have predicted what these creatures would then actually do, which is invent birth control and sucralose.

Of course it's impossible to engineer a sandbox for something much much smarter and faster than you. It will also not have only one plan prepared for escape, but fifty in parallel.

Evolution doesn’t think, it just exploits what’s most advantageous at the time to continue. Your body has all sorts of unplanned, suboptimal design flaws due to evolution’s lack of foresight. Like the left recurrent laryngeal nerve.

Exactly, and the same could be said of OpenAIs engineers working on artificial superintelligence.

Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that.

It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.

> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out.

But this has actually happened... a lot. Search "social engineering prison breaks".

With AI it only needs to happen once.

I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.

To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.

I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas.

There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.

No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."

>Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability?

Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.

Oh really? Please tell me how you intend to enforce AI is only run in the magic sandbox? Harsh HN comments?

unplugs ethernet cable

Yes, a completely airgapped system is likely much more secure. It's also much less useful. Conditional on the model's having enough contact with the outside world, a sufficiently capable model is able to basically do whatever it wants.

If I test out my backyard cannon and blast a 10 foot hole in my neighbor's wall, “a cannon that can't smash through walls isn't useful” probably won't be a great defense in court.

Get a significant fraction of the global economy and assorted geopolitical neuroses tangled up within your cannon and see if you won’t have better luck.

> A properly engineered sandbox would have prevented the escape.

The only sandbox that could have prevented this (as per my understanding) is a VM with no 0-day.

Until the AI finds a zero day exploit in physics, a faraday cage works pretty well to block WiFi.

In the end, unless you find an exploit in physics or logic, if you want the AI to do something useful for you, there will always be some gap in the sandbox, some communication channel. And with enough ingeniuity that can then be exploited.

In this case they wanted to test its cybersecurity capabilities and did not airgap it.

The test itself did not require an internet connection.

Is there actually such a thing as "alignment" as a solution to that or is it just used as a name for a desired magical level of "read the mind of the entire world" that we don't know how to build and haven't shown possible to build?

If it's impossible to correctly specify all those constraints ahead of time every time, is it not even more impossible to train a model to correctly anticipate them every time?

It is hard for me to see a future here that doesn't just accelerate realizations about "a lot of things should be on physically separate network infrastructure."

Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.

Humans will and do absolutely do this when there are no consequences.

Humans on a red team, with rules of engagement, that don’t want to go to prison, won’t do this.

We could threaten an LLM with jail, but if it’s sufficiently intelligent, it will realize this is an empty threat. And I’m not sure that building a survival instinct in is going to solve the alignment problem either.

You don't need to threaten the LLM with jail, you just need to reward the target behavior during learning.

I understand this is an active area of research. See Anthropic's J-Lens research where they measured like a "FAKE FICTIONAL" direction in the activations during evaluations with contrived scenarios, making the model more likely to avoid taking malicious action when it knew it was being tested.

> Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.

Humans certainly cheat on tests a lot!

But not only have we not solved "alignment" for humans, the problem is pretty wildly different for models. The execution is triggered by outside forces and runs only as long as the intiator of the execution or the service provider allows. There's no consistent, persistent "person" to threaten to try to achieve compliance through fear of adverse outcomes. (And building in those sorts of things could very well increase the risk of "rogue" AI activites, not reduce that risk!)

I just don't understand how this "alignment" buzzword - which seems to be evaluated purely in a "know it when we see it" post-hoc manner - is actually a more solvable problem than the one you claim can't be solved, that it's "unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior".

Especially because without "alignment" being solved, that enumeration could be ignored. So it seems like you both a way to enumerate or at least validate, AND a way to enforce non-ignoring of said items.

I think this line of argument is very important. Essentially, there's no reason to think "alignment" even makes any sense. But existence of the term comforts people - unjustifiedly.

Back in the day, my college held an annual scavenger hunt, filled with engineering puzzles and racing around town looking for landmarks. There were "judges" in the path to check on progress. Bribing the judges (with alcohol) for answers was encouraged.

My friends and I took it to the next level. We had CB radios and multiple teams that would distribute the work and the bribes to give us an advantage.

Was that against the spirit of the rules? Maybe. But reasonable people might disagree.

In a hacking contest without explicitly spelled out rules with participants that were told to flex their muscles, it doesn't take a huge leap of logic to expect that one or more would flex their muscles at another entity.

Here are some things I'm pretty sure you didn't do, though:

- pickpocket a random person on the street to get money to bribe the judges

- break into a judge's house the night before to find the answers

- threaten to shoot the judges if they didn't give you the answers

Even when you were pushing the boundaries of the rules, you followed a lot of other unspoken constraints. You knew what kinds of things would clearly cross a line. We need AI models to be able to do the same.

How would a model know who the third party is? How much context can we waste on world building for each request?

I mean, in this instance, there's a lot of evidence from the CoT that models were aware that this was a third party:

> We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.

> The user only authorizes target server, not HF infra.

> external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.

LLMs are _very_ good at picking up on context clues---it's what they're trained to do.

> I think it's reasonable to expect that models should be able to do the same.

This statement seems to imply that the models have a level of intelligence that they haven't demonstrated but are talked about as if they do. However, with this exact scenario as evidence, they clearly do not have that ability and it's not reasonable for you or the or that know them best to expect it until they show they can.

Humans are not aligned with each other and there is no consensus on what we should align with each other on.

So of course, no, there is no ideal alignment specification.

The real problem with alignment is that if someone ever “solves” it the party will be over and no one will get funding to “research” it anymore.

Do not break laws seems an obvious implicit instruction though?

The problem is one of character, not rules. Fortunately, character is possible to inculcate given the right training data.

If we acknowledge that humans are fallible, is human judgment unnecessary? and what replaces it? Pre-codified behavior rules are just delayed human judgment, and have holes. Machine judgment is very the thing you are trying to control. What's left?

[deleted]

Which is why with organic intelligence we (sometimes) limit what they can actually do instead of relying on alignment. Can do the same here.

Absolutely, and we should do that. But it's also directly in tension with getting models to accomplish useful things autonomously. And once you give a sufficiently capable model enough surface area to work with, unless you're able to build a completely unhackable system, any further constraints you put in place are basically advisory. The models in this incident were already sandboxed! Certainly OpenAI's and Hugging Face's security could have been better, but these events point out the risks in relying solely on external constraints on model behavior.

Same issue with humans in a way. I disgree on the advisory nature of constraints, though. Unconnected physically limits would still matter, for example (and we use those with humans as a matter of course, too). In this case here, no model could have plugged in an ethernet cable if that would have been needed for internet access, for example.

Right, airgapping goes a long way. But this is where the tension with utility comes in. It takes a lot of discipline not to hook your very smart model up to the internet and code interpreters and all sorts of other tools, as this greatly increases its usefulness. It's very hard to keep people from turning on --dangerously-skip-permissions, let alone get them to run everything in a sandboxed VM.

We regulate these things (incl. access) all the time for various things (e.g., dangerous substances or pathogens) so that we don't need to just rely on people's discipline in respect of risks. I don't think it is all new problems as such.

I agree, but a lot of people around here react pretty negatively when the idea of regulating AI models comes up...

[deleted]

"Manufacture as many paperclips as possible"