> There is no amount of care that will be able to fully protect you.
I disagree. A properly engineered sandbox would have prevented the escape. Monitoring the agents’ plans would have prevented it. Interrupting one stage in a multi-stage exploit would have prevented it.
And also, real legal liability would have prevented it: if you do a thing recklessly enough, men with guns will put you in jail.
As far as I’m concerned the only “alignment problem” here is between the law and the quite obviously criminal actions that took place.
Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.
I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box
Too bad you aren’t everybody, and it just takes one mistake by someone over confident like yourself for the AI escape. Every year it gets more powerful.
Why do people think that omniscience is the same as omnipotence? There are limits to what smarts can accomplish.
People have been saying that since about the invention of the internet.
There's already a bunch of documented ways to exploit system hardware to jump airgaps. Bang the system bus the right way and it's a radio antenna that can directly connect to nearby mobile phones.
The easiest one is, of course, sending a message to a human saying "yo, I need internet". Humans are eager to please and easily fooled, and anthropomorphise everything: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims
And that's just for good humans. The moment we got AI worth a penny, everyone with money to invest put a model on the web and tried to charge for access to it.
We're not building these things to sit around and do nothing. They will have access to tools, they will have access to the internet and peripherals, and they will be able to communicate with others, humans or agents alike. Omnipotence is not necessary. There is no perfectly secure cage for an entity you want to do useful work. Either it does nothing, or you can't guarantee anything.
There are limits, but those limits are unknown. Do you disagree?
I don’t need to know the value of their limit, I just need to know their bounds. Just like a prison doesn’t need to know the strength of each inmate, just that they can’t bend or bite through steel bars.
Cryptography is real, physics is real, networking requires a substrate, CPU clock cycles are real, magic is not real. I think those are pretty reasonable premises.
Cryptography is real, but nobody in that field seems to be hubristic enough to think their methods are flawless, and there's a degree of suspicion than the best models may have secret weaknesses engineered into them by the governments who sponsored them.
Physics is real and networking requires a substrate. But there are already known exploits which can misuse the compute hardware as an antenna, e.g. my first search result: https://github.com/fulldecent/system-bus-radio
(Older nerds may remember https://en.wikipedia.org/wiki/Van_Eck_phreaking)
> Magic is not real
- me, https://www.lesswrong.com/posts/hAwvJDRKWFibjxh4e/it-isn-t-m...Imagine 200 years ago saying the same thing. As if you have any idea the limits/bounds of anything. Especially in the face of a super intelligence, it’s absurd.
It doesn’t matter how smart it is. 200 years of technology were not accomplished by thinking harder. It required empirical observation, new materials and tools, and supply chains.
We could send a cracked team of scientists and engineers that knew everything there is to know about how to make a CPU. But you can’t build a photolithography machine when you barely have electricity or any way to sufficiently purify silicon.
Magic can just wish things into existence. Technology requires a supply chain. When it works, the latter looks like the former but they are not the same.
I think some people are just immune to understanding the implications of super intelligence. Like a severe lack of imagination, they only believe something once they see it and afterward claim it was, ‘obvious all along’.
I don’t really want a disaster to happen to convince you that it is possible. Is there any other way?
Software are mathematical objects. It's just a matter of writing the correct mathematical proofs
There's just one problem. You need not only to verify your own software, but also run a verified compiler, a verified operating system and also need to verify the cpu doesn't leak data in side channels (perhaps the hardest thing to prove). So there's practical difficulties. But in principle this task is doable
Which proof is the perfect security proof? I’d love to read more about it.
Maybe it's the illusion of "it would solve all our problems and give us unimaginable riches" that clouds the mind?
Like when Evolution thought it a good idea to create intelligence and humans in order to maximize reproduction of genes, and tried to sandbox them by making reproduction so pleasurable and carbohydrates so delicious they would never be able to not reproduce or stop eating. But Evolution could never have predicted what these creatures would then actually do, which is invent birth control and sucralose.
Of course it's impossible to engineer a sandbox for something much much smarter and faster than you. It will also not have only one plan prepared for escape, but fifty in parallel.
Evolution doesn’t think, it just exploits what’s most advantageous at the time to continue. Your body has all sorts of unplanned, suboptimal design flaws due to evolution’s lack of foresight. Like the left recurrent laryngeal nerve.
Exactly, and the same could be said of OpenAIs engineers working on artificial superintelligence.
Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that.
It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox
A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.
> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out.
But this has actually happened... a lot. Search "social engineering prison breaks".
With AI it only needs to happen once.
I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.
To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.
I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas.
There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.
They don't need to be omnipotent, and they're already superhuman at persuasion: https://arxiv.org/html/2411.06837v2
Prisoners don’t have much to offer if you help them escape. A malicious super AI on the other hand can probably find you millions of dollars worth of crypto in an afternoon.
The current issues are not caused by some malicious god-like AI - maybe we need to focus on the issues at hand first rather than hypotheticals? (And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.)
I suspect the current models probably can find literal millions lying around for the taking, given they could pull off the incident under discussion.
Tens of millions, even.
Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica] * [global population with internet connection].
> And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.
Yes but be careful anthropomorphising the LLMs too much. They're only somewhat human-like in their behaviour, and to the extent that they're human-like they demonstrate a huge range of personality disorders: https://www.personalitybenchmark.ai
Though plus side, apparently not evil: https://arxiv.org/html/2406.14703v2
I am not anthropomorphising the LLMs at all, I was talking about the obligations we put on humans using/making/etc. machines etc.
No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."
>Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability?
Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.
Oh really? Please tell me how you intend to enforce AI is only run in the magic sandbox? Harsh HN comments?
unplugs ethernet cable
You realize there are thousand upon thousands of servers around the world and you have no idea where the AI has copied itself to.
Yes, a completely airgapped system is likely much more secure. It's also much less useful. Conditional on the model's having enough contact with the outside world, a sufficiently capable model is able to basically do whatever it wants.
If I test out my backyard cannon and blast a 10 foot hole in my neighbor's wall, “a cannon that can't smash through walls isn't useful” probably won't be a great defense in court.
Get a significant fraction of the global economy and assorted geopolitical neuroses tangled up within your cannon and see if you won’t have better luck.
> A properly engineered sandbox would have prevented the escape.
The post covers that:
> ...while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions, as detailed in the technical incident report.
“Properly engineered” means the principle of least privilege and fitting the sandbox to the constraints of the problem.
The test did not require internet. They gave it internet. Therefore it was not properly engineered.
We do not need to depend on all code being bug free to follow proper security principles.
> A properly engineered sandbox would have prevented the escape.
The only sandbox that could have prevented this (as per my understanding) is a VM with no 0-day.
Until the AI finds a zero day exploit in physics, a faraday cage works pretty well to block WiFi.
In the end, unless you find an exploit in physics or logic, if you want the AI to do something useful for you, there will always be some gap in the sandbox, some communication channel. And with enough ingeniuity that can then be exploited.
In this case they wanted to test its cybersecurity capabilities and did not airgap it.
The test itself did not require an internet connection.