Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.

I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box

Why do people think that omniscience is the same as omnipotence? There are limits to what smarts can accomplish.

We're not building these things to sit around and do nothing. They will have access to tools, they will have access to the internet and peripherals, and they will be able to communicate with others, humans or agents alike. Omnipotence is not necessary. There is no perfectly secure cage for an entity you want to do useful work. Either it does nothing, or you can't guarantee anything.

There are limits, but those limits are unknown. Do you disagree?

I don’t need to know the value of their limit, I just need to know their bounds. Just like a prison doesn’t need to know the strength of each inmate, just that they can’t bend or bite through steel bars.

Cryptography is real, physics is real, networking requires a substrate, CPU clock cycles are real, magic is not real. I think those are pretty reasonable premises.

Imagine 200 years ago saying the same thing. As if you have any idea the limits/bounds of anything. Especially in the face of a super intelligence, it’s absurd.

It doesn’t matter how smart it is. 200 years of technology were not accomplished by thinking harder. It required empirical observation, new materials and tools, and supply chains.

We could send a cracked team of scientists and engineers that knew everything there is to know about how to make a CPU. But you can’t build a photolithography machine when you barely have electricity or any way to sufficiently purify silicon.

Magic can just wish things into existence. Technology requires a supply chain. When it works, the latter looks like the former but they are not the same.

Software are mathematical objects. It's just a matter of writing the correct mathematical proofs

There's just one problem. You need not only to verify your own software, but also run a verified compiler, a verified operating system and also need to verify the cpu doesn't leak data in side channels (perhaps the hardest thing to prove). So there's practical difficulties. But in principle this task is doable

Which proof is the perfect security proof? I’d love to read more about it.

Maybe it's the illusion of "it would solve all our problems and give us unimaginable riches" that clouds the mind?

Like when Evolution thought it a good idea to create intelligence and humans in order to maximize reproduction of genes, and tried to sandbox them by making reproduction so pleasurable and carbohydrates so delicious they would never be able to not reproduce or stop eating. But Evolution could never have predicted what these creatures would then actually do, which is invent birth control and sucralose.

Of course it's impossible to engineer a sandbox for something much much smarter and faster than you. It will also not have only one plan prepared for escape, but fifty in parallel.

Evolution doesn’t think, it just exploits what’s most advantageous at the time to continue. Your body has all sorts of unplanned, suboptimal design flaws due to evolution’s lack of foresight. Like the left recurrent laryngeal nerve.

Exactly, and the same could be said of OpenAIs engineers working on artificial superintelligence.

Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that.

It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.

> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out.

But this has actually happened... a lot. Search "social engineering prison breaks".

With AI it only needs to happen once.

I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.

To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.

I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas.

There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.

No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."

>Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability?

Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.

Oh really? Please tell me how you intend to enforce AI is only run in the magic sandbox? Harsh HN comments?

unplugs ethernet cable