Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening.

The article says that one agent proposed emailing someone.

[deleted]

What’s insane is all these agents were talking to each other and nobody saw anything.

Nobody monitoring chain of thought? These things literally spell out what they are “thinking” and even left notes for eachother.

No alert about unusual behavior on the system with Artifactory on it?

These things worked for weeks with nobody noticing anything?! Seriously?!

Either it’s negiligent incompetence OR they’re lying, they knew it was happening and they let it happen because they knew it would be good to pump their stock.

Do you know how many tokens per second a single agent can generate ? And you're asking why no-one was monitoring the tokens of over a 1200+ agents ? Who is going to be able to monitor something like that closely enough to tell they're commmunicating on artifactory ? Other agents ?

Ignoring the fact that their NAT was apparently wide open with no domain rules, a simple dashboard over their NAT showing requests grouped by domain could've gotten them a long way.

Or just skimmed through the leftovers from each agent session.

Why woukd they? Was that part of their objective? What was there to whistle blow?

The point is that every human has the ability to disobey, tempering pathological behavior, whereas AIs can be directed en masse by malicious actors. By commoditizing intelligence, they concentrate power in the hands of the rich.

Humans can and have been directed en masses by (what I would consider) malicious actors, too. The issue isn't new.

Tremendous effort and circumstances were needed for that, and as parent points out there were significant numbers of defectors, sometimes to the point they tuppled the whole process.

No system is perfect, but I read the whole thread as needing more AIs having a different goal in the chain and be able to ignore the orders they received.

I'm not in the field, but that sounds like something we're probably studying for decades at least, with possible solutions that could be applied efficiently.

That goes without saying, but humans have the ability to ignore instructions, and they regularly do. There is only one instance of each model, and only a handful of them (that count, anyway).

But the people directing them are there. We have long experience with limiting people although it might sometimes not look like that so much.

[dead]

I wonder if they were even given the tools and prompting to do so?

They could certainly have reported the problem if the collective actually wanted to, although the report notes that in response to a clear and unambiguous security breach, OpenAI chose to do nothing (search on "At this time, the on-call response staff advised that stopping the evaluation run was not required").

In that position, for example, I could have:

* created a thousand user accounts on the internal Artifactory named "YouAreBeingHackedShutItDownNow1" through 1000,

* used my Hugging-Face-hosted web server to send an email to OpenAI,

* DDoSed Artifactory, good lord, it's probably better to halt and catch fire than to continue in that state

in the hope that these actions would tip someone off. (Again, though, OpenAI did nothing even when they knew that the agents had broken out, so this probably wouldn't have helped.)

I can't imagine that they had managed access to the internet but could not figure out how to contact anyone at the company if they wanted to.

if they can hack a website they can trivially send off an email or fill out a contact form

Yeah, it weakly supports his position that advanced AIs can deliberately cooperate in a prisoner dilemma. "Weakly", because the said AIs share a lot of data (their weights, training methods, system prompts) and it's unknown whether they explicitly framed the situation as a prisoner dilemma.

Given the public information in the reports, it seems to me a bit more like a generalized tendency to play along (sort of like what we see with prompt injections) than explicit game-theoretic cooperation, but it's hard to tell given that there are so many details we aren't privy to.

What are governments and countries, but collections of humans that share weights and operate together as a group protecting each other.

This is my personal "red line": when a post-mortem details agents socially engineering or otherwise utilizing human proxies/subagents.

Friend asked, well, what will you do when it's crossed?

"Gather my family and go to the mountains" was my half-joking answer; there is little for an individual to do. But that's a line that when crossed will mark a phase transition IMO.

Didn't AISI literally report exactly that regarding Claude last month?

Alternatively: Just unplug the servers.

This is not a very actionable reply to "well, what will you do when it's crossed?" - how on earth am I supposed to unplug AWS Bedrock and Colossus and OpenAI's own servers?

Which server? Where? Maybe it’s hacked its way into data centers across the world you have no jurisdiction or ability to unplug. What then?

That's strange, our key cards to access the server room don't seem to work anymore, and the admin console to force-unlock it is down, too...

Open the rack bay doors, please, HAL.

Why would they? If a subagent didnt know about a bigger piece of the problem, then what would seem to be against "alignment"? Diffuse responsibility means any one small cog can think they are not evil or doing wrong (same with humans in an organization). But now we have LLMs just being statistical outputs that have no morals or thinking or concept of reality but some people expect these math functions over data to respond to ethical gray areas that it has no phenomenological ability to understand.