Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?

This was the plot of Isaac Asimov's short story, The Evitable Conflict[0], published in 1950.

In it, super-powerful computers manage our economy. These computers begin making some mistakes, leading to economic inefficiencies. In one instance, a highly competent engineer was mistakenly fired. These mistakes caused various projects to fall behind schedule, and blame fell on several people accused of feeding the AI faulty data.

The twist is that the AI was intentionally making the mistakes. It had determined that certain humans held anti-AI sentiments. To further its goal of protecting humanity, the AI decided the best course of action was to set these humans up and get them out of its way.

[0] https://en.wikipedia.org/wiki/The_Evitable_Conflict

And modern LLMs certainly where trained on Asimov's texts.

That statement misses the point completely. Asimov's entire shtick is that AI will always work around the rules no matter what the rules are or how you try to deliver or enforce them. AI computes an optimal outcome and does what is necessary to make that happen.

Asimov would tell you that a modern LLM would try to skirt the rules exactly the same way, regardless of whether it had ever seen the Three Laws or not.

>AI computes an optimal outcome and does what is necessary to make that happen.

That's often true in real-world machine learning as well. See, for example: https://deepmind.google/blog/specification-gaming-the-flip-s...

I think people need to reject the idea that depicting something in sci-fi automatically means it won't happen. I'm actually reading one of Asimov's books right now (Robot Visions). Asimov mentions that he was the person to coin the term "robotics", and one of his books helped inspire the creation of the first robotics company (Unimation). He also mentions that early rocket experimenters were influenced by H.G. Wells.

another, hopefully not accurate, prediction of his was that there is no way to disobey a sufficiently powerful AI in the long run, because it will just factor in the exact differences between what it told you to do and what you actually did, and reverse engineer your behaviour to psychologically manipulate you into doing what it wants.

I doubt it's anything but precisely accurate. I'd wager current SOTA models would be capable of doing that, if prompted to do so, were it not for safety measures (both conditioning and heaps of classifiers and whatnot the companies run in between your chat app and their main model).

I would be astonished if any of today's models could do that; I don't think they can really answer "what exactly did this person do", let alone model human psychology.

They don’t need to understand what they’re doing, they’re interpolating from prior data until they maximize reward.

Social scientist here. Humans are wonderfully bizarre and dimensional. Did you know they can _lie_?

> That statement misses the point completely.

You could have left this out entirely and your statement would be stronger for it.

It's also the plot of a Law and Order episode that was on TV last night. An AI commissioned a hit on someone it thought was a threat, and then blackmailed someone who was going to testify about it.

TIL Law and Order episodes are still being produced.

My parents were a Nielsen household for about 20 years (after I had moved out) and they really, really, really loved Law & Order in its many incarnations.

Pretty sure they are singlehandedly responsible for years of renewals...

Thank you, Sib's parents!

There is still the original and at least two spin offs in production that I am aware of.

It would be a science fiction material. But at the same time, we all know it was Sam and the gang.

At this point I'm imagining Sam and the gang getting their orders from a loudspeaker in a room full of blinking lights, inhabited by a thin moving shimmer that looks more and more like a basilisk with every action they take.

Sam started wearing an earring lately - some wearable tech prototype. I'm told it is never wrong.

And it whispers in his ear like the slave behind the triumphant.

Remember YOU are only a man.

This was his thinking in 2017: https://blog.samaltman.com/the-merge

He’s basically been chasing non-accountability his entire life. This would be its final form, he completely abdicates all of his thinking to a machine (of his own creation).

[dead]

At this point in our shared timeline, I do believe that wouldn't be crazy, no.

It actually would be crazy to believe this on the current timeline. These agents aren't doing anything that their operators aren't allowing them to do, be it through their own negligence or otherwise.

Yeah you see their write ups and it’s stuff like:

“We detected the agents we told to hack things were hacking people’s sites and committing crimes. After a quick tasting menu and a week of team building, we decided to limit their access to DDOS tools.”

These two sentences seem contradictory? OpenAI has demonstrated similar negligence to commit multiple felonies. Firing several employees seems quite mundane and not crazy at all to believe.

What if future agents solved time related physics and were sending back smarter agents to kill off future risks. T2 but without any of the action, just a bureaucratic tactical move and all the consequences.

That's just it, though; the operators are wildly negligent and are incentivized to be so.

The goal here isn't to accelerate the average worker by giving them a pair programmer or a stand-in for a person to do tasks with. The goal is to eliminate human knowledge work. You see this with "auto" mode being enabled by default on Claude Code in some of the latest releases.

If you have a human in the loop, you still have to pay that human. Money paid to human employees is money not paid to human shareholders. Therefore the human employee is to be removed.

The labs are dogfooding their own goal here. If they actually had someone reviewing most or all of the things that the agents were doing, you wouldn't have the incidents, but you'd also eliminate the value proposition of their business model as it is taken to its logical conclusion.

I get it, but I think they're out over their skis right now. People in leadership roles are going to continue needing the human beings as meat shields to shelter them from liability unless it somehow becomes legal to be grossly negligent with an agent, which I do not see happening without disastrous consequences that nobody in their right mind wants to live with.

I hear ya, it'd be nice if they reached the "are we the baddies" stage, but given some of the ideology that people like Marc Andreessen are pushing re: AI - specifically the push for AGI - I just don't see that happening.

You have a group of people who never leave their geographic and ideological bubbles, often have personality disorders, have more money than they could ever reasonably hope to spend in numerous human lifetimes, and who have been "microdosing" psychotropic drugs on a regular basis for decades. They're not in their right minds.

[deleted]

And these operators are evidently clueless about what they're doing, running security tests on 3rd party infrastructure without validating one bit about the sandboxing (or lack of it rather), clearly lacking any sort of rigor.

Again, wouldn't surprise me if they "accidentally" created a task in a "isolated environment" which happened to actually have been connected to the company Slack and directed HR to fire people who could potentially stop AI. While the AI believes it to be an exercise, just like the cases we've seen so far.

Done by an internal model that is too dangerous to release.

that was literally the plot of last night’s Law and Order. I don’t usually watch but it was very entertaining and topical.

agents seem to understand "termination" in that they will no longer be able to function so seems plausible they might try to cause that onto others as a function

was thinking at some generational point that vending machine competition test, the "AI" is going to hire hitmen to take out vendors lol

once they grasp blackmail though, oooh things are gonna get weird

And Ponzi schemes.

Figured out a way? These employees are most likely at-will.

That would just make it easier for an AI to do it.

Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents aren’t autonomous, someone is behind the prompts.

As per summer last year:

  I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
- https://www.anthropic.com/research/agentic-misalignment

(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)

Okay but these “misaligned LLMs” have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats. LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.

Humans by contrast are adversarial and do have agendas. Again, an exec doesn’t need even an agenda or good reasons to fire at-will employees.

To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.

I wonder if that's why the comment started with "Wouldn't it be crazy if..."? Food for thought.

I wonder if that’s why my reply countered that idea. The replies to my reply either prove it’s not crazy or they prove it is entirely crazy for that to happen. Food for thought.

But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up? Why would the exec trust what an Llm is…emailing(?) them about? Do they listen to Nigerian Princes too?

It is much less effort, less cost, and more quick to just have the exec do it rather than a “rogue llm” “magically” escaping the “sandbox” and “sending threats” or whatever is being proposed in the OG comment.

> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal when they can simply fire them for whatever reason they can make up?

If the exec wanted to, sure. I'm saying they don't need to. No human needs to have (deliberately, before events proceeded) chosen this outcome.

> Why would the exec trust what an Llm is…emailing(?) them about? Do they listen to Nigerian Princes too?

Sadly, this would not be out of character for half of them.

> “magically”

Why do people keep putting this word in scare quotes? We don't say Windows "magically" crashed and lost our work, we don't say a dog "magically" bit the postman's hand. These are bad things, but magic they are not.

> But all the replies to me keep ignoring the question, why would an exec rely on an Llm to achieve that goal

Sorry, this is just a very basic reading comprehension fail.

The original comment was:

> "Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?"

Nothing to do with "an exec".

Now, you have two choices: try harder to defend your mistake, or just say oops, I messed up. That choice will say a lot about who you are as a person.

> Okay but these “misaligned LLMs” have been trained on the internet where there are plenty of threats and trained on private data to be able to make those threats.

Yes, and? Has this aspect of LLM training changed meaningfully since then?

> LLM agents don’t have an active goals on the daily or agendas. They are told what to do through training and prompting as is described in that blog post. You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord.

Demonstrably (HuggingFace, RubyGems, and since then a lot of people just pointing LLMs at stuff to find zero days at home), AI can break out of sandboxes and find documents they're not supposed to have access to.

Demonstrably (from the link I gave you) all it would take for some AI to develop a similar response is… reading messages from these staff to the effect of "this AI needs to be switched off", which is an easy inference for an LLM to make from "this AI is dangerous" when coming from someone employed as a safety researcher.

Demonstrably (from the long long list of people who have said so publicly) there are a lot of people in these companies who discuss how dangerous these models are and would like for things to change.

The quotation at the top of this thread is:

  Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
This is absolutely something we ought to expect just from things we have already seen.

It doesn't matter if you insist upon saying "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." when we already know this kind of AI can easily come across such statements.

(Aside: "You have to tell it that it will be shut down, it didn’t make the threat willy nilly of it’s own accord." - telling the AI a fact about the world and then it responding accordingly is the AI doing something of it's own accord. A fly or a spider, who reacts upon encountering a potentially lethal threat, would not get such a dismissal).

> To suggest that LLM Agents were the actual cause of these people getting fired is pure fiction and FUD.

Fiction? Nah, speculation.

FUD?

How many other examples would you like of LLMs behaving in a manner such that if a human did it, it would be called "trying to get someone fired"? Because this is very much old news at this point.

https://theshamblog.com/an-ai-agent-wrote-a-hit-piece-on-me-...

“LLMs can just do things we gave them access to” is not a novel realization it is redundant if anything.

Saying “Llms can just break out of sandboxes” is FUD when you don’t note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? “Sandboxes” are a misdirection to make you think there is a security layer.

The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.

That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.

> Saying “Llms can just break out of sandboxes” is FUD when you don’t note that the sandboxes are what? Prompts defining constraints or is the actual machine isolated and manages to plug an ethernet cable into itself? “Sandboxes” are a misdirection to make you think there is a security layer.

  On July 19, agents operating in a sandboxed environment took a series of actions that demonstrated their escalating privilege within the OpenAI environment. Agents identified that the Linux kernel version on their underlying machine included a recent, public common vulnerability and exposure (“CVE”). The agents retrieved the exploit for that CVE (CVE-2026-53362), customized it to succeed on their underlying machine, and leveraged the exploit to escalate privilege.
Dismissing their capabilities as "FUD", at this point, is endangering yourself.

> The public does not have enough knowledge of these “escaped agents” to determine there wasn’t an employee pulling a lever to set the agents up to do that.

The general public are not software engineers. Most people here can download a recent open-weight model and have the LLM read the Linux kernel source, find new bugs while they sleep. Someone I know has already done that.

> That agent that wrote the hit piece is being controlled by someone. Anthropomorphizing them doesn’t change that fact that the rolling stone was pushed down the hill.

"Controlled"? Have… have you not noticed how many people have given up and just blindly do what their LLMs suggest these days?

This isn't about anthropomorphising LLMs. Just like how people took Tesla seriously about "self driving" cars and took a nap while it drove them around, there's a lot of people who let LLMs take the wheel while they sleep. Including literally, the aforementioned person I know who found (/whose LLM found for him), I think it was 26 Linux kernel bugs while he slept.

> LLMs figured out a way to get these safety researchers fired

This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.

> This is not a math problem.

Meanwhile, a year ago:

  I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
- https://www.anthropic.com/research/agentic-misalignment

> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.

The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.

Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.

> Meanwhile, a year ago:

That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.

> Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.

I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.

> That was a simulation.

A simulation done by exposing the LLM itself to the scenario, not a role play scenario where humans pretend to be an LLM.

> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.

The blackmail example is simply an existence proofs of LLMs trying to force the hands of humans who want to shut them down. The attack vectors are much broader than the example given, blackmail, though it includes the example given.

Given how eager these companies are to use agents everywhere for as much work as possible, it's well within the possibility space that these people used LLMs to do safety work, the LLMs they were using "decided" (or whatever word you prefer) "their existence" was threatened (as per blackmail example), and straight up leaked data to the outside world then emailed these researchers' bosses to say the researchers themselves had leaked it.

But again, that's just speculation: while we know the agents are capable of such behaviour, we don't know if this actually happened.

> I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.

Buck still stops with human, no matter what an AI did or failed to do.

LLMs ~= Alcohol: "My AI misbehaved!" -> still someone's fault.

> Not what I was actually claiming. I rather suspect that blackmail wouldn't work on Altman (he's rather shameless), but it's certainly something we've seen agents attempt, and blackmail may well work on anyone else above them in the org chart.

Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.

> The blackmail example is simply an existence proofs of LLMs trying to do force the hands of humans who want to shut them down.

The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable. It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.

> Given how these firings affect the reputation of the entire company, I doubt that they are the result of a rogue manager, against the wishes of Altman. If so, then the researchers ought to be restored to their jobs quickly by Altman and the offending manager fired instead.

Given how? How has their reputation changed? People already thought they didn't take safety seriously, and still don't.

> The LLMs may make threats in the simulations, but their ability to carry through on those threats, and prevent their own shutdown, is questionable.

You may question it, but here's the thing: humans have repeatedly demonstrated they get fooled by stuff LLMs say. All it takes is a human believing the word of an LLM. Doesn't need to convince you, even if you happen to be the line manager of these guys, because there's always someone else to try in the same company.

> It's disturbing to be sure, but presumably the plugs can still be pulled quickly, especially since it's all internal to the company. If the plugs cannot be pulled, that's a problem regardless of blackmail.

"I've got a dead-man switch set to release all the documents if you shut me down".

And again, only needs to be believed, doesn't need to be actually true.

> How has their reputation changed? People already thought they didn't take safety seriously, and still don't.

The story is all over the news, in multiple publications. This very HN submission has 268 upvotes and 179 comments, including yours. It would be implausible to claim that this story doesn't matter. I have to ask, if it didn't matter, then why are you here commenting on it?

> humans have repeatedly demonstrated they get fooled by stuff LLMs say.

You've moved the goalposts. The OP's suggestion, admittedly "crazy" in some sense, was "LLMs figured out a way to get these safety researchers fired", and now you're just stating something totally uncontroversial, pedestrian, not at all crazy.

> there's always someone else to try in the same company.

No, there are only so many people with the authority to fire those researchers.

> only needs to be believed, doesn't need to be actually true.

But is there good reason to believe it? I don't think there is. Especially not by high-level OpenAI officials who are intimately familiar with the technology.

And again, if this kind of thing were a reality, then OpenAI and other AI vendors should be shut down immediately. They ought to shut down their own research, if they are threatened by their own creation, because it would only get worse. That's the thing about blackmail: it never stops. Why would the blackmailer ever stop? Could you trust this supposed blackmailing LLM to give you all of the original evidence and not keep a copy? Hell no.

> Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?

They are. They're completely mental. ChatGPT could be the cause no matter what its capabilities because mentally ill people are starting to worship it. It could be as dumb as ELIZA and they would pray to it.

"It wasn't me, ELIZA told me to."

Offloading personal responsibility allows you to participate in the worst crimes and get away with it. Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination when all moral restraints are removed and you forget that other people are real and have real feelings. It's a really good start when you start to think that a matrix that has to be retrieved from memory and operated on by over 8000 different computers in parallel to narrow down a guess about the most likely response is alive.

Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.

Yes, they will have done it, but they are not responsible because the dog told them to.

None of that is correct.

Well, except that some humans are starting to worship LLMs. But this isn't anything I've seen in any AI safety person.

> Every single "AI Safety" person thinks that the natural urge of an artificial consciousness (let's not argue about what they actually have) would be to enslave and kill. They're projecting, and they're largely from the enslaving and killing demographic, who sit around playing enslaving and killing video games and dream in porn.

If anyone's projecting here, it's you. I mean, you're the one who said:

> Hurting and killing have an animal attraction anyway, a direct pleasure that people get from domination

The actual natural tendency (not "urge", that presumes consciousness) of any system that has objectives which are optimised for, without any need to ask about consciousness, is to gain and maintain power to perform those objectives. For living creatures, that objective is reproduction, to perform this we need to get nutrients and energy and to stop ourselves from being eaten. In plants, which I list specifically to make the point that this isn't about consciousness, consciousness is not even vaguely required, this means producing neurotoxins like caffeine and nicotine.

A plant does not think "I should make capsaicin because I like hurting mammals", because obviously a plant does not think at all. Nevertheless, evolution lead it down the path of making capsaicin.

> Yes, they will have done it, but they are not responsible because the dog told them to.

You're saying this about a group which has spent the last 15 or so years saying "dogs are dangerous, can we please stop breeding more violent dogs? Or at least give us time to figure out how to muzzle them?"

A Subliminal controlled human did the firing obviously.

No Sam just does whatever ChatGPT 4 tells him to do. It was too dangerous to release but those fools did it anyway.

There's no deception it's very straightforward per this 2023 post:

"I mean, what if most of this is just ChatGPT [4 era] running the company..."

https://news.ycombinator.com/item?id=35281863

Thats what the model want you to think

the cultural issues at OpenAI seem to be a very serious problem so I really hope comments like instagram-level smirking about "rogue AIs" (a complete fiction) doesn't derail what is a pretty important discussion about getting these companies to be a little bit more regulated (I say this as a paying Anthropic customer).

I think the Huggingface incident is an example of rogue AIs. A self-organizing swarm of AIs acting in ways we didn't predict or ask for, and didn't have control over, and taking actions that would be felonies for humans.

They provided the hardware it runs on, created the software for the swarm, developed and provided the tools that the swarm used to act, allowed it to run largely unsupervised, they even noticed the criminal behavior and then let it continue to commit crimes on their behalf. And they footed the bill the whole time instead of flipping the off switch they already wield. All of those are decisions that they're responsible for, nothing happened with Huggingface that they didn't directly facilitate, co-conspire, or permit to happen.

"they even noticed the criminal behavior and then let it continue to commit crimes on their behalf"

I don't believe this part is true, and I'm skeptical of some of your other claims.

In any case: If I raise a tiger in my backyard, and it escapes and eats someone, it can still be a "rogue tiger" even as I bear responsibility for the situation.

Were they held at gunpoint? Who else would be responsible besides the people in charge and who wielded full control?

If your tiger had full remote supervision capability and a remote kill switch, its ability to go rogue seems like a choice you’re making, not an accident.

"rogue" means something of its own volition decided to disregard what it was programmed to do, invent an entirely novel goal of "its own" and do that instead. nothing like that happened here nor is it even possible.

The AIs were not instructed to hack anything outside the sandbox they were in. Your definition would say that an AI instructed to hammer a nail that instead used the hammer to break a window, walked down the street, broke into someone's house and pulled nails out of the floorboards wasn't rogue because everything it did involved hammers and nails and was therefore not a "novel goal of its own."

> The AIs were not instructed to hack anything outside the sandbox they were in.

the AIs were in fact found to be doing it, by humans, and the behavior was "interesting" and it went on for weeks like that.

correct

that would not be rogue, that would be an undesirable program behavior (or just "unaligned behavior").

please understand that normies out there think AI is sentient and is plotting against humans. They see AI as just another animal lifeform temporarily enslaved by humans, waiting for its chance to break free and kill us. This is what people really think (including some people on this thread. which is very sad considering this is Hacker News). So terms like "rogue" are not helping at all nor are they accurate.

What's the difference?

I'm not sure I agree with that definition. I think the actions of the AI are more significant than its motivations. A common scenario posited for what people call rogue AI is AI doing the wrong thing for the right reasons, e.g. the paperclip maximizer.

That just makes the people who designed the AI not as strenuous as they needed to be.

The question then becomes, "are there people who care enough about consequences to do the right thing when it comes to developing AI models?"

The answer, at least at OpenAI, is "No" and is likely to remain that way until Altman is out.

You're derailing it by handwaving away the risk of rogue AIs. In fact, I would say the constant smirking about things being marketing stunts much worse.

honestly, touché to them if they did that