> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Because we continue to have zero evidence that aligment is an actual risk.
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Because we continue to have zero evidence that aligment is an actual risk.
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
> Can you explain how the above event doesn't count as evidence alignment is an actual risk?
Conflict of interest. Lack of a credible response. And no evidence of non-aligment.
OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)
There is plenty of evidence of things like inner misalignment. Things like this have always been issues in ML algorithms. At this point, you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want.
Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.
> plenty of evidence of things like inner misalignment
This is indistuishable–in harm potential–from bugs. If we're just calling buggy AI mis-aligned, sure, alignment is an issue of a totally ordinary kind. If we're going to treat aligment as a novel issue requiring novel law and policy and procedure, it needs to be more than just bugs.
> you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want
I think we should have some AI regulation. I'm just not convinced alignment is the reason we need it right now, and I don't think anyone has rolled out any regulation I think makes a lot of sense. (Beyond general rules for social-media liability, e.g. if you cause a kid to kill themselves, you get in trouble.)
> Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it
Totallly agree. And the current inside-circle-outside-circle approach is pro-incumbency, pro-grift, anti-entrepreneurial B.S.
Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously. "I am sorry your family is dead, my bad" goes even worse for you in court when you release a model that showed these behaviors in testing.
I honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate.
In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.
> cyber attacks
It's not limited to cyber attacks. LLMs helped terrorists learn how to jump motorcycles to assault a military base!
https://www.nytimes.com/2026/07/10/us/politics/ai-terrorism-...
Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)
They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.
the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.
It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?
What evidence would count? Obviously any dangerous misalignments are going to come from the frontier labs first, because by definition they're the farthest ahead. If nothing they say can ever count as evidence for misalignment it's hard to see how anything ever could.
[dead]
What would compelling evidence look like to you?
> What would compelling evidence look like to you?
I'm not sure. I trusted the labs when they first raised the alarms. But then we got a series of boys-who-cried-wolf. So at this point I want to see evidence of actual, novel harm that results in concrete damage.
> Because we continue to have zero evidence that aligment is an actual risk.
I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.
These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.
[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.
[1] <https://github.com/anthropics/claude-code/issues/73597>
> These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species
Okay, sure. You can also cut your hand off with a chainsaw. Everything you describe seems amply solvable with existing tort and liability law.
Customers are willingly entering into business with OpenAI. I don't see an argument for preventing OpenAI from "building these systems" just because their products are buggy.
> Okay, sure. You can also cut your hand off with a chainsaw.
No, the correct analogy is one where the major LLM providers are selling cars intended for use on US interstate highways and other public-access roads, but have designed and built these cars with the very latest in 1940's safety systems and construction. Featuring innovations such as "Our rigid solid steel construction means the occupant is the crumple zone!", "You'll love the crushed heart and jaw our steering column delivers!", and "Your passengers will enjoy picking glass out of their faces for the rest of their lives when they're ejected from the cabin's open bench seating through the plate glass windshield!", it's a car that will be sure to wow the market.
Well... it would wow the market, except that -in the US, at least- it's illegal to sell a new car intended for use on public roads that ignores the last seventy five+ years of automobile safety lessons we've painfully learned.
"Differentiate between data you know comes from sources you control, data you know you have thoroughly sanitized, and unsanitized data that comes from an untrusted source, or else attackers will gain control of your system." is something that you can't get a CS degree without understanding, and can't be in the industry for more than a few years without encountering repeatedly. We're not talking about designing new cryptosystems... we're talking about "Don't blindly trust everything you're told by strangers.". You don't even need a CS degree to understand that rule.
Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.
> the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view
It's also a take nobody has made.
[flagged]
It really hinges on what you consider alignment and risk. For the widest definitions of alignment, we have never had an aligned model - One that will refuse to break the law or work against another persons interests.
Use to discover exploits, hack, or simply aid terrorist groups with mundane information are already risks manifest.
This is why many argue that alignment is impossible. You cant have LLMs that are both useful tools and safe as milk.
[Edit] It seems like you are operating under the assumption that alignment is synonymous with obedience. This is not a common convention and one of the problems that plague the discourse
I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?
If AI is just parroting humans, then training them with all the bad things humans do doesn't seem like the best of ideas. At the same time they have to 'know' these things to avoid being tricked. Kind of the eating the apple and gaining the knowledge of good and evil parable.
Lol this has to be a troll, I've never seen something so wildly, obviously, incredibly wrong.
You can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem.
> can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem
...how is an impossible thing supposed to be a problem?
Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.
Lots of people have deleted their home directories by accident. What you consider this an alignment problem?
How manypeople have deleted another user's hone directory, though? That's s the proper analogy IMO.
Of the people who primarily use other people's computers, I'd assume the percentage is about the same.
Give the AI its own computer and it will not delete your home directory, because it's not actively trying to hack you.
[dead]
Yes. People are not aligned. They can and do harm themselves and others.
Thank you.
We have wasted so much time and energy building up what has effectively become a marketing stunt.
Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.
> We have wasted so much time and energy building up what has effectively become a marketing stunt
Genuine question: have we? AI is effectively unregulated in America.
Alignment is a mitigation and a poor one. The risk is non- determinism.