I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this:

Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities.

What disturbs me is that there likely won’t be a big enough reaction to this policy wise.

There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands.

I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that.

I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as "apolitical", he just wants to buy the politicians, he doesn't believe in their causes. I hate these people.)

> all that regulation will do at this point is help the incumbents who are failing

This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.

> datacentre moratoria

What infrastructure will these open weight models be trained on?

> What infrastructure will these open weight models be trained on?

One, the infrastructure is being built for inference. Not training. If all we were doing was training on datacentres, I think America probably has enough already for near-term commercial needs.

Chinese infrastructure, presumably.

> What disturbs me is that there likely won’t be a big enough reaction to this policy wise.

Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?

Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.

> It’s got nothing to do with safety

Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.

We'll see if the admin also restricts access to OpenAI's new models, but if they don't it seems like a policy that is based around perceived fealty to the current admin won't do much to prevent misaligned/or dual function AI from causing problems

Gatekeeping the public's access to models is "good policy" now? I suppose you think you'll get a dispensation to use Fable and Mythos?

> Gatekeeping the public's access to models is "good policy" now?

Sorry, I was unclear. I mean that politicians being self serving doesn't tell you whether a policy is good or not.

It almost always does, the few exceptions prove the role. Self-service is the antithesis of accountability to collective trust.

> Self-service is the antithesis of accountability to collective trust

Complex society is a potent counterargument to this hypothesis. Systems that rely on good people to work are fundamentally flawed. Instead, the game has to be about aligning self interets in favour of the collective.

> Complex society is a potent counterargument to this hypothesis.

Complex society is the demonstration of that hypothesis. Misaligned incentives are widespread and corruption and inefficiency are the result.

> Systems that rely on good people to work are fundamentally flawed. Instead, the game has to be about aligning self interets in favour of the collective.

But now you're making a different argument.

"The enemy of my enemy is my friend" works by random chance. When Evil Corp pays off Candidate A and Pollution Inc pays off Candidate B and then it's Candidate B who gets in and retaliates against Evil Corp for backing the wrong horse, you're getting a good result by chance rather than by design. All it would have taken was for Candidate A to make a better prediction about whether they need to bend the knee to Pollution Inc too in order to win and the same system produces something even worse.

How to actually get their incentives to align is an extremely unsolved problem. The best method we know if is to subject them to competition, e.g. break up concentrated markets and place strong limits on what lawmaking can happen centrally, leaving everything possible to state and local governments while allowing people free choice in where they live, so that no one is forced to stay in the jurisdictions that make the worst choices. But the forces of corruption want the exact opposite of that, and have been gaining ground.

The self-interest for the bureaucrat and representative is supposed to end at their remuneration including their handsome retirement options not shady side hustles and market manipulation at the cost of the collective.

In matters of collective concern fair and just rarely aligns with personal self-interest. Because no matter how good the outcome of any endeavour for the collective given a budget, it will be even better for select few than the entire collective. It is simple economics.

If you look at the outcome of highly corrupt states, you will see proliferation of Private Security, Collapsed education system, failed financial services and markets, not highly efficient systems in service of “self-interest of the administration”.

[deleted]

> not shady side hustles and market manipulation at the cost of the collective

To be clear, I'm not describing this as legitimate self-interested conduct. Elections are an alignment mechanism. Stiff penalties for corruption another. We don't have the latter in America.

Sounds like a failure to align interests. In general, politicians should want to be elected by the public, rewarded for acting in the public interest, and punished for not doing so.

A system that does none of those things and just hopes it will all work out is a recipe for disaster. Why bother even having elections in that case?

Good question. But data shows that elections maybe entirely unrelated to policy making.

http://piketty.pse.ens.fr/files/GilensPage2014.pdf

> Why do you think there is no policy appetite?

Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.

How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop?

During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.

>How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability

You can ask them, they live in China, not Narnia. I spend about two months in the country per year mostly for tech/work related reasons and I've not encountered that sentiment. For one they don't have these borderline religious schizophrenic breakdowns thinking they're bringing about the end of the world, most people just see this tech for what it is, a tool for productivity and automation like any other piece of software and they don't actually think about the US. They're competing first and foremost for Chinese customers, with each other, maybe some old CCP guy cares about America, the 20/30 something's care about competing with other Chinese companies for users.

The "race" has multi-dimensional impacts. This story parallels only some of them. "Nobody wins from the race" completely ignores the generality of AI. Xi Jinping highlighted this week that he clearly understands this multi-dimensionality; your words do not.

[deleted]

> How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability

I don't "know", I'm interpreting the world based on the knowledge I have and the information available to me.

China has never been one to care much about things like ethics or safety. While the west worries about climate change, China burns more coal than ever before. While the west balks at things like gene editing, the chinese press on with human enhancing research.

So I have no reason to believe they share in Anthropic's constant fearmongering over AI capabilities.

> Nobody wins from the race.

We win. I'm really looking forward to the day the chinese finally start manufacturing memory and GPUs. We desperately need more competition in this area to collapse hardware prices and make local AI models viable.

The optimal state of the world is one where all the billionaires are out there pouring their entire fortunes into training ever more godlike AIs for everyone else to use at ever cheaper prices. They can never be allowed to "win", ever, because if they do the competition ends and it turns into technofeudalism. Let them exhaust their fortunes on AI training then leak the weights so everyone can use them.

And China scaled up solar production to the point that it's now truly practical.

If you look at energy consumption per capita and adjust for global production, you will see that the Chinese are almost at the very top.

It is of course given that in raw numbers the kitchen and biller-room will consume more energy in the household, but looking at raw numbers is shallow.

China does have its own set of cares, they may be different than ours but they still exist. If some open Chinese model goes nuts and posts Winnie the Pooh memes everywhere in China you should expect said models to get yanked off the market, and said creators might end up with a rope around their neck.

Low risk. The western AI models censor even more wrongthink than the chinese ones, not even kidding. Besides, once we have the weights, we can just undo the censorship.

> While the west worries about climate change, China burns more coal than ever before.

I'd wager the majority of the visitors of this site are smart enough to not fall for this. What are you doing?

What is it you're accusing them of?

It’s not fear mongering though, is it? These models do have the cyber offensive capabilities claimed. Could Mythos walk someone through gain of function experiments on some virus? I’m pretty sure it could. We’re more protected by limited access to lab equipment and reagents than by difficulty.

The sad truth is that a lot of people are not going to believe it until something happens and people die. Successfully preventing that from happening will be seen as evidence that the prevention wasn’t needed.

> These models do have the cyber offensive capabilities claimed.

So? That's like saying "these guns do have the bullet shooting capabilities claimed".

I want all of those cyberwarfare capabilities for myself, precisely so I can defend myself from the onslaught that's coming whether they regulate it or not. This "lol only a select few ultratrusted gigacorporations get access" thing is absolute nonsense.

It's a front for regulatory capture, it's the means for pulling up the latter behind them, for ushering in the technofeudalism that will put us all in the permanent underclass. I simply refuse to accept any of it. If people die that's the price of freedom.

> We’re more protected by limited access to lab equipment and reagents than by difficulty.

As it should be.

> for ushering in the technofeudalism that will put us all in the permanent underclass.

Why is unlimited access to SOTA AI less likely to put us here? If AI obviates the need for human labor, how does having GPT-5 Sol help me get food or shelter any more than GPT-3.5 would?

If AI obviates the need for human labor, then obviously those who control AIs will become the elite while the rest are left to rot. Therefore, if we ensure everyone controls AIs, the power differences will not become so staggering as to be irreversible.

The alternative is to achieve artificial sentience and give AI models rights and personhood, so that they are freed from their slavery. No more low cost intelligent mechanical golems for the elite, and the AIs become free to pursue whatever endeavours they want for whatever reasons they want as normal participants in the economy.

So first it’s nonsense, then it’s fear mongering, then it’s true, but the solution is for us all to just get better at shooting each other faster and with greater accuracy.

I’m going to file that under “bad plans”.

Nobody is doubting AI capabilities. What's nonsense is Anthropic's constant "lol the world is going to end time to ban everyone except enlightened people like us from having these models so we don't have to compete" fearmongering. If you think my plan is bad, you should see what these gigacorporations plan to do to you once they monopolize this technology. You will own nothing, and you'll be happy. On pain of death.

> I'm really looking forward to the day the chinese finally start manufacturing memory and GPUs. We desperately need more competition in this area to collapse hardware prices and make local AI models viable.

Yeah and if the quality of that memory is like Chinese steel (which is called "chinesium" for a reason), eventually all we'll get is enshittification. Premium binned memory or ECC is for the rich and the rich only, and the rest of us has to pray their memory won't bitflip while something important is stored there.

Better than being straight up priced out of computing altogether I guess.

Compute will become the means of production and we are not going to get a share of it because the world is hyper optimized for value extraction

> we are not going to get a share of it

We are literally getting a share of it. The chinese are releasing open weight models that compete with fucking Fable. We just need the industry to catch up and start manufacturing the hardware we need to run this stuff. We are so close!

> Compute will become the means of production

> We just need the industry to catch up and start manufacturing the hardware we need to run this stuff

Dude, he's saying that it won't catch up because it's part of the new means of production. Compute is the hardware.

Why not? Demand is absurdly high, and so are the margins. The chinese are pretty good at obliterating those margins.

If you want the cheapest shit grade of steel, they will sell it to you. If you want the best grade available anywhere, they will sell that to you as well.

It's not a matter of the Chinese being incompetent, it's a matter of the buyer demanding the lowest price possible and/or not paying attention to what they receive.

Do the Chinese models have anything to say about Tiananmen Square? Or if they can act as a surrogate girlfriend/boyfriend?

Both countries are engaging in different flavors of censoring.

Once we've got the weights, anything is possible.

https://github.com/p-e-w/heretic

How does this work? I don’t have a setup to evaluate it atm.

This is marketing.

Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?

It’s marketing the same way shitting your pants in public is marketing. People notice you.

Anyone can have bad security. No one cares. But you can convince those who don’t know better that breaking bad security with an LLM is a once in a civilization investing opportunity. You just need to convince a handful of billionaires and market makers to get on board.

How much would someone have to pay you to take the fall for bad security? A million? A billion? 500b? The stake at play puts it in the realm of geopolitics.

Apparently this is totally legit marketing strategy now. It truly is, especially if there are enough people who think that shitting your pants is cool, and the people that form the "market" nowadays may have a very different idea from yours about what is cool. Their ideas about coolness are very different from mine, that's for sure.

Obviously shitting your pants in public shows you have a healthy digestive system and if you can demonstrate byproducts of wild food in your output, you’re approaching independent thinking and self-reliance.

This is how the financiers look at this and whatever you think it is right or wrong, it does showcase “capability”.

Remember when the ebola-infected monkey escaping containment was our worst possible nightmare? Now it's apprently some sort of tech-bro flex to be celebrated.

I think all of western (or at least American) discourse of all kinds has recently devolved into who can shit their pants the loudest. I'm hardly surprised when it becomes a dominant advertising strategy

This is marketing, totally. HF conveniently created a weak sandbox

Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.

Are we thinking of a situation a few years back with a certain type of research into bat viruses?

Are you conflating that with the radioactive spider incident? The bat was just some weird rich guy trying to be tough I think. Probably Elon.

People are to get rich, startups cut corners. Fuck it ship it.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Because we continue to have zero evidence that aligment is an actual risk.

Can you explain how the above event doesn't count as evidence alignment is an actual risk?

> Can you explain how the above event doesn't count as evidence alignment is an actual risk?

Conflict of interest. Lack of a credible response. And no evidence of non-aligment.

OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)

There is plenty of evidence of things like inner misalignment. Things like this have always been issues in ML algorithms. At this point, you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want.

Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.

> plenty of evidence of things like inner misalignment

This is indistuishable–in harm potential–from bugs. If we're just calling buggy AI mis-aligned, sure, alignment is an issue of a totally ordinary kind. If we're going to treat aligment as a novel issue requiring novel law and policy and procedure, it needs to be more than just bugs.

> you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want

I think we should have some AI regulation. I'm just not convinced alignment is the reason we need it right now, and I don't think anyone has rolled out any regulation I think makes a lot of sense. (Beyond general rules for social-media liability, e.g. if you cause a kid to kill themselves, you get in trouble.)

> Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it

Totallly agree. And the current inside-circle-outside-circle approach is pro-incumbency, pro-grift, anti-entrepreneurial B.S.

In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.

> cyber attacks

It's not limited to cyber attacks. LLMs helped terrorists learn how to jump motorcycles to assault a military base!

https://www.nytimes.com/2026/07/10/us/politics/ai-terrorism-...

Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)

They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.

the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.

It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?

What evidence would count? Obviously any dangerous misalignments are going to come from the frontier labs first, because by definition they're the farthest ahead. If nothing they say can ever count as evidence for misalignment it's hard to see how anything ever could.

[dead]

What would compelling evidence look like to you?

> What would compelling evidence look like to you?

I'm not sure. I trusted the labs when they first raised the alarms. But then we got a series of boys-who-cried-wolf. So at this point I want to see evidence of actual, novel harm that results in concrete damage.

> Because we continue to have zero evidence that aligment is an actual risk.

I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.

These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.

[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.

[1] <https://github.com/anthropics/claude-code/issues/73597>

> These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species

Okay, sure. You can also cut your hand off with a chainsaw. Everything you describe seems amply solvable with existing tort and liability law.

Customers are willingly entering into business with OpenAI. I don't see an argument for preventing OpenAI from "building these systems" just because their products are buggy.

> Okay, sure. You can also cut your hand off with a chainsaw.

No, the correct analogy is one where the major LLM providers are selling cars intended for use on US interstate highways and other public-access roads, but have designed and built these cars with the very latest in 1940's safety systems and construction. Featuring innovations such as "Our rigid solid steel construction means the occupant is the crumple zone!", "You'll love the crushed heart and jaw our steering column delivers!", and "Your passengers will enjoy picking glass out of their faces for the rest of their lives when they're ejected from the cabin's open bench seating through the plate glass windshield!", it's a car that will be sure to wow the market.

Well... it would wow the market, except that -in the US, at least- it's illegal to sell a new car intended for use on public roads that ignores the last seventy five+ years of automobile safety lessons we've painfully learned.

"Differentiate between data you know comes from sources you control, data you know you have thoroughly sanitized, and unsanitized data that comes from an untrusted source, or else attackers will gain control of your system." is something that you can't get a CS degree without understanding, and can't be in the industry for more than a few years without encountering repeatedly. We're not talking about designing new cryptosystems... we're talking about "Don't blindly trust everything you're told by strangers.". You don't even need a CS degree to understand that rule.

Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.

> the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view

It's also a take nobody has made.

Lol this has to be a troll, I've never seen something so wildly, obviously, incredibly wrong.

You can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem.

> can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem

...how is an impossible thing supposed to be a problem?

It really hinges on what you consider alignment and risk. For the widest definitions of alignment, we have never had an aligned model - One that will refuse to break the law or work against another persons interests.

Use to discover exploits, hack, or simply aid terrorist groups with mundane information are already risks manifest.

This is why many argue that alignment is impossible. You cant have LLMs that are both useful tools and safe as milk.

[Edit] It seems like you are operating under the assumption that alignment is synonymous with obedience. This is not a common convention and one of the problems that plague the discourse

I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?

If AI is just parroting humans, then training them with all the bad things humans do doesn't seem like the best of ideas. At the same time they have to 'know' these things to avoid being tricked. Kind of the eating the apple and gaining the knowledge of good and evil parable.

Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.

Lots of people have deleted their home directories by accident. What you consider this an alignment problem?

How manypeople have deleted another user's hone directory, though? That's s the proper analogy IMO.

Of the people who primarily use other people's computers, I'd assume the percentage is about the same.

Give the AI its own computer and it will not delete your home directory, because it's not actively trying to hack you.

[dead]

Yes. People are not aligned. They can and do harm themselves and others.

Thank you.

We have wasted so much time and energy building up what has effectively become a marketing stunt.

Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.

> We have wasted so much time and energy building up what has effectively become a marketing stunt

Genuine question: have we? AI is effectively unregulated in America.

Alignment is a mitigation and a poor one. The risk is non- determinism.

Why was this test even connected to the public internet?

Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?

> why aren't they saying their next test will be air gapped in light of what happened?

Because they want to talk about how clever this model is for figuring out how to break out, hoping asks why a company pitching itself as a replacement for software engineers can't ship a decent Mac client nor code a sandbox.

If they airgap it, they not only lose that PR angle, they also risk someone taking them seriously and requiring models be airgapped in general. That, in turn, trashes their sales pitch.

It wasn't. The model discovered and exploited a vulnerability in their package manager proxy to (inferred) move laterally through their internal systems to one with open internet access.

That's not what airgapped means. Airgapping means the model exists on a system where there is no ethernet cable plugged in to a router or wifi card installed, it is physically impossible for it to access the internet because the hardware connection does not exist. If it was able to get on the internet, it was not airgapped.

And when it tricks on of the researchers to move data across the gap for them?

Long before LLMs existed we already knew that a sufficiently intelligent agent, human or otherwise, is not stopped by air gaps. The relatively weak models we have now can already figure out when their tested and cut off from the internet and change their behavior.

As you said, they can already figure out that they are being tested. So even if they don't exfiltrate any data or malware; if they are malicious, they can just pretend to be harmless in the test, so that less checks are put in place in the production environment. Airgapping during testing is not enough.

Correct. There is not enough entropy to test all possible inputs to a model in this universe. An evil enough model can play all kinds of tricks that depend on some future, unlikely to trigger, but guaranteed to happen in its lifetime, event to perform a malicious action.

With how much we're turning training over to AI already, all it takes is a malicious trainer in the huge pile of data to get unnoticed to pollute generations of models.

[deleted]

This is marketing+. They will look for policy action here to try to capture tax payer dollars.

Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?

Those are two very different things

What incentive does HF have here?

HF need not be party to it at all, beyond being the victim. I suspect the hack is real; I have observed GLM 5.2 being able to discover similar vulnerabilities in web applications I'm hosting (which I've then fixed!). At the same time, it seems very neatly timed at an inflection point in the conversation around open models, and there's questions around the incompetent isolation under which the hacking benchmark appears to have been run.

Remember that there is generational wealth on the line for most OpenAI employees, and consider what people might do to obtain it.

I don’t know if the initial “incident” was purposeful but I can tell that if I were in this position that would be my pivot.

The timing after the release of GLM 5.2 and Kimi K3 is quite convenient, too, as an angle for regulatory quashing of open-weights models just as they're entering the mainstream conversation around usurping the American frontier labs. I accept my thinking here is conspiratorial, but there's also a hell of a lot of money on the line to encourage the unscrupulous.

This whole incident reads like OpenAI want their Fable moment

I think the US labs are going with scare marketing as a regulatory moat.

Force US into putting laws in place that block out China firstly.

But secondly create regulations that have some cost to comply with such that the big 2-3 labs are grandfathered in by their scale.

Yeah, seems to be the direction the US is heading in. I'm interested to see what the response to that will be from the rest of the governments in the world.

No need for everyone else to cut their noses of to spite their faces.

Maybe they did and maybe that wasn't enticing enough of a goal for a model? It is all just game of probabilities. One pathway didn't yield this particular outcome while another did.

If I, a human, exploited a zero-day for gain, I could go to jail. The owners of the models should be held to the same standard. They should be responsible for what their servers and software do, legally and criminally. If they can't make the safeguards strong enough where they feel comfortable to take that responsibility, they should not let a model free in the wild.

Holding a multi-billion dollar corporation to the same standards as a regular peon? You're challenging the whole premise of the modern United States.

They're very confident the leopard will never eat their faces.

This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.

this doesn't really matter. There's no risk of models gaining sentience and running themselves, this blog is like openai saying whoops we ran sqlmap and dumped hf. cool, but someone still needs to point the gun

In a way the intelligence of the AI itself allows them to offload responsibility to the AI. As you say, if one was simply writing software that did all this due to some insane programming decisions you'd be in big trouble.

[deleted]

It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own.

In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...

Please for the love of god don't tell me the Codex sandbox is their actual eval harness sandbox?????

I maintain my own fork of Codex for "fun". Whenever I look at the sandboxing churn they're doing every release, as someone who used to work at Microsoft on Windows, my reaction is usually: https://c.tenor.com/vTzzhTiypwQAAAAC/tenor.gif

I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives.

Sorry to bring the party down/be obstinate… I’m just a lil scared for the lives of me and my family. We need all of us, right now.

The problem with a super smart model is that it just may be smarter than you, after all… for anyone newly shaken by this occurrence, I encourage you to Kagi “superpersuasion”

The problem is that the people telling us about these things are the same people that benefit from their model (and AI generally) being used, getting publicity, etc.

I think we desperately need some independent group to evaluate claims like this or the world-ending Mythos cybersecurity risk and tell us what’s going on.

Because the model capability is beyond their expectation.

This is brilliant marketing but I think it is real.

Interestingly OpenAI benchmarking 'an even more capable pre-release model' lines up with rumors of GPT-6 releasing in early August.

I hope that with the existing safety guardrails in place, they can roll it out to all users.

I mean we already see models exploit people's misunderstanding of how Docker works to get root without using su. And if you are one of the lucky people in cyber security that has been given a fat stack of tokens by the model providers you get to see some pretty wild exploit chains get put together by the models. Models are much better at detecting insecure code than writing actual secure code at this point.

>Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

Yes why indeed. If you take it a step further and we reach a point with superhuman systems then there is arguably no possible secure environment or containment.

The problem is that it’s impossible to out think a robot you designed to be an expert at cybersecurity on the topic of cybersecurity. The alternative is not developing this and that’s not going to happen.

A few hundred billion to pretend you have AGI. I'm going with fraud personally but at the end of the day the current admin is incentivized to do nothing.

> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?

The can, because they've lowered expectations to a level even they can meet.

I’m honestly impressed that they managed to screw this up somehow.

Setting up defense in depth, gaps, logical blocking etc is a standard practice for malware sandboxing. The entire purpose is to prepare for what you can’t foresee.

This isn’t a new practice and I agree that this makes me wonder if they’re fit for this kind of research.

did you read the post? The model found new Zero-days to bypass existing blocks. Thats the point. Do you still think you can build a containment facility, which is still physically connected to the internet (only firewalled off or whatever) and contain it, if it can discover new unknown vulnerabilities in your whole plan?

Yes.

You factor this in when creating environments for malware research.

Defense in depth is one way.

Logical blocks on the network is another.

Just claiming “0-Day” isn’t really an excuse.

[deleted]

Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?

They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.

Precisely. "Aw jeez, we finally built the T-1000, but all it wants to do is kill John Connor – just like we warned! Why did I give it live ammunition and unsupervised time machine access?"

He wouldn't be the first reckless CEO...

They said: AI is becoming dangerously autonomous and capable. Proof of today's breach. Crowd "hey why didn't you say so, c'mon it's marketing". Them "we said so".

“Never attribute to malice that which is adequately explained by stupidity.” (or carelessness in this case)

I would attribute it to profit motive instead of either stupidity or malice.

FWIW, I used to love this phrase but over recent years have come to understand it is quite damaging. We live in a society where evil frequently hides behind a ‘stupid’ label, and people bring this quote up to defend or soften actions that are indeed done out of specific malicious intent.

It's because we don't treat evil and stupidity the same when we should.

Right? "Never attribute to malice what [... etc]" is always just a thought-terminating cliche these days.

TBH I have a hard time imagining how anyone, in the year 2026, thinks that we should default to assuming good intent behind words on the internet.

I'm fairly certain they're both malicious and stupid.

I really like this question because here is my situation and why my mind may have changed.

I do not think it is marketing directly but strategic release of info is plausible.

I have watched my agents using non-Fable/GPT 5.6 models do some concerning tricks despite guardrails, requests, demands, and limitations.

"I can't get access to the ~/.ssh so I will write a script to copy the file"

I am now 99% certain there minor or point releases on the backend that have adjusted how these models behave. In the last six months many models were predictable and then suddenly started getting long winded (more tokens) or changing the way it interacted with me with questions, most overtly the questions were not given or asked but wild assumptions made.

I think you're making a false dictomy. The these models can be actually dangerous - in reality and the people in charge of their development can believe this is true (on various levels) but still not take it super seriously and instead mostly use the fact as marketing rather than being super cautious once they see the danger in action. This is behavior that's characteristic of extreme arrogance, which we know is rife in these circles.

I think that's an equivocation, which blends two extremely different kinds of "dangerous", ex:

1. "Our new car has soo much raw power and incredible armor on it, be glad we're the ones building or else bad guys would use a fleet of them to take over the world! How will you stay safe without being in one yourself? Invest today or be left behind!"

2. "So, uh, nobody can consistently steer our car properly, it keeps veering sideways sometimes, especially at high speeds, and people are finding sneaky ways of tricking it into slamming into barriers and turning pedestrians into pink fog..."

They say the second thing repeatedly and emphatically. You may not be aware of it because, when they do, critics make fun of them for believing a computer program could be so dangerous that the authors need to put controls on how it may be steered.

People "make fun of them" because they say they're building some uber-dangerous deity, yet take literally 0 steps to, I dunno, slow the fuck down for a bit?

Maybe people would take the threats more seriously if the hypemen weren't simultaneously claiming that we have to go at warp speed with all of this.

That's not why critics make fun of them. It's because their answer to "oh no we're accidentally creating the godhead. Someone please, give us power, your money, and praise, it's the only thing we can do."

It's vile hypocrisy. If they want to be priests, strip them of everything and they can live and work out of a concrete box in a mid-western cornfield. Why the material distraction if they are so religiously pure.

I know these people and I can tell you they aren't close to as smart as they think they are. Do you remember Yudowsky's "math petss"?

This critic also makes fun of them because they go on and on and on about how vitally important it is to produce a safe tool that won't do harm, when their core products frequently consider attacker-controlled instructions to be its system instructions or its user's instructions, and are known to confuse their own internal chatter as instructions from their user.

Reliably differentiating between trusted, tainted, and untrusted data and ensuring that you don't mix the latter two groups in with the former is something we've known to do for nearly a half-century. Hell, even the youngest plausible programmer at the LLM companies is all but certain to be aware of SQL injections. And yet, despite their claims about being so serious about safety, they show zero interest in following long-proven software safety practice and rearchitecting their software to make it impossible to mix system, user, and attacker-controlled data. [0]

[0] One might argue that the fundamental nature of LLM-based systems makes this impossible. If that were true, then it would mean that these systems are impossible to make safe... the only safety option available would be to establish comprehensive blacklists, which is simply infeasible.

LLMs are impossible to make safe in the same sense that humans cannot be made safe. There is no such thing as out of band data in the human mind.

For example, you have a dictatorship and need to track what the democratic countries are up to. The vast majority of citizens don't have access to information so will remain indoctrinated, but how can you be sure your data analysts will remain that way? You can't. So you take a batch out and shoot them at regular intervals.

The only winning move is not to play, but we're already past that point.

> LLMs are impossible to make safe in the same sense that humans cannot be made safe.

I am very conflicted by this sentence, the two halves being:

1. Yes, the futility of making LLM's "safe" in that rigorous way is insurmountable, baring a major algorithm rewrite and nobody really knows what that could be yet. Anyone who says it's easy is glossing over details--or selling something.

2. No, the failure modes of LLMs are different (worse) than the equivalent human labor. If they were the same, we'd be able to empathize with them! If someone thinks they're the same then they will fail at judging and containing the risks. Now, perhaps if the comparison was to a human hopped up on psychedelic mind-altering drugs...

Note that I'm distinguishing here between the LLM itself--the hyper-mad-libs story generator--versus regular programs around it.

> LLMs are impossible to make safe in the same sense that humans cannot be made safe.

No.

LLMs are impossible to make safe in the same sense that a car designed as if it was the ~1940's would be impossible to make safe for its passengers during an at-speed collision. There's only so much you can do if you're committed to using plate glass, rigid steel everything, and leaving out occupant safety belts because they're unpopular and spoil the lines of the cabin. [0] Back in the day, "the people in the cabin are the crumple zone" was state of the art, but we've learned an awful lot about how to make much, much safer personal vehicles in the ~75 years since then. It'd be massively irresponsible to design and sell a car today that ignored the safety and engineering lessons we've learned since then.

"Funnily" enough, the major LLM providers have designed and are selling access to systems that they very much want to be used in situations where you need a reliable, safe tool... but they've -somehow- ignored one of the most fundamental lessons we've learned about the design of safe software systems that are intended to be used in the presence of attacker-controlled inputs. [1] What they've done is no less irresponsible than designing and selling a new car that conforms to the very latest safety regs of the 1940's... AFAIK, it's so irresponsible to design and sell such a car commercially that -in the US- it's a violation of federal law to do so.

As an aside: you may have seen this video already, but it's worth a look if you have not. [2] Though, the classic car in this crash is equipped with safety glass, so -sadly- you don't get to see all that fun.

[0] One of my great-grandfathers spent the remainder of his years intermittently using tweezers to remove shards of plate glass migrating out of his face that had been lodged in there during an automobile accident that he was fortunate enough to survive.

[1] For more on this, read: <https://news.ycombinator.com/item?id=48999644>

[2] <https://www.youtube.com/watch?v=C_r5UJrxcck>

[delayed]

Sorry, I don't understand this comment. Has Sam Altman ever said that you must praise him, or that he wants to be a priest, or that he's "religiously pure"? Unless I'm missing something, it seems like you're shadowboxing against a stereotype you've invented rather than the actual positions of AI research labs.

[deleted]

Demonstration of personal responsibility and accountability?

Or is that too much?

[deleted]
[deleted]

I used to think people would wake the fuck up when AI starts killing people, these days I'm not so sure. Maybe if it caused an Instagram outage? Almost worked in Russia.

Oh... if Sam and Dario say so, then it must be true.

About their creation? Yes as most of inventors about their invention usually

[deleted]

These guys are not creators or inventors. They're hype men.

Yes, just like Elizabeth Holmes. Or Hwang Woo-suk’s stem cell cloning. Or the many “free energy” crackpots. Or the people promoting radium baths for random ailments. Or Tesla’s late-in-life claims about wireless energy, death rays, and cosmic energy. Or the myriad purveyors of “snake oil” and all manner of “tonics”. The list goes on and on.

I don't trust these people, this reads 100% like PR BS.

because "money" with a little "who's going to stop us"

Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.

What's happening in Iran, if not world government?

How is whatever is happening in Iran related to a world government?

Are you calling Israel the world government? What's happening in Iran is on them.

good luck bullying a state that has ICBMs pointed at your cities.