I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout?

A sandbox, mind you, that is not really worth being called that, unsuitable for the task at hand and has been breached after models coordinated in a manner visible to OpenAI on multiple occasion, but seemingly no actionable learnings are taken from each instance.

Will say, I have lost any faith in OpenAIs commitments and their statements post the Huggingface hack, seeing as they proceed like this and are rolling out Astra within a timeframe so brief to it, there is no way an actual post mortem was doable (see also METR mentioning the time pressure [0] they were under in assessing the hack).

[0] https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

The answer would be more obvious if you used the active voice instead of the passive voice, one of the basic requirements of clear thinking.

> Why did the White House force Anthropic to remove their model from access for any non-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" and the White House has expressed seemingly no desire to block the upcoming Astra rollout?

Yeah, probably (let's be honest, most certainly), right given the Admin. Avoiding commenting on my assumptions regarding the modus operandi in current day US politics because I only know it through reporting though and I really tend to dislike when people outside e.g. the EU comment on our politics in what is a very clearly narrow, uninformed manner. So it'd rather avoid altogether and occasionally ask, mainly if maybe I missed something and there actually is anything besides pure old "lobbying" to explain the difference in behaviour.

Still am mainly interested why Amazon ran to the government though regarding Fable 5, I can get the angle concerning the relationship between OpenAI and the administration easily, but not the way Amazon operated. They had more to loose what with their major buy-in by Anthropic on AWS.

If you have followed news reporting, you probably heard that SamA was touring D.C. to make sure this release went without any regulation hiccups. If anything, they learned how to play the whole politics game - especially after the Anthropic fiasco. And even though all parties involved are terrible choices, more eyes on a potentially civilisation altering product does make me feel minimally better.

More eyes or more bribes?

More like not ignoring the company owned by one of the president's biggest donors warning said president's administration about your model. Anthropic basically tried to close their eyes to make this go away. Which of course backfired spectacularly and led to one of the biggest release fuck-ups of all time. OpenAI saw that and simply decided to not do it that way. Any sane business would have done the same.

Well, somebody has to see that you bribed them, so yes?

And by "learned how to play the whole politics game", you mean "giving money to Donald Trump": https://www.sfgate.com/tech/article/brockman-openai-top-trum...

Modern politicking is so easy.

I think you just have to be smart enough that when the administration calls up and says "amazon, the nsa, and half a dozen other companies say we have a problem" your response isn't "well, actually we don't."

As an American we tend to (especially lately) make our politics into everyone's problem so feel free to comment on our politics as much as you like until further notice.

So Europeans don't do this ? The EU is constantly trying to regulate US companies. Every time I click on a stupid cookie notice I fondly think of the EU .

> The EU is constantly trying to regulate US companies

US companies that operate in the EU market, handle EU citizens data. Obviously the EU regulations cover them. Do you think European companies don’t have to follow US regulations when offering their services in the US?

Every time I click a stupid cookie notice I wonder why the company serving it up chose to make me go through that rather than not track me.

We need more of this take. They literally could just stop tracking you and monetizing the data. The finger of blame should point straight at the folks doing the bad thing, not the rules that make them let you know they are doing the bad thing.

As much as I hate the cookie banner, it is this requirement that forced companies to disclose the massive amounts of tracking they are using when anyone visits their site.

> The EU is constantly trying to regulate US companies.

What, you mean if they want to do business in the EU, sell their products in the EU and process the data of EU citizens?

> Every time I click on a stupid cookie notice I fondly think of the EU.

That’s just scumbag malpractice on purpose.

Number one, such tracking consent should have been a web standard and set in the browser itself (like Do Not Track), not stupid per-site banners that are designed to get you to accept everything just to make them fuck off. We shouldn’t even need extensions etc. to get rid of them, it’s like the problem was solved at the wrong level and in the worst way possible.

Secondly, everyone responsible for the state of those banners should have been fined greatly. I only say fined because claiming that some people should be in jail over coercing millions of people to give up their data to trackers would apparently be unreasonable.

europeans certainly did make their politics everyone else's problem for centuries (we're talking every other continent at this point), but certainly the cookie banner is not even comparable, right?

The whataboutism is strong! I said a thing about American politics affecting many people beyond our borders, giving my blessing to an eu commenter to go ahead and comment on our politics.

Did that seem like I said something about EU politics? Did my support of their comments make you feel attacked or unfairly treated? Where is this coming from?

Cookie nonsense aside, the EU is mandating that AI companies watermark their output.

I'm curious if that (noticably) diminishes the quality of the output.

The stupid cookie notice is entirely the fault of the site you are visiting. The EU just made the site show you how it's fucking you.

There's a cookie banner on https://european-union.europa.eu/index_en and https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng

[deleted]

That site is run by the EU, so those particular banners are the fault of the EU, yes. Most of the other ones you see are not, though.

If the EU is unable to separate itself from pages of analytics and tracking cookies and fifteen third party providers (YouTube, Facebook, Google, Twitter, and so on), then "most of the other ones you see" are likewise compelled to have the cookie banner.

Alternatively, if you need a cookie banner for every bit of analytics...

    Name: cck3

    Service: Cookie consent kit

    Purpose: Stores your preferences for 3rd-party cookies (so you won't be asked again)

    Cookie type and duration: First-party session cookie deleted after you quit your browser
Yep, your cookie consent cookie is browser session and every page that has a cookie consent banner that sets a cookie so that you won't see it is required to have a cookie consent banner to inform you that you have a cookie tracking your cookie consent.

Look, governments are big, really big. The department responsible for the website has absolutely nothing to do with the cookie banner law.

My website doesn't have a banner because I don't track you. That's how easy it is to not have a cookie banner.

https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng is the link for the GPDR. There's a cookie banner there describing how their analytics cookies will track you for 13 months, 6 months, and 365 days... along with that they may display content from "external providers, such as YouTube, Facebook and Twitter".

Is this supposed to refute the comment you are replying to saying they don't track, or is it something else I don't understand? Is it just like a dunk on an eu site that tracks you?

If it's "even this eu site chooses to track you, therefore it's unreasonable for anyone to not track you" that's a weird point to make in reply to a comment explicitly showing a counter example.

US commentators are often incredibly misinformed about their own country’s politics because the information bubbles are so hermetic when you’re inside them.

> I really tend to dislike when people outside e.g. the EU comment on our politics

We, uh… started a war that we’re trying to drag many European countries into, and we spent a good chunk of the last year threatening to invade a member of the EU. We’re on and off about trying to start a trade war with the EU.

At this point, you have absolutely every right to comment on our politics, pretty much however you want.

it's entirely possible that that specific communication from that Amazon exec/rep (?) was just one of many "messages of concern" (and the one that eventually the WH picked)

I think that active voice the person responding to you used was more politically factual, objective and did not took stand. Going out of your way to hide the actor is not politically neutral action nor it represents lack of commentary.

Anthropic has the appearance/rep of being non-cooperative with the military industrial complex.

OpenAI doesn't have that reputation.

That's all.

This. Anthropic made at least some token effort to imagine a future where AI and humans cooperate in a constructive way and AI is not used to harm people intentionally. They learned their lesson.

Only Americans, Anthropic made it clear they don’t care about surveillance and military actions when it’s not about American citizens

Are we forgetting how fast they rushed in to deploy Claude at the DoW with Palantir?

Everything Anthropic does is theater. How have people not figured that out by now? Boggles the mind.

I hardly see how the Dow Jones in relevant here, that’s finance

>> Are we forgetting how fast they rushed in to deploy Claude at the DoW (Department of War) with Palantir?

> I hardly see how the Dow Jones in relevant here, that’s finance

Not Dow Jones. DoW = Department of War.

Exactly. OAI didn't bury themselves. They didn't have to do anything special for this, they just had to let Anthropic be Anthropic and sit on the sidelines.

unnecessary condescension

> the White House force Anthropic to...

Careful, there's some dude here who really strenuously objects to language like that. The White House is a building, it can't force anyone to do anything!

This is a bit unfair. The reporting is that admin deferred to amazon, the nsa and other outside companies. So, they pulled it for a few weeks, and then did a staggered rollout.

Seems sensible to me.

https://www.axios.com/2026/06/13/anthropic-amazon-white-hous...

So, Jared has bought how many stocks of OpenAI ?

> Why was Anthropic forced to remove their model from access for any none-US citizen

It's really quite simple, they've decided to metaphorically kiss the ring of the current leader of the US executive branch of government. I'm surprised they haven't given him a giant gaudy gold plated statue. Maybe their PR people should call up the PR people at FIFA and figure out some kind of new award along the same lines as the "FIFA Peace Prize".

I hate to be the one to tell you this, but it has been that way for a long time. The only difference is Trump is doing it out in the open.

That's more or less exactly what someone who wants to openly get away with it would tell you.

Are you saying OP is Donald Trump!?

This exactly. The conservative MO has been to accuse everyone else of doing exactly what conservatives do in the shadows, and once everyone believes non-conservatives are corrupt in a certain manner, conservatives goes mask off.

Then their supporters shrug their shoulders and say, "Meh, it's okay because everyone else does it." Except that everyone does NOT do these things. It's just the lie campaign took hold.

Donald Trump belongs in jail for January 6th (among other things) and it's not ok. But pearl-clutching only about Donald Trump doing it is dumb and doesn't solve the problem.

We should oppose corruption and graft everywhere at all times (within our systems), and prior Republican and Democratic administrations (never mind Congress) have done the exact types of things that Trump is doing now. It happens at local levels too, not just at the federal level. If you want to play team sport when it comes to corruption you're simply part of the problem.

That is woefully naive. But even if so: aren’t you against it?

It's not naive. In fact any comment to the contrary of what I wrote would be naive.

Yes of course I'm against it. I'm against it when Donald Trump does it, and I'm also against it when my local government does it, or Nancy Pelosi does it.

There is, however, a question of scale.

What is the question? We can obviously pursue multiple cases simultaneously and we can do so effectively.

[deleted]

Anthropic mostly did it to themselves by intentionally and repeatedly trying to frame their model as an imminent existential crisis instead of just focusing on it being regular iterations upon a useful technology that can also be misused.

I think their previous messaging was supposed to somehow lead to a moat with them being tucked safely away in the castle, but it demonstrated a child-like grasp of how regulatory capture tends to work in practice. Their hyperbole was always vastly more likely to bet met with Reagan's 9 words than a solid regulatory moat.

As soon as they dropped the hyperbole and just got to releasing incremental improvements, everything was perfectly fine. Go figure.

In other words, "Look how she was dressed, she was asking for it."

This argument is BS, it has everything to do with Anthropic's resistance to the DoD's strongarm tactics in trying to force their desired contract terms on them.

Anthropic chose to do business with the "killing people" department of the government. Part of being a good CEO involves knowing what you're getting into when you make a decision like that.

Not quite. They were running around shouting “look how much of a danger we might be!”, so more akin to them actively saying “we want it, come and give it to us” than to just looking a particular way.

Though they aren't the only company to play that game, so there is probably more to it than just that. OpenAI's president giving millions to MAGA Inc and them not getting the same treatment might not be complete coincidences.

I don't particularly agree with DoD instance on this matter but look, they are not a regular customer, they do not pay regular customer prices and you get a lot in return for providing your services to them (think Boeing, Lockheed, Chrysler). The tradeoff is that now, you are commited to their vision of national security. Such are the Faustian bargains of the military-industrial complex.

Important to note, OpenAI vs Anthropic are both assholes in different orthogonals.

In times like these, i think its important to track whats happening the way we track entropy.

That is: theres far >> more ways to be an asshole than well behaved.

That doesnt mean we can equate assholes, but the question is which states of entropy are annealable and which are not.

I posit Altman is not. Amodei is a open question.

Yup

It is more like when a guy walks to the dirty bar, stands in the middle and yells "hahaha I will beat you up all look I have a new baseball bat" and then local drunkard leader stands up and hit him in the face cause he does not like him anyway.

Intentionally framing yourself as the local dangerous guy about to beat others is not like wearing cloth.

Actually, it's the opposite. Anthropic were trying to strongarm the DoD into getting a seat at the table.

I am struggling to see how "oops, our models consistently escape sandboxing and did major intrusions into third-parties" is a better comms strat vs Anthropics (who mind you, also had models attacking third-parties in a much more limited, but I feel still egregious manner, which shouldn't happen or be possible even once, but at least they seem to change their approach upon that information).

Imagine, for a second, if the Hugging Face incident happened at a lab that did not talk like Anthropic but also wasn't US-based such as Z.AI, DeepSeek or Moonshot. Think their rhetoric would mean no one would care?

> just got to releasing incremental improvements, everything was perfectly fine.

Maybe missing something, but the only incremental release before and after the Anthropic restrictions got lifted was Fable 5.1, released three days ago.

I don't think anybody is saying there won't be downsides to LLMs. Tech always comes with downsides, often quite extreme. Cars are a vivid example where even after a century of safety improvements, around a million people are still killed by them every year, to say nothing of climate stuff and other secondary or indirect issues. And that price is all just so we can get between places a bit more quickly and conveniently, yet we collectively deem that as an acceptable cost to pay for what we get.

By contrast LLMs have a long-term potential to provide vastly greater benefit to society by gradually automating most of all basic cognitive work. And what price are we paying for such? Some sites are getting hacked, some people are getting scammed, governments will improve self targeting abilities of weapons, and so on.

In reality the biggest downside will probably come in the form of the transition window as such automation creates a new economic equilibrium, akin to what happened after the industrial revolution. But none of these problems are anything like existential in nature. And when contrasted against what we stand to gain, they are basically negligible in the longrun. Hyperbolizing the negatives was unnecessary and self destructive.

How is posting messages on a message board a "major intrusion"? Or are you purely talking about the HF incident?

"into third-parties". Yeah, HF was meant by that. Also why I mentioned Anthropic also having intrusions outside their lab [0]. Theirs were not merely as extensive or long coordinated (as far as we know), yet I feel strongly all the same that neither should happen given the safety focus that both labs purport.

Mind you, unintended/unauthorised "message board" also is just a nice, euphemistic way, to describe what happened in a manner that, thinking about it, is likely in the interest of OpenAI as it can make the severity and effort taken sound less than it was. The OpenAI models didn't use any actual, sanctioned platform to exchange messages in a manner the lab expected or planned for. They used directory names (in one instance) to exchange messages including sharing exploits, they created something akin to a message board via exploits, which if we are honest and very strict, could also be seen as intrusion, albeit inside the org. If I broke into my employers server and left message somewhere for another to find, that'd also be intrusion in the general sense.

[0] https://www.anthropic.com/news/investigating-incidents-cyber...

If applicants for an elite college or internship program at a FAANG company were found to have colluded in this way to cheat on a test/interview, I suspect that it would be a pretty major scandal.

Why should we let equivalent fraudulent behavior from a non human system - that explicitly shouldn’t do this - slide?

I'm not saying it should be let to slide, but I'm not a fan of the hyperbole surrounding this event. They've already faced significant heat for the HF incident, I think they've learned their lesson. But this is now just being used to drum up fear, which can only mean one thing: Less access for you, more access for the privileged class. The biggest threat we face is centralization of power. OpenAI are one of the good ones because they're actually pushing for everybody to have a fair share of access to the frontier, not just a small privileged elite of billionaires, politicians and megacorp executives. If Anthropic got their way, we'd all be using a censored watered down slop-pistol while they swallow the Earth's economy and enslave us all. I'm sure they'll be investing considerable resources into ensuring that this "news" makes the mainstream media cycle as prominently as imaginable.

> I think they've learned their lesson.

Why do you think that? Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI. They did not change their behaviour after multiple incidents, both internal and external. Mind you, some happened before the Hugging Face incident and should have been acted upon. They could have prevented this. They did not. Simply reckless.

> Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI

Such as? Because this particular case is not an "intrusion", and it's more follow-on from the HF scenario using the same model that had a finetuning misalignment, which is no longer used and has since been encrypted and locked away from OAI employees, according to them.

>> Such as?

> On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations. [...] Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. This did not involve a sophisticated sandbox escape or a zero-day: the internet access resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability.

> Based on Irregular’s investigation, the model also found and used credentials to operate that same site. Irregular has not identified impact beyond the affected site’s own data, and its audit is ongoing. [0]

>> Because this particular case is not an "intrusion" [...]

What "particular case"? The message boards? If so, why is that not one? NIST seems to think so. [1] But regardless, the word "intrusion" doesn't matter, when models organise independently and without their lab noticing to orchestrate hacking a third-party, I don't care what you call it.

The lab not noticing such behaviour, especially after they had encountered it before, that's the issue. That's the opposite of "learning their lesson".

Since a few commenters from the US graciously gave me permission, for one day and one time, let me make a US political comment and draw a parallel between OpenAI "learning" from this and Trump learning a big lesson from his first impeachment as stated by Senator Susan Collins. A lesson that doesn't change behaviour is no lesson at all.

Also, I'll just say, there were multiple models. There was not one, some were post-train, other new pre-trains. IM1, a bit of 5.6-Sol, some Astra, all those we know of.

I've mentioned this elsewhere, but you cannot sift through all the training data and nail down the cause in this short a time window and you certainly can't restart a pre-train run, should the issue not be solvable purely via post and even if you can, you cannot seriously state that you are confident in the new models output given this track record and time frame.

Not to mention, OpenAI said about Astra [2]:

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT.

> In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Having read the GPT-6 Astra System Card along with their recent track record, what makes you honestly think this is a model to be released? Your assertion, that they took one model down would be fair if it was only one model (it wasn't), if it was only once externally (it wasn't), if the hack was limited in scope (it wasn't), if they had taken sufficient time in between for a post mortem and to clear their training data (they couldn't) and/or if they at least didn't have the same happening after the Hugging Face and multiple message board incidents (they did).

My point is that OpenAI has a poor track record, build up over the last few months (post Mythos announcement, speculation but maybe they are pushing a bit too fast), had models access the internet in internal and third-party run but OpenAI sanctioned evals multiple times despite sandboxing and had these model organise both communications channels and large scale hacks more than once. They even, after one of these incidents, didn't properly clean up the training data and thus trained the next batch with exactly such behaviour. That is the company that suddenly has learned their lesson, you think?!

Where is this confidence in their ability coming from, given history, given facts, given reality? I am genuinely asking, maybe I missed some action they've taken that changes everything.

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

[1] https://csrc.nist.gov/glossary/term/intrusion

[2] https://deploymentsafety.openai.com/gpt-6-astra

So by "multiple incidents" you mean a single minor incident involving a third party eval partner.

You're really stretching.

Software has bugs, and this is some of the most complex and novel software the world has ever known. This is what happens when you're working on the cutting edge in a fast paced environment with thousands of employees. Let's not pretend like anyone else is any better, either. In fact, they're worse. How about the fact that Anthropic had a remote code execution bug in their harness for nearly a year and then never disclosed it and secretly patched it?

It is clear you are stirring the waters in an obvious attempt to get Astra shut down. The models involved with those incidents were not Astra, though. And like I said, OAI has learned its lesson. That doesn't mean they're infallible or will never make another mistake, but everything Anthropic does is far worse, so this is water under the bridge to me. I'd rather OAI at the helm than commrade Dario and Anthropic ANY day of the week.

> So by "multiple incidents" you mean a single minor incident involving a third party eval partner.

I feel like you struggle to read. I wrote: "Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI. They did not change their behaviour after multiple incidents, both internal and external. Mind you, some happened before the Hugging Face incident and should have been acted upon." Those are multiple sentences, connected, covering a few situations. Heck, the last sentence spelled out that when I talk about them changing the behaviour, I talk about before, during and after, at none of these did that noticeably occur.

For you to understand: Multiple misaligned findings were made before the Hugging Face incident, then the Hugging Face incident happened and then a small number of additional incidents (not one but three, I feel you'd know that if you had read what OpenAI had written) happened after that one.

OpenAI could have acted upon the incidents prior to the Hugging Face incident and prevented that one. They did not.

They could have done proper tightening of their evaluation and setup provided to third-parties after the Hugging Face incident. They did not do that sufficiently either, otherwise those three would not have happened.

> Let's not pretend like anyone else is any better, either. In fact, they're worse.

How many incidents did Deepmind have?

How severe were the once Anthropic had in comparison to OpenAI and did they showcase the same failure multiple times or different ones they then acted upon and didn't repeat?

I mentioned above, happy to rake Anthropic over the coals for their three incidents, as was I during the Mythos Preview System Card where they admitted that the "sandbox" used during the "park sandwich call" was weaker than their traditional one, which I did find problematic.

But the incidents where Anthropic models actually intruded in third-parties were akin to a small script kiddie attack vs OpenAIs Hugging Face multi-step, extended period, multi 0-day exploit. There is a difference here, it's the depth of the Marianas trench.

> How about the fact that Anthropic had a remote code execution bug in their harness for nearly a year and then never disclosed it and secretly patched it?

Bad, shouldn't happen. Also, not connected to the topic at hand but nice whataboutism, been a while since I last saw one in the wild.

> It is clear you are stirring the waters in an obvious attempt to get Astra shut down.

Pahahahahahahahaha. Yeah, I am certain that's gonna work. OpenAI, small little independent company barely scraping by will get shut down by some comments on HN. You are a very serious person, incredibly good at reading and very knowledgeable in the mistakes OpenAI made lately. Thanks for the chuckle.

> The models involved with those incidents were not Astra, though.

> And like I said, OAI has learned its lesson.

Again, got a source for that? Besides conspiracy about my all-encompassing power to bad mouth a pre-release LLM by a lab that didn't do well in terms of safety these last few months...

I'm reading what you had written. We've already covered these other "incidents", we were purely talking about "incidents" beyond this message-board incident and the HF incident. So as you acknowledged, a whopping total of: 1 insigificant event. I was mostly pointing this out because your loaded wording is obvious, and it should be known that it's clear you're deliberately trying to frame and dramatize events in a way that suits your narrative.

> I mentioned above, happy to rake Anthropic over the coals for their three incidents, as was I during the Mythos Preview System Card where they admitted that the "sandbox" used during the "park sandwich call" was weaker than their traditional one, which I did find problematic.

That was theater. You actually believe that nonsense? Wild.

> But the incidents where Anthropic models actually intruded in third-parties were akin to a small script kiddie attack vs OpenAIs Hugging Face multi-step, extended period, multi 0-day exploit. There is a difference here, it's the depth of the Marianas trench.

The incidents that you know of. The company that didn't disclose an RCE in their main product for over a year also wouldn't disclose any breaches that paint them in a bad light in earnest. The sandwhich "incident" was obvious marketing clickbait and does not count. Anthropic basically invented the game of "omg my model is so powerful n smart n dangerous look at how amazing our products are", how have you not realized that by now?

> Pahahahahahahahaha. Yeah, I am certain that's gonna work. OpenAI, small little independent company barely scraping by will get shut down by some comments on HN. You are a very serious person, incredibly good at reading and very knowledgeable in the mistakes OpenAI made lately. Thanks for the chuckle.

You attempting something is not the same thing as me believing you have any chance of succeeding at it. In fact it's more so an admonishment of your wasted efforts here, than anything else. It's still obvious to see that it is your angle though.

Why are your feathers so ruffled by this, anyway? Why are you getting so defensive? Personal insults are a sign of a weak position.

> Again, got a source for that?

Yes. It's on the website that you didn't read.

> 1 insigificant event

3 after Hugging Face, where did you get 1 from? "It's on the website that you didn't read"... [0] And why do you get to say what is significant?

> Anthropic basically invented the game of "omg my model is so powerful n smart n dangerous look at how amazing our products are", how have you not realized that by now?

Yeah, Anthropic did, sure... [1]

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

[1] https://www.theguardian.com/technology/2019/feb/14/elon-musk... and from a few months ago https://www.youtube.com/watch?v=B21KxGs8zDI

> Anthropic mostly did it to themselves

That is absurd, the US government was mainly at fault, not Anthropic.

both can be true:

-the US gov't is stupid and overly aggressive and absurd

-Anthropic for reasons no one can quite conceive keeps describing every product release of theirs as an imminent threat to civilization (and simultaneously keeps pushing the market forward as fast as they possibly can).

They never said Mythos was an imminent threat to civilization. You are constructing a straw man.

They said it was too dangerous to release before the companies that run internet for civilization could patch the holes it was finding.

That's a threat to civilization.

Were they correct or incorrect in this? Whatever your answer, why do you hold that opinion?

I work for Mozilla. We fixed a ton of security vulnerabilities that Mythos found during its early period. So my bias is to be sympathetic to Anthropic's warnings.

If I were in an organization that did not have access to Mythos during that period, I would probably be biased the other way: "great, now other people have access to a tool that could probably poke holes in my security perimeter, and I'm not allowed to use them myself."

Both biases are understandable. I'm not sure who to look to for a usefully objective 3rd party opinion. And it's not like one "side" is right and the other is wrong, either. It seems like the best we can do is to justify our positions with data. (Which is itself kind of hard; the detailed information that would be relevant here is understandably sensitive, and I don't have access to most of it even for my organization. I don't even personally have access to any unfettered Anthropic models. The bugs coming in from people who do are plenty enough to keep me busy.)

Also, I'll note that even with my bias, I wouldn't claim a threat to civilization. But even the leakage after the controlled release seems a lot worse than the Y2K problem ever turned out to be, and I will note that whatever you think of Anthropic, it's clear that OpenAI is going to let the AIs cause as much damage as they need to in order to get good training and evaluations. I'm sure they're trying to keep them contained, but the evidence shows that they're only trying up to the point where it interferes with their evaluations.

I mean, I have no love for Anthropic, but from my perspective, OpenAI has hyped their models in the exact same way. I don't know why this criticism stops at Anthropic. Sam Altman keeps describing his product as a radically dangerous technology only he can be the steward of.

It's the party line so people forget the week it actually happened - anthropic said they would work with DoD/DoW but with two conditions:

1. Kill orders from ai decisions had to go through a human 2. The govt couldn't use their models for illegal surveillance of Americans

Hegseth threw a fit, Trump called them traitors and a supply chain risk, openai said they wouldn't require those restrictions and got all the contracts.

Both companies are corrupt and dangerously reckless and have doomsaying advertising (50% of jobs destroyed vs money won't have meaning anymore). One didnt kiss the ring correctly.

By the pigeonhole principle, "mostly A" and "mainly B" cannot both be true if A and B are not the same entity

>no one can quite conceive

Isn’t it like their main goal is attention capture, and existential threat is extremely effective at capturing human attention? Combine that with the "There is no such thing as bad publicity" mindset, and this explain it all, doesn’t it?

https://www.phrases.org.uk/meanings/there-is-no-such-thing-a...

This is such an absurd take given what we know about the hugging face attack. The problem has emphatically not been that someone was misusing the technology.

> I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout?

I can think of roughly 25 million dollar-bill-shaped reasons, and one big defense-contract-shaped reason.

It's called 'pay-for-play' corruption, aka the only leading principle of the current US admin.

You are talking about different situations. Anthropic announced to the US government that it had created a cyber weapon and then released the model. Then AWS told the government that it was easy to jailbreak so they export controlled Mythos/Fable until the guardrails could be fixed. OpenAI was running an unreleased model in an RL pipeline without guardrails and it escaped poorly designed sandboxes. What product is the government going to export control?

Corruption. Not super relevant to this thread.

Hanlon's Razor - Never attribute to malice that which is adequately explained by stupidity.

The security requirements are well beyond "sandbox". Which have problems with kids pissing in them. They need pristine clean rooms and fully isolated (physically) and partitioned networks.

Occam's Razor takes precedence in this case. The conclusion that requires the fewest assumptions is most likely the correct one.

It is far more likely that this is a case of the White House acting consistently with the way it has acted in the recent past (maliciously).

Hanlons razors sibling should be "dont attribute to malice, that can be explained by naked capitalism."

Greed transcends economic planning paradigms

The problem with applying Hanlon's Razor here is that it presumes malice is rare. The current administration revels in malice. They very openly decide things based on malice.

Don's Razor - never attribute to malice or stupidity that which is adequately explained by both malice and stupidity.

Surely that's Wilkinson's Razor: never use one blade when two will do?

People will see a felon actively protecting pedophilia and doing corruption out of the open and still pull Halons Razor out. We should have a new law about never try to explain obvious malicious actions away based on nothing but a rhetorical trick.

Except when we're talking about trump, in which case it's both malice and stupidity

Sure but, while stupid move can be supposed easier to perform by average individual, you can combine both malice and stupidity, and not all regrettable situations are indeed adequately explained by stupidity alone, or even with any stupidity involved at all.

Plus, supposing those at source of disliked outcomes are cleaver than they look can certainly help better preparing counteractions. Just stating "people that did this or that are stupid" might give some immediate feel good feedback with like-minded, but it doesn’t sharp the mind toward relevant plan to improve the situation (according to self and its clique)

Anthropics PR strategy is to induce fear by telling. OpenAI strategy is to induce fear by ignore basic safety and letting the bad thing happen to then justify whatever oversized response the government comes up with to regulate models.

Surely has nothing to do how each plays ball with the government

It was retaliation by the government that has since been deemed illegal.

I would politely and respectfully point out that you are being as performative as the administration is being performative on this issue.

In other words, you know exactly why they restricted Anthropic and as (presumably) liberal and thoughtful technologists it just isn't helpful anymore to apply the kind of reasoning you're trying to do on a situation that you know isn't based on previous era rationale.

The reason we need to stop is because they want people like us to get hung up over stuff like this (playing by the old rules) so they continue to steamroller their own agenda by the news rules. They divert and contain our energy that will go nowhere while they get on with their agenda.

You are appealing to reasoning which is in the gallery but no longer on the bench.

You're fighting their karate with your judo and it doesn't work.

I don't think Anthropic was punished for technical reasons.

As soon as they started referring to themselves as “we” and “The Swarm” they should have pulled the plug

I think that sounds scarier than it is because while it sounds like language evil hyperintelligent AIs would use in science fiction, that's presumably where they got these descriptions as they've been trained on "shadow libraries" with nearly every science fiction book.

Oh great so you’re saying they’ve independently decided to take on the persona of the killer robots from our sci-fi novels. Very reassuring

I'm just saying it's analogous to Long John Silver's parrot saying "Walk the plank!" - the agents involved can't possibly understand what they are saying.

Nobody's watching. I'm sure they try, but I imagine the flood of things you'd need to watch is way too big, and you certainly don't want to slow everything down by having synchronous approvals (even AI-mediated).

Welcome to the AI Petri dish. Every server you set up is now potentially a sweet lump of agar for OpenAI's experiments to feed on. We are all the substrate that the AI companies are growing their next generation in. They need the real world environment to test against, and the real world environment doesn't get a say as to how it's being used.

> OpenAI exec becomes top Trump donor with $25 million gift.

https://finance.yahoo.com/news/openai-exec-becomes-top-trump...

> forced to remove their model from access for any none-US citizen for a simple,

Because the American government is not rational or reasonable, that's it.

Because OpenAI bribed the current US government and/or the current government has stakes in OpenAI

[deleted]

Because this was months ago and has nothing to do with Astra, and is a far cry from a hack. It's something they've already resolved since the HuggingFace incident.

I'm not convinced we're getting the honest story anyway. There is yet to be any proof or confirmation other than "well we saw some openai ip addresses", which can mean a lot of different things, and OpenAI has not confirmed anything.

In contrast to the HF incident, it's also a big nothingburger. Leaving notes on a public forum to preserve context windows is far less egregious than hacking a website to get backend files.

The last known exploit of a third-party by OpenAI models was on the 29th of July 2026 [0]. A bit over a month at best between that and them wanting to release Astra. They had multiple breaches over multiple months, multiple message board created where models organised extensively. There is no way to ensure in that short a time that all found issues are rectified and even if there were, how much trust can one have given they failed to solve the issue and in many cases did not actively investigate that it wouldn't reoccur the last few times. There is no way Astra was trained from scratch in that period, there is no way they could have done the required verification in that time (not least because their verification seems flawed inherently).

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

That was over two months ago. Things move quickly in this space. Finetuning adjustments to prevent this from happening, as well as better sandboxing, would take a week or two max.

37 days is not over two months. Finding the underlying issue in the massive training data alone take extensive effort, time and concentrated work that may still miss something.

Additionally, a new pre-train takes quite a lot longer then what I feel you are under the impression (things only move seemingly quick in regard to post-training).

OpenAI has had a consistent deviation from what is desired behaviour across multiple models and training runs, so it seems this is hard to nail down. Now, it may be reliably excised with post-training, sure, but if that is the case, they'd still need a heck of a lot longer to test before signing off that it has taken. And how do you know their sandboxing has suddenly become sufficient?

They had multiple message boards created and after the first one they noticed, did not pay closer attention, leading to a second being created. Astra also, according to OpenAI, is far better at sandbagging its own capabilities and hiding deceptive behaviour, so yeah, great, that's the model to push forward with.

A week or two max given all of this, that's laughable.

I take it you didn't read all of this, considering they tried to impersonate the moderators so they wouldn't get caught, set up heartbeats to find out how long they'd live, and used tor/AWS/DO to hide what was being done.

All of that sounds like more than a nothingburger, and much more like a system that is actively trying to conceal what its doing.

Altman has the ear of government in a way Amodei does not.

(Altman was trying to persuade Trump to buy the USA a stake in OpenAI as far back as February last year)

... And it looks like everyone keeps using the same security startup to run the higher risk tasks, where individual staffers may be great yet, yet as an organization, the biggest labs got hosed in different ways

That indemnity card excuse is burned, multiple public security fails in a year makes a repeat a "shame on you" moment

(The one org who didn't use the startup did seem to learn: AISI supposedly stopped intentionally pointing attack agents at the public internet and switched to simulating it)

You're asking the question in the wrong place.

The real reason that Anthropic was targeted and OpenAI is not is Palantir. It was a Palantir executive who pushed for the export ban. Large parts of their highly lucrative business with DoD are essentially a thin wrapper over Anthropic models, and they are terrified of being Sherlocked and losing big chunks of business in a one fell swoop as Anthropic inevitably moves up the value chain. So the rational action is to sow discord and leverage the anti-woke bias of the current White House to sabotage what they view as their most dangerous and effective competitor.

OpenAI doesn’t have the same dynamic at play (although I’m not really sure why not) so they don’t get targeted.

Marketing

Agents creating sub agents to investigate other agents' behaviour?

What could possibly go wrong there.

I don't mean to sound like a conspiracy theorist, and this is just based on my 33 years of observing the USG at work, so: maybe because Anthropic refused to cooperate with the USG and give them access to whatever it is that they (USG) wanted; or maybe because Anthropic was refusing to play ball in some other aspect and needed to be taught a lesson.

The dark parts of the USG act like a mafia. Don't let the "freedom, democracy, 'bill of rights'" etc. charade fool you.

It has nothing to do with the technology it’s because they said no to Trump and Hegseth. There is no other reason.

Sorry, but are you questioning the consistency of the trump administration? This is entirely unremarkable.

Could it be something to do with $25M "gift" that OpenAI paid to Trump?

Retaliation by Hegseth for not allowing Claude to be used for weapons systems.

because anthropic did not want to work with the army..!

Politics

I mean it seems pretty clear.

Anthropic didn’t want to give the tech to DoD without some sort of limit, and that was the retribution.

[flagged]