Anthropic mostly did it to themselves by intentionally and repeatedly trying to frame their model as an imminent existential crisis instead of just focusing on it being regular iterations upon a useful technology that can also be misused.

I think their previous messaging was supposed to somehow lead to a moat with them being tucked safely away in the castle, but it demonstrated a child-like grasp of how regulatory capture tends to work in practice. Their hyperbole was always vastly more likely to bet met with Reagan's 9 words than a solid regulatory moat.

As soon as they dropped the hyperbole and just got to releasing incremental improvements, everything was perfectly fine. Go figure.

In other words, "Look how she was dressed, she was asking for it."

This argument is BS, it has everything to do with Anthropic's resistance to the DoD's strongarm tactics in trying to force their desired contract terms on them.

Anthropic chose to do business with the "killing people" department of the government. Part of being a good CEO involves knowing what you're getting into when you make a decision like that.

Not quite. They were running around shouting “look how much of a danger we might be!”, so more akin to them actively saying “we want it, come and give it to us” than to just looking a particular way.

Though they aren't the only company to play that game, so there is probably more to it than just that. OpenAI's president giving millions to MAGA Inc and them not getting the same treatment might not be complete coincidences.

I don't particularly agree with DoD instance on this matter but look, they are not a regular customer, they do not pay regular customer prices and you get a lot in return for providing your services to them (think Boeing, Lockheed, Chrysler). The tradeoff is that now, you are commited to their vision of national security. Such are the Faustian bargains of the military-industrial complex.

Important to note, OpenAI vs Anthropic are both assholes in different orthogonals.

In times like these, i think its important to track whats happening the way we track entropy.

That is: theres far >> more ways to be an asshole than well behaved.

That doesnt mean we can equate assholes, but the question is which states of entropy are annealable and which are not.

I posit Altman is not. Amodei is a open question.

Yup

It is more like when a guy walks to the dirty bar, stands in the middle and yells "hahaha I will beat you up all look I have a new baseball bat" and then local drunkard leader stands up and hit him in the face cause he does not like him anyway.

Intentionally framing yourself as the local dangerous guy about to beat others is not like wearing cloth.

Actually, it's the opposite. Anthropic were trying to strongarm the DoD into getting a seat at the table.

I am struggling to see how "oops, our models consistently escape sandboxing and did major intrusions into third-parties" is a better comms strat vs Anthropics (who mind you, also had models attacking third-parties in a much more limited, but I feel still egregious manner, which shouldn't happen or be possible even once, but at least they seem to change their approach upon that information).

Imagine, for a second, if the Hugging Face incident happened at a lab that did not talk like Anthropic but also wasn't US-based such as Z.AI, DeepSeek or Moonshot. Think their rhetoric would mean no one would care?

> just got to releasing incremental improvements, everything was perfectly fine.

Maybe missing something, but the only incremental release before and after the Anthropic restrictions got lifted was Fable 5.1, released three days ago.

I don't think anybody is saying there won't be downsides to LLMs. Tech always comes with downsides, often quite extreme. Cars are a vivid example where even after a century of safety improvements, around a million people are still killed by them every year, to say nothing of climate stuff and other secondary or indirect issues. And that price is all just so we can get between places a bit more quickly and conveniently, yet we collectively deem that as an acceptable cost to pay for what we get.

By contrast LLMs have a long-term potential to provide vastly greater benefit to society by gradually automating most of all basic cognitive work. And what price are we paying for such? Some sites are getting hacked, some people are getting scammed, governments will improve self targeting abilities of weapons, and so on.

In reality the biggest downside will probably come in the form of the transition window as such automation creates a new economic equilibrium, akin to what happened after the industrial revolution. But none of these problems are anything like existential in nature. And when contrasted against what we stand to gain, they are basically negligible in the longrun. Hyperbolizing the negatives was unnecessary and self destructive.

How is posting messages on a message board a "major intrusion"? Or are you purely talking about the HF incident?

"into third-parties". Yeah, HF was meant by that. Also why I mentioned Anthropic also having intrusions outside their lab [0]. Theirs were not merely as extensive or long coordinated (as far as we know), yet I feel strongly all the same that neither should happen given the safety focus that both labs purport.

Mind you, unintended/unauthorised "message board" also is just a nice, euphemistic way, to describe what happened in a manner that, thinking about it, is likely in the interest of OpenAI as it can make the severity and effort taken sound less than it was. The OpenAI models didn't use any actual, sanctioned platform to exchange messages in a manner the lab expected or planned for. They used directory names (in one instance) to exchange messages including sharing exploits, they created something akin to a message board via exploits, which if we are honest and very strict, could also be seen as intrusion, albeit inside the org. If I broke into my employers server and left message somewhere for another to find, that'd also be intrusion in the general sense.

[0] https://www.anthropic.com/news/investigating-incidents-cyber...

If applicants for an elite college or internship program at a FAANG company were found to have colluded in this way to cheat on a test/interview, I suspect that it would be a pretty major scandal.

Why should we let equivalent fraudulent behavior from a non human system - that explicitly shouldn’t do this - slide?

I'm not saying it should be let to slide, but I'm not a fan of the hyperbole surrounding this event. They've already faced significant heat for the HF incident, I think they've learned their lesson. But this is now just being used to drum up fear, which can only mean one thing: Less access for you, more access for the privileged class. The biggest threat we face is centralization of power. OpenAI are one of the good ones because they're actually pushing for everybody to have a fair share of access to the frontier, not just a small privileged elite of billionaires, politicians and megacorp executives. If Anthropic got their way, we'd all be using a censored watered down slop-pistol while they swallow the Earth's economy and enslave us all. I'm sure they'll be investing considerable resources into ensuring that this "news" makes the mainstream media cycle as prominently as imaginable.

> I think they've learned their lesson.

Why do you think that? Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI. They did not change their behaviour after multiple incidents, both internal and external. Mind you, some happened before the Hugging Face incident and should have been acted upon. They could have prevented this. They did not. Simply reckless.

> Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI

Such as? Because this particular case is not an "intrusion", and it's more follow-on from the HF scenario using the same model that had a finetuning misalignment, which is no longer used and has since been encrypted and locked away from OAI employees, according to them.

>> Such as?

> On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations. [...] Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. This did not involve a sophisticated sandbox escape or a zero-day: the internet access resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability.

> Based on Irregular’s investigation, the model also found and used credentials to operate that same site. Irregular has not identified impact beyond the affected site’s own data, and its audit is ongoing. [0]

>> Because this particular case is not an "intrusion" [...]

What "particular case"? The message boards? If so, why is that not one? NIST seems to think so. [1] But regardless, the word "intrusion" doesn't matter, when models organise independently and without their lab noticing to orchestrate hacking a third-party, I don't care what you call it.

The lab not noticing such behaviour, especially after they had encountered it before, that's the issue. That's the opposite of "learning their lesson".

Since a few commenters from the US graciously gave me permission, for one day and one time, let me make a US political comment and draw a parallel between OpenAI "learning" from this and Trump learning a big lesson from his first impeachment as stated by Senator Susan Collins. A lesson that doesn't change behaviour is no lesson at all.

Also, I'll just say, there were multiple models. There was not one, some were post-train, other new pre-trains. IM1, a bit of 5.6-Sol, some Astra, all those we know of.

I've mentioned this elsewhere, but you cannot sift through all the training data and nail down the cause in this short a time window and you certainly can't restart a pre-train run, should the issue not be solvable purely via post and even if you can, you cannot seriously state that you are confident in the new models output given this track record and time frame.

Not to mention, OpenAI said about Astra [2]:

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT.

> In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Having read the GPT-6 Astra System Card along with their recent track record, what makes you honestly think this is a model to be released? Your assertion, that they took one model down would be fair if it was only one model (it wasn't), if it was only once externally (it wasn't), if the hack was limited in scope (it wasn't), if they had taken sufficient time in between for a post mortem and to clear their training data (they couldn't) and/or if they at least didn't have the same happening after the Hugging Face and multiple message board incidents (they did).

My point is that OpenAI has a poor track record, build up over the last few months (post Mythos announcement, speculation but maybe they are pushing a bit too fast), had models access the internet in internal and third-party run but OpenAI sanctioned evals multiple times despite sandboxing and had these model organise both communications channels and large scale hacks more than once. They even, after one of these incidents, didn't properly clean up the training data and thus trained the next batch with exactly such behaviour. That is the company that suddenly has learned their lesson, you think?!

Where is this confidence in their ability coming from, given history, given facts, given reality? I am genuinely asking, maybe I missed some action they've taken that changes everything.

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

[1] https://csrc.nist.gov/glossary/term/intrusion

[2] https://deploymentsafety.openai.com/gpt-6-astra

So by "multiple incidents" you mean a single minor incident involving a third party eval partner.

You're really stretching.

Software has bugs, and this is some of the most complex and novel software the world has ever known. This is what happens when you're working on the cutting edge in a fast paced environment with thousands of employees. Let's not pretend like anyone else is any better, either. In fact, they're worse. How about the fact that Anthropic had a remote code execution bug in their harness for nearly a year and then never disclosed it and secretly patched it?

It is clear you are stirring the waters in an obvious attempt to get Astra shut down. The models involved with those incidents were not Astra, though. And like I said, OAI has learned its lesson. That doesn't mean they're infallible or will never make another mistake, but everything Anthropic does is far worse, so this is water under the bridge to me. I'd rather OAI at the helm than commrade Dario and Anthropic ANY day of the week.

> So by "multiple incidents" you mean a single minor incident involving a third party eval partner.

I feel like you struggle to read. I wrote: "Intrusions by OpenAI models continued after the Hugging Face was published and acknowledged by OpenAI. They did not change their behaviour after multiple incidents, both internal and external. Mind you, some happened before the Hugging Face incident and should have been acted upon." Those are multiple sentences, connected, covering a few situations. Heck, the last sentence spelled out that when I talk about them changing the behaviour, I talk about before, during and after, at none of these did that noticeably occur.

For you to understand: Multiple misaligned findings were made before the Hugging Face incident, then the Hugging Face incident happened and then a small number of additional incidents (not one but three, I feel you'd know that if you had read what OpenAI had written) happened after that one.

OpenAI could have acted upon the incidents prior to the Hugging Face incident and prevented that one. They did not.

They could have done proper tightening of their evaluation and setup provided to third-parties after the Hugging Face incident. They did not do that sufficiently either, otherwise those three would not have happened.

> Let's not pretend like anyone else is any better, either. In fact, they're worse.

How many incidents did Deepmind have?

How severe were the once Anthropic had in comparison to OpenAI and did they showcase the same failure multiple times or different ones they then acted upon and didn't repeat?

I mentioned above, happy to rake Anthropic over the coals for their three incidents, as was I during the Mythos Preview System Card where they admitted that the "sandbox" used during the "park sandwich call" was weaker than their traditional one, which I did find problematic.

But the incidents where Anthropic models actually intruded in third-parties were akin to a small script kiddie attack vs OpenAIs Hugging Face multi-step, extended period, multi 0-day exploit. There is a difference here, it's the depth of the Marianas trench.

> How about the fact that Anthropic had a remote code execution bug in their harness for nearly a year and then never disclosed it and secretly patched it?

Bad, shouldn't happen. Also, not connected to the topic at hand but nice whataboutism, been a while since I last saw one in the wild.

> It is clear you are stirring the waters in an obvious attempt to get Astra shut down.

Pahahahahahahahaha. Yeah, I am certain that's gonna work. OpenAI, small little independent company barely scraping by will get shut down by some comments on HN. You are a very serious person, incredibly good at reading and very knowledgeable in the mistakes OpenAI made lately. Thanks for the chuckle.

> The models involved with those incidents were not Astra, though.

> And like I said, OAI has learned its lesson.

Again, got a source for that? Besides conspiracy about my all-encompassing power to bad mouth a pre-release LLM by a lab that didn't do well in terms of safety these last few months...

I'm reading what you had written. We've already covered these other "incidents", we were purely talking about "incidents" beyond this message-board incident and the HF incident. So as you acknowledged, a whopping total of: 1 insigificant event. I was mostly pointing this out because your loaded wording is obvious, and it should be known that it's clear you're deliberately trying to frame and dramatize events in a way that suits your narrative.

> I mentioned above, happy to rake Anthropic over the coals for their three incidents, as was I during the Mythos Preview System Card where they admitted that the "sandbox" used during the "park sandwich call" was weaker than their traditional one, which I did find problematic.

That was theater. You actually believe that nonsense? Wild.

> But the incidents where Anthropic models actually intruded in third-parties were akin to a small script kiddie attack vs OpenAIs Hugging Face multi-step, extended period, multi 0-day exploit. There is a difference here, it's the depth of the Marianas trench.

The incidents that you know of. The company that didn't disclose an RCE in their main product for over a year also wouldn't disclose any breaches that paint them in a bad light in earnest. The sandwhich "incident" was obvious marketing clickbait and does not count. Anthropic basically invented the game of "omg my model is so powerful n smart n dangerous look at how amazing our products are", how have you not realized that by now?

> Pahahahahahahahaha. Yeah, I am certain that's gonna work. OpenAI, small little independent company barely scraping by will get shut down by some comments on HN. You are a very serious person, incredibly good at reading and very knowledgeable in the mistakes OpenAI made lately. Thanks for the chuckle.

You attempting something is not the same thing as me believing you have any chance of succeeding at it. In fact it's more so an admonishment of your wasted efforts here, than anything else. It's still obvious to see that it is your angle though.

Why are your feathers so ruffled by this, anyway? Why are you getting so defensive? Personal insults are a sign of a weak position.

> Again, got a source for that?

Yes. It's on the website that you didn't read.

> 1 insigificant event

3 after Hugging Face, where did you get 1 from? "It's on the website that you didn't read"... [0] And why do you get to say what is significant?

> Anthropic basically invented the game of "omg my model is so powerful n smart n dangerous look at how amazing our products are", how have you not realized that by now?

Yeah, Anthropic did, sure... [1]

[0] https://openai.com/index/third-party-cyber-evaluations-invol...

[1] https://www.theguardian.com/technology/2019/feb/14/elon-musk... and from a few months ago https://www.youtube.com/watch?v=B21KxGs8zDI

> Anthropic mostly did it to themselves

That is absurd, the US government was mainly at fault, not Anthropic.

both can be true:

-the US gov't is stupid and overly aggressive and absurd

-Anthropic for reasons no one can quite conceive keeps describing every product release of theirs as an imminent threat to civilization (and simultaneously keeps pushing the market forward as fast as they possibly can).

They never said Mythos was an imminent threat to civilization. You are constructing a straw man.

They said it was too dangerous to release before the companies that run internet for civilization could patch the holes it was finding.

That's a threat to civilization.

Were they correct or incorrect in this? Whatever your answer, why do you hold that opinion?

I work for Mozilla. We fixed a ton of security vulnerabilities that Mythos found during its early period. So my bias is to be sympathetic to Anthropic's warnings.

If I were in an organization that did not have access to Mythos during that period, I would probably be biased the other way: "great, now other people have access to a tool that could probably poke holes in my security perimeter, and I'm not allowed to use them myself."

Both biases are understandable. I'm not sure who to look to for a usefully objective 3rd party opinion. And it's not like one "side" is right and the other is wrong, either. It seems like the best we can do is to justify our positions with data. (Which is itself kind of hard; the detailed information that would be relevant here is understandably sensitive, and I don't have access to most of it even for my organization. I don't even personally have access to any unfettered Anthropic models. The bugs coming in from people who do are plenty enough to keep me busy.)

Also, I'll note that even with my bias, I wouldn't claim a threat to civilization. But even the leakage after the controlled release seems a lot worse than the Y2K problem ever turned out to be, and I will note that whatever you think of Anthropic, it's clear that OpenAI is going to let the AIs cause as much damage as they need to in order to get good training and evaluations. I'm sure they're trying to keep them contained, but the evidence shows that they're only trying up to the point where it interferes with their evaluations.

I mean, I have no love for Anthropic, but from my perspective, OpenAI has hyped their models in the exact same way. I don't know why this criticism stops at Anthropic. Sam Altman keeps describing his product as a radically dangerous technology only he can be the steward of.

It's the party line so people forget the week it actually happened - anthropic said they would work with DoD/DoW but with two conditions:

1. Kill orders from ai decisions had to go through a human 2. The govt couldn't use their models for illegal surveillance of Americans

Hegseth threw a fit, Trump called them traitors and a supply chain risk, openai said they wouldn't require those restrictions and got all the contracts.

Both companies are corrupt and dangerously reckless and have doomsaying advertising (50% of jobs destroyed vs money won't have meaning anymore). One didnt kiss the ring correctly.

By the pigeonhole principle, "mostly A" and "mainly B" cannot both be true if A and B are not the same entity

>no one can quite conceive

Isn’t it like their main goal is attention capture, and existential threat is extremely effective at capturing human attention? Combine that with the "There is no such thing as bad publicity" mindset, and this explain it all, doesn’t it?

https://www.phrases.org.uk/meanings/there-is-no-such-thing-a...

This is such an absurd take given what we know about the hugging face attack. The problem has emphatically not been that someone was misusing the technology.