Go and get one of their models to hack something, it won't do it, why?

They have claimed this happened during a "training run", but why are they training on systems connected to the internet?

That's why people are skeptical.

The public models won’t hack because they have a classifier that shuts down anything that looks like hacking; without the classifier they are perfectly capable of hacking, multiple third-party evaluators have confirmed this.

The models were not trained on systems intentionally connected to the internet; they chained mutliple zero-days (that they discovered) together to get access to the open internet and into huggingface.

> The models were not trained on systems intentionally connected to the internet...

If Amazon connects an AWS Top Secret region to the Internet, it doesn't matter whether or not it's intentional... they're getting nailed to the wall by the US government either way. Frankly, it's way worse for them if it was accidental; deliberate, sophisticated sabotage is a much better story than rank incompetence and/or negligence.

A similar sort of thing applies to the manufacturers of tools that they claim to be dangerous, that have been deliberately built to exceed their authorized access to other computer systems, and are deliberately being tested on how well they can do the thing they've been built to do.

Deliberate, sophisticated sabotage by one or more humans in their employ is much more forgivable than "Whoopsie, we didn't think to make it literally impossible to connect this dangerous automated computer-hacking tool to the Internet.".

It seems that we’re mostly in agreement? I agree that OpenAI has been terribly irresponsible, and that this attack being an accident makes things worse.

What I dispute is that AI agents are simple tools. I think rogue is an accurate word to describe them; I think what OpenAI is doing is more akin to gain-of-function research on a dangerous lifeform. I think this attack would have been prevented by air-gapping, but that wouldn’t solve the fundamental issue which is that they are creating something dangerous that they have no idea how to control