I don't understand why so many comments here are so confident that this is all marketing, that rogue is just hype, that agents are just simple tools, etc. If a bunch of nuclear engineers were going to the news and saying "Our reactor is dangerously close to a meltdown - we need government intervention now!" would your response be that they're just hyping up boring old power generation technology?
These companies are building software. That doesn't work very well. The output it produces does make sense at times, but there are times when it doesn't. And instead of fixing that, or admitting it can't fixed, they started bolting actuators to them, executing actions online (for now) based on the output of their buggy software.
And when this results in actuators executing some bad actions they scream in horror "AI went rogue! It escaped the containment!!! It's going to kill us all!!!"
Go fix your software before you let it do stuff online or IRL. It's not "Terminator", it's just bad QC.
But the thing is, they are going to keep bolting more and more actuators on, and training more and more powerful agents, and we (society, especially the tech industry) are going to keep using them, because they are extremely useful. And I don't see why you're so confident that frontier agents can't get powerful enough to do serious, real, lasting damage to the world; as far as I can tell, AI models have been improving at an accelerating rate, and there is no sign that that is slowing down or will slow down in the near future.
For some reason, people think that it's money that's motivating the billionaires' performances. Imagine.
You wouldn't be suspicious if the nuclear engineers kept on feeding their reactor and building more powerful reactors while they went crying to the news?
In the nuclear case, no, because I’d assume they still have a mortgage to pay.
In the AI case, no, because I think the engineers believe that there is a high probability of enormous upside as well, if it doesn’t kill us all.
I’d roll my eyes if the engineers stated that they didn’t design the reactor to melt down, and that it simply developed rogue meltdown-desiring behavior on its own, and I would also wonder about negligence if they claimed that nobody could have anticipated this (given that, like with botnets and viruses, we have decades of knowledge and experience regarding reactor meltdowns)
I mean sure, negligence is absolutely on the table; but that makes the problem worse, not better! We don’t allow nuclear engineers to be negligent; they can go to jail if they don’t follow strict protocols to make sure the dangerous systems they work on are safe.
Go and get one of their models to hack something, it won't do it, why?
They have claimed this happened during a "training run", but why are they training on systems connected to the internet?
That's why people are skeptical.
The public models won’t hack because they have a classifier that shuts down anything that looks like hacking; without the classifier they are perfectly capable of hacking, multiple third-party evaluators have confirmed this.
The models were not trained on systems intentionally connected to the internet; they chained mutliple zero-days (that they discovered) together to get access to the open internet and into huggingface.