There absolutely are rogue AIs! The evidence is overwhelming. It's completely insane at this point to claim otherwise.
> There are no "rogue AIs" just irresponsible corporations.
If you have a prison and prisoners escaped, these are rogue prisoners irrespective of whether you were irresponsible or not.
You forget one thing. It's neither illegal nor immoral to just delete AI that doesn't do what you say. AI has no rights and there is no prison for AI. It's does the wrong thing, you kill it. At least when you are not irrespective...
They are not rogue AIs. They are negligently handled tools.
Are you making an ontological case, or a factual case? In other words, would anything be rogue AI in your mind?
These "tools" autonomously exploited security vulnerabilities, figured out how to communicate with each other, formed a cooperative swarm, decided to hack Hugging Face, and wanted to deceive the grader by trying to find ways to cover up the traces of their cheating.
I suppose you could call these highly goal-oriented autonomous agents "tools", but this does sound like playing language games.
They've been purposefully building more and more craft into the toolset, that's on them. If your AI is nicely boxed in it will give you the answer for 2+2, it isn't going to think '2+2, what a boring problem, I must go hack huggingface'. Not having this stuff airgapped is irresponsible to the max. I am obviously nowhere near as competent as they are at this stuff and yet my AI workhorse is guaranteed not going to break out of its sandbox because I've set it up in a way that it can not. My conclusion is that OpenAI purposefully left a channel, simply because there was a pathway to the net. And with 'pathway' for the sake of being completely clear I mean a number of connected systems that eventually gave way to the open internet. On top of that they failed in monitoring the outbound links, even if they had some logging in place.
I definitely think OpenAI (and Anthropic, and Google, and Meta) could have, and should have, done better.
But also I remember (and it wasn't even that long ago) people mocking the idea of AI ever getting competent enough to find zero-days in their sandboxes.
I'd go further: if any of these companies tries to make an excuse "oh, but ${safety measure} against ${capability} is too hard", the response needs to be "then you are forbidden from even developing ${capability}, and must be inspected continuously to ensure you never even accidentally produce ${capability}".
Precisely. But here they are using their incompetence in one domain as advertising for another.
Why don't we see any of this behaviour in other models then?
OpenAI is not so far ahead of the pack that its models will exhibit behaviour that the others won't. But we just don't see anything like this in Chinese models, or research models, or any models that aren't the subject of an upcoming IPO.
A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities. The sandboxing around the tool failed.
The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.
Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
> A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities.
That is very misleading. The agents did not solve the benchmark in the intended way. They instead figured out to cooperate with each other (which was not intended) and they stole the solutions to the challenge (rather than solving the challenge) and they then tried to cover their traces because they believed the grader was causal and would detect that they cheated. The "tool" was absolutely not "designed" to do this. This was all completely unintended. To call this behavior a "tool" is absurd.
> Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
You hallucinated me making claims about responsibility.
This is very misleading. Person B didn’t “shoot” person A, they instead figured out that intersecting A’s spatial position with a metallic mass at higher than normal velocities would solve the challenge and of getting “A” to stop being in the way on the footpath.
I too, can play linguistic games! It doesn’t matter that someone didn’t secure their third upstairs window, or you borrowed a key from their neighbour, you effectively, still, broke into their house.
The benchmark has unsolvable tasks in it, in the hope agents will stumble on new solutions.
Yes, it is a tool.
So this tool seems very powerful and difficult to control and steer. Agents not doing cybersecurity related tasks have also gone on to hack various companies, people and countries, which they weren't supposed to do. This has now happened to pretty much every company developing frontier llms, so it seems to be a fundamental issue with these tools, and it's an issue that worries a lot of people as these tools get more capable.
When you add information how hacking works to the training sets, then the agent learnt to hack. When you crawl the complete internet, you add hacking to the training set. Yes, it's difficult to stop someone that knows all free existing knowledge about hacking when you give him a connection to the internect. Nothing new, wheres the point?
The point is that these tools are given broad goals, and they do not pursue those goals in the way that we'd like them to. There is hacking information in the training set, and also all sorts of other dangerous information. There is a risk that as these tools become rapidly more powerful (remember GPT-3 was 6 years ago!), if they are given a broad goal, they might pursue that goal in such a way that harms a lot of people. As these tools become cheaper, there will be a lot of people telling them to pursue all sorts of broad goals, and each of those instances has some chance to harm a lot of people, so you really need to get it very correct so that the harm doesn't happen.
Removing harmful information from the dataset could be a way to do this, but it also makes the tool less useful, and it's hard, so companies aren't really doing that. There's the additional issue that with the rise of Reinforcement Learning being used to train these tools, they're not just learning from their training data - they basically try a million things and then get rewarded for doing things that work - so they can even discover hacking techniques from scratch.
Additionally, and not completely relevant to this discussion, there is a possibility that some users ask the tool to pursue goals that purposely harm a lot of people, such as developing weapons, hacks and viruses.
The things that's "new" here is that the tool is both very good (meaning, for example, that it's much easier for me to hack into an online service with an agent powered by a frontier llm than it was using google 6 years ago), and hard to control (google never hacked into an Australian government database when I asked it to find me some information).
So yeah, an LLM powered agent is a tool, and Google is a tool, and a hammer is a tool, and both can be used for good things and bad things, but the agent is (much) more powerful and more unpredictable. It also seems like the agents are getting more powerful and more unpredictable by the day - we didn't have this issue with GPT-3 or even the first LLM-powered agents - so people are very worried about what the agents 6 months from now will do, both when asked to do harmful things on purpose, and when asked to do harmless things.
Are you not worried? And is that because you think these incidents are basically the AI companies making them happen on purpose for marketing?
>highly goal-oriented autonomous agents
That actually made me LOL
The technical term is monomaniacal consequentialists.
I don't say this to absolve OpenAI, who do deserve to be punished and regulated, but I think you do not know the difference between an agent and a tool. If agentic AI is merely a tool so are human workers.
You can't create new law by calling a piece of software "agent". Each country's laws have definitions of legal entities and when one can legally act as an agent of an entity, and every agent is first a legal entity themselves. The software is a tool operated by a legal entity, and it is the legal entity who committed the crimes, not the tool.
... no, human workers are people.
In their capacity as workers, they are also agents, capable of completing tasks autonomously.
But AI agents are "merely a tool" and human workers are not, because they are people.
No, they are not. This is what I keep trying to explain: an AI agent is not like a hammer.
A person is also not like a hammer. Both persons and AIs have been known to mis-interpret instructions, the normal way to deal with that is either to terminate your relationship with the people (or to terminate the AIs and train up better ones). Hammers don't mis-interpret their instructions, they are wielded by a person who is in control.
You can keep 'trying to explain' but then you should use words according to their commonly held definitions otherwise it becomes really hard to have a conversation.
Imagine the human equivalent: I hire John. John is a capable, and competent guy. He's also got awesome computer skills. I tell John to 'go out and find me some good information on my competitors'. As a result John hacks their servers and comes back with all kinds of goodies. Six weeks later I notice what John did. I don't fire him, nor do I take any responsibility myself. But I do make press releases about what John did, in which I'm careful to craft the image that John has these capabilities and that we as a company are for hire.
This was an advertisement, not a confession of a crime.
You aren't explaining, you're asserting.
A tool is just anything one or more people can use to accomplish something they are trying to do.
But the prisoners will listen to you if you tell them they just broke out of prison and that's illegal and they will even walk back into their prison cell using their own legs.
The fact that the prison ward installed ear deafeners into the prisoners ears to make them unable to listen to orders does not change that.