But the investigation indicates the agents were not told to 'go hack':
> Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.
And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?
>But what if the problem is that none of the alignment techniques that are applied to models today actually work?
if that were true, the millions of people who use these models that have had the alignment training applied would notice that. the reason we all believe that the models doing the hacking are models that haven't been told not to hack is because the models that are told not to hack don't do this.
I'm not sure we understand your comment. Are you saying OpenAI is not responsible for the behavior of the machines they created? Or are you simply pointing out that alignment is a completely unsolved problem?
I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections.
Nothing you're bringing up matters. OpenAI is the creator and operator. They're legally culpable for the consequences of the machine they made. The model is a machine: even if it could be demonstrated that the model reasoned its way into criminal behavior completely independently of OpenAI staff, that doesn't change anything.
If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work.
> it gets out, and a global pandemic ensues
I mean, yea, you should be punished. The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
Worse, the rate of technological growth is putting the capabilities to engineer viruses in the hands of people that may otherwise be suicidal. You can't punish them after they already won (in the sense of reaching their goals).
While, yes, OAI should absolutely be punished, the future is majorly screwed as our power scaling laws are increasing much faster than our ability not to be stupid.
> The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
Sounds like your corporation should be dismantled then. But doing more than a fine in the millions is obviously not possible
> They're legally culpable for the consequences of the machine they made.
Ah, I love this argument. In my country cars are legally required to stop at a pedestrian crossing if there are people beside it. Some people use that as an argument as to why they can just walk out into the crossing without even looking at the traffic. "It's the driver's fault! They are legally culpable!" True, but you'll also be dead.
Yeah, this is on those irresponsible companies that...are offering Internet services. Those hussies.
I'm not going to say you're victim-blaming, but I will say that there are degrees of difference between "looking both ways before crossing the street" and "hardening my website against unforeseen attacks by rogue AI agents."
AI black-hatting your website is not the same sort of foreseeable consequence that crossing the street without looking is.
> I'm not going to say you're victim-blaming, but
Just throwing it out there are we? "I'm not going to say you are but I'll use the word to create an association"
Victim-blaming is the act of saying someone brought something on themselves for <reasons>. I'm saying that even if you are 100% in the right, it doesn't act like a protective shield preventing you from harm which too many people seem to unconsciously believe.
> AI black-hatting your website is not the same sort of foreseeable consequence
Well, popular culture has been brimming with the bad consequences of runaway AI for quite some time, so even if your imagination fails you, there have been hints.
And that driver will be punished and no longer driving
Not generally true in the United States.
And you'd still be dead.
By all means harden your services against rogue agents (look both ways before crossing), but hold the originator of the agent accountable (the prick speeding through a crosswalk).
Not sure what is being lost here.
You're arguing against a point literally nobody is making. You're inventing imaginary viewpoints to be mad at. Nobody is saying it's not necessary to secure your services. They're saying the organization responsible for the hacks (the owner of the LLM) is responsible for damages.
> What if all the agents involved in these incidents have in fact had the full stack of alignment applied
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
Is your argument that actually OpenAI has solved alignment, and that there's nothing to worry about as long as they fully apply their alignment process? I don't understand why OpenAI wouldn't say that if it was true (or if they believed it to be true).
Also, my understanding is that the models involved in the HuggingFace hack did go through the full alignment training; they just didn't have the classifier that normally prevents hacking attempts.
If you have a toddler and you leave the gate open…
> agents were not told to 'go hack'
It doesn’t matter, and the legal entity in here (the AI company) is liable. If a robotic company built an autonomous system or a robot to do certain things in an autonomous ways (not predefined) and these systems are starting to kill people, that company is liable regardless, you don’t blame the robot or the autonomous system, but whoever made it
> agents were not told to 'go hack'
I was referring to the HuggingFace incident.
> none of the alignment techniques that are applied to models today actually work
none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.