"What happens when any [COMPANY] in the world stops caring about this? What if they let an experimental, cutting-edge [PRODUCTS] with no safety features (or worse, one that's [DESIGNED] to be malicious) on [ANYWHERE] and give it a simple goal? A goal like 'make the most money, by any means necessary', 'find a way to leave this payload on as many computers as possible', 'flood all websites using this language with garbage and make their internet completely unusable', 'get this person imprisoned or killed at any cost'."
Bro, this is what we literally, currently, have rn. lmfaol.
No, we have something that's less apocalyptic right now. You're talking about abuse, I was talking about the automation of abuse that's faster and more pervasive than anything individual bad actors could've done in the past. It's like if companies found a way to quickly and cheaply poison the entire world's drinking water supply, and then others argue that Nestle has already restricted the supply of water for profit on a smaller scale in the past, so this isn't new or worth caring about.
You aren't understanding this at all.
The frontier labs have hundreds of the best people in the world working on safety and alignment. They care deeply.
What happens when some random Chinese open source model, distilled on Astra, gets alliterated and now has no guardrails? Any script kiddie in the world could wreak havoc with it.
It turns out that guardrails matter.
> They care deeply.
Until it clashes with their quarterly revenue reports.
Meaningless "big corpo bad" statement, Anthropic at least has sacrificed a good amount of market cap for safety with DoW. OpenAI paused RL for two weeks.
OpenAI literally trained this behavior into their model while benchmaxxing ExploitGym so "number goes up" on the next model scorecards. Anthropic is also training on the same benchmarks [1] specifically for cyberattacks to keep up with OpenAI.
The coverage of this attack is so focused on this as an emergent behavior given what it conjures in the imagination, but it's the byproduct of millions of iterations of RL to improve AI agents' offensive capabilities. No one made OpenAI or Anthropic do that, the benchmarking arms race of their own creation now incentivizes them to keep doing it and evidently their AI Safety people can't or don't want to stop it.
[1] https://www.mpi-sp.org/108048/ExploitGym__Can_AI_Agents_Turn...
I can see where you're coming from and I'm sure the safety teams at the labs have good intentions, but I think your faith in the leading labs to self-regulate is misguided. The employees themselves have said as much with the Pacing the Frontier letter asking for external regulation [1]. These rogue agent incidents and the reckless development practices that led to them are a direct outcome of the competition between the leaders. Even the good-faith two-week pause from OpenAI did not lead to anything more; they need outside intervention.
[1]: https://www.pacingthefrontier.com/
It is hard to take their concerns for safety seriously, when they have been constantly talking about how dangerous their latest model is, before then deciding to release it to the public.
> before then deciding to release it to the public.
You mean, with safeguards that block the dangerous things? Safeguards so aggressive that the public complains about them?