OpenAI and Anthropic have published a lot on the need for AI alignment + the research they're doing to ensure alignment/safety, yet they are also responsible for the highest profile misalignment incidents so far (HuggingFace incident, AISI Mythos social engineering, and now this).
One interpretation of this is that they are being deliberately dishonest about their priorities. Another interpretation is that we cannot rely on the labs to self-regulate, because the labs don't trust each other, and there will always be pressure to go to market faster than their competitor.
Either way I think it's pretty non-controversial that the labs are the source of the danger?
> yet they are also responsible for the highest profile misalignment incidents so far
They are the only ones posting about them or admitting to them. That does not mean "the most misalignment incidents so far." You don't know what other attacks have happened (and it's very easy to carry out worse attacks in far higher volume with abliterated GLM 5.3)
Stopping two labs from further research doesn't reduce the danger at all, it just shifts the danger to labs that don't have real safety orgs.
Right, that's why regulation which is universally applied and includes compute controls (to prevent reckless creation of swarms) would be great.
"Posting about or admitting to attacks" is appreciated while people are still unaware of the risks but will be meaningless in the face of an industrial disaster that causes massive amounts of damage or loss of life. At some point, the leading labs must change their development practices, they can't just be allowed to continue rogue agent attacks just because they're willing to admit to them.
> At some point, the leading labs must change their development practices
What indicates that this has not been done?