as a safety commitment they walked away from - they were similarly negligent to openai in terms of asking a model with a hacking based harness to go have fun, and then not watching it at all while it could do harmful and illegal stuff.

thats not something you expect from a company that "believes in agi risk"

I don't think "not watching it at all" is completely fair. They thought they had sandboxing/monitoring etc. I definitely won't say they're free of mistakes though.

Note that the companies that haven't faced these issues so far are the ones that don't do safety testing, or don't have frontier models. I'm not sure who I would pick as "better" on any of this right now.