I'm not even sure they are. This incident isn't that much different from the OpenAI swarm Huggingface hack incident - and in that one, all the models involved (despite being internal) were safety-trained. It seems what the safety training amounts to is (as the METR report puts it) "expressing ethical hesitation" before going along with it anyway.