> In adversarial settings (where we push the model to evade our monitors)
...why exactly are they training for that?
presumably that's a safety evaluation not a training setting
The whole Huggingface attack happened during training runs
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval
Ah, ExploitGym. Not ExploitBench.
Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
presumably that's a safety evaluation not a training setting
The whole Huggingface attack happened during training runs
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval
Ah, ExploitGym. Not ExploitBench.
Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions