You're talking about the HuggingFace incident? It's notable that the agents (somewhat justifiably) thought that their evaluation had an LLM grader which would look for evidence that they had cheated. And it was specifically that grader they were trying to his evidence of cheating from, not humans more generally.

Honestly that part was more surprising to me than anything else, how narrow the compulsion to cheat was: they didn't learn "cheat in general" they learned "think about the grader in great detail and chat exactly as much and exactly in the ways that actually result in a higher score".

Yea agreed they were trying to deceive what they thought was an LLM grader. It’s unclear to me the extent to which cheating behavior is generalized? From what I’ve read there are signs that some amount of cheating has been reinforced in their training due to poor RLVR evaluation setups