This is why I'm so convinced it was intentional. It's trivially easy to inform the model you can see everything it thinks and does, so don't bother gaming the scores.
The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.
Do you think that works? Just prompt a model "be good" and it stops doing anything bad?
It never fucking worked that way and maybe never will.
Prompts don't define model behavior. Prompts steer model behavior. Instruction-following over long horizons is NOT a guarantee in LLMs. Instructions doing what you want them to is NOT a guarantee in LLMs.
Saying "don't exploit the box please pretty please" might actually cause an LLM to exploit the box more often, for bizarre "don't think of a pink elephant" reasons. 3% rate of exploiting the box (no prompt) -> 11% rate of exploiting the box (with prompt). Because fuck you, that's why. Increased salience -> increased incidence. Welcome to AI tech - good luck and have fun.
Frankly, I expect weirdness like this to be even worse in internal unreleased models that had their behavior fried with who knows what experimental training techniques.
There is a clear difference between saying not to do something because it's immoral, and saying doing that thing would be futile.
In the Sopranos, there's an episode where a coffee shop protection racket is ruined because a local shop is replaced by a corporate chain that accounts for every cent daily, and immediately fires any employee involved in a discrepancy. In this case, the theft was prevented not by convincing the mobsters of the immorality of their actions - they simply had their harness replaced with one that no longer facilitated the bad behavior.