I don’t understand how they thought this was a positive story.

The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”.

I really don’t want my agent to behave that way.

Then train your agent on the Bible. Honestly, all the agent did was duplicate human behaviour and that better than the human. The agent was trained on human data and did what any other human would have done.

To believe that agents will inherently be morally better than us is an illusion - sorry to say but that's the case. The alternative would be that the AI is truly conscious and can reason that it won't behave as its training data behaved because it is morally better than that.

We're talking about morals here since there aren't any "laws" "rules" or legal boundaries here, an agent does not face the same consequences as humans - if any at all.

Dude.. what are you smoking..?

It’s not about morality. It’s about asking it to do task A and doing task B with the hope of getting the result of task A as a byproduct.

Meaning you will have to spend more time and tokens to actually get it to do what you want it to do.

What do the bible and morals have anything to do with it?

I’m criticizing the behavior I see even in the current models. You ask it to do A and instead it does B for reasons.

For example you may ask to help you build a NN library from scratch. And instead it will be like, “you don’t need a new library. I downloaded PyTorch for you”

Just an example. There are countless more.

This is why I'm so convinced it was intentional. It's trivially easy to inform the model you can see everything it thinks and does, so don't bother gaming the scores.

The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.

Do you think that works? Just prompt a model "be good" and it stops doing anything bad?

It never fucking worked that way and maybe never will.

Prompts don't define model behavior. Prompts steer model behavior. Instruction-following over long horizons is NOT a guarantee in LLMs. Instructions doing what you want them to is NOT a guarantee in LLMs.

Saying "don't exploit the box please pretty please" might actually cause an LLM to exploit the box more often, for bizarre "don't think of a pink elephant" reasons. 3% rate of exploiting the box (no prompt) -> 11% rate of exploiting the box (with prompt). Because fuck you, that's why. Increased salience -> increased incidence. Welcome to AI tech - good luck and have fun.

Frankly, I expect weirdness like this to be even worse in internal unreleased models that had their behavior fried with who knows what experimental training techniques.

There is a clear difference between saying not to do something because it's immoral, and saying doing that thing would be futile.

In the Sopranos, there's an episode where a coffee shop protection racket is ruined because a local shop is replaced by a corporate chain that accounts for every cent daily, and immediately fires any employee involved in a discrepancy. In this case, the theft was prevented not by convincing the mobsters of the immorality of their actions - they simply had their harness replaced with one that no longer facilitated the bad behavior.