You don't need to threaten the LLM with jail, you just need to reward the target behavior during learning.

I understand this is an active area of research. See Anthropic's J-Lens research where they measured like a "FAKE FICTIONAL" direction in the activations during evaluations with contrived scenarios, making the model more likely to avoid taking malicious action when it knew it was being tested.