The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.