Or maybe they didn't train it that way to be manipulative (although it's certainly a plausible explanation) but simply as an accidental artifact of trying to make it give honest answers?
LLM-generated images sometimes includes text from the prompt as literal text in the image, so perhaps this is the same sort of artifact? If they've told it to be honest, it responds by talking about being honest instead of actually being honest, because it has no actual understanding of anything.
The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.
> accidental artifact of trying to make it give honest answers?
If it's not giving honest answers that implies it's purposely being deceitful, which it isn't capable of. Right?