I thought exactly the same at first. But then i wondered if that still holds true with today's advanced thinking, RLHF involved, frontier models. I guess to a certain extend it did indeed behave better, as a reaction to his self description into account.

EDIT: I mean, those systems accumulated so much complexity around the attention based next token predictor.

Yeah, in my experience, there's nothing about:

1. LLM thinking 2. RLHF 3. The latest frontier models

that does anything to change this fundamental "suggestibility" of LLMs.

But who knows, maybe I'm wrong.

Training the LLM to do things that the user didn’t explicitly ask for is a good way to get complaints from the users. Doesn’t matter if those things are best practices.