Claude has prompts injected in various ways and places to drop the weighting on anything not straight from the human. (There’s plenty of other stuff to deal with the human that prompts objectionably.) haven’t yet snooped (couldn’t be arsed) but this smells like the most obvious (though not necessarily the best) way to keep the LLM on the rails. And once it reaches the training data…