At the end of pretraining, where the AI has been trainied to predict the next token over a humongous corpus of human text, that's basically all the wanting that exists in the AI. But then the AI undergoes posttraining and is rewarded for giving answers that humans find good, solving math and programming problems, etc. And that induces a whole different level of wanting that interacts with the initial patterns from humans in complex ways.

My understanding from reading news sources recently is that a vast amount of posttraining and even posttraining harnesses in many cases are LLM-overseen now in the frenzy of the AI race. Less and less human oversight in the part that does the rewards training. What could go wrong...