Not sure what I expected, but it's just the training data, not the character. It'd be so cool if such systems had natural curiosity at this checkpoint. Eg:

> Me: "What's semiotic crystallography? > Response: "I don't know, what is it?"

Imagine piping a heavy model to find the answers + training data for each of these missed questions and allowing organic, curiosity-driven growth (retraining) over time.

It would lose knowledge about existing subjects unless it’s continually retrained on those too. It could help inform the next training dataset though.

Good thoughts here. Forgetting is important, but that's too advanced for modern LLMs.