> all its training is to suck up and maximize engagement rather than learning

That's speculative, isn't it

No, that's a direct consequence (intended or not) of how RLHF works.