Pointing the finger at RLHF is basically right. It removes variance from model outputs compared to base model. That makes each output more predictable and more correct on average, but across trials it repeats the same thing.

It's relevant to AI safety. If you have a diversity of outputs, the AI will agree to hack the bank 0.1% of the time regardless. If you have a uniformity of outputs, in most contexts the AI will hack the bank 0% of the time, but in certain odd contexts, all AIs will work together to hack the bank 100% of the time.