Aren't such attractors also problems during training? Presumably alignment constraints would need to apply to continuous learning as well.
Aren't such attractors also problems during training? Presumably alignment constraints would need to apply to continuous learning as well.
I mean yes, but that's far more difficult than one would expect in a layered system. You may be able to keep a concept aligned, but can you keep the meta concept aligned? Continuous learning means there is continous opportunity for a more powerful system of misalignment to form itself and take control of your alignment. Even worse is hidden layers of this misalignment that can hide itself by not using an interpretable language.