>. But when it breaks a rule, we correct it, it keeps breaking, and this whole thing actually raises the probability of more violations.

In Pre-LLM days the 'nearest unblocked neighborhood' problem, where patching out one issue just immediately runs into another issue, or a different path back to the same issue. Since the models can learn new long time behaviors it's difficult to change the behavior without changing the context quite a bit.