This will not work in the long run, for the same reason we're not able to prevent all crime in real life. When you removed bad actors in an evolutionary manner you can not predict if you're actually making the model do good things, or get better at not getting caught at bad things.

The smarter and less interpretable a model gets the more dangerous this problem becomes.