This is a great thought experiment bc it raises the question of WHY humans wouldn’t behave this way. IMO the answer is a lot of socially enforced incentives that are dynamic and would be tough to fully articulate in a prompt.

The white hat has their own liability to consider, and the liability of their employer. Reputation and relationships are a big factor. All these tie into fundamental human incentives: survival, community acceptance, safety and freedom (prison not preferred!).

It’s a good sketch of why alignment is difficult, at least when it’s conceived of as an attempt to match human behavior.