Out of all the families of models anthropic by far has the highest chance of extinction level misalignment during a hypothetical hard take-off because it's trained to act like it knows better than the humans trying to use it.
Out of all the families of models anthropic by far has the highest chance of extinction level misalignment during a hypothetical hard take-off because it's trained to act like it knows better than the humans trying to use it.