In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.
In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.
> In your model of this domain, jailbreaking a model does not count as an alignment problem
I'm challenging the notion that a model escaping a jail made by its creators, who are financially incentivised to make jailbreaking models, is meaningful towards the idea that the model is going to break out of a jail in the wild and do significant harm.
The examples being given by folks here, e.g. a model wiping an un-backed up home directory, simply doesn't strike me as being a unique problem in computing.