Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously. "I am sorry your family is dead, my bad" goes even worse for you in court when you release a model that showed these behaviors in testing.

I honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate.

> Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously

Not really. If I build a special new wine bottle, and call every breakage a mis-alignment problem, it's not the bottle just being fucked in the same way every fucked bottle is fucked, that's marketing. It doesn't change the fundamental form of the problem.

> "I am sorry your family is dead, my bad"

This should be punished. It's a problem that plagues Instagram and OpenAI. It's not inherently one, though, that has to do with AI. Just sociopaths preying on children.

> honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate

Perhaps. I haven't seen someone explain it to me in this thread in a way that seems separate from bugs.

Where I have seen a separate class of problem argued is where it's existential. But in that case, clarity of definition comes at the cost of any evidence for it.