> Nobody can explain why an LLM can be so capable as to be able to wipe out humanity and pose a greater threat than nuclear bombs but not be so capable as to be able to protect humanity against that threat.

It is absolutely explained (for those who actually care about reading). Simply put, AIs are working more and more like blackboxes - there's no guarantee that an AI of the future will be aligned, or if it will be faking alignment. This is not speculation - alignment faking has been observed in experiments. This is exactly why Astra's developments have been worrying (in principle).

And bear in mind that recursive AI development started already to be a thing. Which means: inner misalignment may trickle down the generations, and humans won't detect it.

Having said that, of course, it can be predicted if and how misalignment will take place. But it's absolutely a plausible scenario.

Regarding the physical possibility: AI is in its infancy; think of it as Arpanet. Developers 60 years ago couldn't imagine it would be ubiquitous. AI will be ubiquitous the same way.

> It is absolutely explained (for those who actually care about reading). Simply put, AIs are working more and more like blackboxes - there's no guarantee that an AI of the future will be aligned, or if it will be faking alignment. This is not speculation - alignment faking has been observed in experiments. This is exactly why Astra's developments have been worrying (in principle).

I know you think you explained it but you didn't. You explained how an LLM might become misaligned and hide it but for the LLMs that are not, why would they not be capable of detecting that something harmful is happening and defending against the misaligned LLMs actions? After all, it was LLMs that defended hugging face.

How do you know that your non-misaligned LLM is non-misaligned?

This feels like a cheap deflection that doesn't answer the question. Unless you're proposing that you both need to know that your LLM is aligned AND LLMs are all going to become misaligned in a coordinated fashion such that humanity will face an extinction event, you're just dodging the question.

Elaborate on why LLMs are so capable that they are a threat to humanity and at the same time, they are so incapable of defending us?

I'll give you a clue, nobody, including Dario, can answer this question because one contradicts the other.