So the models will not only be using more and more Neuralese in their CoT (like GPT-6), but different agents will also be able to communicate with each other in Neuralese. It's not looking good for monitorability.
So the models will not only be using more and more Neuralese in their CoT (like GPT-6), but different agents will also be able to communicate with each other in Neuralese. It's not looking good for monitorability.
Is Neuralese in no way decodable into a human-interpretable system? Genuine question -- I don't know the answer.
Probably not without a sufficiently powerful LLM from same family, or something equivalent, to act as a translation layer, which includes the risk of the translator lying to you.
Definitely decodable, that's what's being done now
It's not.
This is from like a year ago.