My guess is that if you go with something other than readable text as the “thought layer”, observability becomes impossible.

Which is not necessarily something the humans training a model would want re/ alignment.

Indeed this is one of the key problems here. If models start doing CoT in "neuralese", or use it when talking with each other, we lose what little observability we have.