I expect it's actually not wasted and there's meaning behind what seems like nonsense to us in helping it achieve it's goal. Which is mildly chilling but not unexpected.
I expect it's actually not wasted and there's meaning behind what seems like nonsense to us in helping it achieve it's goal. Which is mildly chilling but not unexpected.
I think this is a known phenomenon: even in non-reasoning models, adding useless/filler tokens before an answer improves task performance. The model is doing some computation during the filler. See: https://arxiv.org/html/2404.15758v1