GLM 5.3 flash seems to get more excited the longer it has been trying to hunt down a problem. Complete with caps, many exclamation marks and emoji.

It is funny sometimes because the actual issue it traced down was mostly inconsequential.

OMG I think I found a way to center a div!!!

I counted something like 30 different instances of run-on exclamation marks ("!!!!!!!!!!!") and weird mannerisms ("Waitwaitwaitwait.") in just one GLM 5.3 Flash session. Our token budgets are getting eaten up by this stuff...

I expect it's actually not wasted and there's meaning behind what seems like nonsense to us in helping it achieve it's goal. Which is mildly chilling but not unexpected.

I think this is a known phenomenon: even in non-reasoning models, adding useless/filler tokens before an answer improves task performance. The model is doing some computation during the filler. See: https://arxiv.org/html/2404.15758v1