I hated it at first too...
Now though I'm considering all the hidden "thinking" in the models layers that happens for each token output. It is a wild amount of waste! We just can't see it.
This kind of stupid excessive computation is fundamentally how these models are so good.
One day hopefully not so soon someone smart or a foundation model will come up with a more efficient architecture. That's when things get really scary.