It's not obvious if that would help. Some tests have shown that what exactly the thinking tokens are only makes a small difference to the performance of the model, and that the contents of them are sometimes only tenuously related to to what the model actually does after them. It seems like it could be they are more like a kind of "mumbling" and that the underlying mechanism by which they actually help performance is just that it makes more computation available to the model by just giving more passes on the earlier input tokens through the network.