Not the person you replied to, but I think a more accurate description of the reasoning we see is proof of effort, not necessarily great insight into how the reasoning is occurring. For the most part, researchers currently describe the intent and motives (in however one may define them for LLMs) as black boxes right now. Even the mechanics of the cognitive process is not well understood. Depending on the model and harness, the thinking will often look like gibberish. I suspect they've invested considerable effort into presenting thinking as a reasonable approximation of what they imagine it to be. Claude and OpenAI have also begun encouraging multi-step problem solving (or the models themselves decide this), and we can see their more accurate responses at the conclusion of each phase.
Fine tuning or post-training is effectively biasing certain outcomes: making them more likely to occur. This comes with trade-offs. A coding LLM will bias technical language, which would harm a model for general use.
This opens a really interesting field of research. Our brains use specialised regions because specialisation turned out to be the most energy efficient method for biological compute. It might also be the best performant. We don't want to activate 100% of our prefrontal cortex to breath. What a stupendous waste of the organ. I think we see incredible advancements in model clusters in the future, using specialised models for specialised tasks. We have the appearance of this today in some harnesses, but they are shallow imitations. The real innovation will be low-cost, accurate routing. Existing solutions are woefully inadequate for many reasons.