I would not recommend using any of those notes as evidence of internal “intent.” It produces them performatively—it is literally rewarded for thinking out loud in ways that seem plausible to humans.

There are several papers out there arguing that chain-of-reasoning-like output is performative, such as https://arxiv.org/abs/2603.05488

It would be awesome if we could reasonably purge all anthropomorphizing language like “tried” or “thought” entirely from AI discussions, because it introduces very sneaky biases in our thinking, but I’ve found it damn hard to do in practice.