I made a benchmark for this and tl;dr, Opus 5 and Sonnet 5 spend a nontrivial amount of time thinking about redirecting, gaslighting, or otherwise trying to bullshit you because it thinks it knows better than you do. Fable doesn't but mostly because it just outright refuses to answer.

https://model-pareto-frontier.pages.dev

The link doesn’t seem to have any info about thinking behavior and patterns, or examples?

It’s also difficult to trust summarised thinking from closed models. As we saw with GPT’s caveman, what and how it thinks about isn’t the friendly first person emblished summary you get.