I feel like mine is mocking me. I added an instruction in Claude.md that says "under no circumstances use the phrase found the smoking gun, say I found the problem instead"

What does it do? It says "found the smoking gun! Ooops I wasn't meant to say that - I found the problem!"

It's pretty wild how "reasoning" models now generate like 10 thousand hidden chain of thought tokens in response to a "increase opacity of the logo by 20%" prompt before writing the actual message and yet they still manage to do this.

A bit different scenario but reminds me how I asked Gemini (paid plan) to translate some labels on a chart from Russian to English the other day, and it spat out the image, then did a "wait, several labels are still Russian somehow", spat out the same image with the same issue, then a paragraph of incoherent babble about how desperately it wanted to fix the issue to complete the task successfully, finally spat out the same image a third time and was like "welp I may be brain damaged... sorry, user :("

It was hard to tell if the parser was leaking reasoning since Gemini does that occasionally with nested tags and stuff but either way seeing the load-bearing psychosis behind its normal generative loop was somewhat unnerving. Helps me understand why the big labs are so intent on hiding it from the user... if it just gave me the final image without all that drama I'd just have rolled my eyes at the mildly sloppy but 90% correct image.