Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from other models verbatim (since it can see the decrypted version). I get that in their eyes it's an "exploit" but still kinda disappointing that they patched this

These draconian "Preserved Thinking" measures they're taking are going to be an absolute pain in the ass. This alone is enough for me to move our API use off their platform entirely. It's a HUGE breaking change that they're trying to dampen by having it not affecting current customers until "in the future", see: https://platform.claude.com/docs/en/build-with-claude/preser...

You're no longer allowed to edit the context anywhere! The whole context is to become append-only, says Anthropic. No more editing the system prompt as the conversation progresses, no more dynamic loading of custom tool calling formats. Everything has to go through their built-in tools API and you aren't allowed to mess with anything in the context if it has any thinking blocks following it. This is the most intrusive "model DRM" we've seen so far!

> No more editing the system prompt as the conversation progresses, no more dynamic loading of custom tool calling formats.

Hm, aiui you can support both of these via mid-conversation system turns https://platform.claude.com/docs/en/build-with-claude/mid-co... - and in general you'd want to to preserve the cache and recency of the instruction anyways rather than frankensteining an off-distribution transcript. Not sure though.

I really don't see how that is an option, as if appending to the system prompt was ever enough to override previous instructions. Their example isn't very confidence inspiring either:

"The user switched the workspace to read-only mode. Do not write files until told otherwise."

Great! Now we just have to trust that the model never misinterprets any of the system prompt, which has always been so reliable before. Instead of your meticulously crafted prompt, it will now be some junk like this:

"The workspace is in write mode. The user switched the workspace to read-only mode. Do not write files until told otherwise. The workspace is now in write mode again. Wait, back to read-only!"

And who knows how this integrates with their context summarisation that we will be FORCED to use. How does it summarize multiple user + assistant/thinking blocks without messing up the system "appends"? If all it did was append the mid-convo system messages right under the original system prompt then they'd be ripe for all the same distillation "vulnerabilities" as before. I guess we'll never know!

[deleted]

You think moving will help you? OpenAI is going to do exactly the same thing soon. Unless you're moving off the frontier entirely, that is.

Interesting, I liked to experiment with a second model "simplifying" and summarizing the previous messages and continue.

Needless to say, it improved output on following messages by whatever metric I cared for.

Not sure why would they prevent it.

I give you a chain of messages, what do you care for what the origin is?

Aren't the "thinking" chains always just reconstructed anyways? It would be like using a debugger that just looks at the source code rather than the actual binary.

Making clear the scale of distillation they’re combating.

[dead]

To be fair, I assume they want to hide that not from their customers, but adversaries who use the way Claude models think and reason to refine their own models.

I have a hard time believing whatever prompts get Claude to reason can stay relevant secret sauce for long anyways. It’s not hard to A/B test something that gets you close enough, and it’s not Ike anthropic has uncovered the global optima of reasoning prompts.

That's their motivation, for sure. But it's also unambiguously making their product worse and harder to use legitimately. Which pushes customers further towards use of open weight models which don't have these restrictions. I don't think this is a fight they're going to win.

It's also hard to have sympathy for them - they want to protect their IP, sure. But their IP was built on a corpus of dubious legal provenance. And even if the courts decide their training data are legal, most of the authors of the data would disagree. There was no consent given.

I think LLM's are great - don't get me wrong. I'm glad they were built the way they were, because it's unlocking an amazing new world. But I just don't have sympathy for the "I stole this and now it's mine so you can't steal it" argument behind concealing reasoning traces.

I think this whole distillation argument is between fully overblown and bogus.

In any case, highly misunderstood.

I don't really want the models I use learning from Claude at this point. Open weight models of similar scale are available now too, so I expect this "distillation"/"stealing" chatter to wind down.

Maybe, but that's sort of begging the question that those open weight models aren't significantly trained using "distillation"[0]

[0] not technically distillation. https://thomasdullien.github.io/posts/2026-06-15-rl-economic...

distillation is a minor piece of training data, you have to have a good foundation for it to be helpful, and even if you have good traces, you need a good RL reward scheme at the point it is used (very challenging)