The updated version will handle tool calling better by default, but the reasoning quality is no longer preserved and is mutilated quite badly.
The updated version will handle tool calling better by default, but the reasoning quality is no longer preserved and is mutilated quite badly.
So then how do you run it unmutilated?
Log all the calls and run on a periodic cadence (cron or ever N turns) a larger model (like Opus) to read samples of the traces and edit the template to fix observed problems. There are some signs that help find interesting things to look at, errors of course, but also overly long responses, prefix cache misses, tool call errors, etc.