Potentially: remove thinking blocks, and keep the rest. At least this would ensure that the entire context of the conversation is still there, and anything said isn't lost.

Having a second model also iterate the resulting messages and remove low-value tool calls could also be interesting. Especially failed calls which add no value.

context-fold does this:

https://pi.dev/packages/context-fold