The best part is using harnesses like reasonix or whale make cache hit at a rate close to 98%, making requests converge to practically free. And that's with unsubsidized American providers like cloudflare or Digital Ocean.

In long form tasks, across multiple harnesses (Claude Code, OpenCode, Kimi Code, ZCode) my cache rates are typically 96-99%.

I don’t think that’s particularly out of the ordinary. Do people have different experiences with other harnesses? Which ones?

How much time do you spend on your setup vs getting a lot of stuff shipped by paying for fabel 5? For me, not using the best model is a huge opportunity cost since my company can afford it.

> For me, not using the best model is a huge opportunity cost since my company can afford I

But sometimes not using the "best model" is using the best model. Fable isn't the best at everything. Specialization and optimizations might not just be around cost.

I can see that. I’m just afraid to sync too much time into complex routing schemes when I get pretty consistently good results out of got 5.6 or fabel. For code reviews, I’ll try the best flash and grok at the time but they just don’t come close to gpt 5.6 which has been the best review model for me since 5.5.

How can you be hitting cache on what I think are novel LLM prompts …

Not them but my understanding is that the harness will send a simple 'heartbeat' message to keep the cache 'warm', (see prefix caching: https://handbook.modular.com/inference-optimization/prefix-c... ) which can then be edited/changed, which does cause the user to incur a fee, but its much less than the amount they'd pay on a no-cache hit request.

multi turn sessions, they are typically in the high 90% hit rate across all providers without doing much of anything

I’ve never used either of those tools. Do they work well?

Can you expand more on how you are using it?