Wow. So kimi is suddenly half the price for all users until they hit 256k of context? Thats massive.
As far as I understand, no. They're suggesting that smaller context windows are typically cheaper (fewer input tokens over time).
I don't think so. This is a separate model, so I assume that if you just use this and switch to the 1 million context model when you reach 256k, your cache will be invalidated, so you'll re-pay the 256k tokens on the 1 million context model pricing.
The article explicitly says "The current version switching from 256k to 1M does not affect the cache."
As far as I understand, no. They're suggesting that smaller context windows are typically cheaper (fewer input tokens over time).
I don't think so. This is a separate model, so I assume that if you just use this and switch to the 1 million context model when you reach 256k, your cache will be invalidated, so you'll re-pay the 256k tokens on the 1 million context model pricing.
The article explicitly says "The current version switching from 256k to 1M does not affect the cache."