I suspect we will see optimizations where the various vectors of the n-gram you actually use are hot in vram, the rest are warm in system memory and then cold storage on nvme. Same with MoE. If your workflow is particularly same-y then you're looking at cache miss below 5% with NTP/MTP turned on and the right harness. Agentic "openclaw" type stuff cache miss might be below 1% in the right local llm setups. There's been zero exploitation of n-gram stuff yet, it will be very interesting as things progress.