at first this looked like something one could run on CPU with 64GB RAM with a 2-3 bit quant, at possibly half the speed of 3.6 35B-A3B, however the 50B ngram sidecar makes it impossible. and oddly, unsloth's page lists the ngrams as 50GB even though they say it's in 4 bits. should be 25GB according to my math. anyway, the new ngram architecture makes it pretty much unusable for regular folks who cant afford more than 32-64 GB ram in this RAMocalypse.

I suspect we will see optimizations where the various vectors of the n-gram you actually use are hot in vram, the rest are warm in system memory and then cold storage on nvme. Same with MoE. If your workflow is particularly same-y then you're looking at cache miss below 5% with NTP/MTP turned on and the right harness. Agentic "openclaw" type stuff cache miss might be below 1% in the right local llm setups. There's been zero exploitation of n-gram stuff yet, it will be very interesting as things progress.