HBF will probably save the small inference nodes. (Eventually, the initial price will obviously be in the stratosphere.)

Long term, HBF shouldn't be more than 4x as expensive as commodity flash, and Kimi K3 is an interesting target for a system using it given how aggressively it compresses the KV-cache. An inference box with ~1.5TB of HBF and ~48GB of DRAM should be able to be built for less than a couple of grand, and get something near to 100tok/s on full Kimi K3 for a single token stream.