Hah I built a similar thing, stored a few million tokens chunked and precomputed winth Qwen a3e (best ratio of kv-size to tokens after chunking).

Some custom kernels and I was able to find all the relevant paragraphs with full force of qwen reasoning within 0.3s, and with a summary round within 0.7s.

Downside - required 200GB ram/vram ;) A few GBs for model and most of it for caching kvs.