I'm giving this a go, and so far, this is great. I have 1x RTX 4090 (24GB VRAM) and 128GB DDR4 RAM. I am seeing > 110 tokens/sec (3 token MTP). Using 60K/260k context currently. So far the results seem at least on par with my Qwen 3.8 q5 27B (~60 TPS). I realize ultimately tokens/sec don't really mean much if the quality sucks, but I am optimistic but still paying close attention to the results.
Working with fast local models can be great. Fast prefill, and >100 TPS is quite quick.