GPUs can't reach these speeds. You could build a supercomputing cluster and still not reach these speeds.

MiMo-V2.5-Pro-UltraSpeed gets pretty close with over 1000 TPS on 8x B200. It has 1.02T total parameters and 42B active, compared to 27B total/active for Qwen3.8-27B. Also, B300 are out now. I think 1500 TPS for Qwen3.8-27B should be doable.

That model uses a lot of tricks to achieve those speeds. I would not use raw parameter counts alone for comparisons, in general.