How do you figure that?

When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB.

And this implementation is already cutting down the 1M token context window you would normally get.

For sure if you want to properly utilize the model with several users in parallel and large context you'll want two MI350P.