Good point. It's much more of an issue when running dense models with tensor parallelism. In that case, I'd look for an MoE model instead.