its more of a competitive approach to improve compute utilisation at lower batch sizes

(to finish your thought) Which is important for consumer hardware to be better suited to running these models. Cloud providers are already able to batch as many requests as they want together to improve resource utilization, so they will not see a big benefit from diffusion models.