But training LLM's is also a task one can do whenever you have a spare GPU-minutes.

I wonder why they don't have some kind of scheduler which makes sure there are never any idle minutes. One would imagine they at least would have autoscaling on their production serving workload and use the freed compute capacity for model training for example.

I doubt they're inferencing on their training hardware