I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!
I remember being in the room with pretraining day 1 to help monitor the training job launch. Watching this model train from day 1 has been an amazing experience!
What sort of outputs or telemetry is monitored on a large pre-training run?
Outside of ML metrics, you're monitoring the health of every piece of hardware in the system. You need to make sure that you have every GPU, every CPU, the PCIe buses, the networking fabric are all working without any errors. You need to ensure that you can respond as fast as possible to any possible error. One bad component can bottleneck the entire job.
https://github.com/facebookresearch/metaseq/blob/main/projec...
I really enjoyed reading the log book from the training of OPT-175B at Meta… I guess it’s all classified info but it’d be fun to read a blog post about the crazy day to day issues you run into when doing things at this scale :)
GPU failures are frequent enough that at a certain scale, you constantly have workers dropping out. Designing systems that can still keep training is very interesting!