Ok, you are correct. These models don't train noise predictors at all unlike first gen image diffusion