Diffusion language models work with a discrete output space, unlike image models that repeatedly refine a continuous output, so they don't do the noise-prediction thing anyway.

Ok, you are correct. These models don't train noise predictors at all unlike first gen image diffusion