there's still JEPA to be integrated before AGI.

Would DiffusionGemma be suitable candidate for DFlash 2?

DFlash 2 is a diffusion-based speculative decoding head for autoregressive models; there's nothing to accelerate here, because this is already wholly diffusion.

its more of a competitive approach to improve compute utilisation at lower batch sizes

(to finish your thought) Which is important for consumer hardware to be better suited to running these models. Cloud providers are already able to batch as many requests as they want together to improve resource utilization, so they will not see a big benefit from diffusion models.

Has JEPA shown a shred of viability yet?