> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.

I'm confused, videos contain images and audio ...?

That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.

Video contains images (frames), but not every image would reasonably be found in the frames of a video, or interpreted spatially. In the space of world model synthesis, consider blueprints, relationship diagrams, pages of instructions, sheet music, or a boarding pass.

It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.