Lots of words about multi-modal but then this:

> our mission to develop real-world visual intelligence

Visual is mono-modal, isn't it?

its doing video, audio, images and motion. I think that counts as multimodal.

Is this really the value-add comment you’re going with?

Says the one who posts this comment?