Vision functions the same way as language when inference is done. It's a stream of tokens.
It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.
Vision functions the same way as language when inference is done. It's a stream of tokens.
It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.
my understanding is the vision component of the original model is an independent preprocessing step(except for things like Gemma 12b). And it would be possible for engine to have broken that phase of translation before it hits the real model as encoded text. I am not familiar with Strata engine and their claims, but this seems like an interesting way to test models with relatively straightforward inputs.