my understanding is the vision component of the original model is an independent preprocessing step(except for things like Gemma 12b). And it would be possible for engine to have broken that phase of translation before it hits the real model as encoded text. I am not familiar with Strata engine and their claims, but this seems like an interesting way to test models with relatively straightforward inputs.