It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.

The amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data.

It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.

You can't benchmaxx spatial awareness without solving the fully general problem (at least I figure).