Humans are drawing pelicans riding bicycles now. Just google it and you will find 5 or 10 of them in the first few results. Including a t-shirt design.

So it's a pretty much pointless test now.

Yet no LLM can actually do it.

It's quite surprising actually.

What does this even test for? Can I use LLMs to directly generate machine code to replace my compiler? Or maybe I can use LLMs as a bare metal OS / scheduler to replace my machine's operating system and scheduler? It makes zero sense to test for that.

Not only that this so-called "benchmark" isn't economically useful, but that it tests for the sake of testing; and for attention.

To end this obsession with generating SVG pelicans, Quiver AI's [0] model is actually designed to generate SVGs from prompts and has done so for years.

So there is no need to continue with this un-serious pseudoscientific "benchmark".

[0] https://quiver.ai/