As I looked through the images I was unimpressed entirely, at first. But, then I started thinking, these look a little... "childish" to me.

Childish as in... A newish artist who is drawing a concept rather than light / forms (Which is something artists typically do as they understand drawing more and more).

The rose in the vase specifically - some models understood that there was supposed to be shading, reflections, the concept of refraction - others just drew "blue = glass" and "green = stem" and "red = rose".

Really odd to look at, considering if I saw any of these drawings from a human kid, I would say "good job buddy" and put it on the fridge. I'm expecting these to get better as models improve, and perhaps the artistic progression will be there along with it...

Really odd to look at, considering if I saw any of these drawings from a human kid, I would say "good job buddy" and put it on the fridge.

The Grok ones in particular gave me that thought. Most of them really look like what a kid would do when given the same tools, while the other models' output have distinctly more "AI-ness" to them (for lack of a better term.)

What's interesting is that the way in which they're childish is actually extremely human. In fact, one of the ways that you're often taught to draw more realistically is to stop thinking of the concepts as icons you're drawing the outlines of, and instead sort of blur your eyes and see things as they are: hues and values. In other words, become a camera or a printer that has no idea what it's capturing or printing other than a grid of values. That is how you achieve realism.

The fact that it has clearly iconified these concepts in its mind and is tracing the outlines of the things it thinks/expects to go where is very human.

Parrots can also sound extremely human, but it’s only mimicry.

How do we know the difference?

Yes, but Parrots don't mimic the stages of learning speech development like children do. They just memorize a phrase.

These SOTA LLMs aren't trying to mimic existing children's drawings, but interestingly they're following somewhat similar progression that human children do as they develop.

Grok! LOL!

Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?

Grok 4.5 is...something else.

It performs much better than composer2.5 (while being as fast). It's not Opus, but I think it's not that far off. Definitely better than sonnet for what I've been doing.

On the other hand, I think they probably heavily adapted the training data so that it really is extremely focused on code. I just recently ran my personal "poetry benchmark" on it (where I give it ~850 poems I've written over my life and ask it to comment the corpus as a whole), and it's whack. It tries to write in portuguese (most of the poems are portuguese) and code-switches constantly and mixes up words to the point of making what it writes almost unreadable (e.g. it writes stuff like "You can't QoS that that look for beast poems", in portuguese, all messed up). The quality of the analysis is also quite bad (I'd say it's definitely behind Sonnet).

So I really think they either threw away data that wasn't tied to coding so that they could fine-tune it to that, or somehow they've got such an unbalanced dataset that coding ends up dominating either way. To me, its disastrous performance in this drawing "competition" fits this narrative.

"You can't QoS that" sounds like the title of a nerdcore rap song.

They focused a little too much on Grok Imagine.

the starry night one is soo funny

The razor-wire at the bottom for Starry Night was clever, and very Grok. Really shows its military spirit.

Edit: I just don't see the point of redacting the Mona Lisa

I read the balls as “houses”, though the phallic spire emerging from them dead-center is also very on-brand

I must admit, I have a natural tendency to overlook C&Bs, but solid catch -- it's there. Perhaps that explains the Mona Lisa redaction; Grok probably put more effort into that one.

Did Elon tell some poor engineer to give grok a prompt injection for drawing,"make it look like one of my childhood drawings!" Just like the Tesla truck?

I wonder how much better a harness could get for drawing

The Grok ones are so weird that they cross into uncanny and surreal. Truly bizarre and capable of eliciting feelings from me, if only bad feelings…

Why include Grok when it is clearly not even in the same category as the other three?

Would love to see the results if they had used /goal or similar

Article is AI-written.

I think this is just an ad.

If it is, it's a horrible ad. My main takeaway is that all of them were god awful at image generation and Fable was 20x more expensive and still awful.

Well, your takeaway is bad. This isn't image generation, it's tool use with a digital paintbrush.

And image tokenisation.

The difference in cost is pretty incredible.

Rather the difference of cost of Claude is pretty incredible. Then again, I guess they gotta make it rain while they can.

Claude was clearly 'pushing back' on the coziness of the cabin. But I think it did best with the cat. Grok, I fear, is making a case for euthanasia. It's suffering and I think it would be cruel to let it continue. Someone pull the plug. .

Now add Deepseek, GLM and Kimi :-)

Useless They are not image generation models.

That's what makes it an interesting challenge.

Not useless. LLMs are the most general purpose computer algorithms ever created. They are getting smarter and cheaper at a geometric rate. What is a bad idea today could have useful applications tomorrow.

> cheaper at a geometric rate

Citation needed - my company is paying more than ever for code generation. I have no reason to believe (given anecdotes) that anyone finds themselves in the opposite situation.

It keeps getting cheaper and better -- for me.

I get a lot more use (read: cheaper per interaction) and much better quality results from the $20 that I spend on this stuff every month than I did several years ago.

(And several years before that, it was all essentially unobtanium.)

Now this is a cool test. Sol's the best in each one, while Grok looks like the work of that person who "fixed" that Jesus painting..

[flagged]

Tell them to draw LeBron James or Heisenberg. Everyone will refuse.