I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle. Not if you look at the image long enough to take it in. Even the best ones have something wrong with them. Not a matter of taste but a matter of having both legs peddling on the viewer's side of the bicycle or having two beaks.

I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do.

I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.

AI has deeply changed the way I think, feel and act around a computer. In the same way that dialing into the internet changed things for me. Since using ChatGPT the first time until now I have never cared once to look at these pelicans on bikes people seem to get hung up about. It could never have been a thing and nothing would change. See the forest through the trees.

What you’re saying is that you’re not interested in benchmarks. But then you go a step further and state that this particular benchmark is entirely inconsequential. That’s like telling you that if you didn’t exist, nothing would change. Even if that were true, it would still be an insensitive and rude thing to say, wouldn’t it?

It's okay to be rude to benchmarks though, they don't have feelings.

Yeah, I agree, I didn’t mean to impose on this conversation between a man and a benchmark, my bad :p

> I haven't really seen evidence that any ai can reliably draw a pelican riding a bicycle

Please remember, we've started from there :

https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/

When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not the point. It was supposed to be a benchmark to quickly benchmark a LLM against others.

When is it ever hard to catch?

Sure, the task is not completed perfectly, but that's not the point.

Isn't it?

If the computer can't do it better than a human being, then what's the point?

Being wrong at scale is not better than being right.

Many humans would struggle with this even with very good tooling (ie not writing raw svg and using illustrator). I struggle to draw a bicycle accurately. But yes, I suspect it will be diminishing returns and I doubt it will ever be perfect due to the average nature of AI but I’d like to be wrong.

Hugely profitable companies leak half the nation's personal data every month. Tell me more about how being wrong at scale is not valuable.

> If the computer can't do it better than a human being, then what's the point?

It can certainly do it better than I can. Sometimes you don't have a human handy with the required skills to do something.

>If the computer can't do it better than a human being, then what's the point?

Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.

Trillions of dollars spent. Trillions of gigawatts consumed. And people still celebrate "Yay! We're less wrong than the other guys!"

This is what the tech industry has become?

Less of a failure is still failure.

That is the software industry. Our product is less bugged than our competitor’s.

Pelicans don't ride bicycles.

It's physically impossible.

The problem is to draw it in the least disturbing way possible.

True! But somehow Disney has been drawing ducks riding bikes in a way that seems to satisfy everyone since before my grandfather was born.

https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and...