Benchmarks like the pelican thing are about correlation with “how good the model is”. Better models produce better pelicans.

It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance.

Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway.

TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.