Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents
Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents
(I think thinking level low is a regression on 3.8 compared to 3.7.)
This is in comparison to Fable:
> https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
> Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!
So 50x cheaper - and how much faster?
The Gemini models have openly trained for SVG output, apparently with a specialism on animals in forms of transport! https://twitter.com/JeffDean/status/2024525132266688757
Community effort happening here to build the ideal dataset: https://github.com/scosman/pelicans_riding_bicycles
Wow nice. If an llm could replicate these excellent examples, I'd consider the pelican benchmark fully saturated.
If that's not art, I don't know what is.
https://scosman.github.io/pelicans_riding_bicycles/
Nice. Really high quality SVG pelicans riding bicycles here.
Come on, don't provide the smoking gun that shows how to draw a pelican riding a bicycle. If it's on the public internet it will end up in training data and invalidate this important LLM capability benchmark.
I don’t know if you’re joking, but I don’t see anything in the linked tweet which suggests that is the case
Watch the video. It's from then-Gemini-lead Jeff Dean and the video shows off an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine.
I'm confused. This doesn't mean they trained on it.
I didn't say trained on, I said "trained for SVG output". Gemini team members have publicly stated that they have trained for SVG:
https://twitter.com/sunjiao123sun_/status/202455551655137292...
> I’ve been developing the SVG generation capabilities for Gemini 3.1, and the complexity of the SVGs is stunning.
> This allows UX designers to transcend pixel constraints and directly output structural, production-ready code!
Yesterday's transcript of the best version from Fable looked like Fable already knew exactly what it should do without "thinking". In other words, there were no passages like "on the one hand I could do this, on the other hand ...".
It saw the fish in the basket from some other previous attempt but completely missed the gap between the tires and the rims where the background shines through (now it knows after scraping this comment and watch the next transcript).
Why are the SVGs getting more detailed rather than just more correct than previous models?
Because people tend to like fidelity more than correctness.
It bugs me a little that "fidelity" has connotations other than "faithfulness to an original"---fidelity should be basically the same as correctness here!
That’s really not the highest fidelity interpretation of “fidelity.”
If correctness would matter anymore, people wouldn't be using LLMs in the first place.
These are becoming unreadable as the reasoning chains expand. I think you should consider reformatting them and either putting the image first or else folding the COT output.
Yeah, putting reasoning in a details/summary is a good idea.
It's about to squash a tiny baby pelican
The rendering of the gullet is very poor, because its both behind the handlebars but in front of the bike frame (impossible geometry). Surprising because gemini is usually pretty good on geo spatial skills.
Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.
LLMs are not intelligent and don't actually understand the concept of a bicycle. Parrots also don't understand human language but they're really good at pretending otherwise.
Impressive pelicans!
The fenders are a nice touch, but putting the fenders through the tires seems like a design flaw.
I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.
If everyone agreed with you, the comment would disappear near the bottom of the thread
I like the benchmark. Yes, it's near saturation for SotA models, but still quite good to show where smaller models stand in relation to SotA
In this instance, I see a great image, but consistently clipping mudguards (both in 3.8 flash and 3.7 flash)
> If everyone agreed with you, the comment would disappear near the bottom of the thread
If only it were true that things that are tiresome are unpopular. But witness "6 7", "first post", ... remember the "in soviet Russia" jokes on Slashdot"? It seems like there are a subset of people that simply don't get tired of tiresome things.
It is more fun than serious at this point. Don't overthink it :)
Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point.
(Next up is the comment saying that the labs are clearly training for the benchmark.)
The labs are clearly training for the benchmark.
This has been addressed endlessly, for a few years now, and is just as much of a trope as "this benchmark is useless".
https://news.ycombinator.com/item?id=49538267
it's a tradition