https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level.
Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I think these are the worst I've seen, at least in some time. It's a silly benchmark though, not sure what to make of it
it's not truly tested until it plays a match or ten in Brood War imo
If you think this is bad, look up mistral.
I think the result is fine. The benchmark is silly to the point of being useless nowadays.
It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.
So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":
Does the user want the least lines of code to make it functional, or the best looking version?
The user at the very least expects the bike to have bike geometry; Grok seems to struggle with that
The default reasoning level seems better than the high reasoning level:
Has a shadow
Better shaped beak
Leg position more realistic for bicycle riding
Better feathers
Here are some somehow standardized pelican tests but for 3d scenes in threejs at threejseval.com
Grok 4.7 generations: https://threejseval.com/models/grok-4.7-high
Also go vote on https://threejseval.com so you can help evaluate how Grok and other model performs compared to each other!
Wow, thanks for sharing, fun benchmark!
I tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it's not very token efficient though: https://artificialanalysis.ai/agents/coding-agents?agents=co...
Poor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.
Are there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
What is the default reasoning level?
It’s a bad look to be evaluating models built in the furtherance of fascism and white supremacy. As this is a closed model, any support and training of it ultimately benefits those ends.
Hahaha. I remember when musk and his merry band of sycophants were bragging about grok producing the only physically accurate bicycle. What happened?
Yeah, he had opinions: https://twitter.com/elonmusk/status/2023833496804839808