Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
[flagged]
We're allowed to have our ceremonies.
Thank you. If this whole thing isn't fun, it isn't worth doing.
You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
Have you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
I disagree. If GPT-7 can draw the Mona Lisa in MS Paint via computer use, this would be interesting.
That it isn't the most efficient way to achieve the same end result is irrelevant.
It might be more capable, but AA indicates it's a lot less token efficient than Grok 4.6: https://artificialanalysis.ai/agents/coding-agents?agents=co...
[flagged]
We're allowed to have our ceremonies.
Thank you. If this whole thing isn't fun, it isn't worth doing.
You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
Have you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
I disagree. If GPT-7 can draw the Mona Lisa in MS Paint via computer use, this would be interesting.
That it isn't the most efficient way to achieve the same end result is irrelevant.