I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.
I really like this idea. You could expand on this by giving programming tasks and measuring code similarity. Seems like you could develop a pretty detailed understanding of similarities across multiple queries.
> You could expand on this by giving programming tasks and measuring code similarity.
But the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge.
My initial thought would be not so much to see whether they converge, but which ones seem to have the most similarity to each other, particularly along the lines of tasks we know are deliberate training goals.
But your point about competence cuts against my goal because it suggests that competent models would simply cluster on the right or efficient solution, which is of course true. So in a sense you want some task where competence is held constant or off the table in some way, which is what you are saying.
I hope somebody does this. I think there's valuable fingerprinting to be done that might suggest who is distilling whom, or at least who is training from common corpuses.
Worth to mention that with Claude and GPT this can be result of tournament sampling, which is part of text watermarking. Same answer for all Claude models kind of confirm it, imho.
What was your prompt? Most of these seem to be related to metaphors for "ideas" or thinking, or having a bright moment.
"Zephyr" and "breeze" might be related to forgetting everything, starting fresh.
So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these?
I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.
The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.
I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)
Of course one of the biggest problems we still see with LLMs is when you do the opposite. A highly detailed unique prompt is very likely to get terrible adherence or hallucination or both.
Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases.
The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math.
Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt?
I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test.
> The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.
OK well I couldn't resist this one:
llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
I think they're still visually pretty different. The most common shared details are:
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
I recently was testing something, I asked some models to provide me a single random word:
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt:
I really like this idea. You could expand on this by giving programming tasks and measuring code similarity. Seems like you could develop a pretty detailed understanding of similarities across multiple queries.
> You could expand on this by giving programming tasks and measuring code similarity.
But the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge.
Yeah that's definitely true.
My initial thought would be not so much to see whether they converge, but which ones seem to have the most similarity to each other, particularly along the lines of tasks we know are deliberate training goals.
But your point about competence cuts against my goal because it suggests that competent models would simply cluster on the right or efficient solution, which is of course true. So in a sense you want some task where competence is held constant or off the table in some way, which is what you are saying.
I hope somebody does this. I think there's valuable fingerprinting to be done that might suggest who is distilling whom, or at least who is training from common corpuses.
Just tried M365 Copilot with a premium account. Petrichor
Just tried Space Bunny and it gave me the same word...
I got Peregrine out of GPT-6 too. Huh.
Worth to mention that with Claude and GPT this can be result of tournament sampling, which is part of text watermarking. Same answer for all Claude models kind of confirm it, imho.
So not something internal to model thinking.
This feels uncannily like the ancestor of the Voight-Kampff test[0]
0: https://www.youtube.com/watch?v=Umc9ezAyJv0
[dead]
That is a cool idea. That astra gave the same word as claude is highly unexpected.
I saw an interesting matrix that claimed to show which labs were distilling Claude/OpenAI/Gemini models based on these similarities
What was your prompt? Most of these seem to be related to metaphors for "ideas" or thinking, or having a bright moment.
"Zephyr" and "breeze" might be related to forgetting everything, starting fresh.
So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these?
just “a random word” gives you Zephyr in Gemini, and “Lantern” in Claude and ChatGPT.
I got "Marmalade" in Claude (Opus 5.5)
Lantern in Sonnet 5.5
I got pomegranate in ChatGPT
I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.
The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.
I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)
Of course one of the biggest problems we still see with LLMs is when you do the opposite. A highly detailed unique prompt is very likely to get terrible adherence or hallucination or both.
Just tried Mistral Large 4: Serendipity.
Tried this with gpt-5.6-sol. Lantern!
The eqbench creative writing "slop profiles" do something similar. https://eqbench.com/creative_writing.html
Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases.
Muse Spark 1.3: lighthouse
The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math.
Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt?
I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test.
The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.
> The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.
OK well I couldn't resist this one:
Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...The rover honking is pretty silly, opus has a good sense of humor
I'm getting 403 inside the tool for this one. (The pelican bike on top works) This has been happening a lot recently.
That's a GitHub rate limit. Try again now, I just pushed a hopeful fix: https://github.com/simonw/tools/commit/7793fb74c2d37bd613cdc...
It does seem to, thanks!
Big L for mistral in this benchmark. Sorry Europe.
Not so sure, apparently it is the only one that considered that there are no paved roads on Mars.
To be fair, they don't have Armadillos in Europe. of course, you could say the same for Mars...
Gemini wins this one clearly. Honey please!
https://chatgpt.com/s/m_6ac53d4e5b0c8191949050dbf1f402d7
Not sure I'd call it jaywalking exactly but pretty good
Total Recall did promise us that the prostitutes on Mars would be freaks, but I never imagined it would go this far.
I think they're still visually pretty different. The most common shared details are:
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
They aren't. You aren't looking closely. For example, the first image does not have the frame of the bike in the correct shape even.
Also, why are they almost always riding from let to right?
It's been discussed many times. The reason is bikes are almost without exception depicted that way in order to show the drivetrain.
Ever seen a movie chase scene where cars are going right to left?
Everyone is stealing from everyone else.
Because it's a terrible benchmark