Take a random essay and add in a bunch of the phrases that LLMs love like “load-bearing,” “crucial,” structural,” and “woven,” and then submit the original and the edited version to an LLM and ask which is better. It will choose the second one virtually every time. They have ingrained biases that associate those words with good writing and arguments.
This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.
The idea that that LLM reliability or bias can be solved with more LLM is... infuriatingly persistent.
Isn't this basically the mythical man-months LLM edition?
Sometimes I wonder if there’s just one guy somewhere who loved using the word load-bearing, all his papers got trained on, and now he can’t write anything without being assumed to be Claude.
The prose equivalent of Artgerm (a comic cover artist whose style looks to have heavily inspired a lot of AI art).