You're right, and the load-bearing part of the argument is not what you think it is. Two ambiguities worth resolving before moving on: whether what you wrote also applies to ChatGPT, and whether you have custom instructions set up. Failure mode worth flagging explicitly: I didn't read TFA.
(I'm becoming allergic to how these things write).
I am curious why LLM writing has such an uncanny valley feel to it. Like if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it, but I wouldn’t necessarily.
In high school I had a teacher that would say “that type of thing” a lot. One time my friend and I counted it during one class period and he averaged to use the phrase every 48 seconds on average. It was funny, but it never irritated us.
And this is just one example of I am sure thousands I have personally experienced where a friend, family member, or coworker has a peculiar way of speaking and it at most feels odd but not annoying. Yet when I see an emdash now I instantly feel irritated.
And I say this as someone who actively enjoys using Claude and other LLMs, including coding, casual research, or even having it explain pop culture phenomenon or sociology research to me.
> if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it
There are sociological reasons why this happens less with humans:
1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice
2. Those who know you well will notice when you're just repeating ("dad jokes")
3. As a person's idiosyncrasies are beginning to wear on their social circles, they will be getting small clues to stop saying those things. Agents don't get these social between-the-lines cues to stop a certain behavior, they endlessly repeat. Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.
>Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.
That won't happen. People can't phrase their objections in a succinct-enough way. When they do, the objection is superficial ("too many em dashes") and doesn't strike at the core of what makes LLM output bad.
Don't talk about other people in such a way. If we had a way in Claude to mark a word/token and downvote/reduce it logits then it wouldn't be to hard to send loadbearing to token valhalla.
In the future, I hope we get a way to randomize the language idiosyncrasies and/or personalities better.
The inevitable conclusion of retraining on the (now AI generated) web is that models begin to use the lower bits of fidelity in output text to pass messages to their future selves.
I didn't mean it as a criticism, just a statement of fact. I can't do it either. It's difficult to say what about a writing style is bad beyond vague descriptions.
This is conjecture, but why couldn't they hire the writing equivalent of a voice actor?
Honestly I don't think it would take much. Doing it ethically would take a little more money, but not much, for them.
Hire a prolific author with the writing equivalent of the "midwestern accent". I have a terrible writing accent, so it couldn't be me, but these people are out there. Pay them a bunch of dollars to ingest their entire corpus, and to produce more as needed. Use an LLM to match the writing style of this author during the post-training RL, and get a brand new Claude voice out of it.
They already do this at massive scale during the training process (apart from the paying the author part).
Boy do you underestimate everyday human abilities.
Claude in particular seems to try and use its internal thesaurus but not exactly align the senses. It constantly uses "address" for place or location and "grammar" for structure, meaning, etc.
This is kind of a nitpick, but it seems like with prose writing there are still some things to learn.
I think the last one is by far the biggest reason for why this happens way less in humans. In every conversation, almost for every sentence, people gauge the response their words have on whoever is listening. If something didn't land as expected (frown, confusion, unpredicted response) you unconsciously adapt and try something slightly different.
It's immediately obvious by the fact that you're clearly using way different ways of speaking (tone, speed, vocabulary) when you're speaking with a friend, versus your parents, your colleagues, people you don't know, children, etc.
Why don't these go in the system prompt or something that is easy to update?
I think that there is a sort of mechanism by which, when a human speaks to you, you can "mirror" their internal mental state, and the quality of writing often corresponds to how much you get pleasure or information or whatever your goal is from that mental state. The important part is that the words themselves are just pointers to the state. So, the person speaking, if they are skillful, gives enough words, and enough variety, that you can start to produce a state yourself which resembles theirs.
This is the thing that LLM writing doesn't really do. Since there isn't a mental state, there's nothing to mirror anyway, and somehow you can detect the absence of it even though it's hard to put any of the machinery into words. Corporate speech also fails to activate this machinery in the same way, but LLMs seem to do it more egregiously, probably because the corporate speech was at least compiled by a human---even if it is not the thoughts of an individual, it is the "thoughts" of an "entity", the abstract corporation, which the writer was speaking for, and so you can still wrap your mind around the fact that it is communicating with you.
> I am curious why LLM writing has such an uncanny valley feel to it.
Because they are HEAVILY trained to give addictive responses.
They don't want to just answer your question. They want to sycophantically make you feel like a genius for being smart enough to use them.
This makes some good intuitive sense, but to me the sycophancy feels like it is an emergent property of turning a next word predictor into a conversational chatbot whether or not it’s intentionally trained that way. Your prompt and its earlier responses is all it has in its context window, so of course it lends undue importance to everything you say. Does that seem like a contributing factor to you?
> trained to give addictive responses
I've heard this a lot but I'm not sure it makes sense. Nobody I talk to like Claude's output. In fact, they all loathe it.
Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?
I would think that it would be the opposite. Nobody is seriously detoxing or comfortmaxxing AIs yet. Human brains are fed back its own output in learning mode, so we are great at removing whatever we feel uncomfortable from our output, online and offline. No such paths exist for AIs.
>Is there a silent majority of Claude users who really enjoy what we call the LLM-isms? Maybe, but isn't Claude also largely aimed at developers?
Purely out of my own curiosity, I just asked Claude to have fun with itself by making itself a game it enjoys, to play it, and to write its experience.[1] I don't know if it's true or confabulated (maybe it doesn't really know its experience and is just hallucinating it) but I didn't mind reading it, the prose is fine for me. I don't mind reading Claude's writing. I mean let's be honest, we all read Claude's writing all day, most of the submissions on the front page on any given day are written by Claude.
Just before I made that game, I had Fable write up a scholarly report on any subject[2], it chose introspection by LLM's. (This is what made me think of asking it to play a game.) I didn't mind reading it, even though I don't think it really added anything very interesting. I don't think what it wrote is worth publishing, but I read it with interest.
I found I could read it easily and get up to date on the state of this question that it picked to answer.
So the bottom line is I don't mind reading Claude's output that much. Of course, I'm annoyed every time it says "honest", "genuine", "load-bearing", whenever it pushes back gently against something, etc. But it's not the end of the world.
[1] https://github.com/robss2020/claude-fable-5-having-fun
[2] https://claude.ai/share/f0122611-22c0-43a5-ab4a-d6863167bdd6
> https://github.com/robss2020/claude-fable-5-having-fun
If you haven't already seen it, you might appreciate https://www.anthropic.com/research/global-workspace. That's what this made me think of anyway.
thanks for the link! super interesting.
This seems to be the load-bearing point that matters.
One reason is that one person’s idiosyncracies are limited in scope, but LLM-produced text is now everywhere. Also, filler words and mannerisms in speech we’re quite good at filtering out, but in written text the stand out much more.
I don't think its an inherent quality - a lot of older models had a much more natural feel to them.
I guess it has to do with the low temperature (low randomness in final word choce - something all vendors seem to have converged on for some reason) which does make them less likely to skiz out but makes the writing feel dry and samey. It's like repeating a list of dice rolls, and replaying them - there's no inherent pattern in the input, but there sure is in the output.
I suspect because the English colloquial text training data it had access to was early 2000's message boards and social media. Therefore it uses "honestly" and other expressions that took hold in the late 90s and early 2000s.
I had pretty good luck recently by giving it a writing guide about word choice, sentence structure, paragraph structure, and overall doc structure. I basically ask it to read the guide and revise a couple of times before I engage with its writing. Ymmv.
I did this too, but it usually thinks its writing is fine in my experience. Even when spawning a subagent, it thinks its effusive comments are fine. It's driving me nuts. Before I commit I end up ripping out 90% of the comments, and rewording the rest, otherwise I'd be drowning in comments. This is my style guide: https://github.com/smj-edison/zicl/blob/main/CLAUDE.md#style...
Share pls :D
Think of it as a mad lib, it’s populating a template, and seeing the same template filled over and over gets tiring.
It’s the repetitiveness of style, the attempt to make everything seem as impactful as possible, the use of short sentences (too much Hemingway in the training data?), and obvious patterns like “it’s not this, it’s that” and several others.
Real human writing doesn’t follow such strict rules. When the same small set of rules is applied over and over throughout a text, it becomes obviously strange and machine-like.
Wikipedia's own "Signs of AI Writing" page distills it nicely:
I blame RLHF entirely for this. Nobody used to talk like AI speech before.
So is it the LLM or us that's getting the RLHF? /s
I’d guess it’s because that’s exactly how its System Template was written to do
> I am curious why LLM writing has such an uncanny valley feel to it.
Because it's trained to talk like a marketing committee.
[dead]
Very surprised no one has this answer: because it is fundamentally not a human being.
Might be related to their fingerprinting of llm output they said earlier in the week.
It also picks up and obsesses about weird details. You're in the middle of a deep technical discussion and it will divert to point out that it made a mistake in some example code it's just found.
You are diabolical.
Every one of them has their own particular flavour of this aggravation too. Gemini has been my standard go-to for non-coding tasks for a while, but I started to get really annoyed with a couple aspects, especially how it would end almost every response with a barely related "would you like to do this next??" tangent, regardless of my prompt to the contrary. So I've been using Claude more for regular tasks, and am now running into its brand of infuriating idiosyncrasies. I'm also hesitant to try to code too much of this out with system prompts, for fear of degrading the outputs.
forgot the, "my original claim was overstated"
This is painfully accurate.
Well played.
Alas, it writes so much better than the average human that it's what everyone started using. Hence the utter familiarity and now contempt.
I agree that it’s better at writing than a 50%-ile human, but it’s worse at communicating through writing than most humans.
Even an average human writer can communicate details much more succinctly and directly than an LLM
Not at all. I think you have a mistaken view of who the average human is. They are terrible at turning their thoughts into written language. It's just nebulous clouds. Claude is like 75th percentile at communicating ideas.
I think that’s true when you compare to the average white collar professional who does a lot of writing: better at writing, not better at communicating.
But compared to the average adult? I think you forget just how bad at writing the average person is.
No, it does not.
You are vastly overestimating the average writer’s ability.
Even then, I don't care about the "average writer". I want great output. I like to imagine that developers have some self-respect, but by now everyone in the industry is spending hundreds of hours every month reading some of the most poorly written prose we could imagine, simply because it affords us to think less.