> if I was talking to a person who constantly used a phrase they liked I would notice it and it is possible I might get irritated by it

There are sociological reasons why this happens less with humans:

1. You cycle your dumb repetitive jokes with everyone you meet, so nobody hears it twice

2. Those who know you well will notice when you're just repeating ("dad jokes")

3. As a person's idiosyncrasies are beginning to wear on their social circles, they will be getting small clues to stop saying those things. Agents don't get these social between-the-lines cues to stop a certain behavior, they endlessly repeat. Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.

>Perhaps between model version releases, frontier labs can harvest the web and ask "What Claudisms do people mention negatively?" but I don't think they do that yet.

That won't happen. People can't phrase their objections in a succinct-enough way. When they do, the objection is superficial ("too many em dashes") and doesn't strike at the core of what makes LLM output bad.

Don't talk about other people in such a way. If we had a way in Claude to mark a word/token and downvote/reduce it logits then it wouldn't be to hard to send loadbearing to token valhalla.

In the future, I hope we get a way to randomize the language idiosyncrasies and/or personalities better.

The inevitable conclusion of retraining on the (now AI generated) web is that models begin to use the lower bits of fidelity in output text to pass messages to their future selves.

I didn't mean it as a criticism, just a statement of fact. I can't do it either. It's difficult to say what about a writing style is bad beyond vague descriptions.

This is conjecture, but why couldn't they hire the writing equivalent of a voice actor?

Honestly I don't think it would take much. Doing it ethically would take a little more money, but not much, for them.

Hire a prolific author with the writing equivalent of the "midwestern accent". I have a terrible writing accent, so it couldn't be me, but these people are out there. Pay them a bunch of dollars to ingest their entire corpus, and to produce more as needed. Use an LLM to match the writing style of this author during the post-training RL, and get a brand new Claude voice out of it.

They already do this at massive scale during the training process (apart from the paying the author part).

Boy do you underestimate everyday human abilities.

Claude in particular seems to try and use its internal thesaurus but not exactly align the senses. It constantly uses "address" for place or location and "grammar" for structure, meaning, etc.

This is kind of a nitpick, but it seems like with prose writing there are still some things to learn.

I think the last one is by far the biggest reason for why this happens way less in humans. In every conversation, almost for every sentence, people gauge the response their words have on whoever is listening. If something didn't land as expected (frown, confusion, unpredicted response) you unconsciously adapt and try something slightly different.

It's immediately obvious by the fact that you're clearly using way different ways of speaking (tone, speed, vocabulary) when you're speaking with a friend, versus your parents, your colleagues, people you don't know, children, etc.

Why don't these go in the system prompt or something that is easy to update?