You can imagine that as people get used to working with Claude, they defer to its judgement. So the people choosing which RL path is better may say "yes, Claude, that was a good refactor!" because it did something hard that it may have been able to superficially justify. Actually the change was unnecessary and complicating.
The Claude trainers, as they themselves adapt to Claude's output, are collapsing in their own distribution, so even "new" from-human data is already contaminated.
Would more blame this on the LLM companies, they think they are on the verge of automating all work, I don't think they care about how you feel about the writing style of the Deus Ex Machina, it's not going to get fixed because to them Claude is already above a staff engineer and soon going to smarter than any human that will ever live. All the money will be going into improvements relevant to improving long context operating and correctness, they could fix the writing style but they are disinterested in that for frontier models, maybe some other companies are but they don't have as much money.