That’s what this is: it’s mechanical proofing.

As a "proof" of their poor judgement, take a paragraph you like. Ask the LLM to rewrite it to make it better (which I think we both agree will not make it better), and then in a fresh session ask it which it thinks is best. It'll almost always pick its own writing, even when it sucks.

Right, I agree: that’s exactly not what I’m recommending.

You are specifically recommending asking the model which of two versions is better (the quote in my top-level comment).

We both agree that they are poor "make it better" machines, but I also believe they are bad A/B testers and I'm using the former to demonstrate the latter.

You are writing both paragraphs. You’re specifically not asking a model to make a better paragraph. That would be a load-bearing debacle.

It doesn't matter who writes what, what matters is that LLMs have a preference for LLM-shaped writing. By A/B testing against an LLMs opinion, you are optimizing in the direction of LLM prose even if the LLM never writes any of the prose itself.

LLM style is not literally anticorrelated with quality. There are some things that they tend to do poorly, but you're not going to do those because you're writing the text yourself. If you have it judge your writing, and are careful to avoid the failure modes that the post goes into, it can be helpful by serving as a competent editor that doesn't share your blind spots.

I don't think it's a useful substitute for actual proofreading though, unless you're truly in a time crunch. When I proofread I don't just look for mistakes, I try to put myself into the position of a prospective audience member. LLMs seem to have absolutely no concept of "theory of the reader's mind". If I'm going for extra high effort, I have friends read it, usually asking them to identify anything that was unclear or hard to follow.