It seems to me like he started out mad and looked to justify it.
I'm skeptical that anybody generating LLM text is really all that concerned about optimal word choice. Or even particularly good prose. But let's pretend that person exists.
If that person tried, say, an open model and that same model with watermarking applied, I'd be eager to hear their thoughts on the prose quality. Especially if they built an experiment harness and rated a few hundred blinded examples and found a measurable difference.
But getting this upset in advance of any demonstrated problem? It really seems to me like the point isn't the point
> It seems to me like he started out mad and looked to justify it.
that's been his thing since it was just a blog about apple product speculation and update. It's always been tedious.
Google has A/B tested watermarking on millions of responses. They say they observed no difference in user behavior.
It says they observed no difference in people clicking thumbs up or down. There are loads of other behavior that they didn't observe; like, say, switching to a different LLM.
How exactly do you propose they should keep track of quality, then, if not by A/B testing?
The question here is not to what product management advice to the Gemini team. The discussion here is whether the watermarking is noticeable.
The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.
> It seems to me like he started out mad and looked to justify it.
Yes, but that's neither surprising nor a reason to dismiss the anger. People get angry about DRM schemes in video games, even if the slowdown these cause is practically imperceptible. They're angry -- and Gruber acknowledges that factor too -- because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner.
It might be another instance of consequentialism vs. honor ethics. Many consequentialists don't seem to understand that something that doesn't have demonstrable consequences can still have moral implications.
> People get angry about DRM schemes
Yes but thats a thing that degrades something in a catastrophic way, as in I can use the thing one day, and not the next.
A different randomisation system on something that is a text generator which is designed to be unperceptable sounds like the people who are annoyed at FLAC vs MP3[1]
Done right you won't know the difference, done badly and you will.
[1] ex audio engineer, try me.
> the people who are annoyed at FLAC vs MP3
What’s the story there? I didn’t know that was a thing and I’m curious to learn more.
MP3 is lossy. FLAC is lossless. So obviously a certain type of people are going to make a religious war out of it.
> try me
Good luck explaining that one
> People get angry about DRM schemes in video games, even if the slowdown these cause is practically imperceptible.
Any potential "slowdown" doesn't even come close to making the list of top reasons people get upset about DRM.
Even if honor were real, LLMs do not have honor, and people using LLMs to write without disclosing that fact (or indeed at all, to some purists) do not have honor either.
> because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner.
It's LLM output! It's not your domain, it's the LLM owner's!
I assume they are just passing off AI prose as their own and don't want anyone to be able to tell. Which is surprising for someone who's been blogging for a thousand years. But I don't really see any other reason for this amount of heat and FUD.
Considering Gruber's always comically butthurt about regulation, especially EU regulation, your theory seems accurate.