Do people not realize this will apply to ALL Claude output, not just writing you ask it to produce?
> Marks will apply to output from supported Claude models across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered.
I've never asked an LLM to generate writing I wish to post as my own. I don't understand why we think it is a good idea to fudge all output just so people can't cheat on their homework or generate slop. It won't have any impact on those things because there will always be models that don't do this.
It is deeply misguided regulation and Anthropic should have just said no on grounds of common sense.
The exact words matter to people who people who don’t use it to write for them. A lot of people use Claude as a friend/therapist/romantic partner/etc. That’s who I imagine would be most affected
>A lot of people use Claude as a friend/therapist/romantic partner
People developing a para-social (pseudo-social?) relationship with a corporate robot have far bigger problems than the word-chooser in their robot "friend".
What are you talking about? I understand OP has the perspective of writers, but let’s say you’re asking an LLM to explain a concept and it uses green words that are actually more difficult to understand. Or you ask for an analogy to explain a topic and the analogy doesn’t quite land because the description used “gray” instead of “overcast”.
I mean this is the thing that really comes off hard.
If you want precision and clarity of your writing, then you need to hand write it. Just like when you are optimising, its common to drop to a lower level language because the compiler doesn't express what you want. Sure its hard, but you know, thats kinda the point.
Even if you don't want that, the LLM is an average of the style it was trained to give out. Which is a homogenisation of the language to create a vague padding medium between a few generalised facts. (because a. it makes it less jarring when stuff is wrong, because its smeared over a higher amount of text and b. it looks more 'professional' because American business English is all guff and no meat)
Also yes, two isolated phrases may have subtly different meaning, frankly, the nuance is missed on most people. If you look at the interactions on here, at least 25% of the arguments are caused by people angrily reacting to the things _they_ thought the other person was saying, rather than what the actual person was saying.
So no its not a perversion, the LLM is, if you're gonna be picky about things.
The bigger problem is its effect on the way we use language. The amount of LLM generated or edited media is going to keep increasing and the media that people consume affects their own word choice.
I can't remember any radio determining words or adjusting grammar of the person speaking through it. Neither can I recall there has ever been a moveable type press, laser printer, or inkjet which bastardised the words of authors.
These were mediums _through_ which communication happened. Language models, large or small, are not any such medium.
What an absurd take on a very reasonable concern. As much as I hate “obviously AI” writing, I don’t understand why we have to handicap the tech and prevent it from improving. This manipulation of word choices virtually guarantees AI will always sound like a robot.
Have you compared pre- and post-watermarked text to make that assertion? And I'm not sure how writing your own text if you care deeply about the word choice is handicapping the technology?
I don't think there's been a single version of the major LLM providers that haven't handicapped the tech since the beginning by changing temperature/top-P/frequency penalty/etc. All of those stray from "highest probability" token selection. What's funny is it was done specifically to make it sound less like a robot/deterministic.
The problem is, that LLMs worked very well for me to improve my writing. Especially as I'm not a native speaker, it was a great way to improve the legibility of my work.
I want a tool that helps me improve my writing. A tool I can learn from. Not a tool that switches out "bananas" to "airplanes."
I'm was using Claude Opus and now Fabel extensively for editing my texts and I find the recent updates abysmal. Not sure if it's due to the Text Watermark.
Before Claude was great in sharpening the meaning in my writing, it's now close to unusable.
I'll be interested to see if people can actually pick out which text is watermarked and which isn't once they introduce it. It won't switch "bananas" to "airplanes". It'll switch "I really enjoy eating bananas" to "I love eating bananas" or similar
To that point, given a corpus of writing from person A and another from person B, I have no doubt it’s easy to train a classifier to determine who wrote what. In fact I believe law enforcement agencies already have these classifiers.
The only difference here is that Anthropic is actively trying to make the watermark undetectable.
They're not trying to make the watermark undetectable, that would defeat the point of a watermark. They're making it detectable, but not make the text obviously watermarked
This is an excellent use case. It made learning German much easier. I write what I think is correct, then get a fixed version.
When writing in English though, I use it more like a dictionary. If you want to write past a certain level, an LLM works better as a metaphor and idiom search engine.
I also like to ask it to generate 20 ways to say the same thing. It’s a great way to simplify or smoothen sentences without losing your voice.
Interesting, are you being literal about requesting 20 ways of saying the same thing? It seems pretty excessive and I'm actually impressed that asking an LLM to rewrite a statement 20 times yields output that isn't excessively redundant. Does the LLM do a pretty good job reading your mind, or do you still find yourself manually piecing together pieces from the 20 suggestions into a satisfactory sentence?
Yep, I literally ask it to say it in 20 different ways. Sometimes a sentence structure or a combination of words will just work better. The goal is usually to simplify a sentence without losing meaning.
I don't expect the LLM to read my mind. The unit of work is too small for intent to matter, and I'll just steer the next recommendations in a direction as needed.
Most of the suggestions are crap, but they can contain the seeds of a good sentence.
As a writer, Claude’s metaphors are trite and obvious 90% of the time. It also has a terrible penchant for an immediately recognizable emphatic voice that makes even the best outputs super cringe.
But people do not notice and do not care. In our bubble and John Gruber’s bubble we really overestimate how much people must care. Truth is, we’re in this predicament because statistically speaking the people who hive a fuck are a rounding error.
They really are. This is why I prefer the volume approach. I might not accept any of the ideas it spits out, but it often guides me in a direction I was not considering.
I know that most people don't care, but my online presence is a search query for interesting people, so I care about what I put into it.
Nitpicking here in a way that I would usually avoid, but it is relevant to the conversation being had and it seems like you might appreciate the information... "Readability" would be the more correct word to use here instead of "legibility".
Legibility is close enough for me to know what you mean based on the context, but it really applies to the visual presentation and how easy something is to read at a symbolic level (whether someone's handwriting or font choice is good or bad impacts legibility, whether someone uses good grammar or not impacts readability).
LLMs are no more than pen and paper at this point. Especially for those who aren't trying to create slop. We all would want our pens to accurately reflect the strokes (well in this case thoughts) rather than adding tiny watermarks to identify that it is generated by a particular pen or a user.
Watermarking per model is just the start. The method is cheap enough to distinguish individual users.
That is an insane statement, LLMs generate swaths of text from almost nothing.
If they are adding so little value as to be as transparent as a pen and paper then why use one at all? Transcription doesn't need an LLM so that's not what you're taking about I assume.
LLMs do a whole lot more than writing your thoughts down. They write extra text. If you just want a pen and paper, use Notepad. Or better, a pen and paper
Do people not realize this will apply to ALL Claude output, not just writing you ask it to produce?
> Marks will apply to output from supported Claude models across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and wherever Claude is offered.
I've never asked an LLM to generate writing I wish to post as my own. I don't understand why we think it is a good idea to fudge all output just so people can't cheat on their homework or generate slop. It won't have any impact on those things because there will always be models that don't do this.
It is deeply misguided regulation and Anthropic should have just said no on grounds of common sense.
If the exact words we choose when writing matter so much, then why use a non-deterministic LLM that produces slightly different output on every run?
Additionally, LLM are already rolling the dice on different wordings so I don't see how this watermarking makes it any less precise.
The exact words matter to people who people who don’t use it to write for them. A lot of people use Claude as a friend/therapist/romantic partner/etc. That’s who I imagine would be most affected
>A lot of people use Claude as a friend/therapist/romantic partner
People developing a para-social (pseudo-social?) relationship with a corporate robot have far bigger problems than the word-chooser in their robot "friend".
What are you talking about? I understand OP has the perspective of writers, but let’s say you’re asking an LLM to explain a concept and it uses green words that are actually more difficult to understand. Or you ask for an analogy to explain a topic and the analogy doesn’t quite land because the description used “gray” instead of “overcast”.
Watermarking or not, LLM are already using RNG to pick a variant between different words/expressions.
I mean this is the thing that really comes off hard.
If you want precision and clarity of your writing, then you need to hand write it. Just like when you are optimising, its common to drop to a lower level language because the compiler doesn't express what you want. Sure its hard, but you know, thats kinda the point.
Even if you don't want that, the LLM is an average of the style it was trained to give out. Which is a homogenisation of the language to create a vague padding medium between a few generalised facts. (because a. it makes it less jarring when stuff is wrong, because its smeared over a higher amount of text and b. it looks more 'professional' because American business English is all guff and no meat)
Also yes, two isolated phrases may have subtly different meaning, frankly, the nuance is missed on most people. If you look at the interactions on here, at least 25% of the arguments are caused by people angrily reacting to the things _they_ thought the other person was saying, rather than what the actual person was saying.
So no its not a perversion, the LLM is, if you're gonna be picky about things.
The bigger problem is its effect on the way we use language. The amount of LLM generated or edited media is going to keep increasing and the media that people consume affects their own word choice.
> affects their own word choice.
exactly. in the same way that printed books affected word choice, so did the radio.
I can't remember any radio determining words or adjusting grammar of the person speaking through it. Neither can I recall there has ever been a moveable type press, laser printer, or inkjet which bastardised the words of authors.
These were mediums _through_ which communication happened. Language models, large or small, are not any such medium.
What an absurd take on a very reasonable concern. As much as I hate “obviously AI” writing, I don’t understand why we have to handicap the tech and prevent it from improving. This manipulation of word choices virtually guarantees AI will always sound like a robot.
Have you compared pre- and post-watermarked text to make that assertion? And I'm not sure how writing your own text if you care deeply about the word choice is handicapping the technology?
I don't think there's been a single version of the major LLM providers that haven't handicapped the tech since the beginning by changing temperature/top-P/frequency penalty/etc. All of those stray from "highest probability" token selection. What's funny is it was done specifically to make it sound less like a robot/deterministic.
Reasonable? A concern which is based on no real data?
With all things going on among AI bros and the AI industry as a whole, are you really that surprised there is a widespread aversion against the tech?
One could also use butterflies to write ;) https://xkcd.com/378/
The problem is, that LLMs worked very well for me to improve my writing. Especially as I'm not a native speaker, it was a great way to improve the legibility of my work.
I want a tool that helps me improve my writing. A tool I can learn from. Not a tool that switches out "bananas" to "airplanes."
I'm was using Claude Opus and now Fabel extensively for editing my texts and I find the recent updates abysmal. Not sure if it's due to the Text Watermark.
Before Claude was great in sharpening the meaning in my writing, it's now close to unusable.
I'll be interested to see if people can actually pick out which text is watermarked and which isn't once they introduce it. It won't switch "bananas" to "airplanes". It'll switch "I really enjoy eating bananas" to "I love eating bananas" or similar
That’s my point of the post … I had the feeling that the English text editing skills of Claude went significantly down in August.
I was frustrated at first not knowing what they are doing.
After I read this post, (seeing they introduced it in August) I think it has to do with the watermarking.
Try it on a paragraph … the connections between sentences feel clunky now.
I will play with it more and see if that’s really the case (the watermarking making the text edits worse).
To that point, given a corpus of writing from person A and another from person B, I have no doubt it’s easy to train a classifier to determine who wrote what. In fact I believe law enforcement agencies already have these classifiers.
The only difference here is that Anthropic is actively trying to make the watermark undetectable.
They're not trying to make the watermark undetectable, that would defeat the point of a watermark. They're making it detectable, but not make the text obviously watermarked
Undetectable by a human reader. Come on, give the post a charitable reading.
This is an excellent use case. It made learning German much easier. I write what I think is correct, then get a fixed version.
When writing in English though, I use it more like a dictionary. If you want to write past a certain level, an LLM works better as a metaphor and idiom search engine.
I also like to ask it to generate 20 ways to say the same thing. It’s a great way to simplify or smoothen sentences without losing your voice.
Interesting, are you being literal about requesting 20 ways of saying the same thing? It seems pretty excessive and I'm actually impressed that asking an LLM to rewrite a statement 20 times yields output that isn't excessively redundant. Does the LLM do a pretty good job reading your mind, or do you still find yourself manually piecing together pieces from the 20 suggestions into a satisfactory sentence?
Yep, I literally ask it to say it in 20 different ways. Sometimes a sentence structure or a combination of words will just work better. The goal is usually to simplify a sentence without losing meaning.
I don't expect the LLM to read my mind. The unit of work is too small for intent to matter, and I'll just steer the next recommendations in a direction as needed.
Most of the suggestions are crap, but they can contain the seeds of a good sentence.
As a writer, Claude’s metaphors are trite and obvious 90% of the time. It also has a terrible penchant for an immediately recognizable emphatic voice that makes even the best outputs super cringe. But people do not notice and do not care. In our bubble and John Gruber’s bubble we really overestimate how much people must care. Truth is, we’re in this predicament because statistically speaking the people who hive a fuck are a rounding error.
They really are. This is why I prefer the volume approach. I might not accept any of the ideas it spits out, but it often guides me in a direction I was not considering.
I know that most people don't care, but my online presence is a search query for interesting people, so I care about what I put into it.
> improve the legibility of my work.
Nitpicking here in a way that I would usually avoid, but it is relevant to the conversation being had and it seems like you might appreciate the information... "Readability" would be the more correct word to use here instead of "legibility".
Legibility is close enough for me to know what you mean based on the context, but it really applies to the visual presentation and how easy something is to read at a symbolic level (whether someone's handwriting or font choice is good or bad impacts legibility, whether someone uses good grammar or not impacts readability).
> One could also use butterflies to write ;) https://xkcd.com/378/
This comparison is frankly absurd.
LLMs are no more than pen and paper at this point. Especially for those who aren't trying to create slop. We all would want our pens to accurately reflect the strokes (well in this case thoughts) rather than adding tiny watermarks to identify that it is generated by a particular pen or a user.
Watermarking per model is just the start. The method is cheap enough to distinguish individual users.
That is an insane statement, LLMs generate swaths of text from almost nothing.
If they are adding so little value as to be as transparent as a pen and paper then why use one at all? Transcription doesn't need an LLM so that's not what you're taking about I assume.
LLMs do a whole lot more than writing your thoughts down. They write extra text. If you just want a pen and paper, use Notepad. Or better, a pen and paper
What? The entire reason I use an LLM is to be able to avoid thinking about a topic.
That's their whole damn value prop: outsourcing thinking and producing without understanding.
I don't need to read emails in detail to respond any more.
>LLMs are no more than pen and paper at this point.
Then use pen and paper. It is the same, you say, right?