In today’s column, I examine and bust or straighten out various misleading myths regarding the recently released AI text-oriented watermarking feature of Anthropic. The mainstream news and social media have been making zany and incorrect claims about what this particular technique of watermarking is and does. This, in turn, has tended to create widespread confusion and undue consternation among people who are unsure whether their text will be watermarked or not when using Claude.
Weighty questions on people’s minds include whether text that they submit to the AI for proofreading will end up watermarked, and whether the AI fixing incidental typos will also encompass the implanting of a watermark. To answer those questions, I will briefly lay out how it is that the watermarking actually occurs and explain how the hidden watermarks are infused into text (for my in-depth coverage, see the link here, and for my analysis of watermark detection tools, see the link here). You will end up with a much cleaner understanding of how to judge whether to use the AI for aid in writing and editing of text, and the likelihood of a watermark getting included in the text.
Let’s talk about it. This analysis of AI breakthroughs is part of my ongoing Forbes column coverage of the latest in AI, including identifying and explaining key AI complexities (see the link here).
Watermarking Is Challenging
First, some foundational aspects of the topic of watermarks. We are all aware of watermarking when it comes to paper-based materials and likewise for any tangible artifact that exists in a definitive physical form. A dollar bill can contain a watermark, allowing the naked eye to tell whether it is real or counterfeit. Watermarks can also be hidden from visual inspection, requiring some other means to detect the watermark.
Watermarking for digital photographs and graphical images is more readily accomplished than with text since you can embed all sorts of digital ones and zeros that won’t impact the picture, but that can be detected by inspecting the binary representation. It is possible to use sophisticated mathematical algorithms to populate the bits in a manner that almost no one other than someone armed with the algorithm can later detect as being part of a special pattern.
Trying to watermark digital text is a beast of a different kind. Anything that is done to the text will potentially alter the words we see and impact the meaning of the text. If you had a watermarking algorithm that simply said to replace the word “of” with the word “horse”, the resulting text, which is now presumably discernible as AI-written due to the excessive use of the word “horse”, is going to be nonsensical for human use. Likewise, if the watermark consisted of embedding special characters or the use of emojis, you could quickly find those and remove them easily.
Example Of How It Works
An ingenious way to infuse watermarking is to do so by selecting suitable words that can be viably chosen during the AI writing process. Here’s how that works. Envision that AI is generating a response to a prompt, doing so one word at a time. Each word is carefully chosen by the AI. The choice of which word to use is made from several possible words at each step.
Suppose the prompt was asking the AI how to make a ham sandwich. The AI might start assembling the response word-by-word and could have arrived at these choices: “Place a slice of ham onto a bagel and add mustard.” Each word was selected on a one-at-a-time basis, going from the start of the sentence to the end of the sentence.
When the AI got to the word about the bread, in this instance the word selected was “bagel,” but there were several other options available, such as saying “flatbread” (statistical second choice), “wheat bread” (statistical third choice), “white bread” (statistical fourth choice), and other possibilities. Assume that the word “bagel” was the statistically top-ranked choice overall and therefore chosen accordingly.
Aha, in the realm of watermarking, the AI might opt to intentionally choose the second choice rather than the top-ranked choice; thus, the sentence comes out as “Place a slice of ham onto a flatbread and add mustard.” If the AI consistently keeps picking the second choice for many of the words that are being chosen, this becomes a handy pattern for the AI. A human looking at the sentence doesn’t realize that the second choice is being chosen. They see a sentence that looks completely normal.
Detecting The Watermark
I think you can see that this statistical uplift is going to be quite hard to detect by conventional means. Humans are unlikely to discern the watermark by looking for any patterns in the wording. All the sentences are still going to make sense and abide by whatever the topic at hand is. The subtlety of picking the second statistically viable word on numerous occasions is a hidden way of producing the watermark.
How does an authorized detection tool figure out if the watermark is present?
That is the added trickery. The chances of any usual automated detection method ferreting out the watermark are low. It won’t realize what the watermark method is or how to ferret it out. Meanwhile, a detection tool that is built knowing the method can examine the sentences and compare the word choices to the pattern of word choices that the AI would normally make. If the second word choice is consistently being encountered in the examined text, this is a strong indicator that the AI indeed generated that content.
We can make this method much stronger. Instead of always choosing the second choice, the watermark process does something else. Suppose that 50% of the time the second choice is made, 30% of the time the third choice is made, and 20% of the time the fourth choice is made. This makes things even harder for anyone else to crack and find the watermark. An even better method includes having a secret cryptographic key that guides the watermarking process toward the preferred token patterns.
Breaking The Watermark
Anthropic stated in their recent announcement about the new watermarking that the watermark will persist when the text is copied and pasted somewhere else. This makes abundant sense. A body of text produced by Claude is going to carry the watermark since it has that secret pattern of word choices. If you copy it as is and then paste it as is, the pattern of the words remains precisely as the AI generated it. Ergo, the watermark is still entirely present.
The announcement by Anthropic also noted that the watermark can tolerate some semblance of editing. The question is how much editing can be done to the text before the watermark breaks down and is no longer statistically significant. This is somewhat complicated to identify since we are then relying on statistics and probabilities.
Pretend that I take the watermarked sentence that says to make a ham sandwich with flatbread, and I change the word to a bagel. I have marred the watermark. The AI had explicitly chosen the word flatbread, and it is no longer there. Instead, the word bagel is there.
If the body of text is relatively short, my making that one change could materially undermine the statistical likelihood that the text contains the AI watermark. Perhaps that is the only watermarked chosen word in that sentence. On the other hand, if the body of text is very lengthy, many paragraphs in size, my having replaced one word in one sentence is perhaps like dropping a pebble into the ocean. The watermark remains substantially intact because it is pervasive throughout the rest of the text.
Rules Of Thumb
The larger the body of text that the AI outputs and watermarks, the less destructive to the watermark are my few edits. You see, there will still be a preponderance of text that contains the watermark. The statistical signal of the watermarks might remain at some high percentage after my edits, perhaps 90% to 99%. That is potentially enough to be somewhat sure that the watermark is in there. If the watermarks remain at only 10% after my edits, now things are getting dicey. The detection of the watermark is going to be on thin ice to conclude that the watermark is truly there.
The crux is how much of the text contains the watermarked approach, and how much of the text does not contain the watermarked approach. Suppose I use Claude to generate a few paragraphs for me. Claude produces the text, and it contains the secret watermark due to the words selected for the essay. I then paste that text into a document of ten pages of my own hand-crafted text. At this juncture, suppose that someone wants to know whether the ten pages were handwritten or generated by Claude.
An Anthropic-authorized AI watermark detection tool, which hasn’t yet been publicly made available, will presumably scan the ten pages of text and look to see if the pattern of word selections matches the approach being used by Claude. The few paragraphs will likely get flagged, but the rest of the ten pages are unlikely to get flagged since it wasn’t produced by the word selection method.
A big concern is that the authorized watermark detection tool might simply claim that the entire text was likely generated by Claude, even though only a small portion of the text seems to be so generated. The hope is that the detection tool will be more transparent and offer an estimated likelihood, such as that it is slim, modest, or highly likely that the watermark is present. Worries are that people are going to run with whatever the detection tool says, despite the reality that the watermark might only marginally be present.
Proofreading Of Text
There is confusion in the media about what happens if you ask Claude to proofread a body of text. Some have been warning that the mere act of Claude scanning a body of text is going to somehow magically infuse a watermark into the text. That’s a false or misleading portrayal of the situation, so let’s properly unpack things.
First, assume that I take a body of text that I handwrite and I give that to Claude to proofread. I tell Claude not to change any of the wording. It is only to scan the text and let me know if there are any potential gaffes or factual inaccuracies in my text. Claude proceeds to scan the text and gives me a response that identifies a handful of portions of the text that could use some cleanup. The body of text is still the same as it was when I submitted it to Claude.
Do you think that the text now contains the secret watermark?
I trust that you realize the original text has been untouched in the sense that it wasn’t rewritten by Claude; therefore, the text has not been watermarked. A small twist to keep in mind is that the response about where there are gaffes is indeed watermarked, because that answer was generated by Claude. But the text I submitted to be proofread is not watermarked.
Proofreading And AI Making Edits
Now that you’ve got the gist of things, we are ready to get into more complicated situations. First, I handwrite an essay. I then give the essay to Claude and ask it to proofread the essay. In addition, I tell Claude to go ahead and fix the essay, meaning that Claude can reword the text that I have written.
Keep in mind that we are now allowing Claude to start infusing a watermark. How so? Any of the wording changes that Claude makes are a potential moment for Claude to select a suitable chosen word that is based on the secret approach. I might have a sentence that says the dog jumped over the lazy fox, and Claude changes that wording to say that the dog leapt over the lazy fox. The sentence still has roughly the same meaning, but Claude snuck the word “leapt " into the text and presumably did so as a watermarking insertion.
You will not have any immediate clue whether the wording changes are part of the watermark, though they might very well be. Some of the words that Claude changes could be part of the watermarking, while other words might not be. The main point is that you are giving Claude a chance to make word choices that could allow for the watermarking to occur.
Using the prior rule of thumb, if Claude makes a relatively small number of changes in terms of replacing words with other words, and if the body of text is large, the amount of watermarking is going to be slim. The watermark detection will hopefully not clamor that the text has been generated by Claude. It will presumably find the small portion that is now watermarked and indicate that it is likely that Claude was involved to some extent in the generation of the subset of text. We opened that door to this by allowing Claude to not only proofread but also edit the text.
Fixing Of Typos
One of the flagrant misstatements about the watermarking is in reference to Claude fixing typos. What do you think happens when Claude is asked to fix a misspelled word?
Suppose that I have typed the word “catt” in my essay, and I meant to type the word “cat”. I misspelled the word. I give Claude my essay. I tell Claude to only fix any typos. It is not to do any broad-based editing. Just find and correct any typos. Sure enough, Claude shows me the resulting essay with a few typos that were corrected. The word “catt” is now correctly shown as “cat”. The same goes for the other words that I inadvertently misspelled.
Did this give Claude an opportunity to include the watermarking?
Not if you were explicit about only fixing the in-place words that were misspelled. For example, I had chosen the word “cat” and simply misspelled it. Claude makes the correction and turns the word into “cat”. Claude did not choose the word “cat” at the get-go; I did so. There wasn’t an opportunity for Claude to select some other word; it only made sure my chosen word was correctly spelled.
The twist is this. Suppose that you give Claude some leeway. It finds the word “catt” and determines that the word ought to be spelled as “cat”, but if your instructions were loosey-goosey, Claude might replace the word “cat” with the word “feline”. Perhaps the word “feline” is going to be a word choice by Claude that serves as an element of a watermark. Any moment when Claude can choose which word to include is a chance for Claude to go the route of watermarking.
The World We Are In
Things are going to get quite topsy-turvy once all the other major LLMs implement a form of watermarking, including ChatGPT, GPT-5, Gemini, Copilot, Grok, etc. Each AI maker will potentially adopt an approach that is slightly different from the other AI makers. The odds are that they will use a similar technique of statistically using wording choices as their text-oriented watermark technique, but differ enough that there isn’t one universal means at play (as a side note, this could still be undertaken in a semi-standardized way).
The bottom line is that each AI maker will then need to provide a proprietary authorized detection tool so that people can run text through the tool to find out if the text was potentially generated by the particular AI. I’m betting this will sow confusion. Here’s how. Somebody might have used Claude to generate text, and a person wanting to check it runs the text through an OpenAI watermark detection tool that says the text is free and clear of being composed by ChatGPT. An unaware person won’t realize that they fed the text into a different watermark detection tool and should have used the Anthropic tool instead.
A final thought for now. Albert Einstein famously made this remark: “There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.” The prevailing approaches to having AI watermark text are not really a miracle, though you do need to give credit for the underlying cleverness of human designers. All of us will eventually become used to AI-generated watermarks and take it in stride. Some won’t like it; others will relish it. Time will tell.