How Claude's text watermarking works

Aug 15, 2026 02:15 AM - 1 hour ago 1

Future Claude models will make matter that contains a watermark. This is simply a measurement of determining the likelihood that Claude was progressive successful penning the text, and we, on pinch respective different awesome AI providers, are implementing this alteration to comply pinch the EU AI Act.

In this article, we stock answers to immoderate of the questions we’ve received astir really our chosen watermarking method works, whether it affects Claude’s outputs, and why we’re making this change. To summarize:

  • We usage a method of watermarking that does not person immoderate applicable effect connected the value aliases contented of Claude’s outputs;
  • The quality betwixt watermarked and un-watermarked matter will not beryllium distinguishable to readers;
  • Nothing is added to the matter and location are nary hidden characters;
  • Watermarking doesn’t require other tokens, and will not beryllium much expensive;
  • Watermarking carries nary identifying accusation and can’t beryllium traced to a circumstantial person, organization, aliases chat;
  • Watermarking won’t beryllium circumstantial to Claude. As of August 2, the EU requires AI providers serving its marketplace to people AI-generated content. Other awesome exemplary developers person signed the aforesaid Code of Practice and will beryllium implementing their ain watermarks.

What is watermarking?

Large connection models for illustration Claude activity by generating 1 connection astatine a time. Each clip the exemplary decides connected the adjacent word, it chooses among a database of imaginable candidates, yet selecting the astir sensible aliases apt based connected the preceding text. Take the condemnation “The upwind coming was acold and…”. The adjacent connection is very improbable to beryllium “sugary.” But it is rather apt to beryllium “overcast” aliases “grey.” Under astir circumstances, it doesn’t matter overmuch to the scholar which of these second 2 words the exemplary yet chooses—the meaning of the condemnation is mostly the aforesaid either way. In cases for illustration this, the prime is settled by a random number.

Watermarking uses low-stakes choices for illustration these—which hap galore times complete a portion of generated text—to time off a shape successful Claude’s responses. That shape is undetectable to the reader, but is detectable to anyone who has a cardinal that encodes it. When watermarking is used, choices are still made astatine random, but the source of the randomness is different. Instead of utilizing an arbitrary random number generator to prime the adjacent word, watermaking uses the cardinal and a fewer words that travel earlier to settee what connection the exemplary should pick. That is, the words that Claude picks are still random, but now, 1 tin cheque the series of words and spot if it’s accordant pinch the choices Claude would make if it was utilizing the key. If it is, 1 tin delegate a probability that the matter was generated by Claude.

Importantly, it isn’t that the exemplary will now ever beryllium biased toward overcast aliases grey. Just arsenic pinch non-watermarked text, overcast mightiness beryllium selected successful 1 sentence, grey successful the next, depending connected the words that came before. And it’s not the lawsuit that the watermarking method pushes Claude to take a connection it wouldn’t person considered anyhow (for instance, it wouldn’t make Claude prime a connection for illustration “nubilous”—an obscure1 synonym for overcast aliases grey that Claude almost surely wouldn’t usage nether normal circumstances).

How does impact Claude’s outputs?

Watermarking does not effect the value of Claude’s output. To a reader, a watermarked consequence is indistinguishable from an unwatermarked 1 (in this way, AI watermarks disagree substantially from their namesakes connected banknotes, different beingness objects, and immoderate integer documents, which are visible to the naked eye).

In soul testing, we’ve seen nary effect of watermarking connected the content, level of creativity, aliases readability of Claude’s text. In the SynthID-Text paper, which introduced the method we use, Google DeepMind tested this effect by serving a exemplary that utilized watermarking to a information of their Gemini postulation and comparing thumbs-up and thumbs-down ratings. They recovered nary statistically important differences from the unwatermarked model. And successful a controlled study, quality raters comparing watermarked and unwatermarked answers side-by-side saw nary quality successful quality.

A useful affinity is to ideate you’re playing a crippled for illustration Monopoly. On each turn, each subordinate moves a random number of spaces astir the committee according to the rotation of a die. Suppose that, alternatively of rolling the dice to get this randomness, we decided to usage a book of the digits of pi.2 We commencement from a randomly-chosen digit (say, the 1,012,845th aft the decimal place, which happens to beryllium a 6), and from that constituent connected each subordinate simply uses the adjacent digit successful the series arsenic their adjacent “roll”.

For each intents and purposes, the moves are still random: it makes nary quality to the players—or to the result of the game—whether the randomness comes from pi aliases from dice rolls each time. But if we could spot the series of each the moves aft the crippled (and we knew the worth of pi), we could activity retired whether this was a crippled that apt utilized pi to find its moves. The crippled that utilized pi is, successful a sense, “watermarked”.

It’s the aforesaid for Claude-generated text. Watermarking doesn’t alteration the meaning aliases acquisition for the personification reference it, but if you wanted to cheque aft the truth whether the matter was apt generated by Claude, the watermark allows you to do so.

Which circumstantial method of watermarking do you use?

Claude’s matter watermark is simply a type of the SynthID-Text attack published by Google DeepMind successful a Nature paper successful 2024. It belongs to a family of approaches that spell backmost to a connection by Scott Aaronson successful 2022, each of which stock the aforesaid creation rule that we described above—the watermark only changes the root of the randomness utilized to prime among words.

There are limitations to the effectiveness of watermarking. Using our key, 1 tin only reply the mobility “What is the likelihood this was partially written by Claude?” It doesn’t corroborate whether the matter was human-written, and it can’t show whether the matter was written by a different AI (even if that different AI uses watermarking, it would person a different key; it mightiness besides usage a different watermarking method altogether). Detecting a watermarking besides doesn’t activity good connected mini samples, wherever location are less connection choices and frankincense little accusation to spell on. As a transition increases successful length, assurance astir Claude’s engagement increases too.

Watermarking is sparser connected actual passages wherever location are less choices that tin beryllium made without decreasing the accuracy of the text. For example, return the condemnation “Isaac Newton’s astir celebrated activity was called Principia…”. It really matters whether the adjacent connection is “Mathematica” (it’s the only correct answer), truthful the watermark would person thing to enactment on. The aforesaid is existent for proofreading. If you manus Claude a portion of penning and inquire it to edit only the grammar and punctuation and thing else, the watermark tin only unrecorded successful the fistful of corrections, which mightiness beryllium excessively fewer to register.

What astir cases wherever Claude has proofread aliases edited quality text?

The watermark only applies to words Claude chooses. When Claude proofreads matter written by a person, what it gives backmost has mostly only been lightly edited; because astir each the words are the person’s, there’s very small (if anything) for the watermark to connect to. Depending connected the magnitude of the matter and really heavy Claude has edited it, those changes mightiness not beryllium capable to make Claude’s engagement detectable. The much Claude writes, the much decisions it has to make, and the much abstraction location is for a watermark.

What astir code? 

As we noted above, AI watermarking takes advantage of decisions wherever either prime of a connection would beryllium arsenic good. Where an exact output is required—where location isn’t a choice, and thing would beryllium factually incorrect aliases a portion of codification would break if a different word was chosen—the watermark isn’t applied.

For example, erstwhile the exemplary has written “2 + 2 =”, location is simply a very clear champion prime for the adjacent token (if the exemplary is completing the sum, location isn’t an reply that’s arsenic as bully arsenic “4”; if it’s talking astir George Orwell’s Nineteen Eighty-Four, location isn’t an reply that’s arsenic as bully arsenic “5”). The “nudge” of the watermark wouldn’t beryllium applied here. For the aforesaid reason, code—which successful very galore cases has to beryllium exact—has mostly little watermarking than immoderate different forms of text.

Having said that, successful areas wherever location is an arbitrary prime betwixt peculiar words aliases position wrong the code, the watermark tin beryllium used, specified arsenic comments wrong code. But by definition, it will person a negligible effect connected the existent codification produced.

What does this mean for users?

Does this slow the exemplary down, aliases make it much expensive?


No. Watermarking has a negligible effect connected the velocity of models, and because it produces nary other tokens, the exemplary is the aforesaid value to service and use.

Can a watermark beryllium traced backmost to maine aliases my organization?

No. The watermarking applies to Claude and its outputs. It doesn’t place thing to do pinch individual users. There’s thing successful the watermark, aliases its key, that would let anyone to retrieve immoderate accusation astir the user, their organization, aliases their chats pinch Claude.

Why are you watermarking Claude’s outputs?

We’re implementing watermarking to comply pinch the EU AI Act. Anthropic, on pinch respective different awesome AI exemplary providers and astir 190 full signatories, signed the EU Code of Practice connected Transparency of AI-Generated Content successful July 2026. This requires AI strategy providers to usage methods of “marking” AI-generated text. We’re applying watermarking globally astatine motorboat because we don't yet person a durable measurement to scope it by region. However, we will proceed to measure different approaches, and will stock updates erstwhile we person them.

Other questions

How do I cheque if a portion of matter was written by Claude?

We will soon beryllium offering a watermark discovery API. We’re successful the process of moving retired the specifications of its implementation.

What astir images and different files?

When Claude produces a record of a supported type (such arsenic a .png, .jpg, aliases .svg), it will connect a contented credential successful the shape of a small, cryptographically signed statement successful the file’s metadata, saying that the record was made aliases processed pinch Claude. This is an unfastened manufacture modular called C2PA—the aforesaid utilized by camera manufacturers and successful photo-editing package to grounds wherever an image came from. Any C2PA-aware instrumentality tin publication it; we’ll beryllium providing our ain wherever you tin driblet a record and check.

This metadata explanation is very different from a watermark. Nothing successful the record changes—it is not embedded aliases hidden. As pinch text, the credential only says Claude was progressive successful producing the file; it doesn’t see immoderate identifying information.

Can’t personification conscionable edit the matter to get astir the watermarking?

To immoderate extent, yes. Light editing astir apt won’t region the watermark completely; a complete rewrite wherever each connection is replaced will. In the second case, of course, it’s arguable whether the matter tin immoderate longer beryllium described arsenic AI-generated.

What does a watermark really prove?

A watermark tin only find that Claude was apt progressive pinch the contented astatine immoderate point. It cannot separate “Claude wrote this” from “Claude heavy edited this.”

Do watermarks use to translations?

Yes. A translator produced by Claude carries a watermark, because successful this lawsuit each connection is chosen by Claude.

What astir older Claude models?

The EU rule includes a modulation play for Anthropic models launched earlier August 2, 2026, and we’re moving to adhd watermarking for those models arsenic well. This will beryllium rolled retired complete the coming months.

How does this disagree from AI discovery software, for illustration Pangram?

AI discovery package uses a different method, because the companies that supply it don’t person our key. Among different things, those services look astatine aspects of the matter for illustration the subtle (and not-so-subtle) “tells” that often look successful AI’s phrasing. For example, AI models look to beryllium fond of the building “this isn’t [X], it’s [Y]”, and usage the connection “quietly” a batch much than you mightiness expect. Picking up connected these patterns is fundamentally different from checking for a watermark.

Does this alteration who owns a fixed output, aliases who is legally responsible for it?

No. A watermark only helps trial whether Claude mightiness person produced aliases processed the content. It doesn’t opportunity thing astir ownership aliases authorship, and doesn’t alteration a user’s authorities nether our terms. We only use the watermark erstwhile Claude was progressive successful processing the contented aliases file.

More