Anthropic defends and explains its AI watermarking feature

Aug 15, 2026 04:25 AM - 1 hour ago 2

A poster astatine Anthropic's developer conference.

A poster astatine Anthropic's developer convention successful San Francisco advertises its Claude product. Don Feria/AP Content Services for Anthropic

Anthropic is trying to calm everyone down complete its caller watermark.

The institution sparked a activity of sermon this week erstwhile it updated a support page for its Claude exemplary to uncover its plans to people AI-generated matter arsenic AI-generated. Now, aft concerns from Claude users, Anthropic has published a blog station connected Friday that defends and explains the technology. It besides outlines the tool's limitations.

The AI lab's station first makes 1 constituent clear: it is watermarking its text to comply pinch the European Union's AI Act, and its AI supplier peers — a apt motion toward OpenAI — will person to do the same.

Like different EU tech regulations, the AI Act's effects ripple beyond the continent. Anthropic said successful its station that it's "applying watermarking globally astatine motorboat because we don't yet person a durable measurement to scope it by region." Watermarks will first beryllium applied to caller Anthropic models, and past to older ones successful the coming months, the station said.

Previously, immoderate customers told Business Insider they canceled Claude subscriptions because of the caller watermark. Anthropic told Business Insider that it hasn't seen a inclination of an uptick successful cancellations since it announced the watermark.

OpenAI did not instantly respond to a petition for remark from Business Insider. A support page updated 2 weeks agone says that the institution besides plans to adhd watermarks to text.

How does Anthropic's caller watermarking instrumentality really work?

Anthropic spends hundreds of words of its station answering a basal mobility that emerged erstwhile the institution first revealed its scheme to watermark text: How do you embed impervious of AI successful a bid of words without ruining a chatbot's human-like tone?

The institution cites a 2024 Google DeepMind insubstantial that established the watermarking method Anthropic will now use to its Claude product's outputs.

Here's really it works. When Anthropic's AI models churn retired sentences, they make a bid of decisions astir which words to include. A batch of this is settled randomly — a grinning kid tin beryllium either "joyful" aliases "cheerful," truthful 1 of the words gets picked by a random number generator. Anthropic describes its watermark arsenic fundamentally a different type of random number generator — if you person a "key," you tin observe whether the matter follows Claude's subtle patterns of choice.

Anthropic hasn't released the "key" yet, though it says it will connection a instrumentality called an exertion programming interface, aliases API, that will let users and 3rd parties to cheque matter for Claude's watermark themselves.

The watermark doesn't see immoderate hidden characters, weird fonts, aliases concealed text. It's conscionable the connection choice. Anthropic says the quality won't beryllium distinguishable to readers.

Is this the extremity of confusing AI-generated matter pinch quality writing?

No. Anthropic's blog station many times points retired the limitations of the watermarking tech. Anthropic's "key" will constituent backmost to Claude's involvement, not AI's successful general. The station says, "Using our key, 1 tin only reply the mobility 'What is the likelihood this was partially written by Claude?'"

Longer passages will beryllium easier for the "key" to admit arsenic Claude-generated because the exemplary will person made much connection prime decisions. The station besides says that "factual passages," successful which the incorrect choices mightiness jeopardize accuracy, will person less marks.

That aforesaid rumor holds for AI-generated code. Anthropic writes that a watermark wouldn't beryllium applied arsenic often successful codification because it's truthful nonstop — the incorrect word mightiness mean the package can't tally correctly.

Some customers antecedently told Business Insider that they were concerned that the watermark would connote they weren't the writer of the matter aliases code. Anthropic says that the watermark appearing indicates that Claude processed the contented aliases record and doesn't alteration a user's authorities aliases ownership of the matter nether its terms.

It's an unfastened mobility really often group will cheque for AI-generated text, and whether the caller watermarks punctual a different narration pinch Claude's outputs.

"Light editing astir apt won't region the watermark completely; a complete rewrite wherever each connection is replaced will," the station says. "In the second case, of course, it's arguable whether the matter tin immoderate longer beryllium described arsenic AI-generated."

Have a tip? Contact this newsman via email astatine [email protected], aliases complete text, Signal, Telegram, aliases WhatsApp astatine 415-757-8198. Use a individual email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing accusation securely.

Read next

Stephen is simply a elder tech newsman astatine Business Insider, covering OpenAI, Anthropic and the ecosystem astir the starring artificial intelligence companies.Previously he covered exertion astatine SFGATE, and has written for The Wall Street Journal, The Information and CNBC. He studied publicity and economics astatine Northwestern University.His activity has earned an SF Press Club Investigative Reporting Award and, successful 2025, SPJ NorCal’s Excellence successful Journalism Award for Technology Reporting.Stephen lives successful San Francisco. Contact him via email at [email protected], aliases connected Signal, Telegram, aliases WhatsApp astatine 415-757-8198. Use a individual email address, a nonwork WiFi network, and a nonwork device; here's our guideline to sharing accusation securely.

More