AI company says model-level marking will apply across Claude products and cloud services worldwide, as it moves to meet EU transparency commitments. Watermark-removal programs are sure to follow. 

Throughout the world, students screamed in horror! Anthropic will be adding hidden but machine-readable watermarks to text and images produced by its latest Claude models. Anthropic will be adding this functionality to older models. No more will students be able to “write” their papers with a little help. What’s a kid to do!? 

Well, move to another model, obviously. It’s not that easy, though. The other AI companies are adding watermarking to their services too. For instance, Google is readying its own watermark scheme, SynthID, while OpenAI is integrating both SynthID and Content Credentials (C2PA), an open technical standard that enables publishers, businesses, and others to embed metadata in media for verifying its origin.

This is not to say, however, that these watermarks are proof that a document or image was made with AI. As Anthropic stated:

Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example:

  • Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source;
  • The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.

I see vibe-coded programs to remove watermarks coming in three, two, one…

So, why do it? Well, the immediate reason was to keep the European Union (EU) off Anthropic’s back. The EU’s Code of Practice on Transparency of AI-Generated Content demands that AI companies mark AI-generated and manipulated content. So deepfakes and AI-generated and manipulated text can be easily spotted. 

Rather than confining watermarks to content created in the EU, Anthropic has elected to enable watermarking globally. This is being driven by widespread demands for ways to spot AI-created or edited products, as AI content detectors have proven to be less than reliable. 

A recent academic paper,  “AI Wrote My Paper and All I Got Was This False Negative,” found commercial AI detectors were wildly inconsistent, with false positive rates between 0.05% and 68.6%, and false negative rates between 0.3% and 99.6%. As Patrick Traynor, interim chair of the University of Florida Department of Computer & Information Science & Engineering, said, “These are not reliable or robust tools to use to measure the problem. We really can’t use them to adjudicate these decisions. People’s careers are on the line here.”

True. There’s got to be a better way. 

In Anthropic’s approach, an invisible watermark is embedded directly in model output. The company said the mark will not alter the text’s meaning, quality, or readability, and should remain attached when users copy and paste the material. It may also survive some later editing,

The company is placing the system at the model level rather than building separate mechanisms into individual applications. That means marked output is expected to appear across the Claude consumer service, API, Claude Code, Claude Cowork, Claude Tag, and supported deployments on Amazon Web Services, Google Cloud, and Microsoft Foundry. 

Anthropic will use another mechanism for supported files. Images and other eligible outputs, including SVG, PNG, and JPG files, will receive signed provenance metadata based on the C2PA standard. C2PA is designed to preserve a verifiable record of which tool created or processed a file and whether its provenance information has been altered. 

Keep in mind, though, that Anthropic’s system will not provide a definitive test of authorship. Anthropic cautioned that a watermark indicates Claude processed content, not necessarily that Claude produced it from scratch. A writer who uses Claude to translate, summarize, proofread, or otherwise modify human-authored material could therefore produce a marked result. 

The reverse is also true. A missing mark will not establish that material is human-created: older Claude models are still being retrofitted, and extensive rewriting, paraphrasing, translation, short passages, or unsupported file workflows may make a mark unavailable or difficult to detect. 

Anthropic said it is developing tools that will enable users and third parties to detect the watermark and provenance metadata. It has not yet published the underlying technical details, detection performance metrics, or a timeline for extending marking support to all older Claude models. 

Thus, while watermarking will offer a useful new provenance indicator, it’s not a standalone answer to questions about AI authorship, plagiarism, or authenticity. Still, it’s better than nothing, and we’re still in the early days. Personally, I see a race starting between AI watermarks and tools that will erase and modify them. Let the hacking begin!