Anthropic has announced plans to implement invisible, machine-readable watermarks on content processed by its AI models, including Claude. This move is primarily to comply with the European Union's AI Act, which mandates that AI system providers watermark AI-generated or manipulated outputs. The law, effective for new models released after August 2, provides a grace period until December 2026 for updates to older models.
Moving forward, all new models offered globally will mark AI-generated content from day one. Text outputs will feature embedded watermarks, while other generated files will include digitally signed provenance metadata. Anthropic's approach is to watermark all processed content, even if it only involves minor editing, which goes beyond the EU's specific exemptions for assistive functions like grammar correction.
However, the effectiveness and utility of these watermarks are questioned. The text-based watermarks, which rely on biasing word choices, can be easily bypassed by pasting watermarked text into another AI system that edits it, or by using metadata editing tools for non-text content. Once Anthropic reveals how to identify the marks, removing them could become trivial.
Furthermore, the watermarks are not entirely conclusive, with Anthropic stating they provide a signal that content may have been processed by Claude, but do not guarantee it. The general public may also struggle to differentiate between lightly edited content and wholly AI-generated text, potentially leading to misinterpretations. The company acknowledges that its solution may mark content that the AI Act does not require to be labeled, such as simple proofreading or translation work.
Ars Technica reached out to Anthropic for details on detection timelines and testing results regarding false positives or negatives, and how the watermarks interact with editing exemptions, but did not immediately receive a response. The EU's transparency requirements aim to maintain the integrity of the information ecosystem and prevent misinformation, but the implementation of these watermarks raises questions about their practical application and potential to create confusion.