Key facts
- Anthropic is embedding an imperceptible watermark into text generated by its Claude AI models.
- The watermark is designed to remain readable but may be removed by heavy editing or translation.
- This feature is intended to enhance transparency and combat the misuse of AI-generated content.
- The watermarking aligns with requirements of the European Union AI Act.
- New Claude models will support the feature starting August 2, with plans for older models.
- Anthropic will also make tools available to allow users and third parties to check for its marks.
Anthropic is rolling out a new feature for its Claude AI models that embeds an imperceptible watermark directly into generated text. This development aims to increase transparency and make it more difficult for individuals to pass off AI-created content as their own, a growing concern in industries like publishing and education. The watermark, which does not alter the text's meaning or readability, is designed to travel with the content even when copied and pasted, and may persist through some editing. This move is part of Anthropic's commitment to transparency, particularly in light of the European Union AI Act. Models launched on or after August 2 will support this feature, with plans to extend it to older versions. Anthropic also intends to provide tools for third parties to detect these watermarks. The publishing industry has been grappling with allegations of AI-generated content. While the watermarking offers a new investigative tool for publishers and educational institutions, Anthropic acknowledges potential loopholes. Heavy editing, paraphrasing, translation, or mixing Claude's output with other writing could make the watermark undetectable. Furthermore, the presence of a watermark does not definitively prove original authorship by Claude, as even AI-assisted proofreading can leave a mark. Google DeepMind has also implemented similar watermarking technology for its models.
