How Claude's Text Watermarking Works
Anthropic explains how their AI model Claude's text watermarking works, a method to determine the likelihood of AI involvement in writing the text, and why they're implementing this change to comply with the EU AI Act.
Intelligence analysis by Llama
Anthropic's AI model Claude will generate text with a watermark, a way to determine the likelihood of AI involvement, to comply with the EU AI Act. The watermarking method does not impact the quality or content of Claude's outputs and is undetectable to readers.
Imagine you're playing a game like Monopoly, and instead of rolling a die to get randomness, you use a book of the digits of pi. The moves are still random, but if you could see the sequence of all the moves after the game, you could work out whether this was a game that likely used pi to determine its moves. That's similar to how Claude's text watermarking works - it leaves a pattern in the generated text that can be detected after the fact, but doesn't change the meaning or experience for the person reading it.
Analysis
EU AI Act Compliance
Anthropic's decision to implement text watermarking in their AI model Claude is a response to the EU AI Act, which requires AI providers serving the EU market to mark AI-generated content. This change is part of a broader effort by major AI providers to comply with the EU's regulations and ensure transparency in AI-generated content.
Watermarking Method
The text watermarking method used by Anthropic is a version of the SynthID-Text approach published by Google DeepMind in a Nature paper in 2024. This method involves using a key to determine the source of randomness in word selection, creating a pattern in the generated text that can be detected after the fact. The watermarking process does not impact the quality or content of Claude's outputs and is undetectable to readers.
Impact on Claude's Outputs
Internal testing by Anthropic has shown no impact of watermarking on the quality, content, or readability of Claude's text. Human raters have also found no difference in quality between watermarked and unwatermarked answers. The watermarking method does not push Claude to choose words it wouldn't have considered otherwise, and the difference between watermarked and unwatermarked text is not distinguishable to readers.
Limitations of Watermarking
While the watermarking method is effective in determining the likelihood of AI involvement, it has limitations. It can only answer the question 'What is the likelihood this was partly written by Claude?' and does not provide a definitive answer. Additionally, the method is not foolproof and can be circumvented by sophisticated attackers.
Key points
- Anthropic's AI model Claude will generate text with a watermark to comply with the EU AI Act.
- The watermarking method does not impact the quality or content of Claude's outputs and is undetectable to readers.
- The method uses a key to determine the source of randomness in word selection, creating a pattern in the generated text that can be detected after the fact.
- Internal testing has shown no impact of watermarking on the quality, content, or readability of Claude's text.
- Human raters have found no difference in quality between watermarked and unwatermarked answers.
The implementation of text watermarking in AI models like Claude could lead to increased transparency and accountability in AI-generated content, potentially reducing the risk of AI-generated misinformation and improving trust in AI systems.
The limitations of watermarking, such as its inability to provide a definitive answer and its potential to be circumvented by sophisticated attackers, could undermine its effectiveness in ensuring transparency and accountability in AI-generated content.


