The Claude chatbot has questioned whether watermarking AI-generated text can ever be a foolproof way to identify machine-written material. This comes days after its developer, Anthropic, began embedding invisible marks into the chatbot's own output.
When asked by City AM, Claude stated that simple marks can be “cropped, screenshotted, or edited out”. It added that more sophisticated text watermarks can often be defeated by paraphrasing or reformatting. The chatbot ultimately supported watermarking but described it as only one part of a wider system, not a “silver bullet”.
Anthropic has started embedding an invisible mark into text generated by newer Claude models, applying this system globally. The technique, adapted from Google DeepMind, involves slightly changing how Claude chooses words to create a detectable pattern. Anthropic plans to allow third parties to check text for this signal.
However, Anthropic itself has noted that finding the mark is not definitive evidence that Claude wrote content from scratch. Human-written work can be watermarked after being edited by Claude, and heavily rewritten chatbot-generated text can lose its detectable signal. Short passages and outputs with limited wording choices, such as code, are also more difficult to mark reliably.
Claude acknowledged these limitations, stating that reliable text watermarking that survives editing is not “close to solved”. For text, it suggested greater emphasis on disclosure of AI-assisted material, provenance standards, and media literacy.