OpenAI announced on Monday that it will start adding an invisible watermark to text generated by ChatGPT and Codex in the European Union. This measure is intended to comply with the EU AI Act's transparency rules, which mandate that AI companies mark AI-generated content for identification by other systems.
The watermark functions by subtly influencing the model's word choices, creating a pattern that is invisible to readers but detectable by a specialised system. OpenAI stated that this watermark does not identify the user and did not meaningfully impact the models' performance during testing.
The company also released a technical report for its method, named textGrain, co-written with researchers from the University of Pennsylvania and Yale. OpenAI's tests indicate that editing, such as replacing 10% of words with synonyms, can reduce detection rates. Short passages, mathematical answers, and translated text are also noted as being more difficult to detect.
OpenAI will initially provide detector access only to approved researchers and expert organisations to evaluate reliability and responsible uses. The company cautioned that a missing watermark does not confirm human authorship, as the text could be too short, heavily edited, or originate from another company's AI.