OpenAI has outlined a phased plan for adding machine-readable provenance to generated text to meet requirements under the EU AI Act. The company emphasizes that text watermarking and detection are still early-stage technologies with meaningful limitations, so it is taking a cautious, transparent rollout approach.
What the EU AI Act requires and what OpenAI will do
The EU AI Act mandates that providers of generative AI make generated text identifiable in a machine-readable way. OpenAI described a solution that adds an invisible statistical signal to the model’s word choices; the company calls this method textGrain.
Key steps announced:
- API customers worldwide can opt in starting today to enable watermarking for selected models. Watermarking will remain off by default in the API.
- Over the coming weeks, OpenAI will add an invisible watermark to eligible ChatGPT and Codex text outputs in the European Union.
- Applications for access to the watermark detector are opening now; initial access will be granted only to approved researchers and expert organizations to help evaluate and improve the technology.
How textGrain works and its limits
textGrain embeds a hidden statistical pattern in the model’s choice of words. A detector searches for that pattern to assess whether a passage contains an OpenAI watermark. OpenAI has published a technical report about the method and said it will expand that report in the coming weeks; the company also plans to eventually release the technology as open source.
While textGrain matched or exceeded the performance of other approaches OpenAI tested, the company warns that good results in ideal conditions do not guarantee reliable detection in everyday use. Specific evaluation findings provided by OpenAI include:
- Shorter or less flexible text is harder to detect. At a target false positive rate of 1%, the detector identified watermarks in about 80% of 200-token passages versus about 95% of 400-token passages for content such as psychology. Detection rates were substantially lower for content such as mathematics, where word choice is more constrained.
- Editing weakens the watermark. In tests on 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%; replacing 25% of words reduced detection to 17%.
These limitations are part of the reason OpenAI is initially restricting detector access to approved researchers and expert organizations, who can help assess reliability and responsible use.
What a detected or non-detected watermark does and does not mean
OpenAI listed several important limitations on what conclusions can be drawn from a detection result:
- A watermark does not measure human contribution. It can indicate that an OpenAI system generated or processed part of a passage, but not how much human editing, judgment, or creativity was involved.
- A watermark does not establish ownership or responsibility. It does not determine who owns the text, whether its use was lawful, or who is responsible for it.
- A watermark does not identify the user. It does not link a passage to a person, organization, account, prompt, or conversation.
- A watermark does not verify accuracy. It does not indicate whether a passage is true, misleading, harmful, or presented in the correct context.
- The absence of a detected watermark does not prove human authorship. Text generated with OpenAI tools may be too short, edited, translated, come from an unsupported model, or predate watermarking, any of which can prevent detection.
Rollout details and a layered provenance strategy
- EU rollout: In the coming weeks OpenAI will introduce text watermarking for eligible ChatGPT and Codex users across all plans in the European Union. Watermarking will not be made a global default at launch; the regional approach is intended to allow learning from real-world feedback.
- API opt-in: Starting today, API customers worldwide can choose to enable watermarked outputs for selected models. OpenAI is also working with cloud partners to make watermarking available for model outputs accessed through partner services in the coming weeks.
- Detector access: Approved researchers and expert organizations can apply now. Access will be granted case-by-case, in line with the Code of Practice, to support evaluation and improvement of text provenance. The detector will report whether it finds an OpenAI watermark without identifying users or revealing prompts or conversations. Because of the risk of missed watermarks and false positives, OpenAI is not making the detector publicly available at launch.
OpenAI said it expects to revisit and adapt this approach as the technology, standards, and evidence evolve.
The company is combining text watermarking with its existing image and audio provenance tools in a layered approach: it adds Content Credentials to supported image outputs, conforms to C2PA, embeds invisible SynthID watermarks in supported images and audio, and provides verification through openai.com/verify and the Content Provenance API. These techniques are intended to complement each other—for example, Content Credentials can record a file’s origin and history while invisible watermarks can preserve a signal when metadata is removed.
Continued testing and goals
OpenAI will continue improving detection, studying how watermarks withstand editing and translation, and exploring ways to better distinguish AI assistance from AI authorship. The company plans to expand detector access when it believes results can be interpreted responsibly. For further technical details and user guidance, OpenAI refers readers to its technical report and help center materials.



