When you build with AI tools like ChatGPT or Codex, you need to understand how regulations affect what the models produce. OpenAI’s new text watermarking system, rolling out first in the EU, changes how generated content is identified—but not how it works. Here’s what matters if you’re learning to build with these tools.
How text watermarking works
OpenAI’s system, called textGrain, embeds an invisible statistical pattern in the word choices of generated text. The watermark:
- Doesn’t change the meaning or quality of the output
- Can be detected by specialized tools to identify AI involvement
- Works better on longer texts (400+ tokens) than short ones
- Becomes weaker if the text is edited or translated
Note
Watermarking is off by default in OpenAI’s API. EU users will get it automatically for ChatGPT and Codex outputs in coming weeks.
What this means for your projects
If you’re building with OpenAI’s tools:
- EU compliance: Projects serving EU users must handle watermarked outputs from ChatGPT/Codex
- API choice: Global API users can opt into watermarking for select models
- Detection limits: Short or edited texts may not reliably show watermarks
The watermark doesn’t affect performance—benchmarks show nearly identical results with and without it:
| Benchmark | Unwatermarked | Watermarked |
|---|---|---|
| AutomationBench | 34.09% | 34.86% |
| Terminal-Bench 4.0 | 53.90% | 56.06% |
| HealthBench Professional | 64.27% | 64.60% |
Where watermarking falls short
Text watermarks have important limitations you should understand:
- Not proof of authorship: They don’t show who wrote or edited the content
- No accuracy indicator: Watermarked text can still be wrong or misleading
- Easy to obscure: Simple edits can make the watermark undetectable
Practical implications for builders
-
For EU projects:
- Assume ChatGPT/Codex outputs will soon contain watermarks
- Plan for disclosure requirements if your app generates content
-
For global API users:
- You control whether to enable watermarking (off by default)
- Consider it if your use case needs content provenance
-
For all builders:
- Watermarks won’t help verify facts or authorship
- Edited outputs may lose their watermark signal
Important
Don’t rely on watermark detection alone to verify content origin—the signal can be weak or missing in short/edited texts.
How watermarking compares to other provenance tools
OpenAI uses different methods for different media types:
| Content Type | Provenance Method | Public Access |
|---|---|---|
| Images | SynthID watermark + C2PA metadata | Yes (openai.com/verify) |
| Audio | SynthID watermark | Yes (openai.com/verify) |
| Text (EU) | textGrain watermark | Researcher access only |
What's coming next
OpenAI plans to:
- Expand watermarking to EU ChatGPT/Codex users in coming weeks
- Open-source the textGrain technology for community development
- Gradually widen detector access as reliability improves
For now, text watermark detection remains limited to approved researchers.
Building with watermarked outputs
When your project uses watermarked AI content:
- Disclose appropriately: Follow local regulations for AI-generated material
- Don’t overclaim: Watermarks prove model involvement, not content quality
- Expect edits to break detection: Users modifying outputs may erase the signal
The system works best for:
- Long-form content (400+ tokens)
- Unedited model outputs
- Use cases needing basic provenance
It struggles with:
- Short snippets
- Technical/mathematical content
- Translated or heavily rewritten text
Frequently asked questions
Will watermarking change how my AI tools perform?
No—benchmarks show nearly identical results with and without watermarking across coding, analysis, and creative tasks.
Do I need to do anything if I'm building outside the EU?
Only if you opt into watermarking via the API. The feature remains optional globally except for EU ChatGPT/Codex users.
Can users remove the watermark?
Yes—editing about 25% of words makes detection unreliable. Watermarks aren’t a strong copy protection tool.
OpenAI’s approach reflects both regulatory requirements and technical realities. As you build, focus on watermarking’s practical effects rather than treating it as a complete solution for content verification.
Read next
- What Chatham’s AI-Driven Capital Markets Tools Mean for Beginners Building Software
- What Faster Video Diffusion on TPUs Means for Beginners Building AI
Want to try all of this hands-on? Start with the free Vibe Coding 101 course.
Source
Based on OpenAI’s announcement, “Our approach to EU text provenance rules”. Written for people learning to build with these tools.