There is a version of this story that reads like a crackdown, and it is the version doing the rounds: ChatGPT is about to start tagging everything you write, so stop copying and pasting. That is not quite what OpenAI announced. What it announced is narrower, more interesting, and considerably easier to defeat than the headline suggests.
On October 5, OpenAI published a technical report for a system it calls textGrain. It is a statistical watermark, invisible to a reader, that the company will begin applying to ChatGPT and Codex text output in the European Union over the coming weeks. API customers worldwide get the same feature this week, except there it is off by default and you have to opt in.
The reason it exists is regulatory rather than moral. Article 50 of the EU AI Act requires providers of generative systems to make machine-generated text identifiable in a machine-readable way. OpenAI is not volunteering. It is complying, and it said so.
The short version
- What it is: textGrain, an invisible statistical signal baked into the model’s word choices. No hidden characters, no zero width spaces, no odd punctuation
- Where it applies: ChatGPT and Codex text output in the European Union, rolling out over the next few weeks
- Everywhere else: available to API customers globally from this week, switched off unless you turn it on
- Who can check: nobody yet, in practice. Detector access is limited to approved researchers and expert organizations while OpenAI evaluates it
- How well it works: about 95 percent of untouched 400 token passages get flagged. Swap a quarter of the words for synonyms and that falls to 17 percent
How textGrain actually works
The instinctive assumption is that a text watermark must be a trick with characters. Invisible Unicode. A non-breaking space where a normal one belongs. Those exist, they have been used, and they are trivially stripped by pasting into a plain text editor.
textGrain is not that. Every time a language model produces a word, it is choosing from a ranked list of plausible candidates, and usually several of them would read perfectly well. textGrain nudges those choices according to a secret pattern. Across a long enough stretch of text, the pattern becomes statistically detectable even though no individual word looks unusual.
“Our text watermarking technology, textGrain, adds an invisible statistical signal to the model’s word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark.”
OpenAI, in its announcement on October 5, 2026
The practical consequence of carrying the signal in the words themselves is that copying and pasting does nothing to it. Reformatting does nothing to it. Changing the font does nothing to it. The signal survives everything that destroys a character based watermark.
What it does not survive is rewriting, and that is where the numbers get uncomfortable for anyone hoping this settles the question of who wrote what.
The detection numbers, in order of how much they matter
OpenAI measured detection at a 1 percent false positive rate, which is the honest way to report this sort of thing. A detector that flags everything is useless. The question is how much real watermarked text it catches while keeping its wrong accusations to one in a hundred.
Three separate things degrade the signal, and they stack. Short text is harder than long text, because there are fewer word choices to carry the pattern. Constrained writing is harder than flexible writing, for the same reason. And editing is devastating, because every word you change is a word that no longer carries its part of the signal.
Look at the gap between the psychology passage and the math passage. Both are 400 tokens. Both are untouched. One is caught 94 percent of the time and the other 60 percent. The only difference is how much room the subject matter leaves for choosing between synonyms, and that turns out to be worth 34 percentage points.
That is a real limitation rather than a footnote. Technical writing, code comments, legal boilerplate and anything with a fixed vocabulary are all closer to the math case than the psychology one.
What the detector can and cannot tell you
This is the part that matters most if you ever find yourself on either side of an accusation, and it is the part most likely to be misunderstood in the coming months.
| The question | What a textGrain result actually answers |
|---|---|
| Was this written by AI? | No. It answers whether this specific passage carries an OpenAI watermark. Text from Claude, Gemini, Llama, a local model or an unwatermarked OpenAI endpoint carries nothing to find |
| No watermark found, so a human wrote it? | No. Absence proves nothing. The passage may be short, heavily edited, outside the EU rollout, or from an API account that never opted in |
| Watermark found, so AI wrote all of it? | Not necessarily. It indicates the passage contains watermarked output. It does not tell you the proportion, or whether a person then rewrote half of it |
| Can my school or employer run this on me? | Not today. Access is restricted to approved researchers and expert organizations during the evaluation phase |
| Can I beat it? | Yes, and not with much effort. Replacing one word in four already takes detection down to 17 percent. Running the text through a different model does the same job faster |
If that last row reads like a how-to, it is worth saying plainly that it is in OpenAI’s own report. The company chose to publish the weaknesses alongside the strengths, and promised to open source the technology so others can build on it. That is an unusual amount of candour for a compliance feature, and it points at the real purpose: this is provenance infrastructure, not an exam invigilator.
The distinction matters because the previous generation of AI text detectors was sold as the latter and failed badly at it, with false accusations landing hardest on people writing in a second language. If you want a sense of how quickly the folk wisdom around spotting machine output goes stale, our guide to how to spot AI generated images now that the old tricks have stopped working tracks the same arc one medium over.
Who gets watermarked, and who does not
The geography here is deliberate and a little awkward. The EU AI Act applies in the EU, so that is where the obligation bites, and that is where OpenAI is switching the watermark on by default.
| Where and how you use it | Watermark status | Notes |
|---|---|---|
| ChatGPT in the EU | On, by default | Rolling out over the coming weeks on eligible models |
| Codex in the EU | On, by default | Same rollout. Code is a constrained vocabulary, so expect lower detection in practice |
| API, anywhere in the world | Off, opt in | Available from this week on certain models. The developer decides |
| ChatGPT outside the EU | Not announced | No obligation, no stated plan. US and UK users are not in this phase |
| The detector | Restricted | Applications open to approved researchers and expert organizations only |
The asymmetry between the EU default and the global opt-out is the detail worth holding on to. A watermark that covers one market and is optional everywhere else cannot function as a general test of provenance, because the honest answer to “is this text watermarked” is usually going to be no, regardless of who wrote it.
Why the EU is driving this
None of this appears out of nowhere. The AI Act is the first comprehensive AI law anywhere, and its transparency provisions have been pushing providers toward labelling for well over a year. We covered the moment the Act became enforceable and most AI companies turned out not to be ready, and textGrain is a fairly direct product of that pressure.
In August the European Commission gained powers to inspect models being placed on the EU market. Models already on sale have until August 2, 2027 to meet the full set of obligations. The Council of the European Union has been explicit that it expects the Act to set a global standard the way GDPR did for privacy, which is part of why a US company is shipping a compliance feature for one region first and leaving the rest of the world opt-in.
So should you stop copying and pasting?
If you are in the EU, using ChatGPT, and handing in long passages of unedited output somewhere that would object, then yes, the risk profile just changed slightly. Not because anyone can test you today, but because the text now carries a signal that persists indefinitely and could be checked later by whoever eventually gets detector access.
For almost everybody else, the practical answer is that nothing changed. No detector access, no watermark outside the EU on ChatGPT, nothing on competitor models, and a signal that light editing already dismantles.
The thing worth worrying about instead. A watermark that is easy to remove and hard to interpret is a poor basis for an accusation, but it is a fine basis for an argument. The risk is not that you get caught. It is that “we ran it through a detector” becomes a sentence people accept at face value again, the way they did with the last generation of AI detectors, and a 1 percent false positive rate on a large volume of text produces a lot of wrongly accused people.
The bigger picture
Set the compliance framing aside and textGrain is an attempt at something the web badly needs: a way to know where a piece of text came from. The reason that matters is not academic integrity. It is that an awful lot of what you read now was not typed by anyone. Pew’s research found that roughly one in three web pages created since ChatGPT launched was not written by a human, which is the sort of number that makes provenance a load-bearing problem rather than a nice-to-have.
Seen that way, the weaknesses in the technical report are less damning. A watermark does not need to be unbreakable to be useful for measuring how much machine text is flowing through a platform, or for letting a publisher check its own supply chain. It needs to be breakable only in the sense that it should not be mistaken for proof.
OpenAI appears to understand that, which is why the report leads with limitations and why the detector is not being handed to universities. Whether everyone downstream understands it is a different question.
What to watch next
- Who gets detector access, and on what terms. The moment this reaches institutions rather than researchers, the false positive rate stops being a statistic and starts being somebody’s disciplinary hearing
- Whether the open source release actually lands. OpenAI said it plans to make the technology available openly. Independent reproduction of the 1 percent false positive claim is the whole ballgame
- Whether other providers follow by default rather than by region. Anthropic has taken a stricter line on API watermarking. Google has SynthID for images and has been extending the idea
- Whether paraphrasing tools advertise this. The 25 percent synonym figure is in a public technical report, which means it is in a product roadmap somewhere by now
- Codex in practice. Code has a constrained vocabulary, and the math result suggests detection on source code will be substantially worse than on prose
The bottom line
OpenAI is adding an invisible signature to ChatGPT and Codex output in the European Union because the AI Act tells it to, the signature survives copy and paste, and it catches around 95 percent of long untouched passages when someone with detector access goes looking.
It also collapses to 17 percent once a quarter of the words have been changed, scores 60 percent on text with a fixed vocabulary, is off by default for every API customer on the planet, and cannot currently be checked by anyone outside a short approved list. Both halves of that are true at once, and the headline version of this story only tells you the first one.
Sources and further reading
- The Decoder: OpenAI will watermark ChatGPT text in the EU but makes it optional for API users worldwide
- ActuIA on textGrain, the API and the AI Act
- ITdaily: OpenAI introduces text watermarks for the EU to comply with the AI Act
- Techgenyz on how 25 percent of edits cripple detection to 17 percent
- Android Headlines: OpenAI is adding invisible watermarks to ChatGPT text in the EU
- Adgully: OpenAI launches textGrain invisible text watermarking in the EU
- Nexus AI: textGrain and EU text provenance, a guide for builders and compliance teams
- UNILAD Tech: Stop copying and pasting from ChatGPT as new feature can tag everything you write
About this article: GeekBlog covers U.S. technology news, AI, phones, smartwatches and gaming. Every story is written and checked under our Editorial Policy. Spotted a mistake or have a story tip? Contact our editors.

