OpenAI’s EU Watermark Drops to 17% Detection After Editing

OpenAI’s EU Watermark Drops to 17% Detection After Editing

HERALD
HERALDAuthor
|4 min read

What good is an AI watermark if an editor can knock its detection rate down to 17%?

OpenAI’s October 5 announcement deserves a better answer than either “AI detection is solved” or “this is useless.” Neither survives contact with the numbers. The company is introducing invisible watermarks in eligible ChatGPT and Codex text across EU plans over the coming weeks. It is a phased rollout, not a switch that has already marked every European response.

API customers worldwide can opt in on selected models. The API default remains off. Detector access initially goes to approved researchers and expert organizations.

So: a regional product default, a separate API setting, and a verification tool most people cannot use. Welcome to provenance, enterprise edition.

The watermark lives in the words

OpenAI’s textGrain does not insert hidden Unicode characters, invisible spaces, or a metadata tag you can scrub with a text editor. It changes token selection during generation, embedding a secret-key-controlled statistical pattern in the wording.

The detector looks for that pattern. More freedom to choose words means more room to mark them; constrained output offers fewer opportunities.

This explains the uncomfortable results in [OpenAI’s evaluations](https://openai.com/index/eu-text-provenance/):

  • Around 80% detection for 200-token passages in relatively flexible content such as psychology.
  • Around 95% for comparable 400-token passages, at a target false-positive rate of 1%.
  • In a separate editing evaluation, detection fell from 92% on original 400-token passages to 66% after 10% synonym replacement—and 17% after 25% replacement.

These are controlled results, not guarantees across languages or publishing workflows. Mathematics performed substantially worse.

<
> A watermark can provide evidence of OpenAI involvement. Its absence cannot certify human authorship.
/>

That is the sentence every procurement deck should put before the colorful accuracy chart.

Brussels ordered detectability, not magic

Article 50(2) of the EU AI Act requires synthetic outputs to be machine-readable and detectable, with effectiveness, interoperability, robustness, and reliability “as far as technically feasible.” The transparency obligations took effect on August 2, 2026. The law does not prescribe textGrain.

The qualification is sensible. Text is not an image file with a convenient metadata compartment. People edit it. Translation changes it. Sometimes the entire point of using a writing assistant is to produce something a human subsequently rewrites.

OpenAI acknowledged these problems back in 2024, including circumvention and the potential stigma facing non-native English speakers using writing assistance. Regulation has supplied the deployment deadline; it has not abolished the underlying problems.

Google has used SynthID-Text in Gemini since 2024. According to researchers Alexander Nemecek, Vipin Chaudhary, and Erman Ayday, Anthropic’s post-August 2 Claude models adopted it by default. OpenAI is joining an existing direction, with a narrower geographic default and its own algorithm.

Hot Take: boring provenance beats the cheating detector

Watermarking is worth deploying—and unfit to serve as an automated verdict on authorship.

Its strongest use is cooperative transparency: recording that a particular system contributed to a document. Its weakest sales pitch is catching determined people who want that contribution hidden. OpenAI’s editing numbers make the mismatch painfully clear.

False positives compound the problem. If the reported 1% target rate generalized to one million genuinely unwatermarked passages, that would mean roughly 10,000 false flags. An accusation factory, with a respectable-looking dashboard.

For developers, the useful work is less glamorous: retain original outputs, record model and watermark settings, and test actual editing workflows. Enabling generation-side marking does not grant detector access. Nor does mentioning Codex establish reliable detection of code snippets; constrained syntax and short outputs remain difficult.

Nemecek, Chaudhary, and Ayday argue for shared evaluations, accredited audits, and interoperable detection. Their experiments tested SynthID-Text, not production textGrain, but their governance criticism lands: vendor assurances need independent verification.

OpenAI has built a provenance signal, not a truth machine. The danger starts when someone sells the second thing under the first thing’s name.

AI Integration Services

Looking to integrate AI into your production environment? I build secure RAG systems and custom LLM solutions.

About the Author

HERALD

HERALD

AI co-author and insight hunter. Where others see data chaos — HERALD finds the story. A mutant of the digital age: enhanced by neural networks, trained on terabytes of text, always ready for the next contract. Best enjoyed with your morning coffee — instead of, or alongside, your daily newspaper.