OpenAI’s text watermark loses strength after editing
OpenAI’s textGrain rollout offers a limited provenance signal. Company tests show editing weakens detection, while independent validation remains outstanding.
OpenAI announced textGrain on October 5, offering watermarking for selected API models worldwide and planning it for eligible ChatGPT and Codex output in the European Union over the coming weeks. Detector access initially goes to approved researchers and expert organizations. Its own editing experiment illustrates the central limitation: changing words can sharply weaken the signal. [1]
The regulatory purpose is transparency. European Commission guidance says Article 50 requires machine-readable marking and detectability for covered generated content, while accounting for technical limitations. It distinguishes providers’ marking duties from publishers’ disclosure duties and recognizes exceptions, including standard editing assistance. The guidance also places source code outside the marking obligation’s scope. These distinctions matter because a provenance mechanism and a legal determination answer different questions; this article makes no finding that OpenAI’s implementation satisfies the law. [6]
textGrain operates when a model chooses tokens, the words or word fragments forming an answer. Its technical report describes linking those choices to pseudorandom values derived from a secret key and preceding text. A detector reconstructs the keyed pattern and tests for statistical dependence. It needs the passage and key, rather than access to the generating model. This creates a specific signal to search for instead of asking whether writing merely resembles typical AI output. [2]
Why it matters
The design also tries to preserve variety across answers. It groups vocabulary tokens into blocks and uses mathematical optimization to control the influence of keyed randomness. An entropy budget limits the average sampling randomness removed. The report’s distribution-preservation guarantees involve averaging over keys, however. They do not establish that every fixed key leaves each individual response unchanged. The distinction is consequential for applications seeking several different solutions from repeated prompts, rather than one plausible answer. [2]
OpenAI reports approximately 80% detection for 200-token psychology passages and 95% for 400-token passages at a target 1% false-positive rate; mathematics performed worse. Separately, synonym replacement in 400-token passages reduced detection from about 92% to 66% when 10% of words changed, and to 17% when 25% changed. These conditional results do not establish reliability across everyday editing. [1]
Error rates also need a denominator. As a hypothetical calculation, a 1% false-positive rate across 10,000 unwatermarked passages would produce about 100 flags. This is arithmetic, not a forecast of textGrain’s deployed performance. The editorial implication is that screening at scale requires scrutiny of false accusations as well as missed marks. A threshold chosen in an evaluation cannot automatically settle how persuasive a positive result will be in a different population. [1] [5]
Independent adverse evidence shows why adversarial testing matters. Indian Institute of Science researchers Saksham Rastogi and Danish Pruthi examined fixed and semantic-invariant token preference lists in an EMNLP 2024 paper. With 200,000 tokens of watermarked output, they predicted the lists with F1 scores above 0.8, then used that knowledge to push detection below 10%. Their attack tested different designs. It establishes a vulnerability in those schemes, while providing a concrete challenge that textGrain’s evaluators should investigate. [3]
Other research prevents a blanket conclusion that paraphrasing always eliminates watermarks. An ICLR 2024 study found that surviving fragments could support detection after human or machine rewriting. Following strong human paraphrasing, its watermark required about 800 tokens on average at a false-positive setting of one in 100,000. This favorable result belongs to the system studied there. Its different methods, lengths and thresholds cannot be transferred into evidence that edited textGrain passages are reliably detectable. [4]
A September 2026 Case Western Reserve University preprint adds a separate warning about constrained outputs. Testing an open SynthID-Text implementation on two open-weight models, researchers found little measured prose-quality effect. Code correctness fell by 3.1 percentage points on Llama, while the Gemma difference was small and statistically uncertain. Code detection scored 0.55 and 0.57 on a measure where 0.5 represents chance. These preliminary results concern neither textGrain nor deployed proprietary configurations, and used a raw-mean detector rather than the available trained alternative. [5]
The potential benefit remains useful but narrow. As an editorial inference, researchers or platforms could treat a surviving keyed signal as a lead when investigating content origin, alongside other evidence. The principal harm is turning that lead into a verdict. OpenAI says detection cannot measure human contribution, identify the user or verify truth, and absence cannot establish human authorship. Legitimate assisted writing and heavily edited generated writing therefore complicate any institutional decision based on the result. [1] [4] [5]
Searches for this revision did not identify independent textGrain replication; that is a search finding, not proof none exists. Reliable identification after ordinary editing remains unsubstantiated. Confidence would change with independent tests of the deployed system covering realistic revisions, translation, mixed authorship and short passages at disclosed error thresholds. The editorial standard for public benefit should go further: evidence that investigations improve without disproportionate false accusations. Successful removal or spoofing attacks against textGrain itself would strengthen the cautionary assessment. [2] [3] [5]
Evidence check: Unsubstantiated
Claim examined: OpenAI’s textGrain reliably identifies watermarked text after ordinary editing.
What supports it: The reviewed sources do not establish this broad claim. OpenAI reports specific synonym-replacement tests, rather than a representative evaluation of ordinary editing.
What challenges it: Reported detection fell to 66% after replacing 10% of words and 17% after replacing 25%. Favorable paraphrasing results for another watermark cannot validate textGrain.
What would change our view: Independent deployed-system evaluations using representative edits, multiple languages and document lengths, with disclosed detection thresholds, false-positive rates and uncertainty.
Limits of this reporting
No independent textGrain replication was identified in the searches conducted. Detector access and deployed performance were not tested. Older academic studies and the September preprint evaluate other systems. The Nature paper could not be read during this revision and was excluded. EUR-Lex access did not expose the regulation text; the readable European Commission guidance supplies regulatory context without establishing legal compliance.
Sources & evidence
- Our approach to EU text provenance rules — OpenAI. Published 2026-10-05; accessed 2026-10-06.
- textGrain: Entropy-Calibrated Watermarking for Language Model Text — OpenAI. Published 2026-10-05; accessed 2026-10-06.
- Revisiting the Robustness of Watermarking to Paraphrasing Attacks — Association for Computational Linguistics. Published 2024-11; accessed 2026-10-06.
- On the Reliability of Watermarks for Large Language Models — International Conference on Learning Representations. Published 2024; accessed 2026-10-06.
- Watermarks Without Verification: AI Text Watermarking After the EU AI Act — arXiv. Published 2026-09-09; accessed 2026-10-06.
- Transparency obligations under Article 50 of the AI Act — European Commission. Published 2026-07-24; accessed 2026-10-06.
Source reporting and our analysis are separated in the text. Editorial policy.