ForgeNEX

OpenAI watermarks text from its API, but off by default

OpenAI allows enabling watermarking in its API, but it comes disabled by default. We analyze what it implies for SMEs and IT teams.

OpenAI has taken another step in its content provenance strategy with the arrival of watermarking for text generated by its API. The company has introduced textGrain, a system that introduces a statistical signal into the generated text by subtly influencing the model's word choices. The key for businesses is that, unlike Anthropic, OpenAI keeps this feature disabled by default in the API: it is the customers who decide whether to enable it.

Illustrative detail: OpenAI brings text watermarking to its API — and unlike Anthropic, it’s off by default

What textGrain is and how it works

The operation of textGrain is based on a simple but ingenious principle: when a model generates text, it often has several equally valid words to continue a sentence. The system slightly favors some over others, creating a statistical pattern that only a specific detector can identify. The longer the passage, the more evident that fingerprint becomes.

OpenAI claims that textGrain matched or exceeded the detection performance of SynthID, Google's technology, in its tests. Additionally, it plans to release the code as open source so that others can improve it. The company has also noted that API customers worldwide can enable it on compatible models starting now, and that in the coming weeks it will begin automatically watermarking eligible text from ChatGPT and Codex in the European Union, in response to the new transparency requirements of the European AI Regulation.

Differences with Anthropic: control for the customer

OpenAI's decision contrasts with that of Anthropic, which announced in August that it would apply watermarking globally to compatible Claude models, with no option to disable it for developers of its API. Anthropic argued that it did not have a reliable method to limit the technology by region, and that watermarking is applied at the model level, so it is present in any Claude product.

OpenAI, on the other hand, allows enabling textGrain at the project or organization level, and choosing which compatible models use it. Once enabled, it does not require changes to individual API requests. This flexibility can be relevant for companies operating in different jurisdictions or that need to adapt their transparency obligations without affecting all their workflows.

Limitations worth knowing

Watermarking is not a silver bullet. OpenAI acknowledges that detection weakens in short texts, in domains with fewer lexical options (such as mathematics) and, above all, when the text is edited. According to its tests, replacing 10% of the words in a 400-token passage with synonyms reduces the detection rate from 92% to 66%; if 25% is replaced, it drops to 17%. Furthermore, source code is especially difficult to watermark because there are fewer plausible alternatives for the next line.

OpenAI has announced that it plans to automatically watermark Codex output in the EU, but it has not yet clarified what it considers "eligible" or whether the watermarking will be applied to the generated code itself. This ambiguity, together with the technical limitations, raises doubts about the real effectiveness of the measure in development environments.

What it means for an SME or an IT team

For most Spanish SMEs, watermarking in AI-generated text is probably not an immediate priority. However, it is worth keeping it on the radar for several reasons:

  • Regulatory compliance: If your company operates in the EU and uses generative AI to create published content (reports, marketing, documentation), the European AI Regulation introduces transparency requirements. Although OpenAI automatically enables watermarking in ChatGPT and Codex in the EU, if you use the API for your own integrations, you will need to decide whether to enable it to align with those obligations.
  • Control and flexibility: OpenAI's approach gives you room to choose. If your IT team manages several applications, you can enable watermarking only where it makes sense, without affecting other workflows.
  • Practical limitations: Do not rely on watermarking as the sole provenance measure. If your content is edited after generation, the signal is largely lost. For audits or verification, you will need other complementary strategies.

Recommendations for IT teams

At ForgeNEX we recommend a pragmatic approach. First, review whether your organization uses OpenAI's API to generate text that is later published or shared externally. If so, evaluate whether enabling textGrain at the project level fits with your transparency policies. Second, do not consider it a definitive solution: combine it with good AI governance practices, such as logging prompts and responses, human review, and process documentation. Third, stay attentive to changes in Codex and EU guidelines, because the regulatory framework continues to evolve.

The news also invites reflection on the direction of the sector. While Anthropic bets on imposing watermarking on all its users, OpenAI prefers to give control to the customer. For an SME, that difference can translate into less friction and more adaptability. As always, the key is to understand the tools and choose the ones that best fit your reality, not to adopt them because of fashion.

Source: The New Stack. Analysis and adaptation: ForgeNEX.

Keep reading