NEWS

OpenAI watermarks ChatGPT text, and synonym swaps break detection

OpenAI announced on October 5 that it is applying watermarks to text generated by ChatGPT. The technology is called textGrain.

OpenAI watermarks ChatGPT text, and synonym swaps break detection
Image: Redação iMasters

OpenAI announced on October 5 that it is applying watermarks to text generated by ChatGPT. The technology is called textGrain. The feature starts out restricted to conversations from users in the European Union.

How it works differs from what many people imagine. Instead of invisible characters or odd punctuation, however, the system inserts a statistical signal into the model's word choices.

This way, the response keeps the same quality. Even so, it carries clues about its own origin.

Where the feature is going into operation at OpenAI

In the coming weeks, watermarking will become mandatory for eligible responses from ChatGPT and Codex. That scope, however, covers only the European Union.

The company intends to use this limited reach as a test. This way, it can observe how the identifier behaves among real users.

International expansion remains without a timeline.

For API customers, the logic is different. Activation is also optional and applies worldwide, for compatible models.

Therefore, those who need to comply with local legislation can turn on the feature on their own.

Anthropic took a similar path in August. In that case, however, the watermark was rolled out as mandatory and international in recent Claude models.

OpenAI claims performance superior to SynthID

The comparison appeared in the announcement. According to the company, textGrain matched or outperformed other approaches tested.

Among them is SynthID for text, adopted by Google in Gemini.

Even so, an important caveat came with it. Good performance under ideal conditions, moreover, does not guarantee reliable detection in everyday use.

The numbers that show the method's fragility

Here are the data points that matter most to developers.

Text length weighs heavily. Passages of 200 tokens were detected in only 80% of cases.

With 400 tokens, the rate rises to 95%. Short texts, therefore, escape easily.

Mathematical content also complicates things. Detection for this type of text was substantially lower.

Editing knocks out the signal for good. In 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%.

With 25% of terms swapped, the rate plummeted to 17%.

Notice what this means. A quick pass of rewriting is enough to erase almost the entire trail.

What the watermark fails to answer

The company itself listed the limits. The presence of the signal fails to measure the human contribution to building the text.

It also avoids establishing ownership or responsibility for the content. It also says nothing about who the original human author was.

Truthfulness remains out of reach. The signal fails to assess whether the content is true, misleading, harmful, or presented in the correct context.

There is also the reverse case. The absence of the mark fails to prove that the text came from a person.

The detector remains closed to the public

The detectors were launched the same day. Access, however, requires signing up for a waitlist.

Researchers and specialized organizations can apply. Each request goes through the company's evaluation.

Public release has been left out for now. The reason cited involves the risk of misuse of the tool.

What this changes for product builders

First, treat detection as a weak signal. Any workflow that punishes based on this result will create injustice.

Second, watch out for false positives in short text. Commit messages, comments, and ticket descriptions fall well below 200 tokens.

Third, consider the API option. Markets with regulatory requirements tend to demand this feature contractually.

Also, record provenance in your own system. Internal metadata about content origin works better than statistical inference.

Finally, keep an eye on the C2PA standard for text. Declared provenance and statistical watermarking attack the problem from different angles.

Follow our profile on Instagram!

Translated from the Brazilian Portuguese original · Read the original

More from Redação iMasters
View profile →