NEWS

OpenAI turns on watermarking in its API and leaves the decision to developers

OpenAI announced textGrain, a watermarking system for text generated by its models, on October 5. The move responds to European rules.

OpenAI turns on watermarking in its API and leaves the decision to developers
Image: Redação iMasters

OpenAI announced textGrain, a watermarking system for text generated by its models, on October 5. The move responds to European transparency rules. The announcement, however, included a detail that went largely unnoticed.

API customers can already enable the feature on selected models, in any country. Brazil is on that list.

In the API, however, the option remains off by default. The choice is therefore left to whoever integrates it.

In ChatGPT and Codex, enforcement will become mandatory in the coming weeks, only in the European Union.

How the signal enters OpenAI's text

It's worth understanding the mechanism. During generation, the model picks successive tokens based on probabilities.

textGrain adjusts that process. However, it uses values associated with a secret key and the context of the passage.

The detector, in turn, receives the text and the same key. It then checks whether the choices follow the expected pattern more often than chance.

Notice what's left out. No hidden character, invisible space, or unusual punctuation goes into the output.

The difference from common detectors also matters. Traditional tools look for similarity to writing patterns, whereas here the signal is born during generation.

OpenAI published the numbers, and they call for caution

The tests used a target false-positive rate of 1%. For content such as psychology, the detector identified marks in about 80% of 200-token passages.

With 400 tokens, the rate rose to 95%. In math, however, the result was lower, since there is less freedom in word choice.

Editing quickly breaks down the signal. In 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%.

With 25% substitution, the rate dropped to 17%. The test, however, used responses in English.

These percentages describe specific evaluations. They therefore stop short of guaranteeing performance across any text, language, or situation.

What the watermark fails to prove

In addition, the company laid out the limits clearly. Finding the signal fails to determine how much of the work came from a person.

The watermark also avoids identifying the user, revealing prompts, or establishing ownership and legal responsibility. The accuracy of the information remains out of its reach.

The reverse case deserves equal attention. The absence of the signal fails to prove human authorship, since short, rewritten, translated texts, or ones generated by unsupported models, slip through.

Because of the possible errors, the detector stays out of public reach at launch. Researchers and specialized organizations can request access subject to approval.

OpenAI responds to the AI Act and promises open source

The announcement is part of the response to the European Union's artificial intelligence regulation. The obligations require that generated content be identifiable in a machine-readable format, respecting the technical limits of each medium.

The European Commission separates two things. This labeling differs from the duties of whoever publishes or uses the material.

There is also a relevant exception. Texts that inform the public about matters of public interest are exempt from disclosure when there is human review and editorial responsibility.

The company plans to release textGrain as open source. In addition, Anthropic had already announced watermarks for text from Claude.

What to do with that switch in your product

First, assess your audience. A product with users in the European Union is likely to need the feature turned on.

Second, test the impact on output. The company claims to maintain quality, but it's still worth measuring in your own case.

Third, think about your own record-keeping. Provenance metadata in your database works better than statistical inference after the fact.

Also, be careful with short text in your workflow. Commit messages and ticket descriptions fall well below 200 tokens.

Finally, keep an eye on the open-source release. A public implementation allows the method to be audited and its real cost to be understood.

Follow our profile on Instagram!

Translated from the Brazilian Portuguese original · Read the original

More from Redação iMasters
View profile →