Codex Stored a Year of Drafts, Only OpenAI Could Investigate
For about a year, Codex received all of the drafts by Tristan Buckmaster, a professor at the Courant Institute at NYU. Alongside Levent Alpöge

For about a year, Codex received all of the drafts by Tristan Buckmaster, a professor at the Courant Institute at NYU. Alongside Levent Alpöge, a mathematician at Anthropic, he was tackling one of the Millennium Prize Problems: the Navier-Stokes equations. In September, OpenAI announced a proof for the same problem.
From there, a simple question emerged. Did the drafts sent to the tool help the company's model?
The case was analyzed by columnist Raul Mata in an opinion piece published on September 21. Moreover, his argument goes well beyond mathematics.
Codex, two questions and one answer
On Sunday, September 6, Buckmaster took part in a call with OpenAI researchers. He had just learned of the roughly 100-page proof produced by an internal model.
He then asked two questions. First, he wanted to know whether the model had accessed his sessions in Codex. Then, he asked whether that data had gone into training.
According to the professor himself, the first question was answered. The model, they said, ignores user data. The question about training, however, went unanswered on that call.
It's worth noting his honesty. Buckmaster publicly stated that he had not seen the proof, does not know the model's method, and avoids accusing anyone.
The timeline that explains the controversy
The pair's breakthrough came on August 15. The formal verification in Lean followed on the 22nd.
On September 8, OpenAI made the announcement. It involved about 10,000 agents, 88 hours of runtime, and a cost of several million dollars.
The company ceded priority on the Euler case to the pair. On the other hand, it claimed the Navier-Stokes result and stated it would forgo the prize.
On the same day, OpenAI acknowledged a possibility. Although unlikely, de-identified data derived from product use could have helped improve the models.
Mark Chen, head of research, reinforced the distinction. According to him, no human or agent looked at user data in this effort. However, the company uses feedback and de-identified data to improve ChatGPT and Codex, a common practice among LLM companies.
Codex and the two-month window in the denials
The following statements narrowed the response. However, on September 9, a spokesperson said it was impossible for prompts from the last two months to have influenced the system.
Then, on September 10, the company confirmed this following an internal investigation, including the training hypothesis. Finally, on the 13th, it stated that no input after July 3 could have influenced the model.
The technical explanation makes sense. A model with a training cutoff on July 3 stops ingesting any data after that date. Moreover, the decisive breakthrough on August 15 falls within the covered window.
Even so, the columnist points to the central detail. The project lasted about a year, and the denials cover only two months. That leaves nearly ten months unmentioned in the statements.
There's probably nothing there. Besides, another point raised, however, deals with something else.
Who can audit what left your company
A paying customer asked a question. Answering it required an investigation. And the only entity capable of that was the very company being questioned.
It defined the scope, investigated, and published the conclusion. Meanwhile, the professor had no tool to verify anything.
This setup holds for any AI vendor. The data channel exists and is disclosed. However, the tool to audit what passed through it remains solely with the vendor.
The community's reaction was strong. On September 11, 25 Fields Medalists signed a joint statement pointing to severe misalignment between AI labs and mathematicians. The text cites serious issues of attribution and plagiarism.
The same risk with much more common data
Now swap a proof draft for a margin spreadsheet. The scenario becomes familiar to many companies.
However, the study The Value of AI: Brazil 2026, by SAP with Oxford Economics, brought clear numbers. According to the survey, 82% of Brazilian companies report using third-party AI without formal approval, at least occasionally. The global average stood at 69%.
In addition, 43% cite data leaks or intellectual property exposure as a concern. IBM estimates the average cost of a breach in Brazil at R$7.19 million.
The columnist sums up the cause well. Employees paste documents into chat because the tool gets the job done, and the company has failed to offer an equivalent secure path. In other words, shadow AI signals missing architecture.
The zero-retention clause and the court precedent
The contractual protection exists and is called zero data retention. Under it, the vendor commits to stop logging inputs and outputs.
However, this clause usually comes as an addendum. Anyone who opens a regular tab is left without it.
The New York Times' case against OpenAI showed the practical effect. In May 2025, Judge Ona T. Wang ordered the preservation of output logs, including conversations deleted by users. Enterprise, Edu, and API customers with zero retention were excluded from the order.
Brazil has already put this into public contracts
Ordinance SGD/MGI No. 5.921 has been in effect since September 1. It covers about 250 agencies within SISP, the federal government's IT system.
In infrastructure and service contracts, the rules on AI use by the contracted company must appear in the Terms of Reference. Thus, the vendor's use of AI becomes a contractual clause.
Three questions to answer this week
First, where your context lives. Map the physical and contractual path between the internal document and the model.
Second, what crosses the boundary. Log what actually goes out, with your own output log.
Third, what happens if the vendor changes its mind. Price, terms, and usage policy change, and a replacement route avoids dependency.
In practice, this requires your own pipeline, a defined boundary, and audit capability. Training your own model is not on the list.
What to watch going forward with Codex
Watch for further clarifications from OpenAI about the period before July. Also follow the formal review of the proof.
In the meantime, review your team's AI contracts. The zero-retention clause and the output log determine who can answer when someone asks.
Follow our profile on Instagram!
Translated from the Brazilian Portuguese original · Read the original
Windows Zenith targets the dev who currently chooses Linux
Windows Zenith emerges as Microsoft's bet on a system built for development and local AI. The proposal circulated on September 21.






