Did Buckmaster and Alpöge’s use of Codex contribute to OpenAI’s Navier–Stokes result? Initially, OpenAI said it had not accessed their specific user data, but added that it could not rule out “de-identified data derived from their usage of our products” having helped improve its models. This is basically saying that if you put unpublished research into the system, some of what they learn from it may subsequently improve a model available to your competitors.
Presumably OpenAI realised how serious the commercial implications of this statement were and a couple of days later it updated its blog with the much stronger and much more specific statement that Buckmaster’s Codex prompts from the previous two months “could not have influenced the system in any way, including through training.”
Other researchers have raised similar concerns about research information leaking into later models; for example, mathematician Andreas Thom asked OpenAI’s Mark Sellke whether the conversations he had with ChatGPT about his unpublished research were part of the training data, or accessible during the solving process, which fed into OpenAI’s non-sofic-groups proof. Sellke’s answer was “that did not happen”, which could be read as a no to direct access, or a no to any influence at all, including via training. On the All-In podcast, David Friedberg described a handful of occasions when his team had discussed what he regarded to be novel scientific ideas with an AI model, then later used a different account and newer model and received essentially the same idea back, concluding that the earlier conversations must have found their way into the newer models. Whilst it’s possible that a newer model might independently get to an idea that an older model could not, or that a later prompt might have contained enough information to steer the answer, it’s also possible that the models are learning a lot more from users’ interactions than we realise.
I assume that most researchers or their institutions have opted out of allowing their data to be used to train models, but opting out of training doesn’t stop OpenAI from collecting data about how its systems are being used. For example, OpenAI’s Enterprise Privacy documentation says business data may be processed by automated classifiers, including “to better understand how our services are used”. The resulting classifications, it says, are “metadata about the business data” and do not contain the business data itself. This is not particularly unusual and is the kind of data web analytics platforms routinely collect – how many people opened a page, chose one link over another or how search terms were used. However, an AI system generates much richer information about which tools have been invoked, classification of tasks, sequences of agent actions, and the internal reasoning traces while working on a problem.
OpenAI’s Services Agreement defines Customer Content as inputs and outputs, but it’s unclear how they classify visible reasoning traces and whether those count as output delivered to the user (hidden traces are controlled by OpenAI). OpenAI says it monitors reasoning traces during safety research and that they provide “valuable signals during both training and deployment,” but it’s less clear from its published terms what it can subsequently do with reasoning traces.
So, would you trust OpenAI, or any of the other model providers, with cutting-edge unpublished or commercially sensitive research, given that ownership of the steps the model takes is unclear?
[AI use: the key points of the article were dictated into ChatGPT after reading various posts about OpenAI’s Navier–Stokes solution and the change of wording in OpenAI’s blog post, and listening to the All-In podcast at the weekend to create a first draft. ChatGPT was used to help with research into OpenAI’s terms and conditions, and Claude was used for fact-checking. The text has been heavily human-edited. Putting this together made me realise that whilst OpenAI might not access your inputs or outputs, there’s an awful lot happening in the middle which is probably of much more value and whose ownership and usage is unclear in my view.]



