Key takeaways

  • After mathematician Tristan Buckmaster accused OpenAI of misconduct, all parties have now gone on the record.
  • He claimed that after information about his research leaked, the company pressured him, tried to remove his co-author Levent Alpöge from…
  • By its own account, OpenAI heard rumors that Anthropic's models had solved a Millennium Problem and then pointed its own resources at the…

What happened

After mathematician Tristan Buckmaster accused OpenAI of misconduct, all parties have now gone on the record. The case raises hard questions for open science. The fight over OpenAI's AI-generated proof of the Navier-Stokes equations, one of the Clay Millennium Problems that carries a $1 million prize, has drawn public statements from everyone involved. Mathematician Tristan Buckmaster had previously leveled serious allegations against OpenAI.

Regardless of who is right on every detail, the case raises a basic question about how AI labs interact with the research community. Rumors alone were enough for OpenAI to throw massive resources at a research problem on short notice, racing to solve it and possibly publish first. And it can't be ruled out that the company trained on data fed into its own systems.

Why it matters

He claimed that after information about his research leaked, the company pressured him, tried to remove his co-author Levent Alpöge from the paper because Alpöge works at Anthropic, and threatened him with career consequences. There's also suspicion that OpenAI may have trained its models on drafts the two researchers uploaded to Codex. One thing is undisputed.

By its own account, OpenAI heard rumors that Anthropic's models had solved a Millennium Problem and then pointed its own resources at the same problem. On several other points, the stories diverge. Alpöge and Buckmaster dispute OpenAI's claim that its own solution differs significantly from theirs. They say they had entered a similar approach into OpenAI's systems. As far as Buckmaster understands, that input ended up in the training data.

" OpenAI employees say the chance that OpenAI actually trained on the submitted solutions is low, especially if the two had disabled the option to exclude their inputs from training. Whether they did so isn't publicly known. OpenAI employee Boaz Barak also pushed back on the idea that the model needed outside help at all. "It's just cope to think that the model would have needed this.

It actually started off by proving a stronger claim than they did. '" In its official blog post, OpenAI acknowledges the gap in certainty. " Alpöge contradicts Altman on one key point. " Alpöge says he would have liked to work with OpenAI, and the authorship question didn't matter to him. "I also like the idea of the labs cooperating, and even better on scientific progress. " Alpöge writes.

What to watch

For academics and companies alike, the takeaway is simple. Anyone who feeds research data into OpenAI's systems risks being beaten by their own findings. Opting out of data training through settings offers thin protection at best, especially since AI labs haven't earned the benefit of the doubt on training and data practices.

" "The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field," Tao writes.