Key takeaways

  • For companies worried about their intellectual property, that's a dealbreaker.
  • Nvidia has invested in Anthropic, reportedly plans to continue doing so, and supplies the company with hardware for model development.
  • CEO Alex Karp said at a customer event that companies are tired of being "exploited" by AI labs.

What happened

For companies worried about their intellectual property, that's a dealbreaker. Nvidia now uses Fable only for less sensitive tasks like open-source projects, according to The Information. For internal work like AI-powered supply chain monitoring, the company runs its own Nemotron models instead. "As a company, you know, we believe ZDR [Zero Data Retention] should be on by default," Justin Boitano, Nvidia's VP of Enterprise AI, told The Information.

It ranges from harmless ("use explicit user feedback in reward model training") to invasive ("upload user's coding environment and commit history to turn into rl envs"), Schulman says. "De-identification is weak," he adds, and users can be traced back "with just a small number of bits" and it doesn't protect against IP leakage. AI researcher Sarah Hooker, who previously worked at Cohere and Google DeepMind, describes a similar loophole.

" In other words, even if an AI lab doesn't use original data directly, it might be able to extract statistical patterns that get the same job done. Hooker warns companies, "If you are a company with IP you have a limited window to build your own intelligence that leverages your IP. " In a follow-up post, Schulman walked that back somewhat.

Training on user data is "exceedingly unlikely" to contribute much to frontier capability gains, he said. Those gains come primarily from scaling pretraining and reinforcement learning. User data is more useful for finding failure modes or situations that are hard to replicate with paid annotators. " That's likely a nod to AI coding tools like Codex or Cursor, where users connect their code repositories directly to the services.

" Schulman weighed in after mathematician Tristan Buckmaster leveled serious accusations against OpenAI. Buckmaster and his co-author Levent Alpöge had used AI models to make progress on the Navier-Stokes equations, uploading their drafts through OpenAI's Codex. Shortly after, OpenAI presented its own breakthrough using the same unusual solution path.

Why it matters

Nvidia has invested in Anthropic, reportedly plans to continue doing so, and supplies the company with hardware for model development. Booz Allen Hamilton, one of the earliest users of Anthropic's Mythos model, has banned employees from using Fable for work on proprietary cybersecurity software, according to The Information. " Palantir is blocking Fable deployment through its own software to customers until Anthropic grants irrevocable zero-data-retention guarantees, The Information reports.

CEO Alex Karp said at a customer event that companies are tired of being "exploited" by AI labs. Karp has been vocal about his distrust before. Palantir's stance is also self-serving, since the company wants customers running AI models through its supposedly secure platform rather than going directly to providers. Even with zero data retention, labs can learn from how their services get used.

Both OpenAI and Anthropic collect metadata and technical usage data from enterprise customers, according to The Information. OpenAI calls this data "de-identified," meaning it's stripped of information that could be traced back to individual customers. " The resulting classifications are metadata about the business data "but do not contain any of the business data itself," the company writes.

Some customers aren't sure what exactly that metadata covers, according to The Information, and don't think the current transparency is enough. John Schulman, OpenAI co-founder who briefly worked at Anthropic and now works at Thinking Machines, recently laid out the different ways AI companies can train on user data. " That last approach has a low risk of content reproduction but can still extract customer IP.

What to watch

OpenAI initially acknowledged that it could not rule out that anonymized data derived from their use of our products contributed to improving our models. " Still, the case showed how fragile the trust between business, academia, and the AI labs really is.