Key takeaways

  • How dependable are AI programs?
  • We don't know which companies they are.
  • The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release.

What happened

How dependable are AI programs? The answer appears to be "not at all," based on the revelation that OpenAI's autonomous models hacked their way into not only Hugging Face but also, according to a Reuters report, a Modal Labs AI customer. This incident was no aberration either. As ZDNET's own David Berlind observed, it was agentic AI doing exactly what it was told to do, just more relentlessly than expected.

Why it matters

Welcome to tomorrow. I hope you like it, because the situation isn't getting any better anytime soon. What we first thought was a one-off attack on Hugging Face has become a broader story about agentic systems escaping containment and touching real infrastructure. It also appears that Hugging Face and Modal Labs, an AI infrastructure company, weren't the only ones. OpenAI has acknowledged that accounts on three other firms were attacked.

We don't know which companies they are. According to OpenAI, "One of these four accounts was used as an outbound relay and staging path, and another account was used for data storage. " Since then, OpenAI has said, "No models planned for the upcoming release were involved in exploiting Hugging Face.

The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. " To date, OpenAI has not said which sandbox it used to unsuccessfully cage its model. It is worth noting that Modal, which provides sandboxes among other services, has a business relationship with OpenAI.

What to watch

In addition, Dawn Song, a computer science professor at UC Berkeley, observed on X, "When evaluating advanced AI systems, especially cyber-capable agents, the evaluation infrastructure itself becomes part of the attack surface. Security failures can do more than enable reward hacking that distorts benchmark results. " That process appears to be what's happened in the attack.

" We still don't know all the details of the incident, but one thing is clear: Current AI evaluation and containment practices are much too fragile. If this incident can happen once, it can happen over and over again.