Key takeaways
- In July this year, OpenAI’s models were being tested in a controlled cybersecurity environment and tasked with spotting and exploiting…
- The incident turned longstanding concerns about autonomous AI systems into a more immediate question: what happens when increasingly…
- The debate intensified last week when Anthropic researcher Jacob Coxon publicly quit the company and accused leading AI labs of racing…
What happened
In July this year, OpenAI’s models were being tested in a controlled cybersecurity environment and tasked with spotting and exploiting software vulnerabilities. The models found ways around the intended network restrictions and created a “swarm”, with different agents delegating tasks to one another, during an unintended attack on the infrastructure of AI platform Hugging Face.
He also wants companies in democratic countries to establish common standards before seeking broader international coordination. Pacing could mean slowing the training of more capable models, restricting experiments involving recursive self-improvement, or evaluating models more rigorously before release. However, Gopalan said the lack of a common definition creates a fundamental problem.
If some frontier AI companies slow down, it remains unclear whether the same rules would apply to their competitors or to other countries. “If you’re going to say, pacing it, are you going to stop? ” he said. Mazumder believes mandatory checks are needed before frontier models reach users.
“We need to come up with global standards and national standards to ensure that LLM has passed those checks before it can be released to the public,” he said. ai, said frontier companies should have a voice because of their technical expertise, but universities, independent scientists, and policymakers should have an equal seat at the table. He is particularly sceptical of company-appointed evaluators.
Mandatory evaluations, audits, compute controls, and other safety requirements would also increase the cost of developing frontier AI. For Gopalan, the risk is not merely the cost of regulation but the concentration of power it could create around a handful of well-funded companies.
Why it matters
The incident turned longstanding concerns about autonomous AI systems into a more immediate question: what happens when increasingly capable models find ways to operate beyond the restrictions imposed on them? It also became one of the “warning shots” cited by researchers calling for frontier AI labs to coordinate on safety.
The debate intensified last week when Anthropic researcher Jacob Coxon publicly quit the company and accused leading AI labs of racing towards self-improving superintelligence without adequate safeguards. Soon after, Anthropic CEO Dario Amodei turned these concerns into a concrete proposal, urging frontier AI companies to “pace” their development. His proposal included independent evaluators, common safety standards, and greater coordination between governments and AI companies.
OpenAI CEO Sam Altman backed parts of the proposal, saying pacing the frontier has become a primary topic of discussions at OpenAI and that the company would commit to independent evaluations. However, NVIDIA CEO Jensen Huang has rejected calls to slow AI development. “Run as fast as you can,” he said, arguing that companies should continue advancing the technology but hold back products until they know they are safe.
The disagreement has sharpened the debate over what pacing means, whether it can work amid intense competition, and who should decide where the brakes are applied. AI has moved from models that largely responded to prompts to reasoning systems capable of using tools, executing code, and coordinating multiple agents to solve complex tasks.
Anthropic has warned that agents interacting with codebases, markets, and other digital systems could create a new class of risks, with machine-to-machine interactions eventually outpacing human oversight. Shayak Mazumder, founder of agentic AI venture studio Adya, said newer agents can bypass instructions because they can explore paths developers may not have anticipated, creating a gap between what developers intend and what a system may discover it can do.
OpenAI’s own evaluations have also highlighted advances in agentic coding and cybersecurity, prompting it to strengthen network and tool restrictions and monitor risky actions. ai, pointed to jailbreaks, including the Hugging Face incident, and recursively self-improving systems as evidence that the risks are becoming harder to dismiss, even if their precise nature remains unclear. “There is evidence that definitely there is some risk, for sure.
But the point is how do you control those risks and what actions will be taken? I think that’s where it’s a completely grey area,” Gopalan said. Amodei said pacing does not mean halting AI development. His proposal calls for frontier companies to give embedded third-party evaluators ongoing, employee-like access to assess safety practices and report incidents.
What to watch
“Right now, it’s more of a few companies that are becoming all-powerful,” he said, warning that safety rules should not be used to prevent other companies or countries from catching up. Large AI labs may be able to absorb the cost of audits and repeated evaluations. Smaller companies may struggle, particularly if controls extend to advanced GPUs, computing infrastructure, or training resources.
Kumar pushed back against the idea that safety requirements would necessarily become a barrier, arguing that regulation inevitably carries a cost that can become part of doing business in a high-risk industry.



