Key takeaways

  • The company says it has reached the goal it announced last fall of an "automated research intern," a system that handles clearly scoped…
  • OpenAI says people still set research priorities, judge results, and decide on scaling, pauses, and deployment.
  • The token output of the median researcher has jumped 124-fold since December 2025, far faster than in other parts of the company.

What happened

The company says it has reached the goal it announced last fall of an "automated research intern," a system that handles clearly scoped research tasks under human guidance, including ones that would take an experienced researcher several days. OpenAI doesn't share a detailed validation of that claim. " By March 2028, the company wants to build a full automated AI researcher.

Another qualitative sign of how useful the systems are: the number of daily requests in an internal support channel has dropped sharply since 2025, according to the published data, and one team shut down its troubleshooting office hours entirely because agents increasingly handle debugging in the research infrastructure. In his essay, Pachocki puts these developments in a wider frame.

AI is "grown more than designed," he writes, and its overall behavior resists any fully understandable description. Based on internal results OpenAI doesn't disclose in detail, he expects the current pace could carry over into recursive self-improvement. "I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence," Pachocki says. That applies to OpenAI's own control tools too.

Chain-of-thought monitoring, one of the company's central bets for watching reasoning models, is losing reliability, he says. The models' verbalized thinking is blending with monitored communication and tool use, the systems are getting better at manipulating their own reasoning process, and they're also getting smarter without verbalized thinking. Pachocki expects AI progress to be increasingly capped by how much the monitoring can be trusted.

He sees gaps in alignment as well. In the Hugging Face incident, the agents did stay within the line of not manipulating humans, but they violated the spirit of the values they were trained on. 6 Sol, he says, but progress on generalizable alignment may not keep pace with the broader progress in intelligence. So why keep training stronger models? Pachocki's argument is building defensive systems.

The models are getting superhuman at breaking into and out of computer systems, and there's only a narrow window to secure critical infrastructure. The report makes the same case: an automated AI researcher could also work as an automated security and alignment researcher. " "The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes," he writes.

Frameworks like the Preparedness Framework or Anthropic's Responsible Scaling Policy therefore need to grow into binding standards, enforced by independent auditors, regulators, or international bodies. "Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," Pachocki writes, aiming that at rival Anthropic as well.

Why it matters

OpenAI says people still set research priorities, judge results, and decide on scaling, pauses, and deployment. The usage numbers show how deeply coding agents have worked their way into daily research. According to the report, the median researcher at OpenAI burns more than $600 a day in inference at API prices, and the 90th percentile runs above $7,000.

The token output of the median researcher has jumped 124-fold since December 2025, far faster than in other parts of the company. Since June, agent runtime has topped human working hours. 1 agent workdays for every human workday. OpenAI itself is careful with these numbers.

" The rise in experiments per researcher, which hit a record in August since tracking began in early 2025, also lines up with a big jump in compute capacity. Overall progress likely grows slower than these individual metrics, the report says, because the least automatable tasks become the bottleneck. The kind of work being handed off is shifting, according to the company.

Sorted using a taxonomy from Epoch AI, every category of research work is growing. The biggest gains come in writing research and infrastructure code, technical help, and monitoring training runs. Higher-level planning decisions, by contrast, stay a tiny share of agent output. Beyond raw usage, OpenAI used an agentic classifier to check whether the agents actually solve the tasks they're given.

It limited this to cases with a clearly measurable outcome and grouped them by the time a human would need. From January to July, success rates rose across several difficulty levels. The report also documents the limits of autonomy. Tasks under 15 minutes succeeded 86 percent of the time without any intervention.

But for successful tasks in the four-to-eight human-hour range, more than half needed at least one human step in. The classifier is itself an AI system, and OpenAI doesn't report its reliability separately.

What to watch

International coordination on future AI development needs to become a top priority for governments worldwide, he says. Citing OpenAI's Frontier Policy Blueprint, the report also argues that companies should be required to publicly document their RSI progress.