Key takeaways
- The context is a call from Anthropic CEO Dario Amodei to slow development at the AI frontier in a coordinated way.
- At the center is an index that sorts all development work at Anthropic onto a scale from Epoch AI, running from AL0 (no AI) to AL5 (fully…
- Epoch AI calls AL4 "AI leads," and Anthropic put that in its headline.
What happened
The context is a call from Anthropic CEO Dario Amodei to slow development at the AI frontier in a coordinated way. For that, the public needs more insight, Anthropic says. The metrics are meant to show how models get built, and they complement capability tests that measure what models can do.
So Claude leads a quarter of the work as measured by the human hours it takes. That says nothing about how many decisions Claude makes or whether it has a say in the research direction. The number helps the company two ways, of course: As a performance claim for the product, and as an argument for Amodei's call to slow down.
Anyone who accuses AI firms of using safety warnings to shape regulation in their own favor will find fresh material in how "leads" is defined. On Anthropic's most-used internal platform, about 30,000 agents run at the same time, according to the report. A real-time monitor checks every action and is meant to block things like copying model weights. 002 percent.
A second monitor flags about 100,000 logs a week after the fact, and roughly 50 reach humans. Anthropic admits it has only run this monitoring for a few months and isn't sure it catches every behavior pattern. During a sample week in July, about six percent of the compute for AI research went into safety work.
Why it matters
At the center is an index that sorts all development work at Anthropic onto a scale from Epoch AI, running from AL0 (no AI) to AL5 (fully autonomous). As of August 2026, 26 percent of the work sits at AL4, up from under one percent in February. More than 90 percent reaches at least AL3. Claude hits AL5 nowhere.
Epoch AI calls AL4 "AI leads," and Anthropic put that in its headline. But anyone who ties "lead" to fully autonomy has it wrong. That's what AL5 is for. An example from Anthropic shows where AL4 sits: An engineer hands Claude a bug report, and Claude analyzes, fixes, and tests it without asking questions, but it isn't allowed to ship. A human reads the report and decides.
The task and the direction still come from the human. The difference from AL3 ("collaborates") is mainly that Claude no longer stalls when it runs into a snag. Claude did the scoring itself. Agents gathered evidence from Slack and internal documents, and another Claude model assigned the levels. Anthropic admits this "judge" could make the same mistakes as the system it's checking.
Where "collaborates" ends and "leads" begins isn't clear even to Anthropic. A cross-check by the company shows this: Employees were asked to judge how automated their own work area is. When two people rated the same area, they landed on the same level only about a third of the time. The Claude model that assigned the official scores matched the human judgment 59 percent of the time.
Almost always, in 97 percent of cases, human and model were at most one level apart. " Whether a task counts toward the 26 percent is often a matter of interpretation, it seems. Then there's what the 26 percent even measures. Anthropic listed how its employees spent their work time in July and scored each activity, with tasks that eat up a lot of time counting for more.
What to watch
Anthropic cares about this metric mainly because compute is the input that outsiders can check most easily. If the industry agreed to slow down, it would be a possible lever, for instance through voluntary commitments on the share going to safety research. The low number doesn't necessarily mean little safety work, Anthropic says.
Safety work is mostly about designing experiments, which costs researcher time but few chips, while a training run for a new model burns enormous capacity. Anthropic also didn't count work that advances safety and capabilities equally as safety. The company warns that the line to capability research is blurry, and every vendor is tempted to draw it generously. The burden of proof, it says, should sit with the developer.



