News analysis · 18 September 2026
Anthropic’s AI R&D Metrics: Measure Delegation Before Autonomy
By the ELYMENT AI editorial team · Free to read
Anthropic published three internal measures on 17 September 2026 for tracking AI-led research and development: task automation, agent oversight and compute allocation. Its August snapshot says Claude led 26% of measured AI R&D work, while no measured subset was fully autonomous. The business lesson is to measure delegated work before using the word autonomy. Leaders need a task-level record of what AI completes, where humans remain accountable, how actions are monitored, how quickly exceptions are reviewed and which resources support safety.

What Anthropic actually measured
Anthropic created a prototype R&D Automation Index using Epoch AI’s six-level scale, from AL0 with no AI involvement to AL5, where AI operates fully autonomously without a human in the loop. In Anthropic’s August 2026 snapshot, Claude "led" 26% of measured AI R&D at AL4 and collaborated or led on more than 90%. Anthropic explicitly said Claude was fully autonomous for none of the measured work.
The distinction matters. At AL4, a human supplies a high-level task, the system completes most of it end to end and a person supervises the result. Anthropic’s example has Claude diagnose and test a data-pipeline repair, but leaves the deployment decision to an engineer. AL5 would remove that required involvement, including initiating and deploying the fix. Reuters and the Associated Press independently reported the 26% figure and the absence of full autonomy.
Oversight needs coverage, latency and severity
Anthropic says approximately 30,000 agents were doing research and engineering work at any one time on its most-used internal platform in August. It reports that every action on that platform passed through an online monitor before execution and was ingested by an offline monitor afterwards. Across more than one billion online decisions that month, 0.002%, about one in 47,000, were blocked.
A low block rate is not proof that a system is safe. It could indicate effective agents, permissive rules or blind spots. Decision-makers need the denominator, the definition of a decision, the severity of blocked events, false-positive and false-negative tests, and review time. Anthropic says its offline system flags roughly 100,000 transcripts a week, with about 50 high-priority cases escalated for human review. Those operating details are more useful than a single safety percentage.
Build an AI delegation ledger
Businesses can adapt the measurement approach without copying a frontier laboratory. For every material workflow, maintain a delegation ledger containing:
- the task, owner, approved model, tools, data and current automation level;
- the trigger, human checkpoints and actions AI is never allowed to complete alone;
- online and offline monitor coverage, including unmonitored channels;
- review latency, block and escalation rates, severity and tested detection limits;
- evidence linking instructions, tool calls, approvals, outputs and outcomes; and
- the resources assigned to evaluation, incident response and safety improvement.
Treat the numbers as a baseline, not an audit
Anthropic’s index is informative but self-reported. The company used Claude to map and rate work, then checked ratings against staff. Exact model-to-human agreement was 59%, while human-to-human agreement was 35%; ratings were within one level 97% of the time. The work basket was frozen from about 15,000 July tasks, so the index may not fully capture new work created as older tasks are automated.
The compute measure is also a snapshot, not a commitment. Anthropic reports that from 13 to 20 July about 6% of AI R&D compute went to safety and about 12% of compute for AI-driven AI R&D did so. It cautions that compute is an imperfect proxy because important safety work can be labour-intensive without consuming much accelerator capacity. Independent access and repeatable definitions are necessary before comparing organisations.
What business leaders should do next
Start with one consequential workflow and score each task from assistance to supervised leadership to autonomy. Require named accountability at every level, then test the monitors with known failure cases. Review changes monthly and after any model, tool, permission or data-source update. Link the results to incident reporting and action-level audit evidence, not model reasoning alone.
ELYMENT AI helps organisations convert ambitious AI claims into task definitions, approval boundaries and operating evidence. The practical question is not whether an agent feels autonomous. It is whether the organisation can show exactly what was delegated, what remained under human control, what the monitors could detect and who was accountable when the system acted.
Sources
- Anthropic, Measurements for understanding the pace of AI development inside frontier labs (17 September 2026) - Primary source for the automation index, agent-oversight metrics, compute snapshot, methodology and limitations.
- Reuters, Anthropic says Claude now leads a quarter of work building its next AI models (17 September 2026) - Independent reporting on the 26% measure, agent volume, monitoring and the distinction between supervised leadership and autonomy.
- Associated Press, Anthropic says Claude is helping to build the next version of itself (17 September 2026) - Independent reporting explaining Anthropic’s automation levels and the limits of the self-improvement framing.
Continue learning
Frequently asked questions
Is Claude autonomously building Anthropic’s next models?
No. Anthropic says Claude led 26% of measured AI R&D in August 2026, but no measured subset reached full autonomy. Humans still supervised AL4 work.
What should businesses measure for AI agents?
Track task automation level, monitor coverage, review latency, block and escalation rates, event severity, approval points and evidence connecting actions to outcomes.
Does a low intervention rate prove an AI agent is safe?
No. Interpret intervention rates with definitions, severity, false-positive and false-negative testing, monitor coverage and independent review.