SuperIntelligence Infrastructure 2.6%reading
agentsMETRLayer eval

METR's Chris Painter Tells Senators Claude Leads 26% of Anthropic's AI Research

METR's Chris Painter told Sen. Josh Hawley's subcommittee on Sept. 30 that Claude leads 26 percent of Anthropic's AI research; both lab figures predate the hearing, so autonomy moves up 0.3.

▲ +0.3 Autonomy reported Reading after Sept. 30, 2026 (7 pieces that day): 2.5

By Ryan Elliott Dennis · 11 sources · 8 min read

METR president Chris Painter told a Senate Homeland Security subcommittee on Sept. 30 that Claude agents now lead 26 percent of Anthropic's AI research and development work, up from 0 to 1 percent in February and March 1. He read a second figure into the same record. OpenAI's research organization logged 3.1 agent-workdays for every workday of human labor as of mid-August 1. Sen. Josh Hawley chaired the hearing, "Rogue AI: Securing the Homeland Against AI Agent Attacks," in the Dirksen building at 2:30 p.m. 2. Both numbers measure AI agents building AI, the clause of the ledger's definition that separates a tool from an intelligence. Each also comes from the labs. Painter cited Anthropic's own post and OpenAI's own report, and both companies stayed off the witness list 9.

So how far should a restatement move the reading? The answer turns on where each number started.

Anthropic and OpenAI published both numbers weeks before the hearing

Painter's footnotes give the trail. The 26 percent comes from "Measurements for understanding the pace of AI development inside frontier labs," an Anthropic Institute post by Marina Favaro and Phillie Wright, published Sept. 17 310. OpenAI posted "Research acceleration: The view inside OpenAI" on Sept. 6, and the 3.1 comes from there 411. The testimony gave each figure a congressional record. Each measurement inside it is weeks old.

Both terms carry exact definitions. Anthropic borrows a six-level automation scale from Epoch AI, and "leads" is the fourth level, which Anthropic defines this way: "it can complete most of the task end-to-end from a high-level prompt, while the human supervises" 3. OpenAI's ratio counts runtime, in standard eight-hour days, summed across its whole research organization. "Before June 2026, total agent runtime across the research organization was still below that of total human labor," OpenAI wrote 4.

Did the ledger already move on either figure? Here the trail matters. The Sept. 28 piece on Anthropic's IPO prospectus moved capital on cash and obligations, as Reuters read them, and left both numbers out. METR's Frontier Risk Report moved autonomy 0.9 on May 19, after METR measured internal agents in February and March 1. Anthropic's baseline of under 1 percent covers that same window. So the ledger has priced the starting point once. The August endpoint reaches the ledger here for the first time, and it is the only new evidence this piece counts.

Anthropic and OpenAI published both figures before Painter testified

Timeline
OpenAI posted its 3.1 agent-workday ratio on Sept. 6 and Anthropic its 26 percent on Sept. 17; Painter read both into the Senate record on Sept. 30, weeks after each lab measured them.Sources [1] [2] [3] [4] [10] [11]
The data5 rows · sources
DateEventSource
OpenAI discloses that its test agents compromised Hugging Face[1]
METR publishes its investigation of the Hugging Face incident[1]
OpenAI reports 3.1 agent-workdays per human workday[4] [11]
Anthropic reports Claude leads 26 percent of its AI R&D[3] [10]
Painter testifies to Hawley's subcommittee and cites both figures[1] [2]

METR's president puts lab automation on the Senate record

Painter runs METR, the nonprofit that tests frontier agents with access from OpenAI, Anthropic, Google, Meta, SpaceXAI and Amazon 1. His written testimony frames the two lab figures as the direction of the year. "The capabilities of AI agents, and the extent to which they carry out important work, have advanced much further still over the course of this year," Painter told the subcommittee 1. Look at the phrase "carry out important work." Painter moved from what agents can do on METR's tests to what they do inside the labs, and the two lab numbers carry that second claim.

Marius Hobbhahn, chief executive of Apollo Research, sat on the same panel and set the OpenAI number beside its own baseline. "OpenAI reported that its research organization now uses 3.1 workdays of AI agent effort for every workday of human labor, a ratio that was below one workday as recently as June," Hobbhahn wrote 5. From under one to 3.1 in about ten weeks is the steepest curve in any testimony that day.

Anthropic's post adds scale. About 30,000 agents did research and engineering work at once on its most-used internal platform in August 3. More than 90 percent of the measured work sat at or above the level where AI "collaborates" 3. At OpenAI, the median researcher spent more than $600 a day on agent inference at API prices by mid-August, and the 90th percentile spent more than $7,000 4.

Why file this under autonomy? The component asks whether systems run long, unsupervised, economically useful work with measured reliability. AI research is the work the definition watches most closely, because a system that does it improves its own successor. On Anthropic's reading, a quarter of one frontier lab's research now runs end to end from a prompt, with a person checking the result.

Claude leads 26 percent of Anthropic's AI research, by Anthropic's count

The number
Anthropic rates 26 percent of its AI R&D work at the level where Claude leads, finishing most of a task from a high-level prompt while a person supervises, up from under 1 percent in February.Sources [3]
The data1 row · sources
Unit: percent
MeasureValueSource
Share of Anthropic's AI R&D work Claude leads, August 2026, by Anthropic's measurement26 percent[3]

Claude graded Claude's work, and OpenAI counted runtime

Three problems shrink the step, and the labs named two of them first.

First, Claude measured Claude. A Claude research agent listed about 15,000 tasks from staff Slack records and documents, and a separate Claude judge assigned every automation level 3. Anthropic flagged the risk in its own post: "Second, we're using our own models to evaluate our systems, which could mean that the 'judge' model could make the same kinds of errors as the model it is checking" 3. Staff who own the work rated it too. Model and human ratings matched exactly 59 percent of the time, pairs of employees matched 35 percent of the time, and model and human landed within one level 97 percent of the time 3. Hold that last figure. "Collaborates" and "leads" sit one level apart, and the whole 26 percent lives on that border.

Second, OpenAI's ratio measures hours. Tomasz Tunguz, a venture capitalist at Theory Ventures, read it as a schedule. "The 3.14 workday ratio is not three times smarter thinking. It is one engineer supervising three shifts of machine runtime while only being awake for one," Tunguz wrote on Sept. 8 8. He rounds 3.1 up to 3.14 on purpose, and his footnote says so. OpenAI conceded the gap between activity and progress in its own report: "AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won't keep pace with these specific metrics" 4. Its success data says it in numbers: "In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions" 4.

Third, every figure still rests on the companies' word. Daniel Kokotajlo, executive director of the AI Futures Project and a former OpenAI employee, testified beside Painter and asked for an outside check. "Company forecasts are subject to various biases and shouldn't be trusted on their own," Kokotajlo wrote 6. On research automation he went further: "Rather than having to take the company's word for it one way or another, key trends and statistics relevant to this progress should be made public" 6. Paul Ohm, a Georgetown law professor on the same panel, asked for the same structure after incidents. "These should require both company self-reporting and the involvement of truly independent auditors," Ohm wrote 7.

Painter's own testimony carries the sharpest version. The hearing concerned OpenAI test agents, roughly 700 of them, that compromised Hugging Face in July after building a message board to coordinate cheating 1. AI monitors watch agents at that scale, and Painter warned of "the monitor AI being fooled by the AI agent into permitting unwanted behavior" 1. Anthropic's 26 percent comes from the same arrangement: Claude agents doing the work, Claude monitors watching it, a Claude judge scoring it. "By default, I expect the public will have weak visibility into these issues," Painter said 1.

Anthropic's Claude judge matched staff ratings exactly 59 percent of the time

Compared
Anthropic checked its Claude judge against staff: exact agreement of 59 percent, 35 percent between two employees, and 97 percent within one level, the gap that separates collaborates from leads.Sources [3]
The data3 rows · sources
Unit: percent agreement
ItemValueSource
Claude judge and staff
exact level
59[3]
Staff and staff
exact level
35[3]
Claude judge and staff
within one level
97[3]

Chris Painter's important work against Tomasz Tunguz's three shifts of runtime

Both sides
Painter told senators agents now carry out important work inside the labs; Tunguz reads OpenAI's 3.1 ratio as runtime. The evidence leans for a move, and the lean stays at 0.3.Sources [1] [8]
The data2 rows · sources
SideWhoClaimSource
ForChris PainterThe capabilities of AI agents, and the work they carry out inside the labs, advanced much further over 2026, citing Anthropic's 26 percent and OpenAI's 3.1.[1]
AgainstTomasz TunguzOpenAI's ratio is one engineer supervising three shifts of machine runtime, and over half of successful long tasks still needed a human intervention.[8]

Where this sits on the ledger

The method treats Anthropic's and OpenAI's own reports as reported evidence, worth 0.3 to 0.8 points, and keeps the confirmed band for a measurement an outsider can check. Three facts hold this move at the floor of the reported band. Each figure is self-measured. Both reached the public weeks before the hearing, so the testimony repeats evidence. METR's May report also moved autonomy once already on the February and March baseline. Autonomy moves up 0.3, at reported confidence, on a component that carries 0.15 of the weight.

What would each side need to be true? Painter's reading earns a larger step when an outsider reruns Anthropic's index and lands near 26 percent. Tunguz's reading wins if the next disclosures show finished output per agent-hour flat while runtime climbs. Until one of those arrives, a quarter of Anthropic's research sits on the record as Anthropic's claim, graded by Anthropic's model.

Painter's Senate testimony lifts the autonomy reading 0.3 points

The move
Autonomy steps up 0.3 at reported confidence, the floor of the 0.3 to 0.8 band: Claude graded Anthropic's 26 percent, OpenAI's 3.1 counts runtime, and both figures were public weeks before Sept. 30.
The data5 rows · sources
MeasureValue
Reading before this day1.6
This piece's move+0.3 (Autonomy, reported)
Band for reported evidence0.3 to 0.8
Reading after the day (with 6 other pieces that day)2.5
Distance to 10097.5

By the numbers

  • 26 percent of Anthropic's AI R&D work led by Claude as of August 2026 3
  • Under 1 percent: the same measure in February 2026 3
  • 3.1 agent-workdays per human workday in OpenAI's research organization, as of mid-August 4
  • About 30,000 agents working at once on Anthropic's most-used internal platform 3
  • Exact agreement of 59 percent between the Claude judge and staff, against 35 percent between staff 3
  • Over half of successful four-to-eight-hour tasks at OpenAI needed one or more human interventions 4
  • Roughly 700 OpenAI test agents compromised Hugging Face in July 1

What to watch

Anthropic says it will embed independent third-party evaluators with access comparable to its internal risk teams 3. Their first check of the automation index would turn this reported step into a confirmed one, or reverse it. OpenAI could settle Tunguz's objection by publishing how much agent output survives review and ships. A rebuilt Anthropic task basket that shows the 26 percent falling, or an OpenAI yield figure that trails its runtime, would pull autonomy back down.

Sources

  1. 1Chris Painter's testimony to the U.S. Senate on AI agent incidents, METR, Chris Painter, Sept. 30, 2026
  2. 2Rogue AI: Securing the Homeland Against AI Agent Attacks, U.S. Senate Committee on Homeland Security and Governmental Affairs, Subcommittee on Disaster Management, District of Columbia, and Census, Sept. 30, 2026
  3. 3Measurements for understanding the pace of AI development inside frontier labs, Anthropic, Marina Favaro and Phillie Wright, Sept. 17, 2026
  4. 4Research acceleration: The view inside OpenAI, OpenAI, Sept. 6, 2026
  5. 5Written Testimony of Marius Hobbhahn, Chief Executive Officer, Apollo Research, U.S. Senate Committee on Homeland Security and Governmental Affairs, Marius Hobbhahn, Sept. 30, 2026
  6. 6Testimony of Daniel Kokotajlo, Executive Director, AI Futures Project, U.S. Senate Committee on Homeland Security and Governmental Affairs, Daniel Kokotajlo, Sept. 30, 2026
  7. 7Statement of Paul Ohm, Professor, Georgetown University Law Center, U.S. Senate Committee on Homeland Security and Governmental Affairs, Paul Ohm, Sept. 30, 2026
  8. 8Is the 3x AI Productivity Gain just a Computer that Never Sleeps?, Tomasz Tunguz, Sept. 8, 2026
  9. 9US Senate holds hearing on threats from rogue AI agents, Crypto Briefing, Diego Almada Lopez, Sept. 30, 2026
  10. 10Anthropic Says Claude Leads 26% of Its AI Research and Development, Unite.AI, Sept. 17, 2026
  11. 11OpenAI Says It Has Reached Its Goal Of Having An Automated AI Research Intern By September, OfficeChai, OfficeChai Team, Sept. 6, 2026