Back to Discover

OpenAI shows agent research-acceleration metrics and where humans still steer

Official OpenAI card art: Research acceleration: The view inside OpenAI

On Sept. 6, 2026, OpenAI said its measurements show it reached an automated research intern goal, with 3.1 agent-workdays per human workday by mid-August. The Decoder notes no independent review; humans still set priorities and intervened on more than half of successful four-to-eight-hour tasks.

OpenAI said on Sept. 6, 2026, that its own measurements show it has reached an automated research intern goal for well-defined, human-directed tasks that can take a skilled researcher days. The Decoder reported the disclosure and said OpenAI offered no independent review or detailed external validation.

By mid-August, OpenAI reported 3.1 standardized agent-workdays of effort for every human workday across its research organization, after total agent runtime had still been below human labor before June 2026. The company said the median researcher ranked by agent usage was using more than $600 per day of inference at API prices, and the 90th-percentile research user more than $7,000 per day.

OpenAI frames AI research as a multi-step loop in which researchers design improvements, write evaluations and infrastructure, catch bugs and unsafe or misaligned behavior, and integrate winning ideas into a core training run. It classified agent tokens using Epoch AI's six-phase R&D taxonomy (Decide, Design, Build, Run, Analyze, Communicate) and said every category grew from January to August 2026, while high-level planning remained a minimal share of agent output tokens.

Using an agentic classifier, OpenAI said task-success rates generally rose from January to July across several difficulty buckets, but agents still need significant human steering as complexity rises. In the last six months, more than half of successful four-to-eight-hour tasks involved one or more interventions. The Decoder separately preserved that limit and noted that the classifier's reliability was not reported separately. People still set research priorities, judge which ideas to pursue, and decide whether to scale, pause, or deploy systems.

August 2026 experiments per active experimenter reached an all-time high since tracking began in January 2025, OpenAI said, correlated with Codex adoption and with substantially higher available compute. Its methods appendix says activity indicators are relatively easy to gather but hard to interpret because their relationship to research progress is uncertain, and that coding-agent metrics cover most but not all usage as tools evolve. OpenAI sets March 2028 as its target for an automated AI researcher.

What remains open is how outside reviewers will validate the intern milestone, whether rising runtime and experiment counts track research progress once less-automatable steps become the bottleneck, and how far human steering stays required as task horizons lengthen. OpenAI's self-reported ratios should be weighed separately from independent coverage until those points are settled.

  • @BinaryVerseAI Video

    Outsider walkthrough of OpenAI's automated research intern claim and the 3.1 agent-workday ratio (further-watching; not primary evidence).

    Watch on YouTube
  • @AgenticFrontier1 Video

    Agentic Frontier video on OpenAI's mid-August report that agents log 3.1 research workdays per human day.

    Watch on YouTube