Skip to content
Slow Variables

Ask the data

Answers come from the same store as the site; every number is checked against the record it cites.

Return arrow · Model

Agent-workdays per human workday in frontier research (self-reported)

emerging30/95 confidence, limited or vaguegrade Dleading

1What this measures

Agent effort used by a frontier lab's research organisation, in standard eight-hour workdays, per workday of human labour. Today only OpenAI reports it, about itself.

Why it matters. The return arrow made numeric. If deployment inside the lab is feeding invention at more than one agent-day per human-day, the loop Narayanan and Kapoor draw but do not instrument is running; whether it substitutes for external bottlenecks is the thesis question.

Proxy types
deployment, behaviour
Unit
ratio
Cadence
quarterly
Valve
return arrow

2How we track this

  • series openai_blog.openai.rsi_agent_workdays_per_human.pt
  • source OpenAI research posts · default tier 7 · OpenAI; short quotation
Normal band
≤ 1.00×
Fast band
≥ 3.00×
Falsifying

Below one agent-day per human-day, agents are tools inside a human-run process (normal); at three or more, most research effort is agent effort, the regime the AI 2027 'R&D multiplier' assumes (fast). Between is emerging. Tier 7 evidence caps the status at emerging whatever the number says.

Applied to openai_blog.openai.rsi_agent_workdays_per_human.pt.

3Tracker interpretation

OpenAI says 3.1 as of mid-August 2026, having crossed 1.0 in June; Anthropic says its overall progress multiplier is below 2x. Lab statements about themselves, not outcome measures; the number to watch is an independent replication, which METR is funded to build.

4Evidence

Latest point
3.10×as of 2026-08-15
openai
Value the bands apply to
3.10×as of 2026-08-15

1 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.

5Status and reasoning

emergingsince 2026-09-10 · evaluate

OpenAI's 'Research acceleration' post (6 Sep 2026, via the Internet Archive snapshot because openai.com blocks the fetcher): 3.1 agent-workdays per human workday as of mid-August 2026, having crossed 1.0 in June. Above the fast band (3 or more), but a lab's statement about itself is tier 7 and caps the status at emerging. METR's modelling note puts Anthropic's researcher uplift from coding agents at over 2x; Anthropic's own system card says overall R&D uplift is well short of a doubling. No independent measurement exists yet. Initial seed.

6Timeline notes

  • 2026-08-15 openai · 3.10×as of 2026-08-15

7Counterevidence

What cuts against this reading

Self-measured and self-published by the lab whose timeline it supports; 'effort' counts agent time, not delivered research; the same post says over half of successful four-to-eight-hour tasks still needed human intervention. Anthropic's September 2026 system card reports staff self-estimating about 4x productivity uplift while the company puts its overall progress multiplier below 2x and observes no sustained AI-attributable 2x acceleration. METR's technical-worker survey finds a self-reported 1.4–2x value uplift, and METR's own RCT measured developers slower with AI.

8Update history

  1. 2026-09-10unmeasured to emergingconf 30 · evaluate

    OpenAI's 'Research acceleration' post (6 Sep 2026, via the Internet Archive snapshot because openai.com blocks the fetcher): 3.1 agent-workdays per human workday as of mid-August 2026, having crossed 1.0 in June. Above the fast band (3 or more), but a lab's statement about itself is tier 7 and caps the status at emerging. METR's modelling note puts Anthropic's researcher uplift from coding agents at over 2x; Anthropic's own system card says overall R&D uplift is well short of a doubling. No independent measurement exists yet. Initial seed.

9Confidence

30 / 95 — limited or vague evidence

Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.

10Related