Human interventions on 4–8 hour agent tasks (self-reported)
1What this measures
Share of successful four-to-eight-hour agent tasks inside a frontier lab's research organisation that needed at least one human intervention. Today only OpenAI reports it, about itself, and only as 'over half'.
Why it matters. The reliability side of the return arrow. Agent-workdays can rise while humans still steer every long task; the loop only substitutes for people when this rate falls.
- Proxy types
- deployment, behaviour
- Unit
- share
- Cadence
- quarterly
- Valve
- return arrow
2How we track this
- series
openai_blog.openai.rsi_intervention_rate_4_8h.pt - source OpenAI research posts · default tier 7 · OpenAI; short quotation
- Normal band
- ≥ 50.0%
- Fast band
- ≤ 25.0%
- Falsifying
- —
Humans steering most long tasks (half or more need intervention) is the normal-technology picture; a quarter or fewer is the regime where long agent runs are autonomous. Between is emerging. The reported value is a floor ('over half'), stored as 0.5.
Applied to openai_blog.openai.rsi_intervention_rate_4_8h.pt.
3Tracker interpretation
Over half of successful long tasks still needed a person in the loop in the first half of 2026; agent effort is up, autonomy on long tasks is not yet.
4Evidence
5Status and reasoning
Same OpenAI post: over half of successful 4-8 hour tasks in the last six months involved at least one human intervention, stored as the floor 0.5. That sits at the edge of the normal band (half or more), and tier 7 caps the status at emerging regardless. Humans still steer most long agent runs. Initial seed.
6Timeline notes
- 2026-07-31 openai · 50.0%as of 2026-07-31
7Counterevidence
What cuts against this reading
Self-reported and rounded; counts successful tasks only, so failed runs that needed no intervention are excluded; 'intervention' is defined by the lab.
8Update history
- 2026-09-10unmeasured to emergingconf — → 30 · evaluate
Same OpenAI post: over half of successful 4-8 hour tasks in the last six months involved at least one human intervention, stored as the floor 0.5. That sits at the edge of the normal band (half or more), and tier 7 caps the status at emerging regardless. Humans still steer most long agent runs. Initial seed.
9Confidence
30 / 95 — limited or vague evidence
Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.
10Related
- Agent-workdays per human workday in frontier research (self-reported) emerging
- 50%/80% horizon ratio consistent with normal
- Bottlenecks #4, #23, #35, #65, #86 (Narayanan & Kapoor's list)