Skip to content
Slow Variables

Ask the data

Answers come from the same store as the site; every number is checked against the record it cites.

Products · Deployment & application

Developer productivity uplift (METR RCTs)

consistent with normal70/95 confidence, good evidence, some ambiguitygrade Aleading

1What this measures

Measured speed change for experienced developers using AI tools versus not, from METR's randomised trials. Negative means slower. The band applies to the newest estimate (Feb 2026, newly recruited developers).

Why it matters. The single best-controlled test of whether frontier capability turns into product-level uplift. The paper's claim is that benchmark gains do not translate quickly; this is where that shows.

Proxy types
product, deployment
Unit
share
Cadence
annual
Valve
invention to product

2How we track this

Normal band
≤ 20.0%
Fast band
≥ 40.0%
Falsifying

Normal = uplift of at most about 20%, setting-dependent (the range of controlled studies to date); fast = broad uplift above 40%, the scale the productivity-boom claims need.

Applied to metr_blog.rct_2026.dev_speed_change.pt.

3Tracker interpretation

Two trials, two years, and the best estimate is still near zero. Capability is not the constraint; integration and verification are.

4Evidence

Latest point
-4.0%disputedas of 2026-02-24
rct_2026
Value the bands apply to
-4.0%as of 2026-02-24

2 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.

5Status and reasoning

consistent with normalsince 2026-09-10 · claude-initial-seed

METR's Feb 2026 redesign estimates a -4% speedup for newly recruited developers (CI -15% to +9%; METR calls it very weak evidence); the 2025 RCT found developers 19% slower. Both are far below the 20% ceiling of the normal band. Initial seed.

6Timeline notes

  • 2026-02-24 rct_2026 · -4.0%disputedas of 2026-02-24
  • 2025-06-30 rct_2025 · -19.0%as of 2025-06-30

7Counterevidence

What cuts against this reading

Small samples; the 2026 redesign has selection effects METR itself calls 'very weak evidence'; experienced open-source maintainers on their own repos are the hardest case for uplift.

8Update history

  1. 2026-09-10unmeasured to consistent with normalconf 70 · claude-initial-seed

    METR's Feb 2026 redesign estimates a -4% speedup for newly recruited developers (CI -15% to +9%; METR calls it very weak evidence); the 2025 RCT found developers 19% slower. Both are far below the 20% ceiling of the normal band. Initial seed.

9Confidence

70 / 95 — good evidence, some ambiguity

Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.

10Related