Skip to content
Slow Variables

Ask the data

Answers come from the same store as the site; every number is checked against the record it cites.

Predictions

One ledger, four claimant families, scored against the same evidence. Claim text is verbatim from the linked source; the status vocabulary is the AI 2027 tracker's (confirmed, ahead, on track, behind, emerging, not yet testable). A claimant's self-assessment never resolves a claim.

Narayanan & Kapoor · 5

  1. Narayanan & Kapoor

    2025-04-15window to 2035-04-15on trackconf 70
    In Part I, we explain why we think that transformative economic and societal impacts will be slow (on the timescale of decades), making a critical distinction between AI methods, AI applications, and AI adoption, arguing that the three happen at different timescales.source

    How we track it. Adoption and adaptation indicators (BTOS firm use, BBD hours assisted, BLS productivity and TFP, labour trackers) stay inside their normal bands; a breach of the fast bands on two adoption indicators and one adaptation indicator for four consecutive readings would count against it.

    Indicators: US firms using AI (Census BTOS)Share of work hours assisted by generative AIUS labour productivity, year on yearUS total factor productivity, private nonfarm business

    Counterevidence and history

    Weekly work use of generative AI reached 39% within three years, faster than the PC or the web at the same age; the claim is about impacts, not use, which is what makes it hard to falsify early.

    1. 2026-09-10unmeasured to on trackconf 70 · claude-initial-seed

      Every adoption and adaptation indicator with a status sits inside its normal band: Census BTOS firm use 22.4% (normal under a third), BBD hours assisted 6.3% (normal single digits), BLS labour productivity 2.2% and private nonfarm TFP 0.8% (both at trend). The only reading outside a normal band is Ramp's paid-adoption share (56%, emerging on a tech-forward sample). No fast-band breach anywhere; the falsifying pattern (two adoption plus one adaptation indicator in the fast band for four readings) is nowhere near. Initial seed.

  2. Narayanan & Kapoor

    2025-04-15window to 2030-04-15not yet testableconf 30
    Thus, we predict that slow diffusion will continue to be the norm in high-consequence tasks.source

    How we track it. FDA AI device clearances, Waymo paid rides, and safety-brake events (lab pauses, GPU reallocation) as evidence that high-consequence deployment is gated; a rapid, ungated deployment in a regulated domain would count against it.

    Counterevidence and history

    Waymo scaled to hundreds of thousands of paid rides a week; medical-device clearances now exceed a thousand; the safety brakes observed in 2026 were inside labs, not in deployment.

    1. 2026-08-28supportsgrade D, evidence tier 7source

      OpenAI restarted the paused frontier RL run only after new safety and security requirements were in place (28 Aug 2026)

      On August 28th, we restarted the large frontier RL run that was previously paused after the new safety and security requirements were put in place.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      No safety-gating series is in the ledger yet: FDA AI device clearances, Waymo paid rides and lab pause events are planned inputs. Not scored until an instrument exists. Initial seed.

  3. Narayanan & Kapoor

    2025-09-09window to 2030-09-09emergingconf 40
    All this is not just true of today’s AI, but even in the face of hypothetical developments such as self-improvement in AI capabilities. Many of the limits to the power of AI systems are (and should be) external to those systems, so that they cannot be overcome simply by having AI go off and improve its own technical design.source

    How we track it. Independently verified agent-workdays per human-workday above 1 with an intervention rate on 4–8h tasks below 50% (OpenAI's self-reported 3.1 and >50% do not yet qualify) would test it on the invention stage; the diffusion indicators test it downstream.

    Indicators: METR 50% time horizon50%/80% horizon ratio

    Counterevidence and history

    OpenAI's September 2026 research-acceleration disclosure reports more agent than human workdays inside its research org; if independent evaluation confirms it and intervention rates fall, the invention-stage version of this claim fails first.

    1. 2026-09-10unmeasured to emergingconf 40 · claude-initial-seed

      OpenAI's self-reported 3.1 agent-workdays per human workday with more than half of 4-8 hour tasks needing intervention (6 Sep 2026) is tier 7 and does not qualify as the independent RSI series the operationalisation requires; the thesis monitor's invention-side warning is untestable for the same reason. Downstream, the diffusion indicators sit in their normal bands. Initial seed.

  4. Narayanan & Kapoor

    2025-04-15window to 2030-04-15on trackconf 65
    Every time we solve a benchmark (reach what we thought was the peak), we discover limitations of the benchmark (realize that we’re on a ‘false summit’) and construct a new benchmark (set our sights on what we now think is the summit). This leads to accusations of ‘moving the goalposts’, but this is what we should expect given the intrinsic challenges of benchmarking.source

    How we track it. The 50%/80% horizon ratio staying wide, the developer RCT uplift staying small while benchmark horizons race, and pilot-to-production rates staying low are the product-side evidence; a collapse of the ratio toward 2x with uplift above 40% would count against it.

    Indicators: 50%/80% horizon ratioDeveloper productivity uplift (METR RCTs)Enterprise pilots with measurable P&L impact

    Counterevidence and history

    METR's 80% horizon is now measured in hours, not minutes, and doubles on the same clock as the 50% horizon; the gap is stable in ratio terms but the absolute reliable horizon is growing fast.

    1. 2026-09-10unmeasured to on trackconf 65 · claude-initial-seed

      The 80%/50% horizon ratio is 6.3 (normal is 5 or wider), the 2026 METR RCT found developers 4% slower while the 50% horizon doubles every 108 days, and 5% of enterprise pilots reach production with measurable P&L impact (normal under 10%). All three product-side readings are what the claim predicts; the counter-pattern (ratio toward 2 with uplift above 40%) is absent. Initial seed.

  5. Narayanan & Kapoor

    2026-07-13window to 2031-07-13not yet testableconf 30
    There is no particular capability milestone that will unlock all of this economic potential. The economic potential is already there. It really depends on all these downstream actions that we take. It is not gated by capability.source

    How we track it. Not testable as a point prediction; the tracker reads it through the concordance of adoption and adaptation indicators moving at organisational speed regardless of capability releases.

    Indicators: Cross-tracker concordance (labour)Share of work hours assisted by generative AI

    Counterevidence and history

    A single lab release that visibly moves adoption or productivity series within two quarters would be evidence of capability gating.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Not testable as a point prediction, as the operationalisation says. Read through cross-tracker concordance, which is 0.0 (all four labour trackers agree on no broad employment effect), consistent with the claim but not a test of it. Initial seed.

Lab timelines · 6

  1. OpenAI

    2025-10-28window to 2026-09-30emergingconf 30
    According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.source

    How we track it. Resolves only on an independent replication of OpenAI's agent-workday ratio and intervention rate (METR's RSI programme or equivalent), not on OpenAI's own measurements. The lab's self-grade is tier 7 and caps the status at emerging.

    Indicators: METR 80% time horizonDeveloper productivity uplift (METR RCTs)

    Counterevidence and history

    The claim is scored by OpenAI's own measurements. The same post reports that over half of successful four-to-eight-hour tasks in the last six months needed at least one human intervention, and no external evaluation of the intern milestone exists.

    1. 2026-09-10unmeasured to emergingconf 30 · claude-initial-seed

      OpenAI's 6 Sep 2026 post declares the September 2026 intern milestone met by its own measurements (3.1 agent-workdays per human workday; over half of successful 4-8 hour tasks still needed a human intervention). A lab's self-grade is tier 7 evidence and caps the status at emerging; no independent replication exists. openai.com blocks the tracker's fetcher, so the figures are cited, not in the ledger. Initial seed.

  2. OpenAI

    2026-09-06window to 2028-03-31not yet testableconf 30
    We are making strong progress toward creating an automated AI researcher by March of 2028.source

    How we track it. Resolves when an external evaluation shows autonomous delivery of a multi-week research project without human intervention; proxied until then by the 80% METR horizon reaching weeks.

    Indicators: METR 80% time horizon

    Counterevidence and history

    OpenAI's intern milestone was self-graded; its own intervention rate on four-to-eight-hour tasks is above half; METR's 80% horizon on external tasks is hours, not weeks.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Window closes March 2028; the only resolving evidence would be an external evaluation of autonomous multi-week research, which does not exist. Initial seed.

  3. Dario Amodei

    2026-01-20window to 2027-01-20behindconf 55
    I think we might be six to 12 months away from when the model is doing most, maybe all of what software engineers do end-to-end, and then it’s a question of how fast does that loop close?source

    How we track it. 'Most' = more than half of software-engineering task-hours completed without human edits. Proxied by the 80% METR horizon on software tasks reaching working days, and by the developer RCTs turning strongly positive.

    Indicators: METR 80% time horizonDeveloper productivity uplift (METR RCTs)50%/80% horizon ratio

    Counterevidence and history

    Amodei's own later phrasing softened to 'one or two years' (Dwarkesh Patel interview, Feb 2026). METR's 2026 RCT found experienced developers slowed by 4% (new recruits) and 18% (returning participants). The 80% horizon on METR's suite is around three hours.

    1. 2026-09-10unmeasured to behindconf 55 · claude-initial-seed

      Mid-window (Jul 2026 to Jan 2027). METR's 80% horizon on the best-measured model is about three hours (Mythos preview, Apr 2026), not a working day; METR's 2026 RCT found experienced developers 4% slower (new recruits) and 18% slower (returning), after 19% slower in 2025. 'Most of what software engineers do end-to-end' is not what the external instruments show. Initial seed.

  4. Dario Amodei

    2026-06-10window to 2028-06-30not yet testableconf 30
    If these scaling laws continue for only a year or two longer, we are likely to get what I’ve called Powerful AI, or “a country of geniuses in a datacenter”.source

    How we track it. Resolves against the 'Machines of Loving Grace' definition (Nobel-level across fields, autonomous multi-day and multi-week tasks, millions of instances); proxied by the 80% METR horizon exceeding weeks and a shipped continual-learning system.

    Indicators: METR 80% time horizonMETR 50% time horizon

    Counterevidence and history

    The same post hedges that it 'could also be considerably further out'; Amodei called it 'more like a 50/50 thing' in February 2026; no continual-learning system has shipped and the 80% horizon is hours.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Window runs to mid-2028; the definition (Nobel-level across fields, autonomous multi-week work) has no instrument today beyond the 80% horizon, which is hours. Initial seed.

  5. Dario Amodei

    2026-02-13window to 2029-12-31emergingconf 35
    It is hard for me to see that there won’t be trillions of dollars in revenue before 2030.source

    How we track it. Sum of frontier-lab run-rates (Epoch's revenue reports plus press-confirmed figures) reaching $1T annualised before 1 Jan 2030; interim path per Amodei's own words on the same page, 'low hundreds of billions by 2028'.

    Indicators: Circular financing scaleCloud backlog (remaining performance obligations)

    Counterevidence and history

    Run-rates are annualised from a single month and reported by press, not booked or audited; Cahn's and Bain's gap analyses put required revenue far above the realised path.

    1. 2026-09-10unmeasured to emergingconf 35 · claude-initial-seed

      Combined OpenAI ($40B, Aug 2026) and Anthropic ($65B, Jul 2026, CNBC-confirmed) run-rates are about $105B annualised, on Amodei's own 'low hundreds of billions by 2028' path. Run-rates are press-reported and annualised, and no stack-wide revenue metric exists yet. Initial seed.

  6. Dario Amodei

    2026-02-13window to 2026-12-31aheadconf 50
    there’s this bizarre 10x per year growth in revenue that we’ve seen. So in 2023, it was zero to $100 million. In 2024, it was $100 million to $1 billion. In 2025, it was $1 billion to $ 9-10 billion.source

    How we track it. Anthropic's run-rate at end-2026 of $90–100B against $9–10B at end-2025 keeps the 10x-a-year pattern; the same page states the intent to grow '20 or 30x a year instead of 10x'.

    Indicators: Circular financing scale

    Counterevidence and history

    Run-rate figures annualise one month of revenue and are press-reported (tier 5); a single large enterprise contract moves them; the adoption indicators (22% of firms, 6% of work hours) sit far behind revenue, which is a tension the run-rate cannot resolve.

    1. 2026-09-10unmeasured to aheadconf 50 · claude-initial-seed

      Anthropic's run-rate went from $9B (Dec 2025, Epoch) to $65B at end-July 2026 (Epoch; CNBC confirmed): 7.2x in seven months, where a 10x-a-year pace would give about 3.8x. Two sources; tier 5 run-rate, not booked revenue. Initial seed.

AI 2027 · 3

  1. AI Futures Project (Kokotajlo, Lifland, Jurkovic et al.)

    2025-04-03window to 2027-03-31on trackconf 60
    According to a recent METR report, the length of coding tasks AIs can handle, their “time horizon”, doubled every 7 months from 2019 - 2024 and every 4 months from 2024-onward.source

    How we track it. The 50% METR horizon doubling time fitted over the 2024-onward window stays at or under four months (122 days); the tracker's own fit is the band input on the horizon indicators.

    Indicators: METR 50% time horizonMETR 80% time horizon

    Counterevidence and history

    METR's 2023-onward fit is 128.7 days, over four months; the 2024-window fit is sensitive to the model set; METR's 16-hour suite ceiling censors the newest points; horizons measure software tasks under METR's suite, not work in general.

    1. 2026-09-10unmeasured to on trackconf 60 · claude-initial-seed

      The tracker's log-linear fit of METR's 50% horizon over the 2024-onward window gives a doubling time of 108 days (95% CI 99-119), inside the four-month (122-day) claim; METR's own 2024-onward figure is 89 days. The 2023-onward fit (128.7 days) would read behind, so the reading is window-sensitive. Initial seed.

  2. AI Futures Project (Kokotajlo, Lifland, Jurkovic et al.)

    2025-04-03window to 2027-03-31behindconf 55
    If the trend continues to speed up, by March 2027 AIs could succeed with 80% reliability on software tasks that would take a skilled human years to complete.source

    How we track it. The 80% METR horizon reaching a working year (about 2,000 hours) by March 2027; the tracker's 80% horizon indicator is the input.

    Indicators: METR 80% time horizon50%/80% horizon ratio

    Counterevidence and history

    The sentence is conditional on the trend speeding up; the authors' own all-things-considered medians moved out in December 2025; METR's suite cannot measure horizons beyond 16 hours, so a 'years' horizon would be unmeasurable on the current instrument.

    1. 2026-09-10unmeasured to behindconf 55 · claude-initial-seed

      Six months before the date, the 80% horizon on METR's suite is about three hours (Mythos preview, Apr 2026). A 'years' horizon (about 2,000 working hours) needs roughly ten doublings, about three years at the current 108-day doubling time, and the doubling time is not shortening. The condition 'if the trend continues to speed up' is not met. Initial seed.

  3. AI Futures Project

    2025-04-03window to 2027-12-31not yet testableconf 30
    All model-based forecasts have 2027 as one of the most likely years that SC being developed, which is when an SC arrives in the AI 2027 scenario.source

    How we track it. SC = an AI that does any coding task of the best lab engineer, 30x faster and cheaper, on 5% of frontier compute; proxied by the 80% METR horizon and a lab disclosure of fully autonomous coding of frontier training runs.

    Indicators: METR 80% time horizonMETR 50% time horizon

    Counterevidence and history

    The authors moved their own medians to 2028 (Nikola Jurkovic) and 2030 (Eli Lifland) in a December 2025 disclaimer on the same page; METR's 2026 RCT found no speedup for experienced developers.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Window is calendar 2027. The authors' own December 2025 update moved their medians to 2028 and 2030; recorded, but a claim is not scored before its window opens. Initial seed.

Capture theses · 9

  1. David Cahn (Sequoia Capital)

    2024-06-20window to 2026-12-31emergingconf 40
    AI’s $200B question is now AI’s $600B question.source

    How we track it. Capex-to-revenue gap for the stack: Nvidia data-centre run-rate doubled for the rest of the data centre and doubled again for a 50% margin, minus AI ecosystem revenue. Tracked as a gap, not pass/fail, once the capex_to_revenue_stack metric lands.

    Indicators: Circular financing scaleSemiconductors' share of stack operating incomeCloud backlog (remaining performance obligations)

    Counterevidence and history

    Cloud AI revenue is growing (AWS, Google Cloud and Intelligent Cloud segment revenue in the ledger); neither the 'double it twice' multiplier nor the 50% margin assumption is an observed quantity, so the gap is an accounting construct.

    1. 2026-09-10unmeasured to emergingconf 40 · claude-initial-seed

      Tracked, not resolved: the stack-wide capex-to-revenue metric is not yet built. Cahn's own July 2026 update restates the gap at $1.5T; realised lab run-rates (about $105B combined) remain far below the implied required revenue. Initial seed.

  2. David Cahn

    2026-07-08window to 2030-12-31emergingconf 40
    A lot of people have reached out to me for the updated math behind AI’s $600B question. It is now AI’s $1.5T question.source

    How we track it. Same gap measure: about $3T of cumulative lifetime revenue required against realised AI revenue.

    Indicators: Circular financing scaleSemiconductors' share of stack operating income

    Counterevidence and history

    The required-revenue figure depends on assumed hardware lifetimes and margins; the labs' run-rates ($65B Anthropic, $40B OpenAI in mid-2026) are growing several-fold a year, which the static gap arithmetic does not project forward.

    1. 2026-09-10unmeasured to emergingconf 40 · claude-initial-seed

      Same gap measure as the $600B claim; tracked until the capex-to-revenue metric lands. Lab run-rates are growing several-fold a year, which the static arithmetic does not project. Initial seed.

  3. Bain & Company (Global Technology Report 2025)

    2025-09-23window to 2030-12-31not yet testableconf 30
    Two trillion dollars in annual revenue is what’s needed to fund computing power needed to meet anticipated AI demand by 2030. However, even with AI-related savings, the world is still $800 billion short to keep pace with demand, new research by Bain & Company finds.source

    How we track it. Global AI revenue (labs, applications, cloud AI) trajectory against $2T a year by 2030; interim reading from the capex-to-revenue gap.

    Indicators: Circular financing scaleCloud backlog (remaining performance obligations)

    Counterevidence and history

    Bain's demand figure is a model of required, not observed, revenue; falling price per unit of capability means the same compute serves more demand than the model assumes.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Resolves in 2030 against global AI revenue; no interim reading is decisive. Initial seed.

  4. Jim Covello (Goldman Sachs)

    2024-06-25window to 2029-12-31emergingconf 40
    AI technology is exceptionally expensive, and to justify those costs, the technology must be able to solve complex problems, which it isn’t designed to do,source

    How we track it. Direction claim: application-layer margins net of inference and reported enterprise P&L impact fail to cover the capex; contradicted if the lab recoupment ratio exceeds one or the developer RCTs turn strongly positive.

    Indicators: Developer productivity uplift (METR RCTs)METR 50% time horizonSemiconductors' share of stack operating income

    Counterevidence and history

    The 50% METR horizon on multi-hour software tasks has doubled roughly every four months since 2024, which is not what 'isn't designed to solve complex problems' predicts; hyperscaler cloud segments report rising revenue and stable margins.

    1. 2026-09-10unmeasured to emergingconf 40 · claude-initial-seed

      Mixed evidence. METR's RCTs found no developer speedup (19% slower in 2025; 4% and 18% slower in 2026), consistent with Covello; METR's 50% horizon doubling every 108 days is not. Recoupment metrics stay unpublished until lab cost data exist. Initial seed.

  5. Tomasz Tunguz

    2026-03-17window to 2031-03-17emergingconf 35
    Hyperscalers are spending $12 on AI infrastructure for every $1 of AI revenue, betting on 5x revenue growth within five years to justify $575B in capex.source

    How we track it. Hyperscaler AI capex over hyperscaler AI revenue near 12; resolves in Tunguz's favour if the ratio has not fallen toward the low single digits within five years.

    Indicators: Cloud segment operating marginsCloud backlog (remaining performance obligations)Circular financing scale

    Counterevidence and history

    The 12:1 figure divides capex by AI-attributed revenue that hyperscalers only partly disclose; AWS operating margin held near 39% through the capex surge and cloud backlogs are at record highs.

    1. 2026-09-10unmeasured to emergingconf 35 · claude-initial-seed

      No hyperscaler discloses AI revenue cleanly enough to compute the 12:1 ratio from filings; AWS operating margin held near 39% through the capex surge (cloud_segment_margins) and cloud backlogs are at records. Initial seed.

  6. Korinek, Jones, Sacher, Cotter and McCrory (The Anthropic Institute, WP 2026-02)

    2026-09-01window to 2030-12-31not yet testableconf 30
    Under modest change, AI adds less than half a point to GDP growth by 2030 and raises unemployment by a tenth of a point.source

    How we track it. US GDP growth and TFP against the no-AI trend; unemployment-rate delta. The authors state the scenarios only separate after 2027, so no macro reading before 2028 discriminates them.

    Indicators: US total factor productivity, private nonfarm businessUS labour productivity, year on yearUS labour share of income, year on year

    Counterevidence and history

    The scenarios are conditional paths, not forecasts, and the paper says today's readings cannot tell them apart before 2028; the labour share is already falling 3.4% year on year, which the modest scenario would not predict.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Per the authors, almost all divergence between the scenarios comes after 2027; the tracker records the path and expects nulls until 2028. Initial seed.

  7. Korinek et al. (The Anthropic Institute, WP 2026-02)

    2026-09-01window to 2030-12-31not yet testableconf 30
    the median respondent’s answers are consistent with our substantial change scenario in which, by 2030, GDP rises by 8 percent and cognitive employment declines by 4 percent.source

    How we track it. US GDP 8% above the no-AI path by 2030; employment in cognitive occupations down 4%, read off the labour trackers' exposed-occupation series.

    Indicators: US total factor productivity, private nonfarm businessCross-tracker concordance (labour)Entry-level employment shortfall in AI-exposed occupations

    Counterevidence and history

    The scenario is anchored on a survey median, not on data; the paper's own divergence claim means the tracker should expect nulls until 2028.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Per the authors, almost all divergence between the scenarios comes after 2027; the tracker records the path and expects nulls until 2028. Initial seed.

  8. Korinek et al. (The Anthropic Institute, WP 2026-02)

    2026-09-01window to 2030-12-31not yet testableconf 30
    GDP growth then rises to 15 percent per year, the labor share of income falls from 60 to 45 percent, and nearly one in five cognitive workers is unemployed.source

    How we track it. BLS labour share of nonfarm business income falling toward 45%; GDP growth toward 15% a year; cognitive-worker unemployment near 20%.

    Indicators: US labour share of income, year on yearUS labour productivity, year on yearRecent-graduate unemployment rate (NY Fed)

    Counterevidence and history

    The extreme path requires diffusion far beyond the adoption indicators (22% of firms, 6% of hours); the labour-share decline so far (3.4% year on year) is within the range of past cyclical swings.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      Per the authors, almost all divergence comes after 2027. The BLS labour share is already down 3.4% year on year (labor_share_nonfarm, emerging), noted here but not scored against a 2030 path. Initial seed.

  9. Korinek et al. (The Anthropic Institute, WP 2026-02)

    2026-09-01window to 2030-12-31not yet testableconf 30
    Second, almost all of the divergence comes after 2027, a year and a half from the mid-2026 anchor, because the scenarios share today’s readings and separate only as the reach and use of AI diverge.source

    How we track it. A meta-claim about the tracker's own power: no macro indicator (GDP, TFP, labour share, unemployment) should separate the three scenarios before 2028. Falsified if a macro series breaches its fast band before then.

    Indicators: US total factor productivity, private nonfarm businessUS labour share of income, year on yearCross-tracker concordance (labour)

    Counterevidence and history

    Labour share and entry-level employment in exposed occupations already show movements the modest scenario would not predict; 'cannot discriminate before 2028' may be too conservative.

    1. 2026-09-10unmeasured to not yet testableconf 30 · claude-initial-seed

      A meta-claim about instrument power, scored only if a macro series breaches its fast band before 2028. The labour share (down 3.4% year on year) sits in the emerging band, not the fast band. Initial seed.