Methodology
One tracker, two lenses. The diffusion lens follows Narayanan and Kapoor's AI as Normal Technology: five stocks, valves between them, normal and fast bands with falsification thresholds. The capture lens measures who keeps the surplus, layer by layer. They are the same question asked from two ends, so they share one data store, one indicator object and one discipline, adapted from the AI 2027 tracker.
Three layers, strictly separated
Observation (raw, sourced, dated) → Derived (a formula over observations, defined in the semantic layer) → Indicator (proxies, status, confidence, evidence, counterevidence). No derived number exists without a traceable path to observation ids, and the site never computes a number: every figure was exported with the ids behind it.
Evidence tiers
| Tier | Name in the data | What it covers |
|---|---|---|
| 1 | benchmark | Benchmark or independent evaluation (METR horizons, Epoch benchmarks) |
| 2 | model_release | Model release details (system cards, technical reports) |
| 3 | product_behaviour | Public product behaviour (list prices, usage data from a product) |
| 4 | official_filing | Official filing or policy (SEC filings, government statistics, enacted law, regulator lists) |
| 5 | credible_reporting | Credible reporting (Reuters, Bloomberg, FT, WSJ, CNBC, The Information) |
| 6 | published_analysis | Published analysis with a transparent method (surveys, research notes, analyst baskets) |
| 7 | actor_statement | Statement from an actor about itself (intentions and self-measurements, not outcomes) |
The tier is stored; the letter grade is derived for display. A = tier 1, or tier 4 audited or government statistics. B = tiers 2–3, tier 4 company statements (8-K), tier 6 with a transparent method. C = tier 5, or tier 6 estimates. D = tier 7 or single-source. Tier 7 evidence never moves a status above emerging. Audited is reserved for tier 4 and a 10-K; a run-rate reported by the press or by the company is never audited, whatever its tier.
Status vocabularies
Flow indicators (diffusion): consistent with normal · faster than normal · slower than normal · emerging · not yet measurable, each against a normal band, a fast band and an optional falsifying band with a written rationale. Capture indicators: concentrating · dispersing · stable · unclear, from a trend over N periods with a dead-band. Every indicator also carries a timing tag: leading (moves before capture shows in filings), coincident (moves with it) or lagging (confirms it after the fact). Predictions: confirmed · ahead · on track · behind · emerging · not yet testable. Not yet measurable means the thing cannot be measured yet; an indicator whose connector has not landed is simply unpublished.
Bands
Each diffusion indicator names the single series or derived metric its bands apply to (shown as “applied to” on the indicator page), a normal band anchored on the pace of an earlier general-purpose technology at the same age, a fast band anchored on the pace the AI 2027 scenario needs, and where one exists a falsifying threshold. The rationale is written next to the numbers. Bands change only through a reviewed pull request that states why; the evaluator never moves them.
Confidence, 0–95, independent of status
90–95 multiple strong independent sources; 70–89 good evidence with some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.
Discipline, enforced in code
- Every observation carries a URL that resolved at ingestion, its HTTP status, a content hash, the retrieval time, the tier, an audited/company-stated/reported/estimated flag, the extraction method and the raw snippet.
- Figures quoted from documents enter through a manual connector that fetches the page and asserts the snippet is a verbatim substring of it. Nothing is typed into an indicator.
- No single-source status change unless the source is a benchmark, a model release or an official filing. Non-empty counterevidence before anything publishes.
- Status lives only in status events, each with a reason, evidence ids and an author. The changelog is generated from them. Observations are superseded, never deleted; disputed figures carry their dispute text everywhere they appear.
- The evaluator proposes; it writes its own reason only for a first reading inside a band, and a band crossing waits for a human. Bands and formulas change only through a reviewed pull request.
- Rebuilt nightly at 06:17 UTC into a pull request; merging that pull request is the approval that publishes new observations and statuses.
- Fetched HTML passes through a prompt-injection scrub; anything addressed to an AI agent is logged and never stored.
Query layer
The Ask button and the console on /query run against the same store the site is exported from, on a separate read-only service. A question goes to a model with four tools (read-only SQL over the semantic layer, a metric lookup, an indicator lookup, and status changes since a date). Every number in an answer must be followed by the citation token of the record it came from; a post-check extracts the numbers and verifies each against that record's value, confidence bounds, band edges or verbatim snippet. An answer that fails is revised once and otherwise shown as blocked with the unverified numbers marked. The service cannot write observations or statuses, has a daily spend cap, and logs only a hash of each question. The model and prompt version are recorded with every answer.
Weekly memo
Each Monday a job assembles what changed since the last memo (status events, new observations by indicator, watchlist posts, staleness, the thesis monitor, and any crosswalk pair whose two sides moved in opposite directions) and asks a model to draft the memo and the two lens sentences under the same citation rule. If no key is configured or the check fails twice, the memo is the deterministic digest of the same facts. Either way it opens a pull request; merging it publishes the memo under /memos. Nothing in a memo can move a status.
X posts
X is never scraped. Posts enter through a curated List on the X API, or through a weekly manual drop of URLs fetched via X's public embed endpoint. Every post is tier 7 and recorded as a watchlist row; when a post links to a tier 1–6 artifact, the artifact is fetched and entered as an observation in its own right. Posts never move a status.
Colour
Two accents mark direction (faster or concentrating; slower or dispersing) and always ship with an icon and a word. Nothing is red or green: fast is not good, and concentrating is not good.
Credits and conflicts
Method after the AI 2027 tracker (independent, not affiliated). Disclosure: the maintainer is involved with BetterBrain, a firm in the deployment-services sub-layer; that sub-layer is listed and its indicators are held to the same rules.
Data credits, generated from the source registry: METR, Measuring AI Ability to Complete Long Tasks (Horizon v1.1); SEC EDGAR, XBRL company facts API; METR, Time Horizon 1.1 (29 Jan 2026); Epoch AI, 'Data on AI Companies' (CC BY 4.0), epoch.ai/data (CC BY 4.0); U.S. Bureau of Labor Statistics, Productivity and Costs (PRS85006092/93), Total Factor Productivity (MPU4910012); Bick, Blandin & Deming, Real-Time Population Survey, via FRED; Brynjolfsson, Collis, Eggers, Kazinnik & Nguyen, 'What is Generative AI Worth?' (2026); MIT NANDA, 'The GenAI Divide: State of AI in Business 2025'; U.S. Census Bureau, Business Trends and Outlook Survey; Menlo Ventures, '2025: The State of Generative AI in the Enterprise'; HSBC Insights, 'The billions of AI consumer surplus' (10 Aug 2026); Fortune, 'MIT report: 95% of generative AI pilots at companies are failing' (18 Aug 2025); FRED Blog, 'Does generative AI save time at work?' (27 Aug 2026); Bick, Blandin & Deming, Real-Time Population Survey, retrieved from FRED; Canaries Dashboard, a project of the Stanford Digital Economy Lab and ADP Research; Revelio Labs, AI Labor Market Tracker; California Policy Lab and California Employment Development Department, California AI-Unemployment Tracker; California Employment Development Department, AI and the economy; Anthropic Economic Index, open dataset on Hugging Face (CC BY 4.0) and the 2026 reports (CC BY 4.0 (dataset)); Federal Reserve Bank of New York, The Labor Market for Recent College Graduates; SEC EDGAR, inline XBRL filings (segment disclosures); SEC EDGAR archives; NVIDIA Newsroom; Official Microsoft Blog; Microsoft Investor Relations; Anthropic newsroom; About Amazon; CNBC; TechCrunch (quoting Sam Altman on X); SiliconANGLE (relaying the WSJ report); Investing.com via Yahoo Finance; CDS data ICE Data Services; Ramp AI Index, Ramp Economics Lab; OpenAI; METR; Anthropic, Claude Fable 5.1 & Claude Mythos 5.1 System Card (1 Sep 2026); Jamin Ball, Clouded Judgement; OpenRouter list prices; SEC EDGAR Form D data sets; Epoch AI, Notable AI Models (CC BY 4.0) (CC BY 4.0); Epoch AI, Machine Learning Hardware (CC BY 4.0) (CC BY 4.0); Epoch AI, LLM inference price trends (CC BY 4.0) (CC BY 4.0); Epoch AI, Benchmarking Hub and Epoch Capabilities Index (CC BY 4.0) (CC BY 4.0); Epoch AI, AI Chip Sales (CC BY 4.0) (CC BY 4.0); Epoch AI (CC BY 4.0) (CC BY 4.0); FDA, Artificial Intelligence-Enabled Medical Devices list; NCSL, Artificial Intelligence 2025 Legislation; rl-list.com; Our World in Data, AI investment (CC BY 4.0), from the Stanford AI Index (CC BY 4.0 (OWID; underlying data AI Index / Quid)); Google DeepMind; arXiv; Baseten; Thinking Machines Lab; Cursor (Anysphere); Trajectory; Harvey; Engram; Decagon; Google Research; NBER; MIT Sloan; People Matters; Automation Anywhere; AI Futures Project; AI 2027 Tracker; The Budget Lab at Yale; Humlum & Vestergaard; Chicago Booth; CEPR VoxEU; WashU Olin Business School; PwC 2026 Global AI Jobs Barometer; Outsource Accelerator; Dealroom; Dwarkesh Podcast; Electrek; FDA 510(k) database; STAT News; Innolitics; PYMNTS (relaying Bloomberg); Fortune; X; OpenRouter; Artificial Analysis; Stripe newsroom; NVIDIA Blog; AWS Big Data Blog.
Reuse, citation and corrections
Code is MIT licensed. The compiled dataset (observations, derived values, status events) is CC BY 4.0; each observation also carries its upstream source and licence, which govern that row. Cite the site by URL and the observation ids you rely on. Corrections: open an issue or pull request on the repository; every change to a number or a status leaves a dated event in the changelog.