Indicators / series
metr.claude_mythos_preview_early_inspect.horizon_50.pt
Source: METR time horizons (v1.1) (METR) · METR; analysis code MIT at https://github.com/METR/eval-analysis-public · CSV
METR, Measuring AI Ability to Complete Long Tasks (Horizon v1.1)
| id | value | grade | as of date | published date | tier | audited vs reported | extraction method | review status | retrieved at | http status | content hash | supersedes id | url | flags |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| a3cd6bcee336eaa3 | 17.4 h (509–3304) | grade A, evidence tier 1 | 2026-04-07 | 2026-09-10 | 1 | reported | api | approved | 2026-09-10T00:51:13.090038Z | 200 | aae31902b051 | — | source page | ⚑ disputed: METR: measurements above 16 hours are unreliable with the current task suite |
Raw snippets and dispute text
- a3cd6bcee336eaa3 {"claude_mythos_preview_early_inspect": {"p50_horizon_length": {"ci_high": 3304.261235, "ci_low": 508.876789, "estimate": 1044.780145}, "release_date": "2026-04-07"}} — METR: measurements above 16 hours are unreliable with the current task suite