Safety brakes inside the labs (events, trailing year)
1What this measures
Deliberate slowdowns announced by a frontier lab and recorded as dated evidence rows with a verbatim announcement: paused training runs, gated releases, confinement of a model under a preparedness framework. Derived metric safety_brake_events_12m counts them over the trailing twelve months.
Why it matters. The paper's safety speed limits are argued from regulated domains; the tracker also watches for them on the invention stage, where a brake inside a lab slows the whole chain before any regulator acts.
- Proxy types
- policy, model_release
- Unit
- count
- Cadence
- quarterly
- Valve
- adoption to adaptation
2How we track this
- derived metric
safety_brake_events_12m(formula in the semantic layer)
- Normal band
- ≥ 1
- Fast band
- ≤ 0
- Falsifying
- —
Normal = at least one announced brake a year (the limits bind); fast = none, the pace the AI 2027 scenario assumes on the invention stage. Lab self-reports are tier 7, so the status is capped at emerging until an independent record exists.
Applied to metric:safety_brake_events_12m.
3Tracker interpretation
Three brakes in five weeks of 2026, all inside OpenAI and all self-announced; the count is a floor, since labs announce selectively.
4Evidence
3 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.
Derived rows (3)
| as of | dims | value | inputs |
|---|---|---|---|
| 2026-09-01 | 3 | obs:2ab8fec4obs:67d75927obs:82685d1e | |
| 2026-08-18 | 2 | obs:67d75927obs:82685d1e | |
| 2026-08-07 | 1 | obs:67d75927 |
4Evidence log
Pachocki: voluntary slowdowns should become commonplace until shared safety bars exist (6 Sep 2026)
“I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.”
Astra judged to meet the Critical cybersecurity threshold; release gated behind safeguards (1 Sep 2026)
“We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.”
two-week pause in RL training on deployment-bound models after the Hugging Face incident (Jul-Aug 2026)
“This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems.”
Astra internal activities paused pending strengthened security controls (7 Aug 2026)
“We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.”
5Status and reasoning
Three announced brakes inside OpenAI in the trailing year, all with a verbatim source via the Internet Archive: internal Astra activity paused pending stronger security controls (7 Aug 2026), a two-week pause in RL training on deployment-bound models after the Hugging Face incident (18 Aug), and Astra judged to meet the Critical cybersecurity threshold with release gated behind safeguards (1 Sep). At or above the normal band (one or more a year), but every row is the lab describing itself (tier 7), so the status is capped at emerging. Initial seed.
6Timeline notes
7Counterevidence
What cuts against this reading
Every row is the lab describing itself; a brake announced is not a brake measured, and the same lab restarted the paused run within a month; other labs' brakes are not yet recorded.
8Update history
- 2026-09-10unmeasured to emergingconf — → 40 · evaluate
Three announced brakes inside OpenAI in the trailing year, all with a verbatim source via the Internet Archive: internal Astra activity paused pending stronger security controls (7 Aug 2026), a two-week pause in RL training on deployment-bound models after the Hugging Face incident (18 Aug), and Astra judged to meet the Critical cybersecurity threshold with release gated behind safeguards (1 Sep). At or above the normal band (one or more a year), but every row is the lab describing itself (tier 7), so the status is capped at emerging. Initial seed.
9Confidence
40 / 95 — limited or vague evidence
Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.