Indicators / series
cl_ladder.sdft_self_distillation.rung_research.pt
Source: arXiv abstract pages (manual rows) (arXiv) · arXiv; abstract quotation · CSV
arXiv
| id | value | grade | as of date | published date | tier | audited vs reported | extraction method | review status | retrieved at | http status | content hash | supersedes id | url | flags |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 5fb2b23a62ecf025 | 5.00 rung | grade B, evidence tier 6 | 2026-01-27 | 2026-01-27 | 6 | reported | manual | approved | 2026-09-10T06:54:23.589211Z | 200 | 70720b844357 | — | source page |
Raw snippets and dispute text
- 5fb2b23a62ecf025 In sequential learning experiments, SDFT enables a single model to accumulate multiple skills over time without performance regression, establishing on-policy distillation as a practical path to continual learning from demonstrations.