Indicators / series
cl_ladder.sdpo.rung_research.pt
Source: arXiv abstract pages (manual rows) (arXiv) · arXiv; abstract quotation · CSV
arXiv
| id | value | grade | as of date | published date | tier | audited vs reported | extraction method | review status | retrieved at | http status | content hash | supersedes id | url | flags |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 8d2af92daf94cb81 | 5.00 rung | grade B, evidence tier 6 | 2026-01-28 | 2026-01-28 | 6 | reported | manual | approved | 2026-09-10T06:54:23.630962Z | 200 | 949147a99524 | — | source page |
Raw snippets and dispute text
- 8d2af92daf94cb81 Finally, applying SDPO to individual questions at test time accelerates discovery on difficult binary-reward tasks, achieving the same discovery probability as best-of-k sampling or multi-turn conversations with 3x fewer attempts.