Skip to content
Slow Variables

Ask the data

Answers come from the same store as the site; every number is checked against the record it cites.

Products · Model

Agent reliability across repeated runs (tau-bench pass^k)

emerging55/95 confidence, mixed or hard to operationalisegrade Bleading

1What this measures

Share of tau-bench tasks an agent passes on every one of k attempts. The 2024 paper's gpt-4o figures (pass^1 61%, pass^8 under 25%) set the baseline; the 2026 point is the leaderboard leader's pass^1 and pass^4 as tabulated by Automation Anywhere.

Why it matters. The 70% problem in one number - solving a task once is not solving it reliably. Deployment needs pass^k, not pass^1, and the gap between them is where the normal-technology speed limit lives.

Proxy types
benchmark
Unit
share
Cadence
quarterly
Valve
invention to product

2How we track this

  • series automation_anywhere.tau_bench_leaderboard_1.pass_hat_4.pt
  • series automation_anywhere.tau_bench_leaderboard_1.pass_hat_1.pt
  • series arxiv.tau_bench_retail_gpt4o.pass_hat_1.pt
  • series arxiv.tau_bench_retail_gpt4o.pass_hat_8.pt
  • source arXiv abstract pages · default tier 6 · arXiv; abstract quotation
  • source Automation Anywhere blog · default tier 7 · Automation Anywhere; short quotation
Normal band
≤ 70.0%
Fast band
≥ 90.0%
Falsifying

Normal = under 70% of tasks pass all four runs, so the best agent still fails one attempt in three; fast = 90% or more at k of four or higher, the reliability a workflow can be built on. Between is `emerging`.

Applied to automation_anywhere.tau_bench_leaderboard_1.pass_hat_4.pt.

3Tracker interpretation

Reliability still collapses with k - 70% once, 56% four times running - two years after the paper that named the problem.

4Evidence

Latest point
56.2%as of 2026-05-18
tau_bench_leaderboard_1
Value the bands apply to
56.2%as of 2026-05-18

4 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.

5Status and reasoning

emergingsince 2026-09-10 · evaluate

First scoring. The tau-bench leaderboard leader passes 70.2% of tasks once and 56.2% four times running (Automation Anywhere table, 18 May 2026), inside the normal band; the 2026 point is vendor-tabulated (tier 7) so the status is capped at emerging, with the 2024 paper (gpt-4o pass^1 61%, pass^8 under 25%) as the second source.

The tracker's prior expectation was consistent with normal; the evaluator reads emerging. The evaluator wins until a reviewed override.

6Timeline notes

  • 2026-05-18 tau_bench_leaderboard_1 · 56.2%as of 2026-05-18

7Counterevidence

What cuts against this reading

The 2026 point is vendor-tabulated and the model is unnamed; tau-bench domains are narrow; pass^k penalises benign variation.

8Update history

  1. 2026-09-10unmeasured to emergingconf 55 · evaluate

    First scoring. The tau-bench leaderboard leader passes 70.2% of tasks once and 56.2% four times running (Automation Anywhere table, 18 May 2026), inside the normal band; the 2026 point is vendor-tabulated (tier 7) so the status is capped at emerging, with the 2024 paper (gpt-4o pass^1 61%, pass^8 under 25%) as the second source.

9Confidence

55 / 95 — mixed or hard to operationalise

Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.

10Related