Claims and their evidence
Each number in the working paper's abstract and contributions, what kind of evidence it rests on, and where it is recomputed. Live means agents run with the speculator; replay means a speculator replayed against τ²-bench's environment; record means a replay or count from published trajectories alone (pilot/check_flow.py --trace), whose savings are exact and whose detours are lower bounds. Intervals are 95%, from bootstraps over tasks.
| Claim | Number | Evidence | Recomputed by |
|---|---|---|---|
| Counting the right event calibrates the probability of use | expected calibration error 0.01–0.08 in every τ²-bench domain, against 0.06–0.16 for the next-step probability | Replay, nine agents | scripts/calibration.py (results) |
| The speculator takes most of retail's ceiling | 86.4% of it [80.2, 93.1], 10.2 points more than the next-step speculator [6.7, 14.2] | Replay, nine agents it never saw | results, the paper's Table 3 |
| It works live with no model of its own | GLM-5.3 took 27.9% fewer LLM turns (19.1–35.9%) on 28 retail and airline tasks; 21 passed, against 24 | Live, paired with the recorded baseline | results and its episodes archive |
| It works on frontier models | Claude Sonnet 5 took 20.5% fewer LLM turns (16.5–24.4%) and Claude Haiku 4.5 22.4% (15.8–29.4%), fewer on 25 of 28 tasks each; passes 65 → 69 and 58 → 62 of 84, pass^3 0.679 → 0.679 and 0.464 → 0.536 | Live, pre-registered: three trials of 28 retail and airline tasks, arms side by side | pilot/analyze_trials.py (results) |
| A prompt does not get what it gets | Anthropic's sample prompt for parallel tool calls saved Sonnet 5 3.4% of turns (−1.2–7.8%), Haiku 4.5 5.9% (0.9–10.9%) and GLM-5.3 7.6% (−1.8–15.0%); with the prompt in both arms, the flow saved Sonnet 5 22.9% (19.5–26.1%) and Haiku 4.5 17.2% (10.2–24.0%) | Live; GLM-5.3 against its recorded arms | pilot/analyze_trials.py, pilot/analyze_recorded.py (results) |
| It costs less | the agent's cost at list prices, with prompt caching, fell 11.8% (Sonnet 5) and 8.7% (Haiku 4.5), and its input tokens 17.7% and 20.6% | Live, as above | the same |
| It holds live beyond τ²-bench | on AgentDojo's Slack and travel suites, where the replay found reads to take, GLM-5.3 and Claude Haiku 4.5 took 10.1% fewer LLM turns (5.8–14.0%; 14 pairs fewer, 1 more), passing 27 of 34 in both arms; 6.0% (2.1–9.9%) over all 41 AgentDojo tasks, where the replay projected 7.2%. BFCL, projected at 1.6%, shows no effect (−1.6%, −6.3 to 4.2) | Live, two models, one run per arm, in the benchmarks' own environments and scored by their own checks | pilot/bench/analyze_bench.py (results and its episodes archive) |
| The ceiling is the domain's | 3.5% of turns in WorkBench to 47.1% in AgentDojo's travel suite; 29.0% in τ²-bench | Record, 89 more agents on six benchmarks | scripts/ceiling.py, scripts/bench/replay.sh ceilings (results) |
| Reading the request is a model's job | of the 3–7 points the user's words would add, a small model that picks the value among them takes 30–53% and a pattern 0–17% | Record, with Jev's answers (8,532 questions) | ceiling.py --questions, scripts/model_questions.py, ceiling.py --model-answers (results) |
| It gains where agents read ahead | τ-bench retail +1.4 points (0.8–2.1); BFCL +0.85 (0.35–1.45); DTap-Bench, the agent's own sessions, +1.3 (1.1–1.5), and in its legal domain, replayed separately, +0.8 (0.7–1.0) | Record | scripts/bench/replay.sh, scripts/bench/tables.py (results) |
| It stays out where it cannot help | no lookup in WorkBench's 13,869 turns; in DTap-Bench's medical domain both speculators make the same lookups | Record | the same |
| It carries across harnesses | learned in other harnesses, 59% of what the agent's own sessions save (64% of their detours); from its harness-mates, 79% | Record, DTap-Bench's three agent SDKs | the same (results) |
| It learns fast | ten of an agent's own sessions give 96% (retail) and 93% (airline) of what all of them do | Replay, three agents, three orders | scripts/learning_curve.py (results) |
| Where no user speaks, compile once | the procedure alone passes 35 of 40 held-out solo telecom tasks with no model; with GLM-5.3 on its four hand-backs, 39 of 40 at 0.48 LLM turns per ticket, against 15.9 for GLM-5.3 alone | The procedure run in τ²-bench's environment; the hand-backs live, two trials | scripts/telecom_workflow.py, stretto-procedure (results, workflow) |
What the evidence does not show, in brief (the paper's §6 has the rest):
- Replays assume the agent skips what a lookup already answered. GLM-5.3 did live. In the record, agents repeat 0.8–2.6% of their reads with no write between on five benchmarks, and a few models many more (gpt-oss-120b 23%, Qwen3.5 Flash 17.5%), whose replayed savings are upper bounds.
- Detours from the record are lower bounds. Replay from the record found 55–72% of the environment's on τ²-bench. Replayed from the same agents' own no-flow episodes on AgentDojo, it found 3 of the 52 the flows made live: the record cannot follow a chain through tools the agent never called. Its projected savings held there: 11 turns, against 10 saved live in the pairs where the agent used a lookup (results).
- Few agents live.
- On τ²-bench, the use-before-write speculator ran live with GLM-5.3, once per task, and with Claude Sonnet 5 and Claude Haiku 4.5, three trials each (results). In airline, the Claude models' savings (5.3% and 8.5%) are not distinguishable from none.
- On AgentDojo and BFCL, it ran with GLM-5.3 and Claude Haiku 4.5, one run per arm (results).
- An earlier flow also ran with Claude Haiku 4.5 on ten retail tasks and with Claude Sonnet 5 on three (results).
- Every other agent is replayed.
- Live cost is at list prices. On τ²-bench, three trials with the arms side by side put the flow's saving in the Claude agents' cost at 11.8% and 8.7%, less than in their input tokens, since nine tenths of the input is read from the prompt cache. The subscription those runs used bills no tokens. On AgentDojo and BFCL, with one run per arm, cost varied from run to run as much as it differed between the arms.
- Pass rates are underpowered, and τ²-bench's users are LLMs. With 84 pairs per Claude model, a harm of about ten points would show; none did.