Frontier models and a prompting baseline, live
Every live τ²-bench result with the reach decider so far had GLM-5.3 as the agent. This round runs Claude Sonnet 5 and Claude Haiku 4.5 on the same 28 retail and airline tasks, three trials each, with and without the flow. It also tests the flow against the prompt a skeptic would try first: Anthropic's own sample prompt for parallel tool calls. The plan was written and pushed before any episode of the design ran; the deviations from it are listed below.
Findings
The flow saves turns on both Claude models. With it, Claude Sonnet 5 took 20.5% fewer LLM turns (95% interval 16.5% to 24.4%): fewer on 25 of the 28 tasks, more on 3. Claude Haiku 4.5 took 22.4% fewer (15.8% to 29.4%): fewer on 25, more on 1. By the plan's rule both savings are established.
- In retail the savings were 26.2% and 26.8%.
- In airline they were 5.3% and 8.5%, with intervals that include no change. There the models make more calls at once (Sonnet 5 1.80 per tool turn), and a flow spares a turn only when it makes all of that turn's calls.
Asking the agent to batch its reads does not get what the flow gets. We added Anthropic's sample prompt for parallel tool calls to the system prompt.
- Sonnet 5 took 3.4% fewer turns (−1.2% to 7.8%), which is not established.
- Haiku 4.5 took 5.9% fewer (0.9% to 10.9%).
- GLM-5.3, against its recorded baseline, took 7.6% fewer (−1.8% to 15.0%).
The Claude models already make most of their independent calls at once. The prompt moved Sonnet's calls per tool turn from 1.51 to 1.50, and Haiku's from 1.42 to 1.49. GLM-5.3 changed more, from 1.04 to 1.22 in retail and 1.46 in airline, and still saved about a quarter of what the flow saved it.
The flow saves on top of the prompt.
- With the prompt in both arms, the flow cut Sonnet 5's turns by 22.9% (19.5% to 26.1%) and Haiku 4.5's by 17.2% (10.2% to 24.0%).
- The prompt added nothing to the flow for Haiku 4.5 (+0.4%, −6.4% to +7.3%). For Sonnet 5 it added 6.2% (1.6% to 11.1%), in a comparison across the hour between Sonnet's two runs.
A flow makes reads whose arguments come from results the agent has not seen yet. A prompt can only batch the reads the agent already knows it needs.
Passes did not fall. No comparison's pass-rate interval lies below zero.
- With the flow, Sonnet 5 passed 69 of 84 episodes against 65 without it. Haiku 4.5 passed 62 against 58.
- Pass^3 was 0.679 in both of Sonnet's arms, and rose from 0.464 to 0.536 for Haiku.
The design can see a harm of about ten points, not smaller ones.
Cost fell less than turns. The agent's cost is at list prices, with prompt caching, as Claude Code reports it.
- The flow cut that cost by 11.8% for Sonnet 5 and 8.7% for Haiku 4.5, and its input tokens by 17.7% and 20.6%.
- Nine tenths of the agent's input is read from the prompt cache, at a tenth of the price of new input. The flow cuts those reads by a fifth. But the tokens written to the cache, the lookups' results among them, fell only 5% and 2%, and the agent's output 14% and 8%.
- The agent's own time fell 8.5% and 7.0%. The whole episode, the simulated customer's replies included, took 5.6% and 5.1% less time.
The lookups were the agents' own.
- Of the flow's 289 lookups with Sonnet 5, 270 were calls Sonnet made itself in the same task and trial without the flow. For Haiku 4.5 it was 273 of 304.
- The agents made 4 and 1 of the flow's lookups again, Sonnet's 4 all in airline.
Setup
As planned:
- Tasks. The live reach round's 28 tasks: 20 retail and 8 airline.
- Agents. Claude Sonnet 5 (
claude-sonnet-5) and Claude Haiku 4.5 (claude-haiku-4-5-20251001), each in Claude Code 2.1.283 throughclaude-agent.sh, on the user's Claude Code subscription. Each got τ²-bench's system prompt and τ²-bench's tools over MCP, behindstretto-proxy. - Customer. Claude Haiku 4.5 played τ²-bench's user simulator in every episode.
- Arms.
baseline: no flow.reach: the reach round's flows, learned from τ²-bench's four 2025 runs, at 0.3, asking no model.batch: the system prompt ends with the sample prompt for maximum parallel efficiency from Anthropic's prompting guide, word for word (BATCH_READSinrun_episode.py).batch-reach: both.
- Trials and order. Three trials of every task and arm, trial by trial, with each task's arms back to back in a shuffled order (
run_trials.py). Sonnet 5 ranbaselineandreachfrom 17:27 to 18:44 UTC, andbatchandbatch-reachfrom 18:45 to 19:59. Haiku 4.5 ran all four from 17:27 to 20:56. - Scoring. τ²-bench's database check.
- Episodes. 672 Claude episodes, none failed or rerun. At list prices they cost $46.92: $33.36 for the agents and $13.55 for the customer. The subscription's windows went from 0.08 (five hours) and 0.66 (seven days) at the first episode to 0.68 and 0.80 at the last, under every cap.
- GLM-5.3's batch arm. 28 episodes with GLM-5.3 as agent and customer, 390.6 Z.ai credits, compared with the reach round's recorded arms.
Claude Sonnet 5
84 pairs per comparison, 60 in retail and 24 in airline. The change is in total LLM turns over the pairs, with its 95% interval from a bootstrap over tasks. The sign test counts the tasks whose mean turns fell, rose or held.
| Arm | pass^1 | pass^3 | LLM turns per episode | Calls per tool turn | Input tokens per episode | Agent's cost per episode | Wall-clock |
|---|---|---|---|---|---|---|---|
| baseline | 0.774 | 0.679 | 9.92 | 1.51 | 100,412 | $0.068 | 80 s |
| reach | 0.821 | 0.679 | 7.88 | 1.22 | 82,655 | $0.060 | 75 s |
| batch | 0.821 | 0.714 | 9.58 | 1.50 | 97,486 | $0.065 | 80 s |
| batch-reach | 0.869 | 0.750 | 7.39 | 1.28 | 78,460 | $0.057 | 69 s |
| Without → with | LLM turns | Change | Tasks fewer / more / same | Input tokens | Agent's cost | Passed |
|---|---|---|---|---|---|---|
| baseline → reach | 833 → 662 | −20.5% [−24.4, −16.5] | 25 / 3 / 0 (p = 3·10⁻⁵) | −17.7% | −11.8% | 65 → 69 |
| baseline → batch | 833 → 805 | −3.4% [−7.8, +1.2] | 19 / 8 / 1 (p = 0.05) | −2.9% | −4.7% | 65 → 69 |
| batch → batch-reach | 805 → 621 | −22.9% [−26.1, −19.5] | 26 / 2 / 0 | −19.5% | −11.5% | 69 → 73 |
| reach → batch-reach | 662 → 621 | −6.2% [−11.1, −1.6] | 16 / 9 / 3 (p = 0.23) | −5.1% | −4.5% | 69 → 73 |
| baseline → batch-reach | 833 → 621 | −25.4% [−29.8, −21.8] | 23 / 2 / 3 | −21.9% | −15.7% | 65 → 73 |
By domain, reach against baseline:
- Retail. −26.2% (−30.6% to −21.4%), fewer turns on all 20 tasks. Passes went from 44 to 50.
- Airline. −5.3% (−12.4% to +2.9%), fewer on 5 tasks and more on 3. Passes went from 21 to 19, with an interval of −20.8 to 0.0 points, which is not a harm by the plan's rule. The two episodes that failed only with the flow:
- Task 6's third trial: the flow made no lookup, so the episode ran as it would have without it.
- Task 32's first trial: the customer asked to upgrade a basic-economy booking. The agent priced the upgrade on the existing route through Miami at $396, which the customer declined, and transferred them. GLM-5.3's reach episode on this task had failed the same way. The other two reach trials of task 32 made the same four lookups and passed.
Claude Haiku 4.5
| Arm | pass^1 | pass^3 | LLM turns per episode | Calls per tool turn | Input tokens per episode | Agent's cost per episode | Wall-clock |
|---|---|---|---|---|---|---|---|
| baseline | 0.691 | 0.464 | 8.81 | 1.42 | 79,376 | $0.039 | 73 s |
| reach | 0.738 | 0.536 | 6.83 | 1.18 | 63,048 | $0.036 | 69 s |
| batch | 0.714 | 0.571 | 8.29 | 1.49 | 74,193 | $0.037 | 67 s |
| batch-reach | 0.774 | 0.607 | 6.86 | 1.14 | 65,405 | $0.036 | 70 s |
| Without → with | LLM turns | Change | Tasks fewer / more / same | Input tokens | Agent's cost | Passed |
|---|---|---|---|---|---|---|
| baseline → reach | 740 → 574 | −22.4% [−29.4, −15.8] | 25 / 1 / 2 | −20.6% | −8.7% | 58 → 62 |
| baseline → batch | 740 → 696 | −5.9% [−10.9, −0.9] | 15 / 9 / 4 (p = 0.31) | −6.5% | −5.9% | 58 → 60 |
| batch → batch-reach | 696 → 576 | −17.2% [−24.0, −10.2] | 21 / 6 / 1 (p = 0.006) | −11.8% | −2.1% | 60 → 65 |
| reach → batch-reach | 574 → 576 | +0.4% [−6.4, +7.3] | 11 / 12 / 5 | +3.7% | +0.9% | 62 → 65 |
| baseline → batch-reach | 740 → 576 | −22.2% [−29.6, −14.9] | 23 / 3 / 2 | −17.6% | −7.8% | 58 → 65 |
By domain, reach against baseline:
- Retail. −26.8% (−34.7% to −20.5%), fewer turns on all 20 tasks. Passes went from 43 to 47.
- Airline. −8.5% (−26.1% to +10.9%). Passes held at 15. With the prompt in both arms, the flow saved nothing in airline (+1.9%). There it made 22 detours in 76 lookups, and its lookups' results raised the agent's input by 13%.
GLM-5.3
GLM-5.3's batch arm is compared with the reach round's recorded arms on the same tasks. Those are the paired run's baseline (in airline, the mean of its two trials; passes from the first) and the live reach arm, as that round compared them (analyze_recorded.py, which reproduces its numbers). They were recorded one and two days earlier, so these comparisons are not concurrent.
| Without → with | LLM turns | Change | Tasks fewer / more / same | Input tokens | Passed |
|---|---|---|---|---|---|
| baseline → reach (the reach round) | 310.5 → 224 | −27.9% [−35.9, −19.1] | 23 / 4 / 1 | −21.9% | 24 → 21 |
| baseline → batch | 310.5 → 287 | −7.6% [−15.0, +1.8] | 13 / 9 / 6 (p = 0.52) | −3.4% | 24 → 22 |
| batch → reach | 287 → 224 | −22.0% (the prompt took 28.1% more turns, +14.2% to +46.3%) | 23 / 4 / 1 | 22 → 21 |
Deviations from the plan
- Sonnet 5's batch arms did not wait for Haiku. The plan had them run once the rest were done. They began at 18:45, when Sonnet's first two arms had finished, and ran alongside Haiku 4.5's under the lowest cap, 0.80 of the seven-day window, below Haiku's 0.82. That kept the plan's order of priority, and no cap was reached. As planned, they ran after Sonnet's other arms rather than among them, so
baseline → batchandreach → batch-reachcompare runs an hour apart.batch → batch-reach, likebaseline → reach, compares arms run side by side. - GLM-5.3's comparisons use the reach round's method (above) rather than
analyze_trials.py, which pairs episodes of the same trial only. The statistics are the same. - After the plan:
analyze_trials.pylearned to read several output directories, to analyze Sonnet's two runs as one design.run_trials.pylearned to set aside an attempt a stopped run left unfinished. No run stopped, so that never happened.- Neither change touches a statistic.
- The published episodes leave out the subscription's usage windows, which the agent's stream reports with each request.
Caveats
- The customer is a model. Claude Haiku 4.5 played the customer in every Claude episode. In a plumbing episode before the design, it asked to match the wrong bottle's color and failed its task. Such errors fall on every arm alike, and they lower pass rates.
- One harness, one host. The agents ran in Claude Code's normal mode, as in the Claude round, which adds a short note to the context (the working directory, the model's name, the date). Claude Code ran both models with extended thinking: thinking made up 31% of Sonnet 5's output tokens and 64% of Haiku 4.5's. Every arm had both.
- List prices are not a bill. The subscription bills no tokens. The costs above are what Claude Code reports at list prices, with its prompt caching.
- Wall-clock time depends on the API's latency when each episode ran. A task's arms ran minutes apart to spread it, but not at once.
- Airline has 8 tasks. Its savings and pass rates carry wide intervals for every model.
- The GLM-5.3 comparisons are not concurrent.
Published
- frontier-2026-09-27.json holds every arm, every comparison and every episode's row, for each model, from
analyze_trials.py, and GLM-5.3's comparisons fromanalyze_recorded.py. - frontier-2026-09-27-episodes.tar.xz holds all 700 episodes, laid out as
run_trials.pywrites them:sonnet/andhaiku/,DOMAIN/trial-K/ARM/task-ID, with each run'sdesign.jsonand a ledger of each episode's cost and time;glm/DOMAIN/batch/task-ID.
Reproduce
The design, one model at a time, in τ²-bench's Python environment. A stopped run resumes where it left off.
cd pilot
python run_trials.py --model claude-sonnet-5 --out runs/frontier/sonnet
python run_trials.py --model claude-sonnet-5 --out runs/frontier/sonnet-batch --arms batch batch-reach
python run_trials.py --model claude-haiku-4-5-20251001 --out runs/frontier/haiku \
--arms baseline reach batch batch-reachOne episode, as the runner starts it (--arm reach and the flow for the reach arms, --batch-reads for the batch arms):
python run_episode.py --domain retail --task-id 38 --arm reach --batch-reads \
--flow ../docs/results/reach-2026-09-26-retail.flow.json --flow-oracle mock --flow-threshold 0.3 \
--agent-cli claude --model claude-sonnet-5 --customer-cli claude --customer-model claude-haiku-4-5-20251001 \
--out runs/one --label batch-reachThe tables, from the published episodes:
tar xf frontier-2026-09-27-episodes.tar.xz
python pilot/analyze_trials.py frontier-2026-09-27-episodes/sonnet --markdown
python pilot/analyze_trials.py frontier-2026-09-27-episodes/haiku --markdownGLM-5.3's comparisons, with the paired run's and the reach round's episodes unpacked beside them:
tar xf paired-2026-09-25-episodes.tar.gz && tar xf reach-2026-09-26-episodes.tar.gz
B=paired-2026-09-25-episodes R=reach-2026-09-26-episodes G=frontier-2026-09-27-episodes/glm
python pilot/analyze_recorded.py \
retail:baseline=$B/retail/baseline retail:reach=$R/retail/reach retail:batch=$G/retail/batch \
airline:baseline=$B/airline/baseline,$B/airline/trial-1/baseline airline:reach=$R/airline/reach \
airline:batch=$G/airline/batch