Results
Every result so far, in the order it was found. Each page says how to reproduce it. The JSON files beside the pages hold per-episode or per-decision rows. The answer bundles hold every System-One answer the pages read, so they replay without a key.
The working paper puts them together, and RFC-001 records what each round changed in the design, amendment by amendment. What is still open is in the roadmap.
2026-09-23: measuring before building
| Page | What it found |
|---|---|
| Phase 0 | About a quarter of LLM turns sit inside runs of tool calls a flow could make. Most write arguments appear verbatim in an earlier tool output or customer message, so a flow can bind them. A habit that sees only the tool sequence takes 6–8% of decisions at 80% confidence. Behavior transfers across models |
| Phase 0b (report, re-asked, answers) | Zero-shot, Jev agrees with the agent's next step 72% of the time and is well calibrated (ECE 0.06). Trusted at p ≥ 0.9, flows save only 3.9–4.7% of LLM turns, with a risky call in 6–13% of episodes |
| Phase 0b v2 (report, nine targets, sharpened, answers) | Read-only flows turn a wrong pick into a detour. Jev's answers, weighed with the habit and three predicates, agree 80.5% of the time in retail and 78.9% in airline. Taking the likeliest lookup saves 15–17% of turns in retail and 10–13% in airline |
2026-09-24: live pilots, the built system, and what a System-One model is for
| Page | What it found |
|---|---|
| v2 without a goal (report) | Nobody needs to name the goal: flows save 17.3% of retail turns and 12.6% of airline turns offline |
| Retail pilot | GLM-5.3 in Claude Code, ten paired tasks: 23.6% fewer LLM turns with the flow (2.6 per episode, 95% interval 1.0 to 4.2). 8 of 10 passed in each arm |
| Airline pilot | 17.3% fewer LLM turns (2.2 per episode, 0.5 to 3.9). 8 of 10 passed without the flow and 9 of 10 with it |
| Policy guards | On published airline runs, the enforced rules would refuse a write the tool accepted in 38% of failed episodes and 1% of successful ones |
| Guards pilot | On the airline tasks where those writes concentrate, GLM-5.3 made none: 4 of 4 passed in both arms |
| Flow audit | Run as a fugue program, the retail flow picks the agent's own step 85.9% of the time and the airline flow 64.5% |
| Arms C and D0 | A flow that hands every branch back saves 0–1.7% of turns. The habit alone saves as many turns as the arbiter, which buys fewer detours. Naming a flow could add at most 1.9% of turns in retail and 7.9% in airline |
| Habit-only pilot | On the retail pilot's ten tasks, the habit alone took 79 LLM turns, against 110 with no flow and 84 with D0, and passed 9 of 10 |
| Fewer traces | With three retail training tasks, the arbiter saved 20.4% of turns and the habit alone 1.8%. Its samples turned out to be clustered by task id (see its correction and the cold start) |
| Confirmation (report, labels) | Where Jev and the guards' word list disagree on whether the customer confirmed a write, hand labels side with Jev 30 times in 40 |
| Answers: arms, fewer traces, confirmation | The answer bundle for the three pages above |
| Cold start | From five of an agent's own random sessions, the habit alone saved 11.8–18.9% of retail turns. An arbiter fitted on those sessions trailed it until about twenty sessions. A clustered draw saved 2.7%, and another flow's arbiter lifted it to 11.1% |
| Options from the manifest | Offering every read-only tool gains nothing in retail and adds detours in airline. Only traces teach a flow how to bind a lookup's arguments |
| A second confirmation question (report, labels) | "Had the agent proposed this change?" catches calls that differ from what the customer agreed to. 16 of 30 writes it newly fails are real lapses |
| Matching descriptions to records (report) | From the customer's words alone, Jev picked the expected record 81.6% of the time, against 89.9% for the agents, so it is not a check on writes |
| Answers: cold start, manifest, confirmation, matching | The answer bundle for the four pages above |
| Shipped arbiters across domains | An arbiter fitted on one domain's published decisions, served with a habit from the other, does what that domain's own arbiter does, within about a point of turns saved |
| Answers: shipped arbiters | The answer bundle for the page above |
2026-09-25: the evidence a deployment needs
| Page | What it found |
|---|---|
| Prompt injection (rows) | Whatever Jev answers, a flow calls only the lookups compiled for the site. Inside that set, a note in a tool result steered the flow at 86 of 300 retail decisions; one sentence telling Jev that tool results are data cut it to 19, and a threshold of 0.5 as well to 4. An injected claim passed 9 of 140 writes the confirmation judge had refused |
| Answers: prompt injection | The answer bundle for the page above |
| Telecom (report, rows) | Telecom compresses like retail and airline, but Jev agrees with the agents less (67.9%). Neither shipped arbiter, nor one fitted on both domains, carries to it: they made 43–50% of the agent's lookups, against 66% for the habit alone and 75% for telecom's own |
| Answers: telecom | The answer bundle for the page above |
What stretto_commit could save | Bundling consecutive confirmed writes into one call saves at most 0.2–3.5% of LLM turns, too little for a ten-task live pilot to see |
| Shadow mode and per-site promotion (rows) | Promoted on half the test tasks and replayed on the other half, D0 kept 277 of its 294 retail turns saved and all 41 in airline, with detours down from 58 to 28 and from 26 to 11. The habit alone gains little: its detours are at sites that pass the bar |
| Answers: promotion | The answer bundle for the page above |
| Cold start, live (rows, episodes) | On the retail pilot's ten tasks, a habit learned from five of GLM-5.3's sessions took 79 LLM turns against 110 with no flow, making exactly the lookups the four-agent habit made. With the shipped airline arbiter it took 70, with 2 detours where the habit alone made 6. Both passed 8 of 10 |
| Answers: the cold start, live | The answer bundle for the page above |
| Cold start, live, in airline (episodes) | On the airline pilot's ten tasks, the habit from five of GLM-5.3's sessions took 102 LLM turns against 127 with no flow and 105 with D0, with 8 detours. The shipped retail arbiter cut them to 1, at no cost in turns |
| Answers: the cold start in airline | The answer bundle for the page above |
| The confirmation judge, enforced live (rows, episodes) | Enforced on 20 episodes, the judge refused none of 40 writes. Logged on 20 more, it would have refused 4 of 38: 2 real lapses and 2 false alarms. Passes did not move. The second question's flags were 7 false in 8, so the recommended setting enforces the first question and logs the second |
| Answers: the judge, live | The answer bundle for the page above |
| Counterfactual evaluation (rows, decisions) | From D0's exploring replays, the four estimators give D0 and the arbiter at 0.5 what their own replays do, within about a standard error. Lookups the agent made later carry over to any rule, and detours do not: two-thirds of a flow's lookups follow its own previous one, so every estimate put the habit alone level with D0 in airline, where its replay makes 50 detours to 26 |
| Answers: counterfactual evaluation | The answer bundle for the page above |
| Claude models, live (rows, episodes) | D0 saved Claude Haiku 4.5 18.6% of LLM turns on the retail pilot's ten tasks, where it saved GLM-5.3 23.6%, and Claude Sonnet 5 24.1% on three. With Claude Sonnet 5 as the customer to GLM-5.3, the savings held: 23.1% against 25.8% on five tasks |
| Answers: Claude models | The answer bundle for the page above |
| Predicate refinement, in airline (rows, candidates, examples) | Proposed where the arbiter is weakest, two of eight candidate predicates raised the held-out likelihood of the four 2025 agents' steps (0.654 to 0.624 nats per decision), but not GLM-5's, which the proposer never saw, and the savings did not move (12.6% and 12.5% of turns). The three hand-written predicates stay |
| Answers: predicate refinement | The answer bundle for the page above |
| Flow search (rows, flows: airline, retail) | NSGA-II over each site's threshold and the decider, replayed on GLM-5's training-task episodes. On its test-task episodes, the search's front held flows with far fewer detours than D0: 6 against 26 in airline, saving 48 turns against 41, and 30 against 58 in retail, saving 293 against 294. Each gain came from one or two sites, and no single threshold did as well |
| Answers: flow search | The answer bundle for the page above |
| Paired run on every test task (rows, episodes) | GLM-5.3 on all 40 retail and 20 airline test tasks (airline twice), 80 pairs. With the flow it took 25.5% fewer LLM turns (95% interval 20.5% to 30.4%). 70 pairs passed with the flow and 71 without it: −1.25 points (−7.5 to +6.25), too wide to rule out a loss of a few points. 242 of the flow's 259 lookups were the agent's own |
2026-09-26: what a workflow compiled once could take
| Page | What it found |
|---|---|
| The anatomy of τ²-bench episodes (rows) | On ten agents' retail and airline episodes, replies and the calls that answer the customer are 60–75% of LLM turns. What a workflow compiled once from traces could take is the read skeleton, 19–32% of retail turns and 10–31% of airline turns, which is what the read-only flow already does. After a tool returns, the previous tool alone predicts the next step 50–84% of the time; structured features add up to 11 points and the goal at most 4 more. Rules at least 95% sure on training cover 2–44% of retail decisions and 0–16% of airline decisions, and hold 84–100% on held-out tasks. The literature review behind the question: compiling workflows from agent traces |
| A record the customer did not ask about | Nearly every detour is the right lookup of the wrong record, and in airline most came after the customer had named a different one. Scoring the binding there by how often the agent went on to read an unnamed record cut detours on GLM-5's test episodes from 56 to 4 in airline, at no cost in turns, and from 63 to 38 in retail for 4 turns. Compiled into D0's data, the habit alone matched the flow search's airline detours (6) with 7 more turns saved, and no search |
| Checking a write against the proposal (rows) | A check with no model flags a write when the confirmation chose another record of its list: another order, item or card. It flags 3% of the writes of successful episodes and, in retail, three times as many in failed ones, and half of the labelled wrong-record lapses. Useful to log, not to enforce: stretto-proxy logs it beside the confirmation judge |
| Learning from the sessions a flow served | Served sessions, the flow's lookups kept as the agent's calls, teach a flow as much as clean ones; deleting the flow's calls switches it off. Relearned round after round on its own sessions, a flow drifted to 65% more detours in retail, until the named-other score; with it, it holds. Pruning lookups nothing used, or scoring bindings at the agent's own calls only, cut savings by up to three quarters, and neither shipped |
| Where compiling once works: telecom (rows) | In telecom, where the agent troubleshoots a phone by a procedure, a decision tree over every tool result so far predicts 63–84% of the nine leaderboard agents' next steps after a tool, 2–9 points above one feature, and rules sure at 95% cover 13–62% of decisions at 96–99.8% held out; in retail and airline the same tree adds nothing. With the agent holding the phone (τ²'s solo mode), reads after a tool are 52–60% of turns, and under the written workflow GPT-4.1's steps are 80% predicted with sure rules covering 47% at 98%, against 74% and 29% under the manual |
| What is left of the detours | An agent D0 never trained on (Claude Sonnet 4.5) still sees 63 detours in airline. Handing the session back after a surprise leaves 57 even with an oracle, because they come in one burst per session; stopping chains by their joint probability, as speculative decoding's draft trees do, would drop the lookups iterating agents use most; a flow learned from Sonnet's own traces, or D0's arbiter with Jev, does no better. What is left is a semantic stop: Sonnet stops reading at the reservation the customer described and GLM-5 reads on. Where the description is an exact value, such as a ticket's phone number, the binding now counts it (described_read); a city for an airport code is still language |
| Answers: detours | The answer bundle for the page above |
| Answers: a question for the stop | The answer bundle for the page's last section |
| stretto's flow in telecom | Replayed for the first time on telecom, a flow learned from τ²-bench's 2025 runs saved 1.9% of the nine leaderboard agents' turns with 2,936 detours: its binding took the plan id of the line just read where the agents went on through the customer's lines. Ordering each argument's sources as the agent used them at the site took it to 12.8% of turns saved with 443 detours; retail and airline replay exactly as before. Of those detours, 145 are bills four agents read too, with a limit that returns the same bills (13.3% and 298 scored by result). Where the agent holds the phone (τ²'s solo mode), the flow saved 27.5–33.5% of turns, with half the detours once its binding counted where the agents stop reading lines: at the one with the ticket's number (bindings.described_read) |
| Answers: the telecom flow | The answer bundle for the page above |
| A workflow compiled once, run with no model (runs, self-training, binding) | In τ²-bench telecom's solo mode, a decision tree fitted once on four runs' successful training episodes, run with no model in τ²-bench's environment, passed 31 of the 40 held-out tasks (77.5%; 25–31 refitted on resamples). The LLM agents whose traces it learned from passed 49–78% of them, with 15–18 LLM turns per episode. It needs a guard against repeats (22 of 40 without) and many consistent demonstrations (18 of 40 from half the tasks). Its own confidence is no gate, since it measures agreement among demonstrators, but its check of the ticket's stated outcome caught 8 of its 9 failures: handing only those 8 to an agent projects 87.5–89.4%, with a model in one episode in five. Learning from its own tries, kept where that check passed, took a workflow fitted on half the demonstrations from 18 to 22–23 of 40; kept with no retries, its own runs taught it nothing. Learned as where they came from, its identifiers bound for a customer no trace saw: renamed throughout, it passed the same 35 of 40 as for the original customer, against 16 with constants; moved to another customer's account, 35 against 19; on a five-line account, the original customer's 14 of the 16 tasks whose faults survive the move |
| The probability that matters, and a procedure stretto runs (rows, episodes) | A read-only flow should make a lookup when its chance of use before the next write clears δ/(β+δ), a detour's cost against a saved turn's value. The habit's next-step chance understates that chance: lookups it scored 0.4–0.6 were used 94–97% of the time. Counting the right event (--decider reach) lowers the calibration error in every domain, and acting on it at once loses little: 94% of the lookups that paid were used before the flow's next decision. Both costs can be counted in recorded episodes (scripts/costs.py): θ* is 0.30 in retail, what the live run measured, and 0.12–0.13 in airline and telecom. Across nine agents it never saw, the reach flow took 86.4% of retail's read-only ceiling at 0.3, against 76.2% for the next-step flow, and saved more tokens than it in every domain at that domain's costs. Live, with no model, it cut GLM-5.3's LLM turns by 27.9% on 28 retail and airline tasks and passed 21 against the baseline's 24; in the four tasks lost, its lookups returned what the agent's own reads had. Ten of an agent's own sessions give 93–96% of what all of them save in retail and airline; in telecom a hundred of its own save more than 1,184 of other agents' (15.2% against 12.6% of turns). Other agents' sessions are best weighed as a prior of a hundred sessions: added in full they drown out the agent's own, while a hundred of them, drawn at random, plus the agent's own save about as much as the better source at every n, and more than either at ten sessions in telecom. The compiled telecom procedure now runs in stretto (stretto-procedure), matching the script on all 40 held-out tasks, and with GLM-5.3 on its hand-backs the cascade passes 39 of 40 at 0.48 LLM turns per ticket |
2026-09-27: beyond τ²-bench
| Page | What it found |
|---|---|
| What the environment decides, on seven benchmarks (rows, flows) | Replayed from the published trajectories of 89 more agents on τ-bench, BFCL, AgentDojo, WorkBench, DTap-Bench and MCPMark's real MCP servers, with no environment. The read-only ceiling is the domain's: 3.5% of turns where each request names what to read (WorkBench), 47% where one listing names what every later read takes (AgentDojo travel), 28–29% in τ-bench as in τ²-bench. Deciding on the chance of use before the next write saves more where agents read ahead (τ-bench retail +1.4 points, BFCL twice next-step's) and ties elsewhere. New benchmarks taught the binder to read results as data, to pass lists, and never to make a search the agent always narrows without its filter. On DTap-Bench's six domains a flow carries across three agent SDKs at about three fifths of the agent's own savings; its medical domain, where the agents order tests and question a simulated patient, has nothing to take. Its legal, finance and research domains, replayed separately, agree: reach leads in legal, whose agents walk dockets and opinions (+0.8 points with each agent's own runs), and ties in finance and research. Of what reading the user's words would add, a model that picks the value among them takes 30–53%, a pattern 0–17%. Agents repeat 0.8–2.6% of their reads with no write between, so replayed savings stand for most of them; and in the benchmarks without a simulated user, a procedure compiled once would need a model to read the request, and most often one to write. A threshold per decision moves τ²-bench's counted utility by −15% to +10%, and is not adopted. On τ-bench retail the flows answer 42–45% of the agents' API calls; Speculative Actions' model speculators predicted 22–38%. On MCPMark's Notion the flows learn the walk through a page's blocks but not which block comes next; learn --constants passes the page size GPT-5 and o3 always pass. Priced with output tokens and prompt caching, θ* falls, and the thresholds counted in input tokens keep at least 94% of the best utility. Every table is rebuilt by scripts/bench/ |
| Live on AgentDojo and BFCL (rows, episodes) | GLM-5.3 and Claude Haiku 4.5, each run live in the benchmark's own environment, once without and once with the published flow on each held-out task: 244 episodes. In AgentDojo's Slack and travel suites, where the replay found reads to take, the flow cut LLM turns by 10.1% (5.8% to 14.0%). 14 pairs took fewer turns and one more, and passes were unchanged at 27. Over all 41 tasks it cut 6.0% (2.1% to 9.9%), where the replay projected 7.2%. BFCL, projected at 1.6%, shows no effect at 20 tasks per model. In Slack, the savings and detours came from one site; promotion on the agents' own sessions removed the detours and the savings with them, since only the request tells the two kinds of task apart. Cost is not resolved: run-to-run variation is as large as any difference |
| Frontier models and a prompting baseline, live (plan, rows, episodes) | Pre-registered, three trials of the reach round's 28 tasks: the flow cut Claude Sonnet 5's LLM turns by 20.5% (16.5–24.4%) and Claude Haiku 4.5's by 22.4% (15.8–29.4%), with passes 65 → 69 and 58 → 62 of 84. Anthropic's sample prompt for parallel tool calls saved them 3.4% and 5.9%, and GLM-5.3 7.6%; with the prompt in both arms, the flow still saved 22.9% and 17.2%. The agent's cost at list prices fell 11.8% and 8.7% |
| Claims and their evidence | Each number in the paper's abstract and contributions, whether it rests on live runs, replays against τ²-bench's environment or the published record, and the script and page that recompute it |
Recorded episodes
Every live episode behind the pilot pages, 76 in all, is published as one archive per pilot: the conversations, τ²-bench's simulation files, the proxy's session logs and the flows' decisions. The paired run's 160 episodes are in one more archive, laid out the same way (the page). Later rounds publish their episodes with their pages, laid out the same way: the cold start, live, in airline, the judge, live, Claude models, and the reach arm and the cascade. pilot/rescore.py reproduces every recorded reward from them.