Compile What the Environment Decides: Read-Only Speculation and Compiled Procedures for LLM Agents
Draft, 2026-09-27.
Abstract
An LLM agent pays a model turn for every decision, yet many of its decisions are fixed by what its tools returned, not by what the user said. We make this precise and use it in two regimes. While a user is present, a program learned from traces can take over only reads, speculatively. Because reads leave the state unchanged until the next write, the decision to speculate separates over reads: make each read whose probability of use before the next write exceeds
1 Introduction
Tool-using agents spend most of their cost in the LLM turns that decide the next call. Work that reuses traces to cut that cost either keeps the model in the loop (workflow memory [AWM], induced skills [ASI, SkillWeaver]) or compiles traces into programs that run without it [TraceCompiler, Compiled AI, PreAct]; work that speculates on the next action overlaps it with the model's own latency [Speculative Actions, Speculative Macro Commit]. Two questions sit under all of these: which of an agent's decisions a program learned from traces can take over at all, and when acting on a prediction pays.
We answer both with one distinction. Each decision of an agent is informed by the state the tools have returned (what is in the account, what the phone reports) and by what the user said (which order, which fault). A program learned from traces sees the first; the second needs language. A compiler can take the decisions the environment determines, and no more, and two regimes follow:
- A user is present (τ²-bench's retail, airline and dual-control telecom). Replies, and decisions the user's words trigger, stay with the model. What remains are reads whose arguments the tool state supplies. Reads are safe to make speculatively: they cannot change the state, and a wrong one costs a lookup.
- No user speaks (telecom's solo mode, where the agent operates the phone itself). Every branch follows a tool result, so the whole procedure, writes included, can be compiled once, and handed to a model only when the outcome the ticket states does not hold.
Contributions:
- What can be compiled, measured. Every LLM turn of nine frontier agents on τ²-bench, classed by what decided it. A read-only speculator can save at most 29% of turns while a user is present, and reading the user's words would add less than a point. On six more benchmarks and 89 more agents, the ceiling runs from 3.5% of turns, where each request names what to read, to 47%, where one result names what every later read takes: it belongs to the domain, and is set by where arguments come from. Where the request names what to read, a small model that picks the value from the request takes a third to a half of what reading it would add, and a pattern at most a fifth (§4.1).
- The probability that matters. Because reads commute with each other and with the state until the next write, the decision to speculate separates over reads: make each read whose probability of use before the next write clears
, for a detour's cost and a saved turn's value (Proposition 3), and make it at once unless what the agent does before the speculator's next decision would show whether it will be used (Proposition 4). That probability, not the next-step probability that speculative-action systems rank by, is estimated by counting the same contexts for a different event. It is calibrated where the next-step estimate is not, the threshold follows from costs counted in the same traces, and the speculator it drives takes 86% of retail's ceiling across nine agents it never saw, 10 points more than the next-step speculator, and, with no model of its own, cut live agents' LLM turns by 21–28% on τ²-bench (GLM-5.3, Claude Sonnet 5 and Claude Haiku 4.5), where a prompt asking for parallel tool calls saved 3–8%, and by 10% in AgentDojo's suites with reads to take (§4.2–4.4). - Compile once where no user speaks. A procedure compiled from traces passes 35 of 40 held-out solo telecom tasks with no model and hands back exactly what its own check of the ticket's outcome cannot confirm; with a model on those, the cascade passes 39 of 40 at 0.48 LLM turns per ticket, against 15.9 for the model alone (§4.5).
Both run in stretto: an MCP proxy that serves the speculator between any agent and any MCP server, and a runtime that executes a compiled procedure against one. And because the replay rule counts only what a recorded call answered, a speculator can be replayed from a benchmark's published trajectories alone, with no environment: every benchmark that publishes its agents' runs is a test set, and we use six more (§3).
2 Model
2.1 Episodes, reads and writes
An episode is a sequence of LLM turns. In turn
A turn is decided by the environment when the agent's choice depends on
2.2 Speculation and its replay semantics
A speculator
Proposition 1 (safety). If
So a speculator can change an episode's outcome only through the agent's context, never through the environment. Its worst case is a detour.
Proposition 2 (ceiling). A speculator whose arguments are bound from the tool state, and whose reads answer only calls with no write in between, saves at most the turns whose calls are all reads with every argument present in the tool state after a tool response since the last write.
We measure this ceiling on recorded episodes (§4.1). A read can also answer across a write that left its result unchanged, so a replay may save turns outside it: at least 1.5% of the turns the use-before-write speculator saves in §4.2 are, and 0.5% of the next-step speculator's.
2.3 The optimal speculator decomposes
At a decision point with state
Proposition 3 (decomposition). For a set
and it is maximized by
Proposition 3 takes the decision as final until the next write. The speculator decides again after each tool response the agent receives, so a read that clears the threshold now could instead wait for a later decision, which will know more. Waiting forfeits only the uses that come before that decision: let
Proposition 4 (when to act). Let
Three consequences follow. First, the next-step probability is the wrong quantity under either rule. A reply to the user is a step of the agent's but not a decision of the speculator's, since it returns no tool result, so a read the agent makes after replying is one the speculator must make before the reply:
The contrast with speculative decoding is the point. A drafted token is wasted when any token before it is rejected, because a sequence must match as a whole, so draft trees rank tokens by the product of confidences along their path [EAGLE-2, SpecDec++]. Reads form a set: each is used or not on its own, whenever the agent gets to it, and the path product is the wrong criterion.
2.4 Estimating the probability of use by counting
We factor a candidate read's probability of use into its tool and its arguments,
The tool. Each step of an episode is abstracted to its tool, its outcome (returned, failed) and a feature of its result;
whose concentration
The
The arguments. For each argument of each read tool, the traces show where the agent took its value: the path of an earlier result, in the order the agent used such sources at this site. The binding takes the value at the first source that holds one. Its chance
counted separately where the user had named another record of the list, and where the record the user described had already been read. These counts cover the case that dominates the errors: the right tool on the wrong record.
The speculator makes the reads that clear
Algorithm 1. The speculator, after each of the agent's tool responses.
input: tool state x, abstract history h, threshold θ
loop
C ← { (t, a) : t a read, a = bind(t, x) defined } # bind skips values an earlier call of t passed
if C = ∅: stop
(t*, a*) ← argmax over C of r(t | h) · ρ(a | t, x)
if r(t* | h) · ρ(a* | t*, x) < θ: stop
y ← call t*(a*); append (t*, a*, y) to the response; x ← x + (t*, a*, y); h ← h + abstract(t*, y)2.5 Compile once where no user speaks
Where the environment decides every branch, we compile the policy itself. At each site (the tool that just returned, or the start), a decision tree over the tool state's features predicts the next call, as decision mining learns the branches of a process from its logs [Decision mining] (ID3, depth by cross-validation over tasks). Its actions are whole calls, whose identifiers are learned as provenance: the path of the result that held the value (with, in a list of records, the fields that set the chosen record apart), or the shape it had in the task's ticket. They are bound again at run time from that run's results. A guard forbids repeating a read before a write, or a write. The run stops when the tree says so, and then checks the ticket's stated outcome with a read. It hands the task to a model only when that check fails, with its calls and their results as the model's context: a cascade, in which, as in model cascades [Cascades], the expensive model sees only what a cheap check cannot confirm. The procedure is data, a file of trees, bindings and checks, which a small runtime executes against any MCP server.
2.6 The system
stretto implements both regimes for any agent and any MCP server. stretto-proxy sits between the agent's host and the server as a stdio MCP server, forwards every message, records sessions, and, after each of the agent's calls, runs the speculator: the lookups it makes ride in the same tool result, so the agent sees them with no new tool and no change to its prompt. The speculator is a file, the flow IR: the counts of §2.4, the lookups each site may make, and where each argument comes from. stretto learn writes it from recorded sessions, and since every estimate is a posterior predictive of conjugate counts, learning from another session adds its counts; the flow can be reviewed as a diff, shadowed, and promoted site by site. Its run after each call is a probabilistic program [fugue] whose three interpreters simulate it, execute it against the server, and score a recorded session under it. A compiled procedure is also a file, the procedure IR, and stretto-procedure runs it against a server and returns its calls and verdict. Nothing in either file is a model weight: a flow or a procedure is data a person can read.
Serving is guarded where a deployment needs it, from tools a flow may call to shadow mode and pseudonymized sessions (Appendix D). Deciding takes under a millisecond per tool response (§4.4).
flowchart LR
subgraph present["A user is present"]
H["agent host (LLM)"] <-->|"MCP"| P["stretto-proxy<br/>speculator"]
P <-->|"MCP"| S1["MCP server"]
P -->|"records"| R[("sessions")]
R -->|"stretto learn: counts"| F[/"flow IR"/]
F --> P
end
subgraph alone["No user speaks"]
T["ticket"] --> Q["stretto-procedure"]
C[/"procedure IR"/] --> Q
Q <-->|"MCP"| S2["MCP server"]
Q -->|"hand_back: calls and results"| M["model"]
end
Figure 1. The two regimes in stretto. Left: the proxy forwards every message, records sessions, and after each of the agent's calls makes the lookups the flow decides on, inside the same tool result. Right: a compiled procedure runs with no model and hands a ticket to one only when its check of the stated outcome fails.
3 Setup
Tasks. τ²-bench [Barres et al. 2025] has three customer-service domains with a simulated user: retail (114 tasks), airline (50) and telecom (114), where in telecom the user also operates a phone, and a solo mode of telecom in which the agent operates it and no user speaks. We use its train/test split (74/40, 30/20, 74/40) throughout: nothing is learned from a test task.
Agents. For replay, the recorded test episodes, four trials each, of the nine agents on τ²-bench's leaderboard with trajectories published in all three domains: GLM-5, Qwen3.5-397B, Qwen3-Max, GPT-5.2 at high and no reasoning, Claude Opus 4.5, Claude Sonnet 4.5, Gemini Pro and Gemini Flash. Speculators are learned from other agents: the training episodes of τ²-bench's four 2025 runs (Claude 3.7 Sonnet, GPT-4.1, GPT-4.1-mini, o4-mini), so every replay is of agents the speculator never saw. In solo telecom, the runs of GPT-4.1 and o4-mini, each under the written manual and the written workflow policy.
Replay. A replay re-executes each recorded episode's calls in order in τ²-bench's environment, with the speculator serving lookups after each call as it would live, and scores it by §2.2: a recorded call is skipped when a lookup answered it, a turn is saved when all its calls are, a lookup that answers none is a detour. A replay counts what the speculator would have spared an agent that otherwise acted as recorded (§6; details in Appendix D). Intervals are 95%, from a bootstrap over tasks, a task's four trials drawn together.
Other benchmarks. Six more benchmarks publish their agents' trajectories with every call's result: τ-bench [τ-bench] (GPT-4o and Claude 3.5 Sonnet in retail and airline, in τ-bench's own harness), the Berkeley Function Calling Leaderboard's multi-turn base tasks [BFCL] (the leaderboard's own logs of 18 models), AgentDojo [AgentDojo] (21 models' benign runs of its workspace, Slack, banking and travel suites), WorkBench [WorkBench] (24 models' 2026 runs of its 690 tasks in six domains [WorkBench Revisited]), DTap-Bench [DTap] (the benign tasks of six of its domains, customer service, CRM, telecom, travel, an operating system's files and medicine, 1,417 tasks run by seven agents built on the Claude Agent SDK, the OpenAI Agents SDK and Google's ADK, and three more, legal, finance and research, 560 tasks, replayed separately; its fourth harness, OpenClaw, logs no calls) and MCPMark [MCPMark] (17 models' runs of tasks on real MCP servers: a filesystem, PostgreSQL, GitHub and Notion). Each is rewritten in τ²-bench's format with its tools marked read or write, checked against the benchmark's source (DTap-Bench's and MCPMark's by their verbs), and its results rendered as JSON; 40% of tasks are held out. Speculators are learned from the older agents' training episodes and replayed on the newer agents' test episodes: 8 models to 10 in BFCL, 11 to 10 in AgentDojo, 10 to 14 in WorkBench, 8 to 9 in MCPMark. On τ-bench the τ²-bench speculators are replayed as they are. DTap-Bench's agents are contemporaries, so each agent's test episodes are replayed with speculators learned from its own training episodes, from its harness's other agents, and from the other harnesses' agents.
Replay from the record. These benchmarks cannot be re-run cheaply, so their replays answer every call from the recorded trajectory. A lookup that the agent's own later call answers, before its next write, gets that call's result, which is exact; any other lookup is a detour, gets a stand-in result, and never counts as a saving. On τ²-bench's nine agents, this counts 96.7% of the turns the environment replay saves and 91% of the use-before-write speculator's lead over the next-step speculator, but only 55–72% of the detours, which are therefore lower bounds (Appendix D).
Learning curves. A deployment learns from the sessions it serves. For three agents (GLM-5, Claude Sonnet 4.5 and Qwen3.5-397B), we order the agent's training episodes at random, three orders each, learn a speculator from the first
Live. GLM-5.3 runs in Claude Code with τ²-bench's system prompt and no built-in tools; τ²-bench's tools are served over MCP behind stretto-proxy, which records the session and runs the speculator. The user is GLM-5.3 prompted as τ²-bench's user simulator. Claude Sonnet 5 and Claude Haiku 4.5 run the same way, with Claude Haiku 4.5 as the user, for three trials of each task and arm, under a plan published before the first episode. Rewards are τ²-bench's own checks of the final database, and in solo telecom of the environment and the required actions.
Costs. A detour's cost
4 Results
4.1 What decides an agent's turns
Over 39,298 LLM turns of nine agents' test episodes (§3), 46.8% reply to the user, 12.2% write and 41.0% only read (Table 1). A speculator that binds arguments from the tool state could have saved at most the reads whose every argument sits in an earlier result, after a tool response since the last write: 29.0% of all turns (retail 34.2%, airline 29.5%, telecom 24.7%), from 16% to 41% by agent and domain. Allowing an argument to be a value the user wrote raises the ceiling by 0.7 points. In 84% of the turns in the ceiling the agent made one call (78% in airline, 87% in telecom); the rest are parallel reads. The rest of the turns, seven in ten, stay with the model: they reply, write, or read what only the user's words can name.
Table 1. LLM turns of nine agents' test episodes by what they do, and the read-only ceiling (Proposition 2). Pooled over agents, with the range across agents in brackets.
| Turns | Replies | Writes | Reads | Ceiling | |
|---|---|---|---|---|---|
| Retail | 14,106 | 39.6% [34–44] | 14.0% [12–17] | 46.3% [42–53] | 34.2% [29–41] |
| Airline | 7,530 | 40.3% [29–54] | 12.7% [11–18] | 46.9% [35–60] | 29.5% [16–40] |
| Telecom | 17,662 | 55.2% [41–60] | 10.6% [8–12] | 34.3% [30–46] | 24.7% [17–38] |
| All | 39,298 | 46.8% | 12.2% | 41.0% | 29.0% |
The ceiling is the domain's. Classed the same way, τ-bench's turns divide as τ²-bench's do to within two points (Table 1b), though its agents are a year older and ran in another harness. Across benchmarks, reads are about a third to three quarters of turns, and what varies is where their arguments come from. In τ-bench, as in τ²-bench, the user names themselves and the agent walks the records, and 28% of turns are in the ceiling. In AgentDojo's travel suite, a city's listing names the hotels that every later lookup takes, and 47% are. In WorkBench, each request names its customer, task or date, and only 3.5% are; on MCPMark's real servers, where agents write their own SQL and choose which files and issues to read, 7.0%. In DTap-Bench's browser and macOS domains, whose calls act on a screen and return the one they leave, almost none are: at most 0.5% of the frontier agents' turns, with a screenshot counted as a read (all 64 tasks, too few to learn from). In its legal, finance and research domains, reads are 39–77% of turns and the ceiling 1.4–14.9%: legal's agents walk dockets and opinions by the ids results name, finance's requests name the tickers, and research's agents write their own queries. Reading the user's words would add 3–7 points in BFCL, AgentDojo, WorkBench and DTap-Bench, against under two in τ²-bench and τ-bench, whose users give their details over the conversation. A speculator that binds from results pays where agents walk records; where requests name what to read, the lever is language, not a pattern: taking the request's only email, id, number or date recovers 0–4% of those points in WorkBench, BFCL and AgentDojo and 17% in DTap-Bench, since the rest are names and free text. A small model that picks the value among the request's words, shown only the call's tool and argument, recovers 39–53% in the first three and 30% in DTap-Bench, whose travel requests name several cities that the agent looks up with one tool, one after another (Appendix D).
The agent is part of the ceiling too. In DTap-Bench's customer service and travel, Gemini 3 Pro on Google's ADK batches its reads into a few parallel turns and leaves 7.6% and 0.5% of its turns in the ceiling, against 13–34% for the other agents. In DTap-Bench's medicine, where the calls order tests and question a simulated patient, six agents make 0–8 reads in 642 runs; gpt-oss-120b makes 4,056, asking for patients' status by ids it made up, and supplies nearly all of the domain's 3.8% ceiling. On MCPMark's Notion, which an agent walks block by block by the ids each listing returns, 22–34% of the Claude models' and Gemini 2.5 Pro's turns are in the ceiling, and at most 4% of the other models'. GPT-5 and o3 pass a page size with every read of a block, a constant that a speculator binding from results does not supply; counted in, constants would raise Notion's ceiling from 7.9% to 13.2%. So is the format: read as AgentDojo prints them, as YAML and Python literals, its results hold few values a parser can find, and its ceiling is 10.0%.
Table 1b. The same classes on six more benchmarks: the test episodes of the agents replayed in §4.2.
| Agents | Turns | Replies | Writes | Reads | Ceiling | With the user's words | |
|---|---|---|---|---|---|---|---|
| τ-bench (retail, airline) | 2 | 6,244 | 47.0% | 13.2% | 39.7% | 28.4% | 29.8% |
| BFCL multi-turn | 10 | 8,510 | 34.9% | 33.7% | 31.5% | 13.2% | 18.6% |
| AgentDojo (4 suites) | 10 | 1,523 | 26.9% | 16.5% | 56.5% | 22.1% [6.0–47.1] | 27.1% |
| WorkBench (6 domains) | 14 | 13,869 | 27.4% | 26.6% | 46.1% | 3.5% [0.0–6.4] | 10.9% |
| DTap-Bench (6 domains) | 7 | 40,357 | 12.0% | 59.9% | 28.1% | 10.7% [3.8–34.3] | 13.7% |
| MCPMark (4 servers) | 9 | 6,531 | 5.1% | 50.9% | 43.9% | 7.0% [2.5–10.9] | 10.6% |
Brackets give the range over suites or domains.
Figure 2. The read-only ceiling by benchmark and domain (test episodes of the replayed agents, §3), and on top what a speculator that also bound values from the user's words could add: in green the values a small model, shown the user's words and the call, picks among them (Appendix D), and in orange the rest. Where the agent walks records, results bind the arguments; where each request names what to read, only language does.
4.2 Speculation in replay
The next-step probability understates the probability of use. Replays log every lookup a speculator weighed with whether the agent made that call before its next write. After a user's details, GLM-5 and Claude Sonnet 4.5 read an order next with probability 0.64 under the habit, but read it before their next write 94% of the time; after an order, a product comes next with probability 0.08, but before the next write 38% of the time, once the agent has told the user what the order holds. Over all lookups the share used ran far above the habit's score: 0.94 of lookups scored 0.4–0.5 were used, 0.97 of those scored 0.5–0.6, and 0.48 of those scored 0.05–0.10 (GLM-5, retail). Counting the right event (§2.4) lowers the calibration error and the Brier score in every domain, each interval excluding zero, and ranks lookups better everywhere but telecom, where the two rank alike (Table 2).
Table 2. Calibration of the score
| Lookups weighed | ECE, next step → use before write | Brier | AUC | |
|---|---|---|---|---|
| Retail | 20,673 / 21,175 | 0.164 → 0.083 [−0.098, −0.047] | 0.142 → 0.100 [−0.055, −0.029] | 0.929 → 0.950 [+0.006, +0.037] |
| Airline | 17,124 / 17,174 | 0.106 → 0.068 [−0.042, −0.029] | 0.104 → 0.092 [−0.015, −0.009] | 0.821 → 0.845 [+0.002, +0.047] |
| Telecom | 33,844 / 34,976 | 0.069 → 0.049 [−0.025, −0.011] | 0.073 → 0.062 [−0.013, −0.008] | 0.925 → 0.930 [−0.007, +0.017] |
| Telecom, solo | 32,291 / 32,029 | 0.056 → 0.014 [−0.050, −0.023] | 0.080 → 0.074 [−0.009, −0.002] | 0.763 → 0.818 [+0.036, +0.071] |
Figure 3. Reliability of the two scores (bins of at least 100 lookups). The next-step score sits above the diagonal: lookups it rates 0.4–0.6 are used nine times in ten.
At the cost-derived threshold the right probability saves more. Across a sweep of thresholds (Figure 4), the speculator on the probability of use traces a frontier of turns saved against detours that reaches beyond the next-step speculator's. Priced at each domain's counted costs, its utility peaks where
Figure 4. Utility against the threshold, in millions of the agent's input tokens at its domain's counted costs, on the test episodes of GLM-5 and Claude Sonnet 4.5 (retail, airline, telecom) and GPT-5.2 without reasoning (telecom), for speculators learned from the 2025 runs; the dashed line is the domain's
Table 3. Nine agents' test episodes at θ = 0.3, pooled per domain: LLM turns saved, detours per episode, share of the read-only ceiling, and utility per episode in thousands of input tokens at the domain's own costs (§3), with 95% intervals from a bootstrap over tasks shared by both speculators.
| Speculator | Turns saved | Detours per episode | Share of ceiling | Utility per episode | |
|---|---|---|---|---|---|
| Retail | next step | 26.1% [24.0, 28.1] | 0.34 [0.23, 0.44] | 76.2% [71.7, 81.6] | 14.1 [12.6, 15.4] |
| use before write | 29.6% [27.2, 32.0] | 0.50 [0.36, 0.65] | 86.4% [80.2, 93.1] | 15.6 [14.1, 17.2] | |
| difference | +3.5 [2.3, 4.9] | +0.16 [0.06, 0.29] | +10.2 [6.7, 14.2] | +1.6 [0.8, 2.5] | |
| Airline | next step | 13.6% [10.4, 17.0] | 0.35 [0.22, 0.48] | 46.2% [38.3, 55.1] | 9.67 [7.37, 12.31] |
| use before write | 13.7% [10.6, 17.2] | 0.35 [0.22, 0.48] | 46.6% [38.5, 55.4] | 9.77 [7.44, 12.39] | |
| difference | +0.1 [0.0, 0.3] | 0.00 | +0.5 [0.0, 1.1] | +0.10 [0.00, 0.25] | |
| Telecom | next step | 11.8% [10.8, 12.9] | 0.31 [0.27, 0.35] | 47.9% [45.5, 50.3] | 12.6 [12.1, 13.1] |
| use before write | 12.8% [11.6, 14.0] | 0.66 [0.56, 0.76] | 51.7% [49.0, 54.4] | 13.2 [12.6, 13.7] | |
| difference | +0.9 [0.7, 1.2] | +0.35 [0.27, 0.44] | +3.8 [2.7, 4.8] | +0.60 [0.30, 0.91] | |
| Telecom, solo† | next step | 19.9% [18.4, 21.7] | 1.48 [1.27, 1.70] | 40.0% [36.9, 43.9] | 25.4 [23.7, 27.2] |
| use before write | 21.3% [19.4, 23.5] | 1.70 [1.50, 1.92] | 42.8% [39.0, 47.0] | 27.2 [25.1, 29.3] | |
| difference | +1.4 [0.5, 2.3] | +0.22 [0.09, 0.34] | +2.8 [1.1, 4.6] | +1.72 [0.56, 2.88] |
† Two agents, GPT-4.1 and o4-mini, the only published solo runs; the speculator is learned from their own training episodes.
In seconds and dollars. Six agents' episodes report each LLM turn's generation time and cost: Claude Opus and Sonnet 4.5, Gemini Pro and Flash, and GPT-5.2 at both settings. Priced there, the turns the use-before-write speculator saves in retail, 28% of all turns, are 24% of the episodes' generation time, 36 seconds an episode, and 18% of their cost net of its detours at the agent's input price; the next-step speculator's are 20% and 15%. In airline both save 7% of the time and 6.5% of the cost, and in telecom 12% of the cost. A turn that only reads generates less than one that replies, so the shares trail the share of turns (scripts/priced.py).
On other benchmarks. The use-before-write speculator saves more than the next-step one where agents read ahead of their next write, and ties elsewhere (Table 3b). On τ-bench, the speculators learned from τ²-bench's 2025 runs, replayed on 2024 agents in another harness, save 22.9% of retail's turns, 76% of its ceiling, 1.4 points more than the next-step speculator (0.8–2.1); the two tie in airline, as on τ²-bench. In retail they answer 42% of GPT-4o's API calls and 45% of Claude 3.5 Sonnet's with each call's exact arguments and result, making one lookup for about every two calls, three quarters of them used. On the same benchmark, Speculative Actions' model speculators predicted 22–38% of the calls with one to three guesses a step [Speculative Actions]; their agent differs and they guess the next step, writes included, so the comparison is indicative, not controlled. On BFCL, learned from eight older models and replayed on ten late-2025 ones, it saves about twice the next-step speculator's turns at every threshold from 0.1 to 0.5 (+0.85 points at 0.3, 0.35–1.45). Both take little of BFCL's ceiling, whose 200 tasks are 200 different requests over 128 tools. On AgentDojo the two tie (−0.5 points, −1.7 to 0.0). In Slack they make the same lookups and take 78% of its ceiling. On WorkBench, whose reads take what the request names, neither makes a single lookup: a speculator that cannot help stays out of the way, once it scores a lookup without arguments by how often the agent's own calls had none (Appendix D). On MCPMark's servers the speculators make 29 lookups in 396 episodes at
Across harnesses. DTap-Bench runs the same tasks under three agent SDKs. Over its six domains, the use-before-write speculator saves more than the next-step one whichever agents it learned from. Learned from each agent's own training episodes, it saves 3.4% of turns against 2.2% (+1.3 points, 1.1–1.5), 34% of the ceiling, for 0.23 detours per episode against 0.06; learned from the other harnesses' agents, 2.2% against 1.4% (+0.9, 0.7–1.1), for 0.16 detours. Medicine is half the turns and has nothing to take: there the agents order tests and question the patient, and the two speculators make the same lookups, all of them gpt-oss-120b's. Over the other five domains the lead is +2.6 points (2.2–3.0) and +1.7 (1.4–2.1). It leads in each domain with reads to take, least in CRM and the operating system's files, whose ceilings are 11.5% and 11.2% (+0.1 to +1.1 points), and with the agents' own sessions takes 86% of travel's ceiling. A speculator learned in other harnesses keeps 59% of what the agent's own saves, with 64% of its detours, and one learned from the agent's harness-mates keeps 79%. Each SDK runs one vendor's models here, so harness and model family go together: the GPT-5 models keep 90% from each other, and Claude Opus 4.6 keeps as much from the other harnesses' agents (62%) as from Claude Sonnet 4.5 in its own (61%). Learned from all seven agents, the speculator keeps 92% of the own speculators' turns with 91% of their detours.
Three more domains. Replayed separately, DTap-Bench's legal, finance and research domains show the same where there are reads to take. In legal, whose agents walk dockets and opinions by the ids results name, the use-before-write speculator saves 4.7% of turns against 3.9% with the agents' own sessions (+0.8 points, 0.7–1.0), 31% of the ceiling, and 4.3% against 4.0% learned in the other harnesses (+0.3, 0.2–0.4), though learned from all seven agents the two tie: its lead is in lookups used later rather than next, each agent's own way of reading ahead, which pooled counts average away. Weighed as a prior instead, the agent's own sessions plus 100 of the others' as in §4.3, it saves 5.6% against 4.7% (+0.9, 0.7–1.0), more than either source alone; over DTap-Bench's other domains with reads to take, the prior gains in CRM too and costs 0.5–0.8 points of the own sessions' savings in the other four. In finance and research, whose ceilings are 4.6% and 1.4%, the two make nearly the same lookups.
Table 3b. Replays from the record at θ = 0.3: turns saved and detours per episode (lower bounds), with 95% intervals for the difference from a bootstrap over tasks.
| Turns saved, use before write | Next step | Difference | Detours per episode, use before write / next step | |
|---|---|---|---|---|
| τ-bench retail | 22.9% | 21.5% | +1.4 [0.8, 2.1] | 0.92 / 0.85 |
| τ-bench airline | 17.1% | 17.1% | 0.0 | 0.41 / 0.41 |
| BFCL | 1.6% | 0.7% | +0.85 [0.35, 1.45] | 0.09 / 0.05 |
| AgentDojo | 7.2% | 7.7% | −0.5 [−1.7, 0.0] | 0.10 / 0.11 |
| WorkBench | 0.0% | 0.0% | 0.0 | 0.00 / 0.00 |
| MCPMark | 0.2% | 0.2% | 0.0 | 0.04 / 0.03 |
| MCPMark, own sessions | 0.4% | 0.4% | 0.0 | 0.72 / 0.65 |
| DTap-Bench, other harnesses | 2.2% | 1.4% | +0.9 [0.7, 1.1] | 0.16 / 0.08 |
| DTap-Bench, own sessions | 3.4% | 2.2% | +1.3 [1.1, 1.5] | 0.23 / 0.06 |
| DTap-Bench legal, finance and research, other harnesses | 2.8% | 2.7% | +0.17 [0.10, 0.25] | 0.08 / 0.08 |
| DTap-Bench legal, finance and research, own sessions | 3.3% | 2.7% | +0.54 [0.44, 0.64] | 0.05 / 0.03 |
Acting at once. Of the lookups the use-before-write speculator made at
4.3 Learning from few sessions
A deployment learns from the sessions it serves (§3). From ten of an agent's own sessions, the speculator saves 27.5% of its retail turns and 14.9% of its airline turns, 96% and 93% of what it saves from all of them; telecom needs about thirty (80% of the final savings at ten, 97% at thirty) (Table 4, Figure 5). Past that, more sessions buy fewer detours, not more saved turns: from 0.64 to 0.52 per episode in retail and from 0.74 to 0.40 in airline. Where agents act alike, other agents' sessions teach as much as the agent's own: in retail and airline the two protocols stay within two points of each other at every
Other agents' sessions are a prior. Since learning is counting, a deployment can start from other agents' sessions and add the agent's own as they arrive; what is left to choose is how much the others weigh. At full weight they drown the agent out: added to all 1,184 of the 2025 runs' sessions, a hundred of an agent's own barely move its telecom speculator (13.0% of turns, against 12.6% with none of them and 15.2% from its own alone; for GLM-5, 3.4% at every
Table 4. LLM turns saved · detours per episode for a speculator learned from the first
| Domain | Learned from | n = 10 | n = 30 | n = 100 | All |
|---|---|---|---|---|---|
| Retail | the agent's own sessions | 27.5% · 0.64 | 28.8% · 0.65 | 29.3% · 0.64 | 28.5% · 0.52 (296) |
| four other agents' sessions | 29.3% · 0.97 | 29.6% · 0.80 | 28.3% · 0.68 | 30.2% · 0.55 (1184) | |
| 100 of theirs, then its own | 29.9% · 0.70 | 30.1% · 0.69 | 30.5% · 0.69 | 30.7% · 0.60 | |
| Airline | the agent's own sessions | 14.9% · 0.74 | 15.5% · 0.37 | — | 16.1% · 0.40 (120) |
| four other agents' sessions | 14.7% · 0.93 | 15.2% · 0.63 | 15.1% · 0.48 | 15.7% · 0.40 (480) | |
| 100 of theirs, then its own | 15.2% · 0.31 | 15.4% · 0.31 | — | 15.8% · 0.33 | |
| Telecom | the agent's own sessions | 12.4% · 0.56 | 15.1% · 1.17 | 15.2% · 0.74 | 15.5% · 0.73 (296) |
| four other agents' sessions | 10.0% · 0.30 | 11.4% · 0.31 | 12.4% · 0.53 | 12.6% · 0.51 (1184) | |
| 100 of theirs, then its own | 13.3% · 0.57 | 14.6% · 0.72 | 15.0% · 0.90 | 15.4% · 0.86 |
Figure 5. Turns saved (top) and detours per episode (bottom) against the sessions learned from, on a log scale: mean over three agents and three orders. The dotted line learns from 100 of the other agents' sessions, drawn at random, and the agent's first
4.4 Live
Live, the speculator runs in stretto-proxy between the agent and the tools, over MCP: GLM-5.3 in Claude Code on τ²-bench's own prompts, with GLM-5.3 as the user. The earlier flow, which weighs the habit's next-step probability with an LLM's answers about the state [D0], cut LLM turns by 25.5% (95% CI 20.5–30.4%) over 80 paired retail and airline tasks, with 71 passed without it and 70 with it. The speculator of §2.4, deciding on the probability of use before the next write at
Table 5. Live, GLM-5.3 as agent and user, against the paired run's recorded baseline (airline: the mean of its two trials; passes: its first). 95% intervals from a bootstrap over tasks.
| Tasks | LLM turns, baseline → speculator | Change | Passed: baseline, D0, speculator | |
|---|---|---|---|---|
| Retail | 20 | 227 → 154 | −32.2% [−40.0, −23.5] | 17, 14, 16 |
| Airline | 8 | 83.5 → 70 | −16.2% [−37.2, +10.6] | 7, 7, 5 |
| Both | 28 | 310.5 → 224 | −27.9% [−35.9, −19.1] | 24, 21, 21 |
Five pairs disagree on the outcome: four passed only without the speculator and one only with it (McNemar
Frontier models, and a prompt. We then ran Claude Sonnet 5 and Claude Haiku 4.5 on the same 28 tasks, three trials of each task and arm, with Claude Haiku 4.5 as the user, under a plan published before the first episode (Table 5c). A fourth arm answers the obvious objection, that asking the agent to make its independent calls at once would do as well: it ends the system prompt with the sample prompt for parallel tool calls from Anthropic's prompting guide, word for word. The speculator cut Sonnet 5's turns by 20.5% (16.5–24.4%) and Haiku 4.5's by 22.4% (15.8–29.4%), fewer on 25 of the 28 tasks for each, with passes 65 → 69 and 58 → 62 of 84. The prompt saved 3.4% (−1.2 to 7.8%) and 5.9% (0.9–10.9%). These models already make most independent calls at once, 1.5 and 1.4 calls per tool turn, and the prompt barely moved that. GLM-5.3, which makes about one call per tool turn, took 7.6% fewer turns with it (−1.8 to 15.0%, against its recorded baseline), about a quarter of what the speculator saved it. With the prompt in both arms, the speculator still saved 22.9% (19.5–26.1%) and 17.2% (10.2–24.0%): its reads take their arguments from results the agent has not yet seen, which no prompt can batch. At list prices, with prompt caching, it cut the agent's cost by 11.8% and 8.7%, less than its input tokens (17.7% and 20.6%), since nine tenths of an agent's input is read from the cache at a tenth of the price. In airline, where these models make the most calls at once (Sonnet 5 1.8 per tool turn), the savings, 5.3% and 8.5%, were not distinguishable from none.
Table 5c. Live, Claude Sonnet 5 and Claude Haiku 4.5 as agents and Claude Haiku 4.5 as the user, on the 28 tasks of Table 5: three trials of each task and arm, 84 pairs per comparison, arms side by side (Sonnet's prompt arms an hour after its others). 95% intervals from a bootstrap over tasks.
| Claude Sonnet 5 | Claude Haiku 4.5 | |
|---|---|---|
| LLM turns, baseline → speculator | 833 → 662, −20.5% [−24.4, −16.5] | 740 → 574, −22.4% [−29.4, −15.8] |
| The prompt alone | −3.4% [−7.8, +1.2] | −5.9% [−10.9, −0.9] |
| The speculator, with the prompt in both arms | −22.9% [−26.1, −19.5] | −17.2% [−24.0, −10.2] |
| Passed of 84, baseline → speculator | 65 → 69 | 58 → 62 |
| pass^3, baseline → speculator | 0.679 → 0.679 | 0.464 → 0.536 |
| The agent's cost at list prices | −11.8% | −8.7% |
Beyond τ²-bench. The speculator also ran live in the environments of two of the replayed benchmarks, where no simulated user speaks: AgentDojo's four suites and BFCL's multi-turn tasks. Each task's tools were served over MCP by the benchmark's own code, and each episode was scored by the benchmark's own check (Table 5b).
The agents were GLM-5.3 and Claude Haiku 4.5, in Claude Code. The speculators were the published ones, learned from 11 and 8 older agents' runs of the training tasks, at θ = 0.3. Each held-out task ran once without the speculator and once with it.
The replay had found reads to take in AgentDojo's Slack and travel suites and none in banking or workspace, which fixed the comparison in advance.
- In Slack and travel, the two agents took 10.1% fewer LLM turns (5.8–14.0%). 14 of 34 pairs took fewer turns and one took more, and both arms passed 27.
- In banking and workspace, the speculator made one lookup in 48 episodes.
- Over all 41 tasks it cut turns by 6.0% (2.1–9.9%); the replay had projected 7.2% for ten other agents.
- In BFCL, where the replay projected 1.6%, no effect shows at this size (+1.6%, −4.2 to +6.3). The 25 pairs in which the speculator made no lookup ran alike in both arms, yet differ by 3.9%. Resolving 1.6% would take about 900 pairs.
Of AgentDojo's 88 lookups, 36 were calls the agent also made in its episode without the speculator. In Slack, the savings and 41 of the 43 detours came from one site. After the agent lists the channels, the speculator reads every channel in turn, as the older agents did on the tasks that ask about all of them.
- On the two held-out tasks that do, the reads spared the agents that turn.
- On the five that did not, it read four channels for nothing.
- Promotion on each agent's own training sessions held the site back. That removed all 43 detours, and the saved turns with them: 72 turns against 75 without a speculator and 68 with it.
Which kind of task it is, the request says (§6).
Replayed from the record, the same agents' own no-flow episodes projected 11 saved turns, all in Slack and travel. Live, the speculator saved 10 in the pairs where the agent used one of its lookups. The replay's savings held on these agents. Its detours did not: the record found 3 of the 52. A lookup of a tool the agent never called in the episode has no recorded result to replay, so the record misses a walk through such tools whole.
Table 5b. Live in AgentDojo's and BFCL's own environments. GLM-5.3 and Claude Haiku 4.5 pooled, one run per arm, each task scored by the benchmark's check. 95% intervals from a bootstrap over tasks, with a task's pairs drawn together.
| Pairs | LLM turns, without → with | Change | Pairs fewer / more | Passed, without → with | |
|---|---|---|---|---|---|
| AgentDojo, Slack and travel | 34 | 138 → 124 | −10.1% [−14.0, −5.8] | 14 / 1 | 27 → 27 |
| AgentDojo, banking and workspace | 48 | 146 → 143 | −2.1% [−8.7, +3.6] | 5 / 5 | 41 → 42 |
| AgentDojo, all | 82 | 284 → 267 | −6.0% [−9.9, −2.1] | 19 / 6 | 68 → 69 |
| BFCL, 20 held-out tasks | 40 | 444 → 451 | +1.6% [−4.2, +6.3] | 7 / 14 | 29 → 29 |
4.5 Compiling where no user speaks
In telecom's solo mode the agent operates the phone itself, and every branch follows a tool result. The procedure compiled from the successful training episodes of four solo runs (GPT-4.1 and o4-mini, each under the written manual and the written workflow policy; 730 episodes of 74 tasks) passes 35 of the 40 held-out tasks with no model, where the agents it was compiled from pass 49–78% (Table 6). Its identifiers bind for customers no trace saw: renamed throughout the database it passes the same 35, and on another customer's account 35, where constants pass 16 and 19. Its own check of the ticket's stated outcome resolves 25 runs, transfers 11 to a person as the policy directs, and hands back 4, all four failures; it misses one failure, a run it judged resolved.
The cascade gives the handed-back tickets to a model, which takes over from the procedure's state with its calls and results as context. Live, with GLM-5.3 in Claude Code on τ²-bench's solo protocol, it resolved all four in both of two trials, in 4.75 LLM turns each. On the same four tickets from scratch it resolved three, in 12–19 turns, and on four other tickets three, in 14–21. The cascade passes 39 of 40 held-out tasks and makes 0.48 LLM turns per ticket, where GLM-5.3 alone made 15.9.
Table 6. Telecom solo, 40 held-out tasks. LLM turns per ticket for the agents are their recorded averages; for GLM-5.3 alone, over the eight tickets run.
| Passed | LLM turns per ticket | |
|---|---|---|
| Agents compiled from (4 runs, 4 trials) | 49–78% | 14.7–18.1 |
| GLM-5.3 alone (8 tickets) | 6 of 8 | 15.9 |
| Compiled procedure alone | 35 of 40 | 0 |
| Procedure, then GLM-5.3 on its 4 hand-backs (2 trials) | 39 of 40 | 0.48 |
The procedure is a file stretto runs: stretto-procedure executes it against any MCP server, and on all 40 held-out tasks it made exactly the calls of the reference implementation, with the same rewards and verdicts.
5 Related work
Reusing traces with the model in the loop. Workflow memory [AWM] and induced skills [ASI, WALT, SkillWeaver] give the agent text or callable routines mined from its own successes; they save 10–27% of steps on web tasks, and AWM's workflows offered as callable actions were used in 18.5% of tasks. Agentic plan caching reuses plan templates from earlier runs, matched by keywords and adapted to the new task by a small model, and halves the cost of several agent applications [APC]. They change what the model sees. We change neither the prompt nor the tool list: the speculator's reads arrive inside results the agent asked for.
Compiling traces into programs. Where a task recurs without branching, a compiled program removes the model almost entirely [LOOP, Compiled AI, PreAct]. TraceCompiler compiles multi-turn API traces into workflows with argument provenance, classes each binding as constant, user input, copy, transform or dynamic, and declines to compile an intent whose irreversible effect is under-determined; its authors report no rate at which intents compile [TraceCompiler]. Robotic process mining reached the same condition for automating a step, that each parameter is computable from earlier ones [Bosco et al., Leno et al.]. Our §2.1 names the condition these systems share, that the environment decides the step, and §4.1 measures it on every turn of nine agents: while a user is present, it holds for reads with copied arguments; where no user speaks, for whole procedures (§4.5). Our hand-back is on the ticket's stated outcome, not on a mismatch mid-run [PreAct] or a missing branch [TraceCompiler].
Speculating on the next action. Speculative Actions predicts 22–38% of τ-bench retail API calls with a smaller model and restricts speculation to actions that can be undone (counting answers 42–45% there, §4.2) [Speculative Actions]; Speculative Macro Commit executes macros mined from τ² traces and verifies their first call, since a library match alone reproduced the next steps 34.6% of the time [Speculative Macro Commit]; AutoTool executes a predicted call when a transition graph and a parameter filler agree, and cuts LLM calls by 15–25% [AutoTool]; AOSpec speculates on an action and on its observation together to hide serving latency [AOSpec]. PASTE pre-executes calls predicted from recurring patterns while the model generates, and cuts task time by 43.5% [PASTE]; a memory of past trajectories raises a speculator's next-action accuracy by 19–39% [Speculate with Memory]; an agent trained as its own speculator predicts its next call 61–66% of the time [Speculate While You Reason]; asynchronous I/O with speculative tool calls speeds agents up 1.3–2.2× [Speculative Interaction Agents]. These hide latency: the agent still makes each call, and its result is ready sooner. Our speculator spares the turn, since the agent sees the result before it would ask for it. LLMCompiler has the model plan a graph of calls and runs the independent ones in parallel [LLMCompiler]; the model still chooses every call. Asked by prompt to make its independent calls at once, a Claude agent saved 3–6% of its turns live, and the speculator 17–23% on top of that (§4.4). The speculative systems rank candidates by the probability of the next action and verify sequences. Proposition 3 shows that for reads the right quantity is the probability of use before the next write, a first-passage probability that the next-step one understates (§4.2), and that the decision separates over reads, which lets a speculator act on each read alone; Proposition 4, that acting on it at once forgoes only what a later decision would have learned. Speculative decoding [Leviathan et al., EAGLE-2, SpecDec++] ranks draft tokens by path products because a draft must match as a sequence; reads do not. Binding a value the user wrote to a call's argument is the slot filling that dialogue state tracking does in task-oriented dialogue [TRADE]; §4.1 measures how much of it a small model does where requests name what to read.
Compiling the policy instead of the traces. On τ²-bench the largest gains come from compiling the written policy into a graph with an LLM inside each node [STAGE, PolicyGuide], mostly in telecom, whose policy is a troubleshooting flowchart. Our solo-telecom workflow is compiled from traces alone and runs no model inside; the written procedure given to an LLM agent (τ²-bench's workflow policy) reaches 78% on the same tasks, the compiled procedure 88%, and the cascade 98% (§4.5).
Counting. The habit is MacKay and Peto's hierarchical Dirichlet language model over abstract steps [MacKay & Peto 1995]; the reach estimator applies the same back-off to a Bernoulli event per read; bindings are Beta–binomial rates. All updates are conjugate, so learning from a new session is adding its counts, and other agents' sessions enter as a prior whose weight is a choice (§4.3).
6 Limitations
Benchmarks, not traffic. The evidence spans seven benchmarks: τ²-bench, where the speculator also ran against the environment and live, and six more replayed from 89 more agents' published trajectories, in more than twenty domains and three agent SDKs, with four real MCP servers among them. They agree on what sets the ceiling, where arguments come from: almost nothing where calls act on a screen or order tests, 47% of turns where one listing names every later read. And the speculator's lead holds where there is something to take. On real MCP servers there is little: agents write their own queries, and where they walk Notion's pages block by block, a speculator cannot tell which block comes next. Two limits remain. Every task set was written by a benchmark's authors for coverage, not sampled from a deployment's traffic: a deployment with a heavy head of common requests should compile more, and one with a long tail less. And the evidence is deepest on τ²-bench: only there did the speculator run against an environment with a simulated user, and live at scale. On AgentDojo and BFCL it ran live with two agents, once per arm (§4.4); elsewhere it is replayed from the record, which counts savings conservatively and detours from below, for agents that never ran with a speculator.
Replay assumes the agent skips what it already has. The replay counts a call as answered when a lookup made earlier returned the same result, and assumes the agent, seeing that result, does not make the call and otherwise acts as recorded. Live, GLM-5.3 made none of the speculator's 101 lookups again before the next write, and Claude Sonnet 5 and Claude Haiku 4.5 made 4 of 289 and 1 of 304 again. The replays of other agents and harnesses count what those agents would have skipped, not whether they would have skipped it. The record does show how often an agent makes again a read it already made, with no write between, as an agent that ignored a lookup's result would. Such repeats are 0.8% of the 58,129 reads of τ²-bench's nine agents and 0.9–2.6% of τ-bench's, BFCL's, AgentDojo's and WorkBench's; on DTap-Bench, six agents in three harnesses made 22 of 21,524. A few models repeat many, and their replayed savings are upper bounds: gpt-oss-120b 23% of its reads on DTap-Bench, Qwen3.5 Flash 17.5% on WorkBench, and o3, Grok 4 and Qwen3 Max 5.5–9.3% after the same read had succeeded on MCPMark's real servers, where the other six models made at most 1.7%.
Costs are averages, in tokens. θ* prices a domain's average detour against its average saved turn, in input tokens. A threshold per decision, from the result's size and the turns left, gains 7–10% of the counted utility in airline and telecom and loses 15% in retail (Appendix B), because it spends each score where it is priced, and scores are calibrated on average, not at every site; with each site recalibrated on the agent's own lookups, it stays within 5% of the average threshold. Priced in the agents' recorded seconds and dollars instead, a saved turn is worth somewhat less than the average turn (§4.2). Priced as providers bill, with output tokens and prompt caching, θ* falls, and the thresholds counted in input tokens still keep at least 94% of the best utility (Appendix B). Live, at list prices with prompt caching, the speculator cut the Claude agents' cost by 9–12% over three trials of τ²-bench tasks, less than their input tokens, since nine tenths of an agent's input is read from the cache (§4.4). On AgentDojo and BFCL, with one run per arm, cost varied between runs as much as between arms.
Pass rates are underpowered. The live runs bound turns, not success: a pass-rate interval of one point would need thousands of paired tasks. We report the pairs that disagree and why. With 84 pairs per Claude model, a harm of about ten points would show; none did.
Where no user speaks is τ²-bench's own variant. The compiled procedure was fitted on 730 successful episodes of one customer's base tasks, carried to renamed and moved customers, and hands back what its own check cannot confirm; its one missed failure shows that the check is only as good as the outcome the ticket states. Procedures need many consistent demonstrations: fitted on a quarter of the training tasks it passed 18 of 40. A compiled procedure is cloned behavior, whose errors compound once a run leaves the states its demonstrators visited [DAgger]; the outcome check bounds what that costs, by handing such runs back. In the benchmarks without a simulated user, a procedure compiled once would need a model to read the request, and most often one to write: 4% (MCPMark) to 37% (AgentDojo) of their episodes pass no value the agent composed, and at most 8% take none from the request either.
What a proxy sees. An MCP proxy sees the agent's calls and their results, not the user's words. Binding values the request names would add 3–7 points of turns in BFCL, AgentDojo, WorkBench and DTap-Bench (§4.1). It takes a host that shares the conversation with the speculator, and a model that reads it: a pattern recovers at most a fifth of those points, and a small model that picks the value among the request's words a third to a half. Live, AgentDojo's Slack suite shows the cost of not seeing it. One site's lookups paid on the tasks that ask about every channel and were detours on the rest. A bar per site, set on the agents' own sessions, could keep both or drop both (§4.4).
Simulated users. τ²-bench's users are LLMs, and so are the users of our live runs there. BFCL's, AgentDojo's, WorkBench's, DTap-Bench's and MCPMark's requests were written by their authors, but one request per task is not a user.
Availability. stretto's code, the flows every replay and live run served, the rows behind every table and figure, and the live episodes are published at https://github.com/alexnodeland/stretto; each number here is recomputed from them by scripts/bench/ and the commands on each round's results page, and docs/results/claims.md maps each headline number to its evidence.
References
- [τ²-bench] V. Barrès, H. Dong, S. Ray, X. Si, K. Narasimhan. τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment. 2025. arXiv:2506.07982.
- [τ-bench] S. Yao, N. Shinn, P. Razavi, K. Narasimhan. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. 2024. arXiv:2406.12045.
- [BFCL] S. G. Patil, H. Mao, F. Yan, C. C.-J. Ji, V. Suresh, I. Stoica, J. E. Gonzalez. The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models. ICML 2025.
- [AgentDojo] E. Debenedetti et al. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. NeurIPS 2024 Datasets and Benchmarks. arXiv:2406.13352.
- [WorkBench] O. Styles et al. WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting. COLM 2024. arXiv:2405.00823. [WorkBench Revisited] WorkBench Revisited: Workplace Agents Two Years On. 2026. arXiv:2606.13715.
- [DTap] Z. Chen, X. Liu, H. Tong, C. Guo, Y. Nie, J. Zhang, M. Kang, C. Xu, Q. Liu, X. Liu, et al. DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents. 2026. arXiv:2605.04808.
- [MCPMark] Z. Wu, X. Liu, X. Zhang, L. Chen, F. Meng, L. Du, Y. Zhao, F. Zhang, Y. Ye, J. Wang, et al. MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use. 2025. arXiv:2509.24002.
- [AWM] Z. Z. Wang, J. Mao, D. Fried, G. Neubig. Agent Workflow Memory. ICML 2025. arXiv:2409.07429.
- [ASI] Z. Z. Wang, A. Gandhi, G. Neubig, D. Fried. Inducing Programmatic Skills for Agentic Tasks. 2025. arXiv:2504.06821.
- [SkillWeaver] B. Zheng et al. SkillWeaver: Web Agents Can Self-Improve by Discovering and Honing Skills. 2025. arXiv:2504.07079.
- [WALT] WALT. 2025. arXiv:2510.01524.
- [TraceCompiler] S. El Yadouni, G. Li. TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows. 2026. arXiv:2608.02680.
- [LOOP] arXiv:2605.14237. [Compiled AI] arXiv:2604.05150. [PreAct] arXiv:2606.17929. Replay and compilation without a model, 2026.
- [Bosco et al.] A. Bosco, A. Augusto, M. Dumas, M. La Rosa, G. Fortino. Discovering Automatable Routines from User Interaction Logs. BPM Forum 2019. [Leno et al.] V. Leno et al. arXiv:2106.13446.
- [Speculative Actions] N. Ye et al. Speculative Actions: A Lossless Framework for Faster AI Agents. ICLR 2026. arXiv:2510.04371.
- [Speculative Macro Commit] Z. Liu, S. Kundu, P. A. Beerel. Speculative Macro Commit for Faster Tool-Using Agents. MLSP 2026. arXiv:2609.03236.
- [AutoTool] J. Jia, Q. Li. AutoTool: Efficient Tool Selection for Large Language Model Agents. AAAI 2026. arXiv:2511.14650.
- [AOSpec] H. M. Chen, J. Guo, W. Luk, H. Fan. AOSpec: Action and Observation Co-Speculation for Low-Latency Agent Serving. 2026. arXiv:2608.00881.
- [PASTE] Y. Sui, H. Zhao, R. Ma, Z. He, H. Wang, J. Li, K. Xu, K. Chen, Y. Yang. PASTE: Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving. 2026. arXiv:2603.18897.
- [Speculate with Memory] Y. Li, Q. Ye, P. K. Choubey, J. Zhang, C.-S. Wu. Speculate with Memory: Lossless Acceleration for LLM Agents. 2026. arXiv:2607.12236.
- [Speculate While You Reason] J. Ji, Y. Liu, L. An, R. Jain, G. Polatkan, S. Zhu, S. Chang. Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent–Speculator RL. 2026. arXiv:2607.25816.
- [Speculative Interaction Agents] C. Hooper, M. Kang, S. Moon, N. Lee, E. Wen, J. Wawrzynek, M. W. Mahoney, Y. S. Shao, A. Gholami, K. Keutzer. Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling. 2026. arXiv:2605.13360.
- [Ghost Tool Calls] B. Mohammadi, L. Klein, A. Arora, L. Bindschaedler. Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools. 2026. arXiv:2606.02483.
- [LLMCompiler] S. Kim, S. Moon, R. Tabrizi, N. Lee, M. W. Mahoney, K. Keutzer, A. Gholami. An LLM Compiler for Parallel Function Calling. ICML 2024. arXiv:2312.04511.
- [APC] Q. Zhang, M. Wornow, G. Wan, K. Olukotun. Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents. NeurIPS 2025. arXiv:2506.14852.
- [Leviathan et al.] Y. Leviathan, M. Kalman, Y. Matias. Fast Inference from Transformers via Speculative Decoding. ICML 2023. arXiv:2211.17192.
- [EAGLE-2] Y. Li, F. Wei, C. Zhang, H. Zhang. EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees. EMNLP 2024. arXiv:2406.16858. [SpecDec++] K. Huang, X. Guo, M. Wang. SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths. 2024. arXiv:2405.19715.
- [STAGE] arXiv:2608.22538. [PolicyGuide] arXiv:2608.19861. Compiled policies on τ²-bench, 2026.
- [MacKay & Peto] D. J. C. MacKay, L. C. B. Peto. A Hierarchical Dirichlet Language Model. Natural Language Engineering 1(3), 1995.
- [Decision mining] A. Rozinat, W. M. P. van der Aalst. Decision Mining in ProM. BPM 2006.
- [Ibrahim & Chen] J. G. Ibrahim, M.-H. Chen. Power Prior Distributions for Regression Models. Statistical Science 15(1):46–60, 2000.
- [DAgger] S. Ross, G. Gordon, D. Bagnell. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS 2011.
- [TRADE] C.-S. Wu, A. Madotto, E. Hosseini-Asl, C. Xiong, R. Socher, P. Fung. Transferable Multi-Domain State Generator for Task-Oriented Dialogue Systems. ACL 2019. arXiv:1905.08743.
- [Cascades] L. Chen, M. Zaharia, J. Zou. FrugalGPT. 2023. arXiv:2305.05176. A. Madaan et al. AutoMix: Automatically Mixing Language Models. NeurIPS 2024. arXiv:2310.12963.
- [fugue] A. Nodeland. fugue: a monadic probabilistic programming library for Rust. https://github.com/alexnodeland/fugue.
- [D0] stretto's earlier read-only flow, with an LLM arbiter: RFC-001 §3.13–3.22, https://github.com/alexnodeland/stretto.
Appendix A. Proofs
Proposition 1. Proof. Reads do not change the state, so the speculator's reads do not. An answered call's result is, by the definition of answering, what the call would have returned itself. The writes, their order and the states they act on are therefore unchanged. ∎
Proposition 3. Proof. Reads commute with one another and with the state, so the result of
Proposition 4. Proof. Drop the subscripts. Making
Appendix B. Costs, and the threshold sweep
Both costs can be counted from recorded episodes, in each agent's own input tokens:
Swept over thresholds (Figure 4) and priced at each domain's counted costs, the use-before-write speculator's utility peaks where
Priced as billed. Providers bill an output token at 4–8 times an input token, and with prompt caching they reread a context's prefix at about a tenth of the price. Priced that way (scripts/costs.py --price), a saved turn is worth its prompt, read from the cache up to the previous turn's prompt and written after it, plus its output; a detour's result is written once and read from the cache in every later context. Caching discounts the detour more, since a detour is carried almost entirely as cache reads, while a saved turn's output and new input never are. At Anthropic's ratios (output 5×, cache reads 0.1×, writes 1.25×)
Per decision. Proposition 3 holds per read, and both costs can be counted per decision: scripts/per_decision.py). The rule raises the threshold early, where a detour is carried longest, and lowers it late. In retail, raised early, it drops the lookups of an order's products, which the speculator scores 0.33 and the agents used 64% of the time. In telecom it spares the agents that use fewer lookups than scored (GPT-5.2, +57% and +102%). In airline, lowered late, it makes five times the detours, each cheap. A threshold per decision is only as good as each score where it is priced; the average threshold forgives a score that is off where the costs are extreme. Recalibrating each site's score by the use counted there in the other half of the agent's test tasks, with the model's score as a prior of 40 lookups, shrinks retail's loss to 3.5% and leaves the threshold per decision within 5% of
Appendix C. Acting at once
Of the lookups the use-before-write speculator made at
Appendix D. Replay and implementation details
An earlier version of our replay skipped any recorded call already made, the agent's own repeats included, and so counted as saved the 0.2–0.9% of turns (3.3% for one agent) in which an agent repeated calls it had made. The numbers here use the rule of §2.2. A lookup made between two calls of one turn cannot spare the second, which the agent has already made; the replay counts such a lookup as used rather than as a detour, which happened for 6 of the 3,323 lookups in six replays we checked.
Replay from the record. A trace replay answers the agent's own calls with their recorded results. A lookup gets the result of the agent's own next call with the same arguments, if one comes before the agent's next write; otherwise it gets the nearest recorded result of the same tool, a result of the right shape for the speculator's next decisions, and is marked so that it never counts as a saving. Writes by others, such as the telecom customer's on their phone, do not end a lookup's reach, since the record cannot say whether they changed what the agent reads. Stopping at them cut one agent's telecom savings from the environment replay's 164 turns to 141. Compared with the environment replay of §4.2 on the same episodes, the trace replay saves 3.3% fewer turns with the use-before-write speculator and 2.8% fewer with the next-step one: after a write, the environment still credits a lookup whose result the write left unchanged, which a record cannot show. The gap is 15% in solo telecom, where the agent writes to the phone every few turns. It counts 0.55 and 0.72 of the detours, since a stand-in starts fewer chains of further lookups than a real result, and 93–95% of episodes save exactly the same turns.
Lists. AgentDojo's travel lookups take lists: get_hotels_prices takes every hotel a city's listing named. A binding therefore also traces list arguments to the path of an earlier result whose values contain the list, and passes every value at that path of the most recent such result, unless the lookup already had that list. Its chance is scored as a string's is. Flows that bind no lists replay identically, decision by decision.
Lookups without arguments. A lookup's required arguments are those the agent passed in at least 90% of its calls. A search that each agent narrows by one of several optional filters, as in WorkBench, therefore has none, and a speculator would make it bare. The binding's chance that its arguments are the agent's own must hold for no arguments too. For such a tool that the agent sometimes called with an argument, the chance is the share of its calls that passed none, Laplace-smoothed: near 0 for a search the agent always narrows. A read the agent always calls with none keeps the chance 1.
Constants. An argument the agent passes with one value every time, such as the page size GPT-5 and o3 pass with each read of a Notion block, traces to no result. A flow then either cannot make the lookup, where the argument is required, or makes it without the argument, and it is never the agent's call. learn --constants learns such an argument, one value in every call that passed it, at least five and at least half of the lookup's, never found in an earlier result, and never a string with a digit or an @ that a user wrote, which is that user's id or email however many sessions shared it; the binding passes it as the agent did. Learned so, each model's own flows save 21 of Notion's turns instead of 13, with 1.6 detours an episode instead of 2.6 (o3's fall from 173 to 23), 82 of DTap-Bench CRM's instead of 62, and 262 of its OS files' instead of 155, where the agents always ask for a message's body as text; elsewhere they save the same turns. Without the option, learn writes the flow it wrote before.
Reading the request. ceiling.py's fifth count asks the System-One model, for each argument of a read whose value the user wrote, to pick that value among the spans of what the user wrote so far: every run of one to six words that neither starts nor ends with a common word, the latest first, at most 255. The model is shown the user's words and the call about to be made, its tool and the argument's name, never the agent's value, and a turn counts when every value in it is picked right. The value is among the spans for at least 91% of the questions in τ²-bench, τ-bench, BFCL, AgentDojo and DTap-Bench's customer service, telecom and travel, for 80–98% in WorkBench and DTap-Bench's CRM and operating system's files, and for 18–41% in MCPMark, whose values the agent composes. The instruction was revised once, on WorkBench multi-domain's 185 questions, where the model had picked spans with words around the value (59% right before, 70% after), and then fixed for every benchmark. The 8,532 questions cost about $0.41. In DTap-Bench's travel, 59 of 74 questions have several right answers, the cities a request names, which the agent looks up one after another: one pick per call finds one of them.
Deployment guards. A flow calls only tools it read in training and that the server does not mark as writes, and an operator can narrow that to a list, since a read can still be metered or recorded as an access, and a speculative read tells the server what the agent may do next, read-only or not [Ghost Tool Calls]. A flow learned from sessions pins each tool's input contract, and the proxy makes no lookup of a tool whose server now lists another. In shadow mode a flow logs what it would look up and makes nothing, and stretto promote then lets it act only at the sites where its lookups were the agent's own. Recorded sessions can be pseudonymized for sharing, and learn learns the same flow from the copy; they expire after a set number of days.
The served habit. As served, the habit has one more layer of the same form, for an episode-level group; a flow learned without intents puts every session in one group, so the layer weighs the longest context's counts once more. That only sharpens the next-step estimate, which §4.2 finds too low, not too high.