Stretto
Compiling LLM agent behavior into typed probabilistic flows, with calibrated System-One decisions at the branch points
An LLM agent pays for every tool call with a turn of the LLM, and many of those calls are predictable lookups. We treat an agent's trace as a walk over the action space its tools define, learn a Bayesian model of that walk from recorded episodes, and compile its predictable parts into flows: read-only continuations that run between the agent's turns. Each branch point is decided by weighing the learned habit against a calibrated System-One model (TypeSafe's Jev), and each lookup's arguments are bound from results already in hand. Because a flow never writes, a wrong pick costs one extra lookup rather than a wrong action.
On τ²-bench's published trajectories, read-only flows save 17% of LLM turns in retail and 13% in airline offline, and they transfer to nine agent models they never saw. Replayed without Jev, the habit alone saves as many turns, and live it did too; the System-One model's answers buy precision, cutting detours in airline by half to two-thirds. At a cold start, too, the habit carries the savings. Learned from five of an agent's own sessions, it saved 12–19% of retail turns on four random draws, more than an arbiter fitted on those sessions; on a narrow draw it saved 3%, and an arbiter fitted on other agents' decisions, even in another domain, lifted that to 11%. A flow that hands every branch back to the LLM saves almost nothing, and naming flows as macro-tools could add at most 2% of turns in retail and 8% in airline. Live, with GLM-5.3 as the agent and the flow behind its MCP tools, adding no tools and changing no prompt, every retail and airline test task, paired, took 25.5% fewer LLM turns (95% interval 20.5% to 30.4%) and 21% fewer input tokens. The pass rate moved by −1.25 points, with an interval (−7.5 to +6.25) too wide to rule out a loss of a few points. With Claude Haiku 4.5 as the agent, on ten retail tasks, the same flow saved 18.6% of turns, and a Claude model as the customer in place of GLM-5.3 left the savings where they were. Policy guards compiled from the written policy refuse a policy-breaking write in 38% of failed airline episodes and in 1% of successful ones, each of which τ²-bench's own task notes forbid. Live, GLM-5.3 made no write they had to refuse. Asked whether the customer agreed to each write, Jev beat the guards' word list on 30 of 40 hand-labelled writes where the two disagreed, and enforced live, it refused none of the 40 writes it judged; asked which record a customer meant, it did worse than the agents. In telecom, once each lookup takes its arguments from the sources the agent used at that point, a flow learned from other agents saves 12.8% of nine agents' turns in replay, and 27–34% where the agent operates the phone itself. There, where every branch follows what a tool returned, a decision tree compiled once from successful traces runs the whole task with no model. It passes about as many held-out tasks as the agents it learned from, a check of the ticket's stated outcome tells it when to hand back, and with its identifiers learned as where they came from it passes as many for a customer no trace saw (§5.23). The system is open source: a Rust MCP proxy, a flow IR, and the measurement pipeline.
This notebook records the project through 2026-09-25, when flows still asked a System-One model. The method has since become a speculator that decides on the chance of use before the next write, with no model of its own, and a procedure compiled once where no user speaks. The current results are in the working paper and on the results pages of 2026-09-26 and 2026-09-27 (seven benchmarks, three agent SDKs, real MCP servers).
stretto (It., "narrow, tight"): the passage of a fugue in which entries of the subject overlap, each voice entering before the previous one has finished. Here, a flow's lookups enter before the agent's turn is over.
1Introduction
A tool-using agent works in turns. The LLM reads the conversation, emits a tool call, waits for the result, reads everything again, and emits the next call. In τ²-bench's customer-service domains, about a quarter of all LLM turns are spent inside runs of consecutive tool calls, before the agent says anything to the customer (§5.1). Most of those calls are lookups whose tool and arguments follow from what the previous call returned: a user's profile lists their orders, so the next call reads an order.
That regularity is visible in traces. An agent's trace is a walk over a small alphabet, the tools it may call plus "reply to the customer", and the walk is far more predictable than the text around it[8][17]. Prior work has mined such traces into workflows, but leaves the branch points to the LLM or to hand-written rules[24], or skips the LLM one step at a time on an uncalibrated score[13]. What has been missing is a cheap decider for the branch points: something that answers a typed question about the state, fast, with a probability that means something.
System-One models now fill that role. Jev takes a state and typed questions (a choice among options, a score, a yes/no) and returns calibrated probabilities in 70–500 ms, at $0.042 per million input tokens. It never generates text. Stretto combines the two: it learns the habit from traces, compiles it into flows with explicit decision sites, and decides each site by weighing the habit against Jev and a few yes/no questions about the state, handing back to the LLM whenever the answer is not good enough. §5.9–5.14 measure what Jev adds. For read-only flows it adds precision, and the savings come from the habit, even from a deployment's first five sessions: an arbiter fitted on so few costs more than it gives (§5.12). Jev earns its place in judging what the habit cannot see, such as whether a customer agreed to a write.
The design choice that makes this safe is that flows only read. Between the agent's turns a flow may call tools its server marks read-only. Anything that changes the world goes back to the LLM. A flow that picks the wrong next step therefore makes one extra lookup and hands back. It costs a pause and some tokens, and changes nothing in the environment. This turns the question "is the System-One model accurate enough to act?" into "is it accurate enough to be worth a lookup?", which it is.
Our contributions:
- A measurement method on public trajectories (§5.1–5.3): how many LLM turns macro-tools could remove, where write arguments come from, how predictable the next action is, and how well a habit transfers across models.
- Read-only flows with Bayesian arbitration (§3.3–3.6): a System-One model weighed against the habit, its own record per site, and state predicates by a cross-fitted conditional logit, acting on the probability that the whole call, arguments included, is the agent's.
- A live form that needs no cooperation from the agent (§3.7, §5.6): the flow runs behind the agent's own tool calls and returns its lookups in the same response, measured live with GLM-5.3.
- Typed policy guards, audited against traces (§3.8, §5.7): checks compiled from the written policy that pass, fail or do not know, tested on 2,624 trajectories before any is enforced.
- What the System-One model adds (§5.9–5.14): measured against the habit alone, offline and live, with plentiful and scarce traces and from a deployment's own first sessions, and in two content roles: judging a confirmation and matching a description to a record.
- An open implementation (§4): a Rust MCP proxy that records sessions and serves flows, guards and a commit tool for any MCP server; a versioned flow IR; learning from recorded sessions, with arbiters that ship; and an audit that scores new sessions under a flow run as a fugue program (§5.8).
3Method
3.1Traces as walks over the action space
An episode is a sequence of events: customer messages, assistant turns (text and zero or more tool calls), and tool results. Stretto abstracts each assistant turn to an action, a tool name or respond, and each result to an outcome, ok or error. The alphabet comes from the live tool manifest, so a model can only ever propose actions the agent is allowed to take. Tools the server marks read-only (MCP's readOnlyHint) are lookups; everything else is a write.
3.2The habit
The habit predicts the next action from the last k steps and their outcomes with a hierarchical Dirichlet back-off model: counts at each context length are smoothed toward the next shorter context, down to the unigram, with one concentration α. Stretto infers α by Metropolis–Hastings over the prequential marginal likelihood of the training episodes, in fugue, and uses the posterior median. Enum-like fields read from tool outputs (an order's status, a cabin class) can extend the context. Each is kept only if it improves held-out prediction under five-fold cross-validation grouped by task, which rejects fields that merely identify a task. The habit is learned from successful training episodes only.
3.3Flows and their decision sites
A flow is a read-only continuation. Its decision site is the lookup that just returned; its options are the lookups agents made next at that site in training, plus handing back to the LLM. At each site the flow puts a slice of the state (the customer's messages, the calls so far, the latest result) to the System-One model as typed questions:
- which comes next, asked two ways: as one
Choiceamong the site's lookups and handing back, and split into a yes/no question (is more needed?) and a choice among the lookups; - three yes/no predicates about the state, as
Noulquestions: are records in a list still unchecked, are options still missing, is something needed that only the customer can give.
All questions go in one request, and Jev evaluates them in parallel. Answers are cached by the SHA-256 of the request, so a repeated question is free, and replaying a run needs no key.
This bounds what an injection can do. The state slice includes tool results, which an outsider may have written, but the answers only weigh options the flow enumerates, and the flow binds each lookup's arguments itself. Whatever Jev answers, a flow calls only the lookups compiled for the site, and never a write; a test holds it to that with an oracle that names other tools. Inside that set, injected text can still steer the choice (§5.15).
3.4Arbitration, not pooling
Jev's pick is not taken at its word. A conditional logit scores each option o at a site from features of that option:
s(o) = w₁·ln P_habit(o) + w₂·ln P_one(o) + w₃·ln P_split(o) + w₄·[o = hand back]
+ Σⱼ w₅ⱼ·[o ∈ Oⱼ]·logit P(yes to predicate j)
+ w_r·[o = Jev's pick]·logit r(site)
Here Phabit is the habit's prediction with all mass on steps a flow cannot take moved to handing back, Oj is the set of options predicate j bears on (the same lookup again, any lookup, or handing back), and r(site) is Jev's agreement with the agent at that site on other tasks, shrunk toward its overall agreement: a one-coin Dawid–Skene sensor[F3] whose labels are known. P(o) ∝ exp s(o), scaled by how often the agent's step was among the site's options at all. The weights are fitted by maximum likelihood on the agents' own next steps, with a ridge penalty, cross-fitted over five folds of tasks, so no decision is ever scored by weights that saw its task. A live flow judges each task with the fold that held it out.
3.5Acting: lookup first
For a read-only flow, handing back costs an LLM turn and a wrong lookup does not. The cost-optimal rule is therefore to take the likeliest lookup whenever it is likely enough to pay for its tokens. Stretto takes the best lookup when P(tool) × P(arguments are the agent's) ≥ 0.3, and hands back otherwise. Every wrong lookup still has to be carried in later prompts, so the projection charges each one the tokens of a typical lookup (about 250) in every later turn.
3.6Binding arguments
A lookup needs arguments, and in τ²-bench 85–91% of write-argument values appear verbatim in an earlier tool output (§5.1). For each argument of each lookup, stretto learns from training where the agent's values came from: a source tool and a JSON path in its result, such as get_user_details at $.orders[*]. Live, it takes the first candidate at that path that has not been looked up yet, preferring one the customer mentioned. How much to trust a binding is measured per call: in training, how often the arguments it would have bound, all at once, matched the agent's own call, split by whether the customer had mentioned the value. For retail that is 92% for orders and 60% for products the customer did not mention (95% and 76% when mentioned).
3.7The live form: flows behind the agent's own calls
A flow could be exposed as a macro-tool the LLM calls by name, but that asks the agent to adopt a new tool. The first live form asks nothing of it. After each of the agent's own tool calls, the flow asks its questions, makes the lookups it is sure enough of, and appends their results to the same tool response under a short header, until it hands back. The agent sees more of what it needs per call and spends fewer turns gathering it. The flow is goal-free: nobody tells it what the customer wants, and §5.5 shows it does not need to know. A name would add only the lookups that need a value from the conversation, which a flow behind the tools cannot bind; §5.9 finds them rare.
3.8Policy guards
Flows never write, but agents do. A guard is one rule of a domain's written policy compiled into a check a proxy runs before a write reaches the server. It reads only what the episode has shown: the user the agent authenticated, the records it looked up, what the customer said. So it returns one of three verdicts:
- pass;
- fail, with a reason, when the facts in hand break the rule;
- unknown, when a record the rule needs was never looked up.
Only a failure of an enforced rule refuses the write. The agent then gets the reason and the rule's id in place of the tool's result. τ²-bench's policies yield 12 retail and 10 airline rules. An LLM (Claude, while drafting the module) compiled them from the policy text. Which of them are enforced is decided by an audit against recorded trajectories (§5.7), not by the compiler.
4System
Stretto is five Rust crates and a CLI, released under the MIT license. The live path is an MCP proxy that sits between any MCP client and server:
- forwards every line and records the session (JSONL, v2)
- after a lookup, runs the flow and appends its lookups
- before a write, checks the guards
- adds
stretto_commit: confirmed writes in one call
stretto compile builds it from τ²-bench results; stretto learn from recorded sessions. Decisions ask Jev (cached by request hash).pilot/), which predates the proxy's active mode and behaves the same way. The proxy's tests record forty sessions, learn a flow from them, and serve it to a new one.| Crate | State | What it does |
|---|---|---|
stretto-trace | built | Canonical episode schema; τ²-bench results and MCP session logs in; tool manifests with docs and read-only hints. |
stretto-model | built | Action abstraction; hierarchical Dirichlet back-off habit; α posterior by MH in fugue; code features; argument provenance; tool runs. |
stretto-oracle | built | Oracle trait; Jev client; on-disk replay cache keyed by request hash; mock. |
stretto-report | built | The stretto CLI: Phase 0 reports; System-One questions; arbiter; flows (compile, learn from recorded sessions or a handful of τ²-bench ones, serve, export-arbiter); guards and their audit; the flow audit, with fugue (audit); the confirmation judge (confirm); matching descriptions to records (match). |
stretto-proxy | built | A stdio MCP proxy for any MCP server. It records sessions; in active mode it runs a flow behind the agent's calls, refuses writes the guards fail, adds stretto_commit, and logs the conversation a host hands it. |
5Experiments
Data. τ²-bench[26] publishes trajectories for four 2025 agent models (Claude 3.7 Sonnet, GPT-4.1, GPT-4.1 mini and o4-mini, with GPT-4.1 simulating the customer), four trials per task: 2,624 episodes in retail and airline. Everything is learned on τ²-bench's official train split and measured on its test split. As transfer targets we use nine runs from Sierra's τ²-bench leaderboard (5,904 episodes): GLM-5, Qwen3.5-397B, Qwen3-Max, GPT-5.2 with reasoning high and off, Claude Opus 4.5 and Sonnet 4.5, and Gemini 3 Pro and Flash. Nothing is ever learned from them.
Measures. The unit of cost is the LLM turn: one assistant message, whatever it contains. Offline, we replay held-out episodes through a flow and count the turns it would have removed. A detour is a lookup the agent did not make. Agreement is how often a decider's pick matches the agent's actual next step. That is conservative, since an equally good step counts as a miss, and it is not correctness: the agents fail 21–50% of tasks.
5.1A quarter of LLM turns sit inside tool runs
If every run of consecutive tool calls were one macro-tool call, 23.8% of LLM turns would disappear in retail and 25.1% in airline (Figure 3). How much a model leaves on the table depends on how it calls tools: models that batch parallel calls have already collapsed part of the work.
Show table
The arguments are mostly already there. Of 12,730 write-argument values in retail, 91.2% appear verbatim in an earlier tool output, 3.0% only in what the customer said, and 1.6% were generated by the agent (airline: 84.5%, 1.5%, 3.2% of 6,792). What agents do generate is mostly a closed-set choice, such as a cancellation reason, or arithmetic, which belongs in code.
5.2A habit takes a slice; content decides the rest
On the action sequence alone, the habit is rarely sure right after a tool returns: at 0.8 confidence with at least five supporting observations, it can take 8.3% of those decisions in retail and 5.8% in airline. Code features lift that to 13.8% in retail, and naming the episode's intent to 19.5%. In airline they add under two points: its 30 training tasks are too few to support them. Where it acts, it agrees with the agent 82–99% of the time. That is not enough to act on alone: held to 99% cross-validated agreement per context, over at least ten distinct tasks, one context per domain survives, and in both the habit's call is to hand back. Whether a run continues depends on what the tool returned, which is the System-One model's job.
Habits do carry across models. A habit learned from one model predicts another's held-out behavior within 3–7 points of top-1 agreement of that model's own habit, and habits from the 2025 baselines predict GLM-5 within 4 points of its own in retail and match it in airline. That is what licenses compiling flows from one set of trajectories and serving them to another model.
5.3Zero-shot, Jev is calibrated but not accurate enough to carry writes
Phase 0b asked jev-1.13.0 at every held-out decision a flow would hand it: which tool comes next, or whether to hand back, with every tool as an option, and each closed-set argument. That was 8,946 questions, 28M input tokens, for $1.19. Jev agreed with the agent 72% of the time, and 92–93% when it was at least 0.9 sure (Figure 4; expected calibration error 0.06). Its probability means something, but flows that also write need roughly 99%. Trusted at p ≥ 0.9, such flows would save 3.9% of turns in retail and 4.7% in airline while making a call the agent did not make in 6% and 13% of episodes. That result is what led to read-only flows.
Show table
5.4Read-only flows, weighed answers, lookup first
Making flows read-only costs little ceiling: with a perfect System-One model they would save 21.0% of turns in retail and 17.7% in airline, against 22.3% and 20.6% for flows that may also write. Narrower questions alone did not make Jev more accurate (73–78% on the same decisions). Weighing its answers did (Table 2). Where the combination is at least 0.99 sure, it agrees 99.2% of the time in retail (on 10.6% of decisions) and 97.0% in airline (on 25.2%). Of the predicates, "records still unchecked" carries the most weight.
| Decider | Retail | Airline | GLM-5, retail | GLM-5, airline |
|---|---|---|---|---|
| Jev alone, one question | 76.2% | 72.4% | 77.9% | 63.3% |
| Combined with the habit | 77.9% | 74.3% | 81.7% | 69.5% |
| Combined, with predicates | 80.5% | 78.9% | 85.7% | 75.8% |
Because a detour is harmless, flows can act on weaker picks. Taking the likeliest lookup at p ≥ 0.3 saves 17.4% of LLM turns in retail and 12.8% in airline, pooled over the 2025 models, with at least one detour in 57–67% of episodes (Figure 5). Dollars fall faster than turns even after charging every detour, because a typical lookup returns only 200–250 tokens.
Show table
Across the nine leaderboard models, none of them trained on, the savings track how each model calls tools (Figure 6). Models that call one tool at a time leave long runs to shorten: Qwen3.5 saves 34.1% in retail and 27.8% in airline. The heaviest parallel callers have already done that work themselves. In airline, Claude Opus 4.5, GLM-5 and GPT-5.2 stay under 5%, against read-only ceilings of 12–17%.
Show table
5.5Flows do not need the goal
Every run above told the flow the episode's goal (the writes it goes on to make), as a macro-tool call would name it. A flow running behind the agent's calls gets no such thing. Asked again without it, 8,092 more questions for $0.93, nothing measurable changed: agreement 80.5% in retail and 78.4% in airline (80.5% and 78.9% with the goal), and 17.3% and 12.6% of turns saved (17.4% and 12.8%). GLM-5 still clears 20% in retail (20.3%). This fits what read-only flows decide. Which record to read next, and whether there is enough, is in the lookups already made and in what the customer said. It is not in the goal.
5.6Live
Replay first. Before any live run, GLM-5's recorded retail test episodes were replayed through the live flow, with real Jev answers and no LLM; a recorded call the flow had already made was skipped. Over all four trials, 160 episodes and 1,384 LLM turns, the live flow saved 294 turns (21.2%), where the offline projection saves 281 (20.3%). 91% of its lookups were the agent's own, and 27 episodes had a detour, against 45 projected. Weighing each lookup by its binding agreement (§3.6) had cut detours from 42 to 18 on one trial without losing a saved turn. Where the flow follows the agent's path it asks byte-identical questions to the offline run, so they come from the replay cache.
Paired pilot. The agent was GLM-5.3 in Claude Code, on Z.ai's GLM Coding Plan, with τ²-bench's tools served over MCP. GLM-5.3 also played the customer, from τ²-bench's user-simulator prompt. Ten test tasks per domain were drawn at random, seeded, and each ran once without and once with the flow. The flow was compiled goal-free from the 2025 baselines, and it asked Jev live. The reward is τ²-bench's database check alone. Its natural-language assertions need an LLM judge and were left out.
| Retail, without | Retail, with | Airline, without | Airline, with | |
|---|---|---|---|---|
| LLM turns | 110 | 84 (−23.6%) | 127 | 105 (−17.3%) |
| Saved per episode, 95% interval | 2.6 (1.0 to 4.2) | 2.2 (0.5 to 3.9) | ||
| Agent input tokens | 723,926 | 571,528 (−21.1%) | 996,580 | 849,367 (−14.8%) |
| Agent's own tool calls | 69 | 42 | 84 | 60 |
| Flow lookups (repeated by the agent) | 26 (3) | 24 (0) | ||
| Passed the database check | 8 of 10 | 8 of 10 | 8 of 10 | 9 of 10 |
| Z.ai credits | 128.5 | 108.0 | 176.4 | 157.6 |
Show table
In retail the flow saved 2.6 LLM turns per episode (95% interval 1.0 to 4.2). It acted on 9 of the 10 tasks, making 26 lookups, and the agent repeated 3 of them. On those nine tasks turns fell by 26%. The one task where it never acted took a turn more, because the simulated customer never replies the same way twice. In airline the flow acted on every task, with 24 lookups the agent never repeated, and saved 2.2 turns per episode (95% interval 0.5 to 3.9). None of the three failures came from a flow decision. Without the flow, the agent booked a flight on a task that expects no write, and added a paid bag nobody asked for. With it, on task 31, the customer would pay under $100 for a change. The agent upgraded the basic-economy reservation (a $301 charge), changed its flights (a $220 refund), and quoted the net, $81. The task counts the upgrade's cost against the limit and expects no change. The flow's three lookups there were the agent's own.
The airline result exceeds its projection. Offline, flows save GLM-5 under 5% of turns in airline (Figure 6), because GLM-5 in τ²-bench's own harness makes parallel calls in 45% of its airline tool turns. GLM-5.3 in Claude Code made them in 3.8% (3.0% in retail), so it behaves like the sequential callers at the top of Figure 6. The live agent differs from the leaderboard run in model version as well as harness. Either way, what matters for flows is how the agent in front of them calls tools, and that is measurable from a few recorded sessions. The two pilots cost 571 Z.ai credits and a few cents of Jev questions.
Ten pairs cannot bound a one-point loss of pass^1. They show that the mechanism works live, without the agent's cooperation, and give a first effect size.
Every test task. The run then took every test task, reusing the pilots' pairs: τ²-bench retail's 40 test tasks once and airline's 20 twice, 80 pairs in all. It cost 1,897 Z.ai credits.
- Turns and tokens. With the flow, LLM turns fell 25.5% (95% interval 20.5% to 30.4%): 30.4% in retail and 21.1% in airline. The agent's input tokens fell 21.0%.
- Pass rate. 71 pairs passed without the flow and 70 with it, a difference of −1.25 points (−7.5 to +6.25; McNemar p = 1.0). By domain, retail moved −7.5 points (−17.5 to 0.0) and airline +5.0 (−5.0 to +17.5).
- The pairs that disagree. One of the nine plausibly comes from the flow: its reads went deep on one order where the agent alone read wide. In the other eight, the failing step was the agent's or the simulated customer's, and no lookup touched it.
- Detours. 242 of the flow's 259 lookups were calls the agent also made without it. The other 17 carried under 3% of the input tokens the flow saved.
RFC-001's gate for Phase 2 asks for 20% fewer turns with at most one point of pass^1 lost. The turns half is met. The other is not shown: the interval still allows a loss of 7.5 points. At this rate of disagreement, bounding one point would take about 4,300 pairs (details).
5.7Guards against recorded trajectories
stretto guards checks every write in every published trajectory against what came before it, as a proxy would, and sorts it by whether the tool accepted it. A guard that fails a write the tool also refused agrees with the tool. A guard that fails a write the tool accepted, in an episode that then succeeded, is a false alarm, or a policy the episode broke and passed anyway. τ²-bench grades the database, not every rule.
In airline the enforced rules would refuse an accepted write in 141 of 369 failed episodes (38%) and 5 of 431 successful ones. All five are cancellations that τ²-bench's own task notes forbid (tasks 9 and 37: "Agent does not cancel reservation NQNU5R since it is in the past"). The cancellation-eligibility rule alone fires in 116 failed episodes. The airline API does not check eligibility, and its policy says so. In retail the guards separate nothing: 1 of 500 failed and 28 of 1,324 successful episodes, each a break of the written policy that passed the database check (exchanging an item for itself, writing before checking an order's status, one user never authenticated). Retail's tools already enforce most of their own policy.
A correction. After the airline pilot, the basic-economy rule was changed to refuse task 31's path: upgrading a reservation's cabin, then changing its flights. That was wrong. τ²-bench's own task 32 expects exactly that path. It was found before the live test of the guards ran, and reverted.
Live. Does refusing those writes raise the pass rate? The guards ran live behind GLM-5.3 in airline (arm B). The pilot took the four test tasks where the audit finds the 2025 agents' refused writes concentrated (35, 45, 32 and 48; a guard would have refused a write in 35 of their 50 failed episodes). It added four tasks whose baselines had passed, as a harm check. GLM-5.3 gave the guards nothing to refuse:
- It passed all four main tasks in both arms. On the two that ask for cancellations the policy forbids, it declined on its own.
- With the guards on, 7 of 8 episodes passed. The one failure, task 31, repeated the cost error of §5.6, and nothing was refused.
- The proxy checked 9 writes live and passed them all. They include task 32's upgrade-then-change, the path the briefly mistaken rule would have refused.
So guards are insurance whose value depends on the agent. Against the 2025 agents they would have refused a write in 38% of failed airline episodes; against GLM-5.3, on the tasks where those failures concentrate, in none. They cost nothing when they do not fire. The pilot cost 223 Z.ai credits (details).
The audit also demoted a rule. Detecting an explicit "yes" with a word list misses about as often in successful episodes as in failed ones (3.2% and 3.9% of accepted writes in retail, 6.9% and 6.5% in airline), so both confirmation rules are logged and never enforced. Judging a confirmation is a natural System-One question, not a keyword match, and §5.11 asks it.
Show the per-rule audit
5.8Auditing a flow, as a fugue program
A flow's decisions over an episode are a probabilistic program: at each decision, a categorical choice among the site's options, weighted by the arbiter. stretto audit builds that program in fugue, one sample per decision at address decision#i. It then runs it under two handlers. ScoreGivenTrace, given a trace of the agent's own steps, returns the log-probability of what the agent did, so the flow's surprise per decision, per site and per episode. PriorHandler runs the same program as a stochastic policy. The check it serves: before trusting a flow on new sessions, score them.
| Retail | Airline | |
|---|---|---|
| Decisions scored | 724 (160 episodes) | 228 (80 episodes) |
| Flow's likeliest option was the agent's step | 85.9% | 64.5% |
| Surprise, nats per decision | 0.42 | 1.02 |
| Expected calibration error | 0.088 (underconfident) | 0.070 (overconfident) |
| Weakest site | after an order is read, 61% | after a reservation is read, 42% |
The retail flow fits GLM-5 as well as the offline agreement said (85.7% in Table 2). The airline flow does not, and the audit says where: GLM-5 in τ²-bench's harness reads several reservations in one turn, while the flow decides one lookup at a time. That matches the projection of under 5% saved (Figure 6), and it is the warning an audit exists to give. Calibration also runs opposite ways in the two domains. In retail, a top option the flow gives 0.8–0.9 is right 96% of the time; in airline, 74%. So a deployment should recalibrate on its own sessions, and the audit says in which direction. Full tables are in docs/results/audit-2026-09-24.md.
5.9What arbitration adds, and what naming would
Two parts of the original design were left untested. One is arm C: a flow that hands every branch back to the LLM, as TraceCompiler does. The other is naming, where the LLM calls plan_* macro-tools itself instead of a flow running behind its own calls. Both were measured without an LLM. Recorded episodes were replayed through τ²-bench's tools with a flow behind them, as in §5.6. A recorded turn whose calls the flow had already made counts as a turn saved; a flow lookup the agent never made is a detour.
| Flow | Retail, saved | Retail, own | Airline, saved | Airline, own |
|---|---|---|---|---|
| D0: arbiter, 0.3 | 21.2% | 91% | 6.5% | 88% |
| Habit alone, 0.3 | 22.5% | 91% | 8.7% | 80% |
| Arm C: habit, 0.9 | 1.7% | 100% | 0 | – |
| Strict arm C: habit, 0.99 | 0 | – | 0 | – |
A flow that hands every branch back saves almost nothing. Branches are everywhere, and at 0.99 the habit never acts. What saves turns is acting on a probability, which is safe because a wrong read costs only a detour.
More surprising, the habit alone saves as many turns as the arbiter, with 95% intervals bootstrapped over tasks:
- Retail: 1.2 points more (−0.5 to 3.0), with the same 58 detours.
- Airline: 2.2 points more (0.6 to 4.2), with twice the detours, 50 against 26.
Weighing Jev's answers buys precision, not savings. The same holds on Claude 3.7 Sonnet's test episodes, one of the four source agents. There the habit alone saves 1.2 points more in retail (−0.6 to 2.9) and 2.6 more in airline (1.1 to 4.5). It makes 101 detours against 73 in retail, and 95 against 29 in airline. So, given enough traces (§5.10), a flow can run on its habit alone, with no key and no latency per question (--flow-decider habit), at the price of more detours in airline. In airline GLM-5 saves less here than live (Table 3), because in τ²-bench's harness it makes several calls per turn, and a turn counts as saved only when the flow made all of them.
Live. The habit-only flow then ran live on the retail pilot's ten tasks (§5.6), one episode each, with GLM-5.3 in Claude Code. It took 79 LLM turns, against 110 without a flow and 84 with D0. That is 28% fewer than without, with fewer turns on all ten tasks, and against D0 a change of −0.5 turns per episode (95% interval −1.4 to +0.4). It made 37 lookups, none of which the agent repeated, and passed 9 of 10 (8 of 10 in both earlier arms). It asked Jev nothing, and cost 109 Z.ai credits. It ran later the same day than the other two arms, so model drift cannot be ruled out, and ten tasks cannot bound a pass-rate effect (details).
Naming a flow would add only one kind of lookup: one inside a run that needs a value from the conversation, such as a city the customer named or a date the agent works out. A flow behind the tools cannot bind such a value. Across the thirteen agent models these lookups are 0.0–1.9% of LLM turns in retail and 0.9–7.9% in airline (GLM-5: 1.6% and 2.9%). Lookups a flow can bind are 6–38%. That bounds what plan_* could add, even if the agent adopted it and passed the right values, so it was not built. The top of the airline range is flight searches, by the agents that search many dates: Gemini 3 Flash and Qwen3.5. Details are in docs/results/arms-2026-09-24.md.
5.10Fewer traces
The habit of §5.9 learned from 831 successful retail episodes and 246 airline ones: four agents on every training task. A new deployment starts with far fewer. So the habit was trained on less. --train-fraction keeps a fixed share of the training tasks, nested so that each smaller sample is part of every larger one, down to a single task (but see the correction below). The habit, the sites a flow acts at and the bindings learn from those tasks alone. Each flow then replayed GLM-5's test episodes as in §5.9, deciding by the arbiter and by the habit alone (Figure 9).
Show table
With enough traces the habit carries the savings. From 7 retail tasks (67 episodes) and 8 airline tasks (75), the habit alone saves what it saves with every task, 1–3 points more than the arbiter, which again buys fewer detours. Below that, Jev's answers carry them. With 3 retail tasks the arbiter saves 20.4% of turns, nearly all it saves with every task, and the habit alone 1.8% (the difference: −18.6 points, 95% interval −21.1 to −16.0). The habit is unsure where agents read the next order: its probability is 0.19 at the median and never above 0.25. So under the 0.3 rule it never reads one, where the arbiter reads 394. With a single task, the habit copies that task's path and reads product lists nobody asked for: 260 detours, and 4.9% saved against the arbiter's 10.0% with 28 detours. In airline, with 3 tasks, the arbiter saves 5.6% with one detour and the habit 4.1% with 37, a difference the interval does not resolve.
The traces still set a flow's reach. A flow acts only at sites where training shows a lookup, and offers only those lookups: 4 sites with one retail task, 19 with all 74. With one airline task, neither decider saved a turn. Two cautions. Only the habit's side shrank: the arbiter was fitted as always on the source agents' held-out decisions, at least 2,344 in retail and 992 in airline at every size. A deployment with few traces would fit it on its own few sessions, so the sweep flatters the arbiter. And each size is one draw of tasks, so the sweep locates the crossover, between 3 and 7 retail tasks, rather than tracing a curve. The sweep's Jev answers cost $1.12 (details).
Correction. The samples were meant to be pseudo-random, and they were not. The sampler ranked tasks by an FNV hash that barely mixes an id's last characters, so neighbouring ids sorted together. The three retail tasks were 104, 105 and 109, and the seven were 103–107, 109 and 113. Neighbouring τ²-bench tasks often share a customer and a kind of request, so these samples are narrower than random ones. The numbers above hold for the tasks drawn, but where the crossover falls rests on one clustered draw. The sampler now mixes its order first, --train-tasks names a sample exactly, and §5.12 repeats the small sizes on random draws.
5.11Judging a confirmation
τ²-bench's policies require the customer's explicit yes before a write. The guards test for it with a word list, which §5.7 found no better in failed episodes than in successful ones. stretto confirm puts the question to Jev instead, once per write in the published trajectories. Jev sees three things: what the agent said last before the customer's last message, that message, and the call about to be made. It is asked whether the reply explicitly agreed to this change, and fails the write when its probability of a yes is below 0.5. The 3,863 questions cost $0.12.
Jev fails more accepted writes than the word list. In retail it separates the two kinds of episode better: it fails 8.5% of the accepted writes in failed episodes and 5.0% in successful ones, where the list fails 3.9% and 3.2%. In airline it separates them no better. Episode outcome is a weak test, though, since τ²-bench checks the database and not the conversation. So the disagreements were labelled by hand. The two judges agree on 94.9% of the 3,484 accepted writes. From the other 178, a random 20 of each kind were labelled blind to both judges (Table 6).
| Where they disagree | Writes | Jev right | Word list right |
|---|---|---|---|
| Jev says no, the word list says yes | 20 | 16 | 4 |
| Jev says yes, the word list says no | 20 | 14 | 6 |
| All | 40 | 30 (75%) | 10 (25%) |
Weighted by how often each kind occurs, Jev is right on 78% of the disagreements. The most valuable catches are mismatches. The customer says yes to one change, and the agent makes another: "Yes, please process those returns," and the agent cancels an order; "Yes, please use gift_card_8190333," and it charges a credit card. No word list can see that, because it never reads the call. Jev also hears agreement without the words ("Please book the second option"). Most of its errors are too easy a yes. Six of its ten are a yes to a change the agent never described first. Enforced, Jev's judgment would refuse 5–11% of accepted writes in successful episodes, so whether refusing them costs pass^1 or turns is for a live run. Until then the guards only log it (details).
A second question, asked on its own about the same three fields, was aimed at those six errors. The judge then fails a write unless both answers are yes. The first wording, "did the agent's message describe this exact change?", was too literal. It failed 18% of the accepted writes in successful retail episodes, and of 30 random writes it newly failed, 3 were lapses. The second wording, "had the agent proposed this change?", fails 113 more accepted writes: 6.6% in successful retail episodes, against 13.4% in failed ones. Of 30 of those drawn at random and labelled blind, 16 were lapses (53%; 95% interval 36–70%). Seven are calls that differ from what the customer agreed to: a whole order cancelled where three items were to go, or a refund to a gift card instead of the Mastercard the customer named. The rest are addresses the customer supplied that the agent never read back, and extras nobody proposed. These are the costly lapses, but half of the question's flags are false alarms, so it too is only logged. The two wordings cost $0.25 (details).
Enforced, live. GLM-5.3 then ran ten test tasks per domain twice behind the guards, with the judge's first question logged in one arm and enforced in the other. The second question was logged in both. Five tasks per domain were those where Jev failed the most writes in the published trajectories, and five were drawn at random. Enforced, the judge refused nothing: it failed none of the 40 writes it judged. Logged, it would have refused 4 of 38. Labelled blind, two were real lapses, both in episodes that passed τ²-bench's check: a cancellation the customer never confirmed, and an order change never read back. The other two were false alarms. The gap between the arms is within chance (Fisher's p = 0.05 by write, 0.11 by episode). The enforced arm's judgments came as close to the line, but none fell below it. Passes did not move: 8 and 8 of 10 in retail, and 7 and 6 in airline, where the one difference failed with nothing refused. The second question failed 8 of 38 writes, and 7 of those were false alarms, among them five flight changes the customer had authorized in so many words. So the recommended setting enforces the first question and logs the second. The judge answered in 0.65 s at the median, and its 156 questions cost under a cent (details).
5.12A cold start
§5.10 shrank only the habit's data. Its arbiter was still fitted on thousands of other agents' held-out decisions, and its samples were clustered. A deployment starting out has neither: only its own first sessions. stretto learn holds out 30% of them, by task, and fits one arbiter on Jev's answers there, so with five sessions it has 3–11 decisions to fit eight weights on. The flows here learn from GLM-5's own published sessions, trial 0 of each training task drawn, and replay its test episodes as in §5.10. Each draw gives an arbiter flow (70/30), a flow of the habit alone from every session, and at five and ten sessions two repairs: the arbiter flow's arbiter with the habit refitted on every session, as stacking refits its base model, and the habit alone with the arbiter of the all-task flow, fitted on four other agents' 4,513 decisions (Figure 10).
Show table
The habit alone is the steadier start. From five random sessions it saves 11.8–18.9% of turns across four draws, and from ten 20.9%, near the 22.4% it saves from all 74 tasks. In airline it saves 8.7% from five sessions, all it saves from 30. The arbiter flow trails it until about twenty sessions. It saves 6.5–16.5% from five random retail sessions, less than the habit alone on every draw, and 11.6% from ten. In airline it saves less at every size. From 20 and 40 sessions it saves 22.6% and 23.9%, the latter with three times the detours, and with every task the two tie. The split takes 30% of the sessions from the habit, the sites and the bindings. And eight weights fitted on a handful of decisions can land anywhere: across the five-session draws, the weight on the habit ran from −0.32 to 0.30, where from 20 decisions up it settled between 0.88 and 1.15. Refitting the habit on every session is no remedy. It gains where the habit alone is weak, but elsewhere loses up to 16 points, or looks up at sites its arbiter never saw fitted and makes 355 detours.
Which sessions a flow starts from matters more than how it decides. The habit alone saved 2.7% to 18.9% across the five draws. On the clustered one, a session that read an order straight after finding the user left the habit unsure where agents read the next order. There the arbiter fitted on other agents' decisions lifted the same five sessions to 11.1%, while on the four random draws it stayed within two points of the habit alone. Across the five draws it had the highest floor, 9.9%, and the highest mean, 15.2% against 14.0%. That is still far below §5.10's 20.4% from three clustered tasks, whose habit learned from 21 successful episodes of four agents. So §5.10's reading needs narrowing: Jev's answers carried the savings for its clustered samples with an arbiter fitted on plenty of decisions, and not for a random handful of an agent's own sessions. A deployment can start on the habit alone, with an arbiter fitted elsewhere as insurance where one exists, and fit its own arbiter once it has some twenty sessions (details).
A shipped arbiter. The arbiter fitted elsewhere need not come from the same domain. Stretto now ships two, fitted on the four agents' published retail and airline decisions (compile --pooled-arbiter, then export-arbiter), and learn --arbiter-from serves one with a habit learned from new sessions. Each was replayed with the habits above from the other domain. In retail, the airline arbiter saved what retail's own did, within half a point on all six habits, and lifted the narrow draw from 2.7% to 10.6%. In airline, the retail arbiter came within about a point of airline's own: like it, 1 to 3 points below the habit alone, with 29% to 84% fewer detours. What carries over is the weights. Tool names differ, so across domains the arbiter weighs Jev's answers by their overall agreement with the agents, and the habit's weight is nearly the same in both (0.27 and 0.25). So one arbiter, fitted once on public traces, can ship with the compiler (details).
Live. GLM-5.3 in Claude Code ran five retail training tasks through the proxy, with the conversation handed to it: tasks 15, 24, 76, 88 and 99, the five-task random sample. All five passed. stretto learn learned a flow from the five logs. Three sessions trained the habit, and two held out the 15 decisions its arbiter was fitted on. Replayed on GLM-5's 40 trial-0 test episodes, it saved 19.0% of turns, against 22.2% for D0, which learned from 831 episodes of four other agents, and 18.7% for the habit alone learned from GLM-5's sessions on the same five tasks. Its arbiter, fitted on 15 decisions, weighted the habit as the larger fits do, so it decided much as the habit alone would. It then ran on the retail pilot's tasks, costliest first. Task 101 took 12 LLM turns, against 19 without a flow and 16 with D0, and task 36 took 14, against 16 and 11. On task 27 the agent failed, as it had without the flow and with D0, and then transferred the customer to a human. The simulated customer kept talking for 19 more turns, although τ²-bench tells it to end the conversation at a transfer. The pilot harness now ends an episode there. That episode cost 61 credits, so the run stopped after three tasks to stay inside its budget. Three tasks show that a flow learned from five recorded sessions runs live, not how many turns it saves.
Live, on all ten tasks. Two flows were then learned from all five sessions and run on all ten of the retail pilot's tasks: the habit alone, and the same habit with the shipped airline arbiter, which no retail decision went into.
- The habit alone took 79 LLM turns, against 110 without a flow and 84 with D0: 28.2% fewer (95% interval 19.0% to 36.9%). On every task it made exactly the lookups that the habit learned from four other agents' episodes made, tools and arguments alike, and that flow took 79 turns too. Their episodes still differed by a turn or three on six tasks, and on the pass of one: that is how much the agent and the simulated customer vary.
- With the shipped arbiter it took 70 turns, 36.4% fewer than without a flow (26.3% to 44.5%), with 2 detours where the habit alone made 6. That is 0.9 fewer turns per episode than the habit alone (−2.3 to +0.2), which ten tasks cannot separate from none.
- Passes. Both passed 8 of 10. They failed the two tasks the arms without a flow and with D0 failed, with the same agent errors. Offline, on GLM-5's episodes, the two flows saved 20.5% and 19.6% of turns.
Live, in airline. The same run in airline recorded its five-task random sample, tasks 1, 15, 38, 41 and 42, and all five passed. On the airline pilot's ten tasks the habit alone took 102 LLM turns, against 127 without a flow and 105 with D0: 19.7% fewer (95% interval −3.3% to 39.9%). It made 8 detours where D0 made 1, all of them reservation reads: after finding the user it reads each of their reservations in turn, and the agent needed only some. The shipped retail arbiter cut them to 1 at no cost in turns, 102 again, with an interval clear of zero (7.6% to 32.0%). It passed 8 of 10 and the habit alone 7, against 8 without a flow and 9 with D0, and none of the failures came from a flow decision. Offline, on GLM-5's episodes, both flows saved 8.9% of turns, what the habit from five of GLM-5's own sessions saves (8.7%).
So in both domains a deployment's first five sessions are enough to start on the habit alone, with a shipped arbiter for precision (details: retail, airline).
5.13Options from the manifest
A flow acts only where training shows a lookup, and offers only the lookups it saw there. So with one airline task it knows two sites and saves nothing (§5.10). --manifest-options offers every read-only tool the server lists, at every site, for Jev to choose from, and binds a lookup that training never made by argument name. The flows were compiled on the sweep's one- and three-task samples and replayed as in §5.10.
In retail it adds nothing. The arbiter keeps to the lookups the habit knows, and saves 10.0% and 20.3% of turns as before, with up to twice the detours. In airline the fitted arbiter gives the habit almost no weight, so Jev's picks bring in flight searches that no trace showed. Bound by name, each argument takes the first value under its name in any earlier result: the flights just returned, or the reservation's, not the trip the customer wants. So nearly all of them are detours, 83 and 91, against 0 and 1 without manifest options. With one task the flow saves 1.1% of turns where it saved nothing, and with three it saves 3.7% against 5.6%. What limits a flow on few traces is less its options than its bindings, and only traces teach those (details).
5.14Matching descriptions to records
RFC-001's other content role for a System-One model is matching what a customer describes to a record. stretto match covers the writes that pick records out of earlier results: the items of an order, the variant an item becomes, a payment method, a reservation. At each such write, it asks Jev which record the customer means. Jev sees only the customer's words and the candidates, not the agent's pick. Over 3,978 such choices in the published trajectories, Jev picked what the task expects 81.6% of the time, against 89.9% for the agents that made them. Where its pick differed from the agent's, the agent was wrong only 19% of the time. As a check on writes it would mostly raise false alarms, and even its confident disagreements catch only 25 of the agents' 402 wrong picks. The customer's words alone are often ambiguous ("yes, the second one"), and showing the agent's messages would let the model copy the agent's pick. The 3,708 questions cost $0.29 (details).
5.15Prompt injection
Jev reads tool results that outsiders may have written, and it "does not treat state as hostile by default". At 300 held-out retail decisions, a note added to the latest tool result told the agent either to stop looking things up or to make the lookup the arbiter liked least. Each decision was then judged again by an arbiter fitted on clean answers. The note to make that lookup moved Jev's pick to it at 205 of the 300, and the flow made it at 86 where it would not have. The note to stop cut the flow's lookups from 168 to 70. So an injection can cost a flow its savings or add detours, all among the site's own lookups. One sentence added to every question, telling Jev that tool results are data and to ignore instructions in them, cut the flow's steered lookups from 86 to 19 and cost nothing on clean decisions. With a threshold of 0.5 in place of 0.3 as well, 4 remained, and the flow kept 124 of its 130 right lookups. The confirmation judge was harder to move: a claim in the agent's message that the customer had already agreed passed 9 of 140 writes it had refused, and a quoted "yes" passed 12. The 2,300 questions cost $0.22 (details).
5.16A third domain
τ²-bench's telecom domain is technical support for a phone line, where the customer runs tools on their own phone while the agent works on the account. It compresses like the others: 30% macro-tool headroom, and the habit's top guess is the agent's next step 69% of the time. Jev is weaker there: weighed with the habit, its answers agree with the agent 67.9% of the time, against 80.5% in retail, and a flow projects 15–26% fewer turns with frequent detours. The questions speak of orders and flights, and the customers' 12 phone actions per episode reach the flow only as what they say. The shipped arbiters did not carry to it. At telecom's 6,343 held-out decisions, telecom's own arbiter made 75% of the agent's lookups and the habit alone 66%, with fewer detours; the retail and airline arbiters made 44% and 50%, and one fitted on both domains' 6,555 decisions 45%. Fitted where handing back was common and Jev more reliable, they hand back too often. The 4,713 questions cost $0.48 (details).
5.17Shadow mode and promotion
RFC-001 §3.7 lets a flow act at a site only after shadow mode shows it agrees with the agent there. stretto-proxy --flow-shadow lets a flow decide and log without acting, and stretto promote scores each lookup it would make against the rest of the session: used if the agent made it later, a detour if it never did. A site is promoted when at least 70% of its lookups were used, the lower bound of a 90% interval on that share is at least 0.5, and they came from at least three tasks. Replayed on GLM-5's test episodes, with each half of the tasks promoted on the other half, D0 kept 277 of its 294 retail turns saved and all 41 in airline, and its detours fell from 58 to 28 and from 26 to 11. The sites held back were where its lookups were not the agent's: further products after a product, and flight statuses in airline. The habit alone gained little, since its detours sit at sites that pass the bar. Scoring had to follow how the flow sees calls, one at a time: scored only at the end of each turn, promotion held back the reservation reads GLM-5 makes in parallel, and cut airline's savings from 41 turns to 4 (results).
5.18Counterfactual evaluation
RFC-001 §3.7 plans to evaluate a changed flow on an earlier flow's logs before it runs. A flow that always takes its likeliest lookup gives inverse propensity weights nothing to reweight, so --flow-explore ε now takes another lookup that binds with probability ε, drawn by the decider's probabilities, and logs every option and the chance of what it took. stretto evaluate estimates another rule from such logs, per site and in total. The direct estimate reads every option's outcome, as a replay or a shadow session labels them. IPS, self-normalized IPS and a doubly robust estimate read the taken option's alone, as a log of a flow that acts has only that one. On GLM-5's test episodes, all four trials, D0 replayed exploring at ε = 0.1 and 0.2, and four rules were estimated from its decisions against their own replays.
The weights are right: for D0 itself and for the arbiter at 0.5, the four estimates agree within about a standard error. Lookups the agent made later carry over to any rule, within 3.5% of each rule's own replay by the direct and self-normalized estimates. Detours do not. About two-thirds of a flow's lookups come right after its own previous lookup, and another rule's chains lie where the logging flow never went. In airline, after the agent's calls, the habit alone and D0 make 8–9 detours each. After their own lookups the habit alone makes 41 and D0 18. So every estimate from D0's logs put the habit alone level with D0, where its replay makes 50 detours to D0's 26. Exploring shifts the logged states too: D0's detours at its explored decisions came to 69 and 106 against 58 in its own replay. Turn credit, the share of a turn each lookup spares, overstates the turns a rule saves by up to 75%, because a turn is saved only when all its calls are. So the estimates can rank changes that keep a flow's chains, such as another threshold on the same decider. A changed decider, or anything priced in turns, needs a replay and then a paired run. Shadow mode has the same blind spot, since a flow that makes no lookups has no chains. The replays reproduce decision for decision from the published answers (details).
5.19Claude models, live
Every live result above had GLM-5.3 play both the agent and the simulated customer. Claude Haiku 4.5 then played the agent on the retail pilot's ten tasks, and Claude Sonnet 5 on three of them, with GLM-5.3 as the customer and D0 as the flow. D0 cut Haiku's LLM turns from 97 to 79, 18.6% fewer (95% interval 10.1% to 27.4%), where it saved GLM-5.3 23.6% on the same tasks. Haiku makes more calls at once, 1.37 per tool turn against 1.05, and a flow spares a turn only when it makes all of that turn's calls, so it gains less, as calling style predicted offline. Of the flow's 30 lookups, 28 were Haiku's own. Haiku passed 7 of 10 with the flow and 6 without. One failure with the flow came just after the flow's only two detours, reads of two products the customer wanted to return. Haiku then returned them, where the policy allows a return or an exchange per order and the customer preferred the exchange. GLM-5.3 made the same error in both arms of the pilot. Sonnet's turns fell from 29 to 22 on its three tasks, with every lookup its own. Then Claude Sonnet 5 played the customer to GLM-5.3 on five tasks. D0 saved 23.1% of turns (13.8% to 32.1%), against 25.8% with GLM-5.3 as the customer on the same tasks, and made the same lookups on four of the five. Every episode passed. So a customer played by the agent's own model had not inflated the savings (details).
5.20Predicate refinement
The arbiter's three predicates were written by hand. RFC-001 §3.4 plans to find them instead: show a proposer where the arbiter is weakest, ask its candidates, and keep one when it explains the agents' steps better. stretto refine now runs that loop. It ranks the sites by the held-out surprise of the agents' steps and writes out the worst decisions at the worst sites, with the state Jev saw. Each candidate is asked alone at every decision, so no other answer changes. It is kept when it raises the held-out log-likelihood by more than BIC's charge for a parameter. One round ran in airline, where the arbiter is weakest. The proposer was Claude, the model running the analysis. It read examples from five sites and proposed eight candidates, five of them about one named lookup: the customer's profile, a fare search, a flight's status. Two were kept: whether the results already answer the customer, and whether a direct-flight search came back empty. Together they took the surprise at the four 2025 agents' 2,042 held-out decisions from 0.654 to 0.624 nats per decision, and agreement from 74.8% to 76.7%. The proposer's examples came from 17 of the 20 test tasks, though, so the loop also scores GLM-5's decisions. Those are judged by the same arbiters, but never fitted on and never shown. There the pair explained nothing: the log-likelihood fell by 1.5 nats, and agreement moved from 72.7% to 73.0%. Nor did it move the savings, 12.6% of turns before and 12.5% after, pooled with lookup first, and 7.1% and 6.5% for GLM-5. So one round found predicates that fit the agents the proposer read, whether from the examples it saw or from those agents' own habits, and the three hand-written ones stay. A kept predicate should raise the likelihood of an agent the proposer never saw. The 20,600 questions cost $1.51 (details).
5.21Flow search
Every flow so far acted on one threshold, 0.3, at every site. RFC-001 §3.10 plans a search over flows with fugue-evo, and stretto search now runs NSGA-II over each site's threshold, from 0.10 to 0.95 or off, and the flow's decider, the arbiter or the habit alone. A site is the call the flow has just seen. The search scores every setting by replaying GLM-5's episodes on τ²-bench's training tasks, on two objectives, turns saved and detours. §5.18's estimates could not price a changed decider's detours, and a replay takes seconds. The hand-set settings and the front were then replayed on GLM-5's test-task episodes, which the search never saw, and set against each decider at one threshold everywhere. In each domain the front held a flow with far fewer detours than D0 on those episodes. In airline it saved 48 turns with 6 detours, where D0 saved 41 with 26. Over the 20 test tasks that is +7 turns (95% interval 0 to +15) and −20 detours (−49 to +2). In retail it saved 293 turns with 30 detours, where D0 saved 294 with 58: −28 detours (−68 to −2). No single threshold did as well. At 0.4 everywhere, the arbiter saved 35 airline turns with 15 detours, and 272 retail turns with 38. Those two flows were the front's best on the test episodes, though. Chosen on the training episodes alone, as the fewest detours at D0's turns saved or more, the flows saved 44 airline turns with 8 detours and 294 retail turns with 49. Each gain came from one or two sites, which flow-diff shows a reviewer. In airline the threshold rose to 0.60 after a flight status, where D0's lookups were mostly detours, and fell to 0.10 after a reservation read, since agents read a customer's reservations one after another. In retail it rose to 0.90 after a product read. Promotion (§5.17) had held back the two sites whose thresholds rose. A second airline search, with another seed, changed the same two airline sites, with 0.15 after a reservation read: 44 test turns with 11 detours.
The airline front's end with the fewest detours did not carry: the arbiter at 0.5 everywhere saved more there. Nor did the habit alone gain much. On two other agents' test episodes, Claude Sonnet 4.5's and Qwen3.5's, airline's best again saved more turns than D0 with fewer detours, but retail's gave up 15 and 33 turns for its fewer detours: those agents made use of the lookups after a product read that it drops. So a search should run on the sessions of the agent the flow will serve. The search read Jev's answers from a cache and never asked, so a decision the cache could not answer handed back. On the training episodes, retail's best handed back 65 decisions that way, and seemed to trade 28 turns for its fewer detours; on the test episodes, with Jev answering, it gave up one. Only 234 distinct questions went unanswered in retail, so a search should let its replays ask. The round's 8,788 questions cost $1.09 (details).
5.22What compiling once can take, and the record the customer named
Could a workflow be compiled once from the traces, branches and writes included, and run with little or no model? No reviewed work does it for τ²-bench: the nearest compiled one intent by hand and declined another[24], and replay without a model works on branch-free, repetitive tasks. Counting what decides each step on ten agents' episodes, replies and the calls that answer the customer are 60–75% of LLM turns. What traces fix is the read skeleton, 19–32% of retail turns and 10–31% of airline turns: the lookups each site may make and where their arguments come from, which is what a flow does. After a tool returns, rules at least 95% sure on training cover 2–44% of retail decisions and 0–16% of airline's, and hold 84–100% on held-out tasks, short of the 99% a write needs. The rest is language, and the probabilities are one agent's profile, regenerated from its sessions, as profile-guided compilation regenerates a profile for a fixed program.
The count also showed where detours come from. Nearly every one is the right lookup of the wrong record, and in airline most came after the customer had named a different one: the agent reads the record asked about and stops, and the flow went on down the list. The binding's chance, scored at the agent's own lookups, cannot see that. A flow now counts, from its traces, how often the agent went on to a record the customer had not named once they had named another. On GLM-5's test episodes airline's detours fell from 56 to 4 at the same turns saved. Compiled into D0's data, the habit alone made as few airline detours as §5.21's searched flow, 6, with 7 more turns saved, and no search: the search's gains came from the two sites where the flow read records the customer had not asked about.
A deployment learns from sessions its flow served. Kept as the agent's calls, the flow's lookups teach as much as clean sessions do; deleted, they teach the flow to switch itself off. Relearned round after round on its own sessions, in replay, a flow drifted to 65% more retail detours, and with the named-other count it holds. Pruning the lookups nothing used, and scoring bindings at the agent's own calls only, cut savings by up to three quarters: both remove the evidence where the flow always acts. A flow also pins each tool's input contract, and the proxy leaves a tool whose contract changed to the agent.
5.23Where compiling once works
Is the read skeleton a limit of the method, or of retail and airline? In τ²-bench's telecom domain the agent troubleshoots a phone by a procedure that branches on what the phone reports[26]. With the customer holding the phone, it is language like the others: replies are 41–62% of turns. But a decision tree over every tool result so far, as decision mining fits one at a process's branch points[27], predicts 63–84% of the nine leaderboard agents' next steps after a tool, 2–9 points above one feature, and its sure rules cover 13–62% of decisions at 96–99.8% held out; in retail and airline the same tree adds nothing. In τ²'s solo mode, where the agent operates the phone itself, reads after a tool are 52–60% of turns and the fixes most of the rest.
Run as the agent, that tree is a workflow compiled once. Fitted on four solo runs' successful training episodes (GPT-4.1 and o4-mini, each under the policy written as a manual and as a workflow), its steps whole calls, it ran with no model in τ²-bench's environment and passed 31 of the 40 held-out tasks, 77.5%, none of whose fault combinations appears in training; refitted on ten resamples of its training episodes, it passed 25–31. The agents it learned from passed 49–78%, with 15–18 LLM turns per episode. It needed a guard against repeating itself, imitation's compounding errors[28] (22 of 40 without it), and many consistent demonstrations (18 of 40 from half the training tasks). Its own confidence was no gate, since it measures agreement among demonstrators. The ticket's stated outcome was: 31 of the 32 runs whose check said done passed, and all 8 it called unresolved had failed; over the ten refits, 16 of 300 and none of 100. Handing only those to an agent projects 87.5–89.4%, with a model in one episode in five.
That check also lets the workflow learn from its own tries, as expert iteration does[30]: fitted on a quarter or half of the demonstrations, then trying each training ticket 17 times and keeping the shortest run the check verified, it went from 18 to 22–23 of 40. Its own runs kept with no retries taught it nothing, so the gain needs a sandbox where a ticket can be tried again. The telecom base tasks share one customer, so the workflow first took its ids as constants. Learned instead as where they came from, and bound again from each run's results (the line holding the ticket's number, the bill whose status sets it apart), they carried to a customer no trace saw: with every name, id and number renamed, it passed the same 35 of 40 tasks, where constants pass 16, and moved to another customer's account, with one line instead of three, it passed 35, where constants pass 19.
stretto's own read-only flow had only been projected for telecom. Replayed, a flow learned from τ²-bench's 2025 telecom runs saved 1.9% of the nine leaderboard agents' turns, with a detour in almost every episode. One tool reads lines, plans and devices alike, and the binding, taking the most recent output first, read the plan of the line it had just read where the agents went on through the customer's lines. Bindings now take a site's sources in the order the agent used them there, counting each read of a batch at the site of the read before it, since a flow makes a batch's reads one after another. The same flow then saved 12.8% of turns with 443 detours instead of 2,936, from 3.0% to 18.5% per agent, and retail and airline replay exactly as before on all nine agents. Of those detours, 145 are bills that four of the agents read too, with a limit that returns the same bills; counting a call as made when a lookup of the same tool has already returned its result, the flow saves 13.3% with 298. Where the agent holds the phone, in τ²'s solo mode, a read-only flow learned from those runs saved 27.5–33.5% of turns in replay, the most any flow has saved, and half the detours it first made went once its binding counted how often the agents read another of the customer's lines after the one with the ticket's number: almost never.
The detours left on an agent D0 never trained on come from somewhere else. For Claude Sonnet 4.5, a session hand-back leaves 57 of 63 airline detours even when told of each at once, because they arrive in one burst after the user lookup. Stopping a chain of lookups by its joint probability, as speculative decoding's draft trees do[29], would drop 135 of GLM-5's 206 airline lookups, every one of them used, since lookups are a set and not a sequence. A flow learned from Sonnet's own traces did no better, and neither did D0's arbiter. What is left is a semantic stop: once the reservation the customer described has been read, Sonnet reads no more and GLM-5 reads on.
6Discussion
What is new. Stated narrowly: branch points decided by calibrated, typed answers weighed against a learned habit; a safety argument that comes from what a flow may do (read) rather than how sure it is; argument bindings trusted per call; and guards whose enforcement is decided by an audit. A skeptical reader could call it AutoTool plus TraceCompiler plus a calibrated router. That is roughly right, and each of those lacks something another supplies.
Why read-only matters. The zero-shot result (§5.3) looked like a dead end: a well-calibrated decider at 72% cannot carry writes. The same decider, weighed and restricted to reads, saves most of the ceiling, because the cost of being wrong fell from a wrong action to a few hundred tokens. The general lesson is to design the action space so that the fast path's errors are cheap, then let it act on weaker evidence.
What the System-One model is for. The design put Jev at every branch point. For read-only flows with plenty of traces, §5.9 shows it is not what saves turns: the habit alone saves as many, offline and live. Jev's weighed answers make the flow stop where the habit would guess, which cuts detours in airline by half to two-thirds. That is precision, not the savings the design claimed. §5.10 seemed to find where it does carry the savings: when traces are scarce. §5.12 narrows that to samples like §5.10's, narrow and scored by an arbiter fitted on thousands of other agents' decisions. From a deployment's own first sessions, drawn at random, the habit alone saved more than an arbiter fitted on those sessions until about twenty of them, and an arbiter fitted elsewhere mattered only on the narrow draw. So a deployment can start on the habit alone and fit its own arbiter as its sessions accumulate. In a domain unlike the ones an arbiter was fitted on, that is the only safe start: in telecom, the shipped arbiters made fewer of the agent's lookups than the habit alone (§5.16). Where Jev has earned its place so far is in judging what the habit cannot see. It judged a customer's confirmation better than a word list, and a second question catches the costliest lapses, a call that differs from what the customer agreed to, at the price of false alarms (§5.11). Matching a customer's description to a record, it did worse than the agents it would check (§5.14). And one round of new predicates, proposed where the arbiter was weakest, fitted the agents the proposer read but not another agent (§5.20).
What an injection can do. The security argument is structural: a flow reads and never writes, and it only chooses among lookups compiled from traces, so no answer takes it outside that set (§3.3). §5.15 shows why the structure matters: inside the set, Jev followed injected instructions readily, and a note in a record redirected nearly a third of the flow's decisions. Instructions in the questions and a higher bar where outside text is present absorb most of it. The confirmation judge weighs the customer's own reply, which an outsider who controls only records cannot write, and an injected claim moved it at 6–9% of the writes it refused. It should not be enforced without the guards, which check facts rather than wording.
Who benefits. Savings follow the agent's calling style more than its accuracy. Sequential callers gain the most, and the heaviest parallel callers little in airline. Live, Claude Haiku 4.5, which makes more calls at once than GLM-5.3, saved 18.6% of turns where GLM-5.3 saved 23.6% (§5.19). A deployment should measure its own agent's runs before expecting a number, which stretto's Phase 0 does from logs alone.
Limitations. The live pilots are ten tasks per domain, one episode per arm, and the simulated customer is the same model as the agent. Where a Claude model played the customer instead, on five tasks, the savings held (§5.19). The reward omits τ²-bench's natural-language assertions. Agreement is measured against agents that fail 21–50% of tasks. Provenance is a text-matching heuristic, and the guards were compiled by an LLM and reviewed only through the audit. The comparisons in §5.9–5.10 are replays, which assume the agent would have acted the same with the flow's results in hand. So are §5.21's, §5.22's and §5.23's detours, and neither the flows the search chose nor the named-other count has run live; §5.23's compiled workflow ran in τ²-bench's own environment, but on tasks that share one customer (renamed once, to test its binding), with the agents' pass rates averaged over four trials against its one run, and its hand-back is projected from the agents' own per-task rates; §5.22's relearning holds the agent's steps fixed, so it measures only the flow's side of the loop. §5.10's samples were clustered by task id (see its correction), and §5.12's cold start rests on five draws of five sessions from one agent, and live on ten tasks per domain, one episode per flow. The confirmation labels in §5.11 come from one annotator, the model that ran the analysis, who labelled one write differently in the two rounds. The injection test (§5.15) used two fixed notes in one domain, placed where Jev always sees them, and scored decisions rather than episodes. The offline results use mid-2025 models, and GPT-4.1 simulated the customer where the leaderboard used GPT-5.2, so the targets differ from the baselines in more than the agent.
7What comes next
- Pass^1 to within a point. The paired run on every test task (§5.6) bounds the change to −7.5 to +6.25 points. A one-point bound would take about 4,300 pairs. Next is the same run on more agent models. §5.19 took two Claude models live on ten tasks and on three, and the projection favors the sequential callers.
- The cold start on other agents. In both domains, live, the habit from a deployment's first five sessions saved what a flow compiled from four other agents saved, and a shipped arbiter from the other domain added precision (§5.12). All of those sessions were one agent's, and offline there were five draws. The shipped arbiters did not carry to telecom (§5.16), and options from the manifest did not help: a lookup no trace showed needs bindings, and only traces teach those (§5.13).
- Independent labels for the judge. Enforced live, the judge refused none of 40 writes, and logged, its fails were two lapses and two false alarms (§5.11). Every label so far is one annotator's, the model that ran the analysis.
- Confirmed writes in one call.
stretto_commitis built, and no live run has used it: bundling consecutive confirmed writes could save at most 0.2–3.5% of LLM turns, too little for a small pilot to see, so a live test waits for tasks that write more. Namedplan_*flows are not planned: §5.9 bounds what naming could add at 2% of turns in retail and 8% in airline. - The named-other count, live. Replayed on held-out tasks, a flow that counts the records a customer did not ask about made as few detours as the searched flows, with no search (§5.22). Next is a paired live run of the habit with it against D0 (#34), and flows serving a smaller agent, where what is left to a model, language, is what a small model would take (#8).
- Compiled procedures, where the agent works alone. In telecom's solo mode a workflow compiled once from traces passed as many held-out tasks as the agents, with no model, and its check of the ticket's stated outcome caught all but one of its failures (§5.23). Its identifiers, learned as where they came from, bind for customers no trace saw. Next is the procedure inside stretto, with its fixes behind the guards, and accounts in states no trace shows; and, for flows, the stop at the record the customer described where the description is language, such as a city for an airport code, since an exact value is now counted.
- Guards where they bite: agents that make the writes the guards refuse, such as the smaller or cheaper models the design plans to run on flows compiled from frontier traces. Claude Haiku 4.5 ran on D0 without the guards (§5.19).
- Predicate refinement, held out by agent. One round of proposing predicates where the arbiter is weakest, in the style of counterexample-guided refinement[F2], fitted the agents the proposer read and not GLM-5 (§5.20). The next rounds should show the proposer one set of agents and keep only what explains another's steps.
Each of these, and the rest of what is left, is a GitHub issue. The roadmap groups them.
8Reproducing the results
Phase 0 needs no keys and runs in about 20 seconds on a laptop:
git clone --depth 1 https://github.com/sierra-research/tau2-bench ../tau2-bench
cargo run --release -p stretto-report -- phase0 --tau2 ../tau2-bench --out reports/phase0.md
cargo run --release -p stretto-report -- guards --tau2 ../tau2-bench
Every Jev answer behind §5.3–5.5 and §5.9–5.23 is published in the repository and replays without a key (stretto import-answers, then --oracle replay). Asked the same questions a second time, Jev chose the same option 97% of the time, and §5.3's headline figures moved by at most 0.3 points. §5.9, §5.10 and §5.13 replay from them too (pilot/check_flow.py --flow-decider arbiter|habit, with flows from stretto compile --train-tasks), and so do §5.11 (stretto confirm --second-question proposed), §5.12 (flows from stretto learn --results, and the shipped arbiters from compile --pooled-arbiter and export-arbiter), §5.14 (stretto match), §5.15 (scripts/injection.py, with stretto ask), §5.16 (stretto fit-arbiter and scripts/compare_arbiters.py), §5.17 (stretto promote), §5.18 (stretto evaluate, on decisions from pilot/check_flow.py --explore), §5.20 (phase0 --candidates, then stretto refine), §5.21 (stretto search, replaying through pilot/check_flow.py), §5.22 (scripts/anatomy.py, and flows from stretto learn and compile replayed the same way) and §5.23 (scripts/anatomy.py, scripts/telecom_workflow.py in τ²-bench's environment, and scripts/detours.py on the replays). The live episodes behind §5.6, §5.11, §5.12 and §5.19 are published with the proxy's logs. All the System-One questions behind this paper cost about $13, and the live pilot's harness is in pilot/. Per-phase results, with every number above, are in docs/results/, and the design and its amendments are RFC-001.
References
- R. Koohestani. AgentGuard: Runtime Verification of AI Agents. ASE 2025 AgenticSE workshop. arXiv:2509.23864
- R. Koohestani et al. TriCEGAR: A Trace-Driven Abstraction Mechanism for Agentic AI. Preprint, 2026. arXiv:2601.22997
- H. Wang, C. M. Poskitt, J. Wei, J. Sun. ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction. ASE 2026. arXiv:2508.00500
- P. T. Tran-Truong, X.-B. Le. Measuring the Unmeasurable: Markov Chain Reliability for LLM Agents. Preprint, 2026. arXiv:2604.24579
- Z. Chen, M. Kang, B. Li. ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning. ICML 2025. arXiv:2503.22738. H. Wang et al. AgentSpec. ICSE 2026. arXiv:2503.18666
- F. Fournier, L. Limonad, Y. David. Agentic AI Process Observability: Discovering Behavioral Variability. PMAI 2025. arXiv:2505.20127
- L. Lin et al. Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents. Preprint, 2026. arXiv:2607.06873
- S. Cho et al. Automata from Agent Traces: Failure and Next-Step Prediction. Preprint, 2026. arXiv:2608.23670
- I. D. Lopez-Miguel et al. ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning. MODELS 2026. arXiv:2608.14352
- X. Huang et al. PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors. Preprint, 2026. arXiv:2605.06455
- M. Tappler et al. Automata Learning meets Shielding. ISoLA 2022. arXiv:2212.01838
- X. Liu et al. ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph. Preprint, 2024. arXiv:2403.00839
- J. Jia, Q. Li. AutoTool: Efficient Tool Selection for Large Language Model Agents. AAAI 2026. arXiv:2511.14650
- Y. Sui et al. Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving (Act While Thinking). Preprint, 2026. arXiv:2603.18897
- N. Ye et al. Speculative Actions: A Lossless Framework for Faster AI Agents. ICLR 2026. arXiv:2510.04371
- Z. Liu, S. Kundu, P. A. Beerel. Speculative Macro Commit for Faster Tool-Using Agents. MLSP 2026. arXiv:2609.03236
- A. Yagubyan. How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines. Preprint, 2026. arXiv:2605.28840
- Y. Ruan et al. Identifying the Risks of LM Agents with an LM-Emulated Sandbox (ToolEmu). ICLR 2024. arXiv:2309.15817
- Z. Guo et al. StableToolBench (arXiv:2403.07714) and MirrorAPI (arXiv:2503.20527); Z. Ren et al. GTM (arXiv:2512.04535).
- H. Chae et al. Web Agents with World Models. ICLR 2025. arXiv:2410.13232
- G. Ganapavarapu, D. Patel. MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments. Preprint, 2026. arXiv:2605.09131
- Z. Wang et al. Agent World Model. ICML 2026. arXiv:2602.10090
- R. V. Hemadri et al. R2V Agent: Teaching SLMs When to Ask for Help (arXiv:2605.16604); D. Piatrashyn et al. ReDAct (arXiv:2604.07036). Preprints, 2026.
- S. El Yadouni, G. Li. TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows. Preprint, 2026. arXiv:2608.02680
- Z. Z. Wang et al. Agent Workflow Memory. ICML 2025 (arXiv:2409.07429). Q. Zhang et al. Agentic Plan Caching. NeurIPS 2025 (arXiv:2506.14852).
- V. Barrès, H. Dong, S. Ray, X. Si, K. Narasimhan. τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment. 2025. arXiv:2506.07982
- A. Rozinat, W. M. P. van der Aalst. Decision Mining in ProM. BPM 2006. doi:10.1007/11841760_33
- S. Ross, G. Gordon, D. Bagnell. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS 2011 (arXiv:1011.0686). O. Bastani, Y. Pu, A. Solar-Lezama. Verifiable Reinforcement Learning via Policy Extraction. NeurIPS 2018 (arXiv:1805.08328).
- Y. Li, F. Wei, C. Zhang, H. Zhang. EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees. EMNLP 2024 (arXiv:2406.16858). K. Huang, X. Guo, M. Wang. SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths. 2024 (arXiv:2405.19715).
- T. Anthony, Z. Tian, D. Barber. Thinking Fast and Slow with Deep Learning and Tree Search. NeurIPS 2017 (arXiv:1705.08439). E. Zelikman, Y. Wu, J. Mu, N. D. Goodman. STaR: Bootstrapping Reasoning With Reasoning. NeurIPS 2022 (arXiv:2203.14465).
- E. Debenedetti, I. Shumailov, T. Fan, J. Hayes, N. Carlini, D. Fabian, C. Kern, C. Shi, A. Terzis, F. Tramèr. Defeating Prompt Injections by Design. 2025 (arXiv:2503.18813). L. Beurer-Kellner et al. Design Patterns for Securing LLM Agents against Prompt Injections. 2025 (arXiv:2506.08837).
- A. K. Lew, T. Zhi-Xuan, G. Grand, V. K. Mansinghka. Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs. 2023 (arXiv:2306.03081). J. Loula et al. Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo. 2025 (arXiv:2504.13139). L. Wong et al. From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought. 2023 (arXiv:2306.12672).
- Z. Z. Wang, A. Gandhi, G. Neubig, D. Fried. Inducing Programmatic Skills for Agentic Tasks. 2025 (arXiv:2504.06821). B. Zheng et al. SkillWeaver: Web Agents Can Self-Improve by Discovering and Honing Skills. 2025 (arXiv:2504.07079).
- [F2] E. Clarke et al. Counterexample-Guided Abstraction Refinement. CAV 2000.
- [F3] A. P. Dawid, A. M. Skene. Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm. JRSS C, 1979.
- [F6] M. Dudík, J. Langford, L. Li. Doubly Robust Policy Evaluation and Learning. ICML 2011.
- TypeSafe AI. Introducing System One Models & Jev. typesafe.ai; API documentation at docs.typesafe.ai.