Changelog
The stretto CLI and stretto-proxy are the product. The library crates (stretto-trace, stretto-oracle, stretto-model, stretto-report) make no stability promise before 1.0. The file formats carry versions of their own (docs/formats.md).
Unreleased
The console
stretto-console, a new crate and binary: the server of stretto's management plane, a local web app over the data directory (~/.stretto,--dataor$STRETTO_HOME). Its JSON API lists and shows the recorded sessions call by call, with each of the flow's lookups and the decision that made it; the flows, asstretto flow-showreviews them and asstretto flow-diffcompares them; a registry of the MCP servers stretto fronts (servers.json), with the host configurationstretto initprints for each and a live test of the connection; and jobs that run thestrettoCLI (learn,promote,audit,redact,doctor). It requires a token (a cookie, or a bearer), a header on every change, and a loopback Host under--no-auth;--read-onlychanges nothing (its README). The UI isconsole/; its API types are generated from the Rust.The console's UI (
console/, Vue 3), whichstretto-consoleembeds:- an overview of what stretto recorded and learned, by domain, with its health checks;
- the servers, with their form, the configuration for each host and the connection test;
- each session as a timeline: the conversation, the agent's calls, and the lookups the flow read ahead after each call, with their probabilities against the threshold;
- each flow's graph at a threshold of your choosing, its sites, bindings and tools, its review, and a comparison with another flow;
- the jobs, run from the page or again from an earlier job's parameters, with their output as it comes.
It has light and dark themes, works from a phone's width up, and has a command palette (Ctrl-K or ⌘K). It follows the data directory as it changes.
make consolebuilds it and runs the console.The console ships with the rest. The release archives,
install.sh,install.ps1and the Homebrew formula installstretto-console, with its UI built in; the installers still install a release from before it. A second image,ghcr.io/alexnodeland/stretto-console(the Dockerfile'sconsoletarget), runs it on port 8080 with a health check, andcompose.yamlruns it over your~/.stretto. docs/console.md is its guide.The console's jobs can be cancelled, from the job's page or with
POST /api/jobs/:id/cancel: a queued job never runs, and a running one'sstrettois killed. A job that ends so iscancelled.stretto init --upstream URL [--upstream-header NAME=VAR]configures the proxy in front of a Streamable HTTP server.stretto doctor --data DIRchecks a data directory other than~/.stretto; the console's doctor job checks the one it serves.stretto_report::review::viewreturns a flow's review as data, whichflow-showrenders from;review::diff's result lists its changes by heading and serializes.
Documentation
- The documentation site is at stretto.alexnodeland.com. Links to the old address, alexnodeland.github.io/stretto, redirect there.
- The documentation site builds for the address GitHub Pages gives it (
SITE_URL, from the Pages workflow), so its base, canonical links and sitemap follow a custom domain. A link from the repository's Markdown to the site itself stays on the site: the roadmap's and the changelog's links to the documentation site opened the research notebook before.
Development
- A
Makefilewith the everyday commands (make help).make ciruns what CI's check job runs;make coverage,make msrv,make bless,make quickstart,make walkthrough,make site,make dockerandmake docker-consolerun the rest. - CI also runs the doctests, builds the API docs with rustdoc warnings as errors, and checks the workspace on Rust 1.88, the
rust-version, with the committed lockfile. A coverage workflow measures line coverage with cargo-llvm-cov and fails under 78%; its LCOV report is an artifact, and goes to Codecov when aCODECOV_TOKENsecret is set. - A Claude Code setup:
CLAUDE.md, which importsAGENTS.md, and.claude/, with permission rules, a hook that runs rustfmt on each edited Rust file, a session-start hook for Claude Code on the web, four skills (check,results,release,steward) and two subagents (reviewer,claims-checker). - A dev container (
.devcontainer/) and VS Code settings (.vscode/). - A debug build of
stretto-consoleserves the UI inconsole/disteven when the UI was built after it (make ci, thenmake console). Before, such a build served only the page that says how to build the UI. - CI's
uijob checks the console's UI: Prettier, ESLint, vue-tsc, Vitest and the build (make ui-check). It also runs Playwright in Chromium on a mock API and on the console serving its test fixtures (make e2e), and fails on any error in the browser's console. On the mock API it also runs axe-core's WCAG 2.1 A and AA rules on every page in both themes, and on the palette, a dialog, the phone layout, first run and sign-in.make typesregenerates the API's TypeScript, and CI's check job fails when it has drifted.
0.1.0 (2026-09-27)
The first release: RFC-001's design, built and measured, with release binaries, installers, a container image and a documentation site. docs/results has every result, and the working paper puts them together. The development log below lists every change on the way here.
Install and first run
- Each release carries archives for five targets (Linux x86_64 and aarch64, macOS on Apple silicon and Intel, Windows x64) with their
SHA256SUMS,install.shandinstall.ps1, which check an archive's checksum before installing it, and a Homebrew formula. A container image,ghcr.io/alexnodeland/stretto, is published for amd64 and arm64; it runs as a non-root user, with/dataas its home (install). stretto init --host claude-code|claude-desktop|cursor|vscode -- <server…>writes an MCP host's configuration for a server behind the proxy, and prints the steps to a served flow.stretto doctorchecks the installation: the binaries on PATH,~/.stretto, whether a key is set (never its value), and the flows and sessions.stretto completions <shell>prints shell completions.examples/quickstart/records, learns, reviews and serves a flow onstretto-mcp-demo, with no key, in under a second; CI runs it.- Building from source needs Rust 1.88.
Measure an agent (stretto phase0)
- How compressible an agent's behavior is, from τ²-bench's published trajectories, with no key: a hierarchical Dirichlet habit of the next tool call, its concentration fitted with fugue, coverage and agreement on held-out tasks, transfer between agents, macro-tool headroom and argument provenance.
- With
--oracle, System-One questions at every held-out decision (Phase 0b):v1, over every tool, andv2, read-only lookups with a state slice, a stop question, yes/no predicates (data/predicates-v2.json) and an arbiter that weighs the answers against the habit. Answers go through a replay cache, and published answer bundles replay without a key (export-answers,import-answers).stretto askanswers questions of your own through the same cache.--domain telecommeasures τ²-bench's third domain.
Flows
stretto compilewrites a flow (the flow IR,stretto_flow: 1) from τ²-bench results.stretto learnlearns one from sessions the proxy recorded: with the habit alone (--habit-only), with an arbiter fitted on held-out sessions, or with a shipped arbiter (--arbiter-from). A flow's lookups bind their arguments from earlier results, including results that list values one per line.stretto serveandstretto flow-serveanswer a flow's decisions over a local port.stretto export-arbiterwrites a flow's arbiter to its own file (stretto_arbiter: 1), andstretto fit-arbiterfits one on the decision logs of several compiles.data/arbiters/ships two, fitted on four agents' retail and airline decisions; they carry between those two domains, not to one unlike both.stretto auditscores recorded sessions under a flow, run as a fugue program: agreement, calibration and surprise per site. A flow without an arbiter is audited with its habit.stretto promotekeeps a flow to the sites where its lookups were the agent's own, scored on recorded sessions or τ²-bench results; a promoted flow hands back after every other call.stretto-proxy --flow-shadowrecords the sessions to promote on: the flow decides and logs, but makes no lookups.stretto flow-showrenders a flow for review: the tools it may call, the lookups it may make after each call and where their arguments come from, and how it decides.stretto flow-difflists what changed between two flows for a pull request, and exits with 1 when a change needs review (reviewing flows).- Deciders. A flow decides with its arbiter where it has one, and otherwise with
reach: the chance that the agent makes a lookup before its next write, counted from traces, times its binding's, against the threshold. It asks no model and needs no key.serve,flow-serve,audit,promote,initandstretto-proxy --flow-decidertakehabit,reachorarbiter. - Bindings. A lookup's arguments come from earlier results, tried in the order the agent used them at the site.
learn --constantslearns an argument the agent always passed with one value, never a string a user wrote with a digit or an @; a list argument takes every value at the path of an earlier result; a search the agent always narrows is not made bare; and a pick of a record other than the one the customer named is scored by how often the agent made it. - Flow format 2 holds per-site thresholds, constants, lists and those counts; this release reads formats 1 and 2. A flow learned from recorded sessions pins each tool's input contract, and the proxy makes no lookup of a tool whose server now lists another.
stretto evaluateestimates what another rule would have done from a flow's exploring decisions (--flow-explore), directly and by IPS, self-normalized IPS and doubly robust estimates.stretto refinekeeps candidate predicates that raise the arbiter's held-out likelihood.stretto searchruns NSGA-II over each site's threshold and the decider.
Compiled procedures (stretto-procedure)
- Where no user speaks, a whole workflow, writes included, compiled once from agents' traces is a file stretto runs: the procedure IR (
stretto_procedure: 1, formats), with a decision tree per site over the run's state, identifiers bound again at run time from where they came from, a guard against repeats, and the checks of each outcome a ticket may state. stretto-procedure --procedure FILE --ticket TEXT -- SERVER…runs it against an MCP server with no model and prints its calls, its check of the ticket's outcome, and its verdict:resolved,transferred, orhand_backwhen the check failed and the ticket should go to a model.
Checks on writes
stretto guards: typed policy checks for τ²-bench's retail and airline domains, audited against the published trajectories.stretto confirm: a System-One model judges whether the customer confirmed each write, with an optional second question.stretto matchasks which record a customer means.
stretto-proxy
- A stdio MCP proxy that forwards byte for byte and records sessions (
stretto_mcp_log: 2).--upstream URLwraps a Streamable HTTP server instead of a command, with headers from the environment that are never logged. - Active mode:
--flowruns a flow behind the agent's calls and appends its lookups to the result.--guardsrefuses writes a policy check fails.--confirm-judge log|enforceadds the confirmation judge, and--confirm-second-shadowasks its second question but only logs the answer.--commitadds a tool for confirmed writes in one call.--contextreads the conversation the host writes. - Jev's key can come from a file:
TYPESAFE_API_KEY_FILEnames it whenTYPESAFE_API_KEYis not set. stretto-mcp-demo, a tiny server for trying it.--retain-days Ndeletes, at start, the logs and cached answers older than N days.- Each run of a flow after one of the agent's calls is a fugue program (
program.rs, RFC-001 §3.2):- The sites. A decision site
decide#icomes before each lookup, and the flow's arbiter decides it. An outcome siteoutcome#icomes after, and takes the server's answer and scores it. - The program is in the flow. The flow IR holds it as
program, in fugue's serializable program format (RFC-001 §3.5), and it is checked when the flow loads. A flow written without one runs the standard run, which draws exactly what the Rust it replaced drew.flow-showprints the program, andflow-difflists a change to it as needing review. - Its distributions. The flow registers two,
DecideandOutcome, from its own statistics: a flat Dirichlet's predictive over what the agent did next at a site, and a flat Beta's over how often a tool succeeded, from fugue's conjugate helpers. They carry their site as metadata (WithMeta). - Three interpreters. The proxy runs the program with fugue's
run_async.PriorHandlersimulates a flow with it, andScoreGivenTracescores a recorded run. - The flow log. Decision lines carry their
address, and each run ends with arunline holding its trace and its surprise.
- The sites. A decision site
--flow-tools TOOLSnames the only tools a flow may call on its own, and--flow-explore EPSILONexplores at read-only sites for counterfactual evaluation.
Privacy
stretto redactwrites a pseudonymized copy of recorded sessions: a value fewer than--keep-sharedsessions contain becomes a salted hash, the same wherever it appears, and--hash-fieldhashes named fields however many sessions share them. A flow learned from the copy matches one learned from the originals, on synthetic sessions, 40 real ones and 5 with their conversation (privacy).stretto learn --sessionsandredactskip the confirmation logs the proxy writes beside the sessions; they failed on them before.
Live runs and research tooling
- The pilot harness runs a Claude model as the agent or the customer (
run_episode.py --agent-cli claude,--customer-cli claude --customer-model, throughpilot/claude-agent.sh), and Z.ai credits count the GLM side only. The guards arm takes the confirmation judge (--confirm-judge,--confirm-second,--confirm-second-shadow), with Jev's key handed to the proxy in a file, and--labelnames an arm's directory. pilot/run_paired.pyruns every test task of a domain in both arms, reusing a pilot's pairs, under a Z.ai credit budget that it checks against a ledger before each episode.pilot/analyze_paired.pyreports on the result: paired pass rates with a bootstrap interval and McNemar's test, turns and tokens saved, detours and their token cost, and a pooled estimate across domains (the paired run). Itsarmscommand compares any arms run on the same tasks (the cold start, live).pilot/run_episode.py --soloruns τ²-bench's no-user mode live, and--handbackhas the agent take over from a compiled procedure.--batch-readsadds Anthropic's sample prompt for parallel tool calls to the agent's system prompt, the prompting baseline.pilot/run_trials.pyruns a paired design with Claude models as agent and customer, resumably, within a Claude Code subscription's five-hour and seven-day windows;pilot/analyze_trials.pyreports pass^k, turns, tokens, cost and time per arm and compares the arms, andpilot/analyze_recorded.pycompares arms recorded at different times.pilot/bench/runs AgentDojo and BFCL tasks live, their tools served over MCP by the benchmarks' own code and each episode scored by the benchmark's check.scripts/holds every analysis behind the results: the ceilings by where arguments come from (ceiling.py), the costs and the threshold (costs.py,per_decision.py,priced.py), learning curves, repeated reads, the replays of six more benchmarks from their published trajectories (scripts/bench/), and the paper's figures.
Documentation
- The walkthrough runs the whole loop on the official MCP filesystem server, and CI runs it.
- Reviewing flows, with example diffs of real flows that a test keeps current.
- Privacy: what each file holds, what is sent to the System-One model, retention and redaction.
- The file formats, the CLI reference (generated from the code, kept current by CI), the design and the roadmap.
- The working paper, published as the research notebook at stretto.alexnodeland.com/notebook.
- The documentation site (
website/, VitePress): a guide, integrations, the reference, the research and the community pages. - The brand kit: the mark, color tokens, an interactive explainer, and the explainer and walkthrough videos, narrated and captioned.
- The working paper as an arXiv-ready LaTeX manuscript (
paper/latex/), built from the Markdown. CONTRIBUTING.md,SECURITY.md,CODE_OF_CONDUCT.md,CITATION.cff, and issue and pull request templates.
Release
- Install a release with
install.sh,install.ps1, Homebrew or the container image (install), or build the tag:cargo install --locked --git https://github.com/alexnodeland/stretto --tag v0.1.0 stretto-proxy stretto-report. It is not on crates.io: it depends on fugue'sprogramfeature, which is not in a fugue release yet. - The shipped arbiters stay in
data/arbiters/; take them from a checkout or from the tag. - A version tag, or the release workflow run by hand, makes a GitHub release with that version's section of this file as its notes (releasing).
Development log before 0.1.0
Every change made before the first release, newest first, as it was logged.
Install and first run.
- Each release now carries:
- archives for five targets (Linux x86_64 and aarch64, macOS on Apple silicon and Intel, Windows x64), with
SHA256SUMS; install.shandinstall.ps1, which check the archive's checksum before installing;- a Homebrew formula.
- archives for five targets (Linux x86_64 and aarch64, macOS on Apple silicon and Intel, Windows x64), with
- A container image,
ghcr.io/alexnodeland/stretto, is published on each version tag, for amd64 and arm64. It runs as a non-root user, with/dataas its home. stretto init --host claude-code|claude-desktop|cursor|vscode -- <server…>writes an MCP host's configuration for a server behind the proxy, and prints the next steps.stretto doctorchecks the installation: the binaries on PATH,~/.stretto, whether a key is set (never its value), and the flows and sessions.stretto completions <shell>prints shell completions.examples/quickstart/records, learns, reviews and serves a flow onstretto-mcp-demo, with no key, in under a second, and runs in CI. docs/install.md and docs/releasing.md cover installing and cutting a release.- Building from source needs Rust 1.88.
strettoruns on a thread with an 8 MB stack. On Windows, whose main thread has 1 MB, a debug build overflowed it parsing the command line. CI now builds on macOS and Windows and runs each binary's--version.
- Each release now carries:
Live runs on other benchmarks (
pilot/bench/).bench_mcp.pyserves an AgentDojo or BFCL task's tools over MCP, running each call with the benchmark's own code.run_bench_episode.pyruns one episode in Claude Code and scores it with the benchmark's own check. It stops with a clear error when the task's server never started.run_bench_paired.pyruns every held-out task in both arms, resumably. It runs every task's first arm and then every task's second, so that two arms sharing a prompt prefix do not share the prompt cache.run_bench_paired.py --split trainruns the training tasks, to record an agent's own sessions forstretto promote.live_to_tau2.pywrites a live run's episodes as τ²-bench results, so the same episodes can be replayed from the record. Replayed there, the no-flow episodes projected 11 saved turns, against 10 saved live where the agent used a lookup. They found 3 of the 52 live detours, because the record cannot follow a chain through tools the agent never called (pilot/trace_env.pynow says so).analyze_bench.pypairs the arms by task, each arm known by its label, so a promoted flow can run as an arm of its own. It reports:- bootstrap intervals over tasks, with a task's pairs for all models drawn together;
- an exact sign test;
- rows pooled over models;
- named sets of domains (
--set); - the pairs in which the flow made no lookup, which measure run-to-run variation.
Results for GLM-5.3 and Claude Haiku 4.5, 244 episodes (results):
- In AgentDojo's Slack and travel suites, 10.1% fewer LLM turns (5.8% to 14.0%), with passes unchanged.
- Over all of AgentDojo, 6.0% fewer, where the replay projected 7.2%.
- No effect in BFCL, whose replay projected 1.6%.
- Promotion, live on Slack: scored on each agent's own training sessions, it removed every detour and the saved turns with them. One site served both kinds of task, and only the request tells them apart. It also kept lookups of calls the agent had already asked for in the same turn (#40).
A flow without an arbiter decides with
reachby default, the chance of each lookup before the agent's next write.stretto-proxyserves such a flow withreachunless--flow-decidernames another. Before, it refused one without--flow-decider habit.stretto promotescores it withreachtoo, andstretto initwrites--flow-decider reachand passes--decider reachto the promote step it prints.stretto flow-showandflow-diffshow what it does at each site withreach.- A flow learned before flows held reach's counts is served with the habit, as before.
stretto auditstill scores a flow without an arbiter with the habit, the model of the agent's next step. - The quickstart and the walkthrough serve with the default, and save as many calls as with the habit.
- With
reach, a hand-back now names the likeliest lookup under the threshold and its probability, instead of "no lookup to make".
stretto auditandstretto flow-showname the reach decider when a flow is served with it, instead of the arbiter.The documentation site (
website/, VitePress), deployed to GitHub Pages: a guide, integrations, the reference, the research and the community pages. The research notebook moves to/notebook/.The brand kit (
brand/): the mark, the wordmark and lockups, color tokens with a contrast check, messaging, the social card, an interactive explainer, and the launch and walkthrough videos with their sources. The videos are narrated, with WebVTT captions, and the explainer reads each step aloud when its sound is on. The voice is Kokoro-82M, an open text-to-speech model run locally (brand/video/narrate.py).CONTRIBUTING.md,SECURITY.md,CODE_OF_CONDUCT.md,CITATION.cff, issue and pull request templates, and.editorconfig.DTap-Bench's legal, finance and research domains (200, 200 and 160 tasks) are counted for their ceilings and replayed separately from the six (
scripts/bench/replay.sh dtap-more,tables.py'sdtap_more_by_source): reads are 39–77% of their turns and the ceiling 1.4–14.9%; in legal the use-before-write speculator saves 4.7% of turns against 3.9% with the agents' own runs (+0.8 points, 0.7–1.0) and ties with flows from all seven agents, and its scores there are better calibrated (an expected calibration error of 0.009 against 0.042). Its lead is each agent's own reading ahead, which pooled counts average away; weighing the others as a prior (the agent's own runs plus 100 of theirs) keeps it, at 5.6% of turns against 4.7%; over the domains with reads to take, the prior gains in legal and CRM and costs 0.5–0.8 points of the own flows' savings in the other four. In finance and research the two tie. Its browser and macOS domains (34 and 30 tasks) are counted only: where calls act on a screen, almost nothing is in the ceiling.dtap_to_tau2.pyreads abrowse(finance'sbrowse_stock) and leaves the Claude Agent SDK'sAgenttool out of reads and writes, asTask; no published label changes.scripts/ceiling.pycounts, per episode, whether every value of every call, writes included, came from a result, a constant or what the user wrote, with none the agent composed: 4–37% of the episodes in the benchmarks without a simulated user, and at most 8% with none from the user's words either, so compiling a procedure once there needs a model to read the request, and most often one to write. Figure 2's per-agent counts of the fifth ceiling are now the single-pick ones its summary and text report (64 rows had come from an experiment that walked the model's ranking).Repeated reads, counted for every replayed agent (
scripts/remade.py --results): with no write between, agents make again 0.8% of their reads on τ²-bench, 1.0% on τ-bench, 0.9% on BFCL, 1.5% on AgentDojo and 2.6% on WorkBench, most of them by a few models (Qwen3.5 Flash 17.5%, Llama 3.3 70B 6.2%), whose replayed savings are upper bounds, as gpt-oss-120b's are on DTap-Bench.scripts/ceiling.pycounts a fifth ceiling, words by model: a value of the user's counts as bindable when the System-One model, shown what the user wrote so far and the call's tool and argument, picks it among the spans of the user's words.--questionswrites the questions,scripts/model_questions.pyasks them throughstretto askand its replay cache, and--model-answersreads the picks back; every other count is unchanged. Of what the user's words add, the model takes 39% in WorkBench, 53% in BFCL, 48% in AgentDojo, 30% in DTap-Bench, 53% in τ-bench and 50% in τ²-bench, where a pattern takes 0–31%; MCPMark's composed values are rarely a span of the request (2%). The picks are published (results).scripts/ceiling.pycounts a fourth ceiling, words by shape: a value of the user's counts as bindable when it is the only one of its shape (an email, an id, a number or a date) in what they wrote so far, as a pattern could take it with no model. Every other count is unchanged. Of the 3–7 points the user's words would add in BFCL, AgentDojo, WorkBench and DTap-Bench, a pattern recovers 0–4% in the first three and MCPMark, and 17% in DTap-Bench.Constants, opt-in (
stretto learn --constants, flow format 2'sbindings.constants). An argument the agent passed with one value in every call of a lookup that passed it, at least five calls and at least half of the lookup's, and that was never found in an earlier output, such as a page size, is learned as a constant, unless it is a string a user wrote with a digit or an @, such as an id or an email (of the published flows, that drops only the repository EasyR1, which Kimi K2's MCPMark GitHub users named, where the flows save nothing either way): the binding passes it as the agent did, so that the lookup is the agent's own call, and a required argument with a constant needs no source.flow-showlists them, andflow-diffflags a new or changed constant for review, since it could be one user's value that every training session shared. Without the flag,learnwrites the same flow as before.scripts/bench/replay.shgainsmcpm-constanddtap-const: each model's own flows save 21 of MCPMark Notion's turns instead of 13, with 1.6 detours an episode instead of 2.6 (o3's fall from 173 to 23, since its lookups now pass the page size it always passes), 82 of DTap-Bench CRM's instead of 62, and 262 of its OS files' instead of 155, where the agents always ask for a message's body as text; elsewhere they save the same turns.scripts/ceiling.pycounts a third ceiling, with constants: an argument that took one value in every call of its tool the agent made, at least five, counts as bindable. Every published count is unchanged. Constants would add 5.2 points of MCPMark Notion's turns (GPT-5's and o3's page size), 2.2 of DTap-Bench's OS files', 2.0 of its CRM's, 0.8 of its customer service's, 0.1 of BFCL's, and nothing elsewhere.MCPMark's Notion server (
scripts/mcpmark_to_tau2.py --service notion). A tool name's hyphens become underscores, anddtap_to_tau2.py's verb rule reads Notion's POST searches and database queries as reads: an HTTPpostis a write only when no read verb follows it (no label of the other benchmarks changes, and the other three servers convert byte-identically).scripts/bench/replay.shreplays Notion with the pooled flows and addsmcpm-own, each newer model with a flow from its own training runs on all four servers. The ceiling is 7.9% of Notion's turns, 22–34% for the Claude models and Gemini 2.5 Pro; the flows save 0.2% of MCPMark's turns at θ = 0.3 pooled and 0.4% with each model's own, since they learn the walk through a page's blocks but not which block comes next.θ priced as billed* (
scripts/costs.py --price OUT,READ,WRITE,--sweep). Prices a saved turn's output tokens and, with prompt caching, its prompt as the previous turn's prompt read from the cache plus what came after it, written; a detour's result is written once and read after.--sweepprices the reach round's threshold sweep at those costs. The default, input tokens alone, reproduces Appendix B. At Anthropic's ratios θ* is 0.29/0.07/0.11 (retail/airline/telecom), at OpenAI's 0.23/0.05/0.08; the thresholds counted in input tokens keep at least 94% of the best swept utility (97% with caching).costs.pyalso findsscripts/bench/replay.sh's replay folders, which name the decider.The benchmarks round, rebuilt from the repository (
scripts/bench/).replay.shlearns every flow and runs every replay behinddocs/results/benchmarks-2026-09-27.md, set by set, from the converters' output in$WORK; a replay that finished is skipped.tables.pyprints and writes that page's tables, andREADME.mdgives the order of steps. Checked against the published rows: every table the page draws from these replays comes out the same, andpriced.py's JSON is identical. Customer service's pooled flows keep 90% of its own flows' turns with 60% of their detours; the page had given the four domains' 94% and 79% as customer service's.priced.pynow skips the solo replays, whose detours are not counted.DTap-Bench, across harnesses (
scripts/dtap_to_tau2.py). DTap-Bench's benign runs of one domain become τ²-bench results, one file per harness and model (--namerenames a domain that τ²-bench also has). Each run holds its request, its turns (consecutive calls, with any message the agent sent with them) and its judged outcome. The harnesses' own calls that list or load tools (List MCP Tools,ToolSearch) are left out with their results. Results are paired with calls in the order the calls were made, since the Claude Agent SDK names some wrongly, and MCP text blocks that the OpenAI Agents SDK logs as Python reprs are unwrapped. OpenClaw's runs, which log only the reply, are skipped. A tool is a read when the first verb in its name reads (existsreads;request, a test ordered for a patient, writes); the Claude Agent SDK's own tools are neither. Six domains are replayed: customer service, CRM, telecom, travel, an operating system's files and medical, 1,417 tasks. A flow learned in other harnesses keeps 59% of the turns the agent's own saves, with 64% of its detours. Medical, half the test turns, is a control: its agents order tests and question a simulated patient, six of them make 0–8 reads in 642 runs, and both deciders make the same lookups, gpt-oss-120b's (results).MCPMark, real MCP servers (
scripts/mcpmark_to_tau2.py). MCPMark's logged runs (OpenAI Responses format) of one server become τ²-bench results: the assistant's message and the calls it makes before their results arrive are one turn, results pair with calls by id, and MCP text blocks are unwrapped as for DTap-Bench.--learn-from-allgives every run a reward of 1, keeping the verdict asverified, sincelearnfits the habit on successful runs and few of MCPMark's pass. On the filesystem, PostgreSQL and GitHub servers the read-only ceiling is 6.6% of turns and the flows save 0.3% (results).Saved turns in seconds and dollars (
scripts/priced.py).pilot/check_flow.pyrows now name the turns a replay saved (saved_at), andpriced.pyprices them at each turn's recorded generation time and cost, net of detours at the agent's fitted input price. In τ²-bench retail the use-before-write speculator's saved turns are 24% of the episodes' generation time and 18% of their cost (results).scripts/remade.py --resultscounts, in any recorded runs laid out as τ²-bench's, the reads an agent made again with no write between, as an agent that ignored a lookup's result would.A threshold per decision, evaluated (
scripts/per_decision.py). Prices each lookup's detour (the tool's mean result tokens times the turns left) and saved turn (the next turn's input tokens) where it is decided, and evaluates that threshold, and flat ones, off-policy over the logged decisions of a replay. At flat thresholds it reproduces the sweep's turns saved to within 5.2%. On τ²-bench it changes the counted utility by −15% in retail and +7–10% in airline and telecom;--recalibrate WEIGHTscores each lookup by its site's own rate of use, cross-fitted over halves of the tasks, which brings every domain within 5%. The served threshold stays the domain's (results).pilot/check_flow.py --explorenow logs each decision's place in the episode (at) and, for each used option, the message of the call it answers (use_at).A search the agent always narrows is not made bare. A lookup whose arguments the agent passes in under 90% of its calls has no required arguments, and a flow made it with none. On WorkBench, whose searches take one of several optional filters, that was every lookup the flows made: a search of all tasks or all customers, which no agent makes, 289 detours for one saved turn.
learnandcompilenow count, for each lookup without a required argument that the agent sometimes called with one, how many of its calls passed none (bindings.bare). Such a lookup is as likely as those calls, Laplace-smoothed as a string binding's agreement is. A read the agent always calls with no argument keeps the chance 1 and is not counted.flow-showshows the chance. A flow with the count is format version 2. Flows without it replay as before.Bindings pass whole lists. A lookup can take a list, such as the hotels a city's listing named, passed to a lookup of their prices.
learnandcompilenow trace a list argument to the path of an earlier result whose values contain every value of the list (bindings.lists), and the binding passes every value at that path of the most recent such result, in order, unless the lookup already had that list. Its chance is scored as a string argument's is.flow-showlists such a source as the path's "every value". A flow with lists is format version 2, since a build that ignored them would make fewer lookups than the flow was learned to. Flows without lists replay exactly as before: GLM-5's retail replay with the paper's flow gave the same decision log, byte for byte. On AgentDojo's travel suite, flows had made none of the lookups its 47% ceiling holds (results).Replay from the record (
pilot/check_flow.py --trace,pilot/trace_env.py). Replays another benchmark's published trajectories with no environment. The agent's calls return their recorded results. A lookup returns the result of the agent's own later call with the same arguments, if the agent made it before its next write, which is exact. Any other lookup is a detour, gets the nearest result of the same tool as a stand-in, and never counts as a saving. On τ²-bench's paper batch, it counts 96.7% of the environment replay's saved turns, 91% of the reach decider's lead, and 55–72% of the detours (results). A domain τ²-bench does not know takes its writes from its checkout'stools.py.Four more benchmarks, as τ²-bench results (and DTap-Bench, above).
scripts/taubench_v1_to_tau2.py: τ-bench's historical trajectories.scripts/bfcl_to_tau2.py: BFCL v4's multi-turn tasks, run on BFCL's own backends as ground truth; with--runs, also the leaderboard's logged runs of any model from a BFCL-Result snapshot, scored by its checker.scripts/agentdojo_to_tau2.py: AgentDojo's benign runs, with results rendered as JSON.scripts/workbench_to_tau2.py: WorkBench's runs, the loggedFinal Answeras the reply.
Each also writes a checkout-shaped folder:
tools.pywith every tool marked a read or a write (checked against the benchmark's source), and a train/test split by task.stretto learn --results,pilot/check_flow.py --trace,scripts/ceiling.pyandscripts/calibration.pyread them unchanged (results).scripts/learning_curve.py --prior OTHER.json...learns from those other agents' training sessions together with the first n of the agent's own, as a deployment that starts from other agents' sessions and adds its own;--prior-n Ktakes K of them, drawn at random, so that they weigh as K sessions do (results).pilot/check_flow.pysplits each replay row's used lookups by when they were used:used_nextbefore the flow's next decision (before any call of the agent's that no lookup answered),used_laterafter one, so that a later decision could have made the lookup in time. At 0.3, 94% of the reach decider's lookups that paid in retail, airline and telecom were used before its next decision (results).scripts/remade.pychecks the replay's assumption on live episodes: whether the agent made a flow's lookup again before the next write. On the reach arm's 28 live episodes, GLM-5.3 made none of the flow's 101 lookups again (results).stretto flow-showsays when a flow carries thereachcounts, which--decider reachserves.scripts/costs.pycounts a detour's cost and a saved turn's value in recorded episodes, in each agent's own input tokens, and so the threshold a read-only flow should use, θ* = δ/(β+δ). On τ²-bench's leaderboard episodes it gives 0.30 in retail, what the live paired run measured, and 0.12–0.13 in airline and telecom, whose contexts are longer and results shorter (results).scripts/learning_curve.pymeasures how a flow's savings grow with the sessions it learns from. For each seed it orders an agent's training sessions at random, learns a habit-only flow from the first n (or, with--pool, from other agents' sessions), and replays it on the agent's own test episodes at--deciderand--threshold: a row per n and seed (results).scripts/paper_figures.pydraws the working paper's figures, the reliability diagrams and the learning curves, from the round's published rows (docs/results/reach-2026-09-26.json).Compiled procedures in stretto (#38). A whole workflow compiled once from agents' traces, writes included, for where no user speaks, is now a file stretto runs: the procedure IR (
stretto_procedure: 1, formats), per site a decision tree over the run's state, calls whose identifiers are held as where they came from and bound again at run time, a guard against repeats, and the checks of each outcome a ticket may state.scripts/telecom_workflow.py --exportwrites one;stretto-procedure --procedure FILE --ticket TEXT -- SERVER…runs it against an MCP server with no model and prints the run: its calls, its check of the ticket's outcome, and its verdict (resolved,transferred, orhand_backwhen the check failed and the ticket should go to a model).stretto_report::procedureis the runtime.pilot/run_procedure.pyruns it on τ²-bench telecom's 40 held-out solo tasks overtau2_mcp.pyand scores each with τ²-bench's evaluator: the published procedure (docs/results/telecom-workflow-2026-09-26.procedure.json) made the same calls as the script on all 40, passed the same 35, and handed back the same 4.Live solo runs and the cascade.
pilot/run_episode.py --soloruns τ²-bench's no-user mode live (telecom): no customer, τ²-bench's solo system prompt with the ticket, the phone's tools and τ²-bench'sdonetool served bytau2_mcp.py --solo, and the episode scored as τ²-bench scores solo runs; as in τ²-bench, a turn that ends with neitherdonenor a transfer is an agent error.--handback RUNS.jsonmakes the agent take over from the compiled workflow (scripts/telecom_workflow.py --json):tau2_mcp.py --prefixmakes the workflow's calls first, with their effects in place, and the agent is told what they returned and that the workflow's check of the ticket's outcome failed.--arm reachserves a flow with--decider reach. The harness now stops an episode whose MCP server failed to start, rather than letting the agent run without its tools.The replay counts only what a lookup answered.
pilot/check_flow.pyskipped any recorded call already made, the agent's own repeated calls included, and so counted as saved the 0.2–0.9% of LLM turns (3.3% for one agent) in which an agent repeated calls it had made. It now skips a recorded call only when a flow lookup answers it: the same call, a lookup no recorded call has used yet, that returned what the recorded call returned (with no write since, the agent's or the customer's, it did).--legacyreplays with the old rule, which reproduces the published numbers.--explore 0labels each optionusedonly when the agent makes the call before its next write, as the rule has it, and alsonextwhen it is the agent's very next call.scripts/ceiling.pyclasses every LLM turn of τ²-bench's recorded test episodes as a reply, reads or writes, each after the customer or after a tool, and counts the turns any read-only speculator could save: every call a read, with its arguments in earlier results, after a tool response since the last write.scripts/replay_stats.pygives replays' turns saved, detours and share of that ceiling with 95% intervals from a bootstrap over tasks, andscripts/calibration.pyhow calibrated a flow's scores were against theusedlabels (binned, with the expected calibration error, the Brier score and the AUC); with--versus, against the same episodes replayed under another decider, with 95% intervals for the differences from a bootstrap over tasks.Deciding on the chance of use before the next write (
--decider reach). A lookup pays off if the agent makes the call at any point before its next write, not only next, since a read's result stays current until then.learnandcompilenow also count, over the habit's histories, how often each action came before the agent's next write (reachin the flow IR; each action its own Beta–Bernoulli, backed off as the habit is), andserve,flow-serve,audit,promoteandstretto-proxy --flow-decidertakereach: of the lookups that bind, the one whose chance times its binding's is highest, when that clears the threshold. A flow learned before has no counts and refuses the decider; a build that does not know the field serves the flow as before, so the format version stays. On GLM-5's retail test episodes, flows learned from the 2025 runs scored lookups with a calibration error of 0.10 against the habit's next-step probability's 0.19, and at the threshold of 0.3 saved 358 turns with 55 detours against 307 with 33.Counterfactual evaluation (RFC-001 §3.7, #15).
stretto-proxy --flow-explore EPSILON, andserve --exploreandflow-serve --explore, explore at read-only sites: with that probability the flow takes a lookup other than its rule's choice, drawn by the decider's probabilities among the lookups that bind. Each decision then logs itspolicy: every option and the chance that the flow took what it took.stretto evaluatereads such decisions, with each option's outcome, and estimates what another rule (the arbiter's or the habit's probabilities at a threshold) would have done, per site and in total: directly from every option's outcome, and by IPS, self-normalized IPS and a doubly robust estimate from the taken option's alone, with each site's effective sample size.pilot/check_flow.py --explorewrites the labelled decisions from replays.Predicate refinement (RFC-001 §3.4, #16).
phase0 --candidates FILEasks candidate predicates alone at every asked next-step decision, with the same state as the next-step question, so every other answer keeps its cache key. Their answers go to--oracle-log, and--weigh IDweighs one in the arbiter.stretto refinereads such a log, fits the arbiter by cross-validation over tasks with each candidate added, and keeps candidates greedily while one raises the held-out log-likelihood of the agents' steps by more than a penalty (by default what BIC charges a parameter). A transfer target's decisions, judged but never fitted on, are scored alongside as an out-of-sample check.--exampleswrites examples from the sites where the arbiter is weakest, for whoever proposes the candidates.A predicate can bear on one named lookup:
"favors": {"lookup": TOOL}(formats).Flow search (RFC-001 §3.10, #25).
stretto searchruns NSGA-II, from fugue-evo, over each site's threshold (0.10 to 0.95, or off) and the decider (the arbiter, or the habit alone). It scores each setting by a replay command that printspilot/check_flow.py'sCHECKtotals, and reports the settings that save the most turns for the fewest detours, starting from the hand-set ones. It counts the decisions a replay could not answer, and says so when a setting on the front has any.--rescorereplays an earlier search's front on other episodes, such as held-out ones.Per-site thresholds in the flow IR (
thresholds): at a listed site a flow acts on that threshold in place of the served one, and above 1 it never acts there. A flow with them is format version 2, which 0.1.0 refuses rather than ignores; this build reads versions 1 and 2.flow-showandflow-diffshow them.flow-difflists a changed threshold without flagging it, as it would the served one, and flags a site switched back on.stretto-proxy --flow-tools TOOLSnames the only tools a flow may call on its own. A read-only tool can still be metered, rate-limited or recorded as an access; without the option, a flow may call every tool it reads as a lookup, as before.Pinned inputs. A flow learned from recorded sessions pins each tool's input contract (
contracts): its arguments, their JSON types and which are required, as the server listed them.stretto-proxymakes no lookup of a tool whose server now lists another, and says so once.flow-showlists them andflow-diffshows a change.A record the customer did not ask about.
learnandcompilecount, per lookup, the picks the binding would make where the customer had mentioned a value at its sources and the binding would pass another, and how many of them the agent went on to pass (bindings.named_other). The binding's chance there is that rate, shrunk towards its unmentioned chance.flow-showandflow-diffshow it. A flow with the count is format version 2. Replayed on GLM-5's test episodes, airline detours fell from 56 to 4 at no cost in turns (results).scripts/proposal_check.pychecks each write of τ²-bench episodes against what the agent proposed, with no model: it flags a write when the confirmation chose another record of the list its value came from (results).scripts/anatomy.pyclasses every LLM turn, every decision after a tool returns and every argument of τ²-bench episodes by where its information came from: copied from an earlier result (one value at its path, or one of several), the customer's words, or made up, after TraceCompiler's binding classes; and how well the previous tool, structured features of the results and the goal predict each decision (results).scripts/anatomy.pyreads telecom too: τ²-bench's read-only tools for the agent and the phone, features from results in text (eachKey: valueline's value, or its|-separated parts), and results without ids paired with their calls in order. It also predicts each decision with a decision tree over the whole state (every tool's last result and the tools called), its depth chosen by cross-validation, as decision mining does at a process's branch points (results).Sources in the order the agent used them at the site.
learnandcompilecount, for each lookup argument with more than one source, where the agent took its values at each site, the tool whose result came last (bindings.site_sources), each read of a batch at the site of the read before it, as a flow makes them. Where a site has enough of them, the binding tries its sources in that order before recency: after reading a phone line, the next line id, not the line's plan id. Replayed on the nine leaderboard agents' telecom test episodes, a flow learned from τ²-bench's 2025 runs saved 12.8% of turns with 443 detours, against 1.9% with 2,936; retail and airline replay exactly as before on all nine agents (results).pilot/check_flow.pyreplays telecom: the customer's calls on their own phone are the customer's turn. A flow with them is format version 2.The record the customer described, already read.
learnandcompilecount, per lookup, the picks where another value of the same list was already looked up and returned a value the customer gave, which another's result did not (the line with the ticket's phone number), and how many of them the agent went on to pass (bindings.described_read). A value the customer gave has a digit and at least four characters, and the customer wrote it or the agent passed it before any result held it. The binding's chance there is scored as withnamed_other. In τ²-bench telecom's solo mode the agents almost never read another line after the one with the ticket's number (7 of 297), and the flow's detours fell from 387 and 430 to 203 and 231 with more turns saved (799 and 647, from 776 and 633); all nine leaderboard agents replay retail and airline exactly as before (results).flow-showandflow-diffshow it. A flow with the count is format version 2.flow-showlists each lookup argument's sources by site, in the order the binding tries them, andflow-difflists a change in the source it tries first, without flagging it.learn --resultstakes τ²-bench's solo runs (no-user), where the agent calls the customer's tools itself: those tools join the flow's, read-only or not as τ²-bench marks them. Only calls that returned count, since an agent that calls the customer's tools when the customer holds the phone is told they do not exist (Claude 3.7 Sonnet did 130 times in the 2025 dual-control runs, which had made a flow learned from them offer the phone's tools).pilot/check_flow.py --soloandpilot/tau2_mcp.py --soloserve them. In telecom's solo mode a habit-only flow saved 26.9–32.5% of LLM turns in replay, and 27.5–33.5% withbindings.described_read(results).stretto-proxy --confirm-judgealso logs, with each judgment, the proposal check (proposal_check): the call's values whose record the confirmation chose another of, with no model. Only values with a digit are checked, as ids carry them; on τ²-bench's 2025 runs that flags exactly whatscripts/proposal_check.py's closed-choice rule flagged. It refuses nothing.scripts/telecom_workflow.pyruns a workflow compiled once from τ²-bench telecom's solo runs (a decision tree per site over the whole state, its steps whole calls) with no model in τ²-bench's environment, and scores it with τ²-bench's evaluator;--surehands back where the tree is not sure,--no-guardlets it repeat itself, and--train-sharelearns from fewer tasks (results).--self-train ROUNDSthen learns from its own runs, as expert iteration does: each round it tries every training ticket once as it stands and--rolloutstimes drawing from its leaves, keeps the shortest run its own check of the ticket's criterion says resolved, and is refitted (results);--bag Naverages N trees fitted on resamples.--symboliclearns each identifier a call passes as where it came from (a result's path, with the fields that set its record apart in a list, or the ticket) and binds it again from the run's results, preferring the record that holds the ticket's value, and keeps a tool's state on the record the ticket names;--renameruns the test tasks for a customer no trace saw, and--as-customer IDfor another customer of the database, keeping the tasks its gold actions still solve and whose faults the account still shows (τ²-bench's flags for the line: active, roaming allowed, data used up). Renamed, the workflow passed the same 35 of 40 as for the original customer, against 16 with identifiers as constants; moved to Michael Lee (one line, not three), 35 against 19 (results).anatomy.tree(counts=True)returns a leaf's calls by count.scripts/detours.py toolslists each lookup tool's detours and uses over tagged replays, by the kind of task, with how many detours the agent made with other arguments (results).scripts/detours.pysays where a replayed flow's detours come from: a session hand-back simulated on the replay, each lookup's place in its chain against its probabilities and use, and, in airline, reads after a reservation the customer had described (results).pilot/check_flow.py --same-resultalso counts a recorded call as made when a flow lookup of the same tool, agreeing on every argument both pass, already returned its recorded result, as an optional argument that changes nothing does (alimitabove the number of bills), and counts that lookup as used, not as a detour. Each such pair is written to the episode'ssame-result.json. Scoring by exact arguments stays the default (results).pilot/check_flow.pynames each recorded episode it replays after its folder, or, when two folders share a name (a task's trials, each in its own folder), after its path below their common parent. Such episodes shared one replay folder before, and with--jobsone trajectory, which the flow reads; the published replays were of τ²-bench results files, whose episodes are named by task and trial, and are unaffected.