CLI reference
Every command and option of the three binaries: stretto, which measures agents' traces and compiles, learns, serves and audits flows; stretto-proxy, which records an MCP server's sessions and runs flows, guards and the confirmation judge on them; and stretto-procedure, which runs a compiled procedure against an MCP server with no model. --help prints the same text.
The sections below are generated from the code. After changing an option, run STRETTO_BLESS=1 cargo test --bins to rewrite them; CI fails while they are out of date.
stretto
Compile an agent's recorded behavior into flows, serve them, and measure them against published τ²-bench trajectories.
| Command | What it does |
|---|---|
phase0 | Measure how compressible an agent's behavior is (no API keys needed), and with --oracle, how well a System-One model takes the decisions flows would hand it (Phase 0b). |
flow-serve | Serve a live read-only flow (RFC-001 §3.13). |
compile | Compile a live read-only flow and write its IR (JSON) to a file. |
learn | Learn a live flow from sessions recorded by stretto-proxy (JSONL logs in a directory) and write its IR, as compile does from τ²-bench results. |
serve | Serve a compiled flow (from compile), as flow-serve does. |
guards | Test the policy guards (typed checks a proxy runs before a write) against recorded τ²-bench trajectories. |
confirm | Judge the customer's confirmation before each write with a System-One model, next to the guards' word list, on τ²-bench's published trajectories. |
match | Match descriptions to records (RFC-001). |
audit | Audit a flow against recorded episodes. |
flow-show | Show a flow as a reviewer reads it (Markdown). |
flow-diff | What changed from one flow to another, as a change list for a pull request (Markdown). |
evaluate | Estimate what another rule would have done on a flow's logged decisions (RFC-001 §3.7). |
refine | Refine the arbiter's predicates (RFC-001 §3.4). |
search | Search a flow's settings (RFC-001 §3.10). |
promote | Promote a flow's sites (RFC-001 §3.7). |
redact | Pseudonymize recorded sessions. |
jev-check | Check that Jev is reachable with TYPESAFE_API_KEY. |
export-arbiter | Write a flow's arbiter to its own file, to ship. |
fit-arbiter | Fit one arbiter, to ship, on the held-out decisions of one or more compiles. |
export-answers | Write every cached oracle answer to stdout as JSON lines ({"key", "response"}). |
ask | Ask a System-One model questions of your own. |
import-answers | Read JSON lines from export-answers on stdin into a replay cache. |
init | Print the configuration that runs an MCP server behind stretto-proxy in an MCP host, recording its sessions, and the steps from there to a served flow. |
doctor | Check the installation. |
completions | Print the completion script for stretto in bash, zsh, fish, powershell or elvish to stdout. |
stretto phase0
Measure how compressible an agent's behavior is (no API keys needed), and with --oracle, how well a System-One model takes the decisions flows would hand it (Phase 0b).
Usage: stretto phase0 [OPTIONS] --tau2 <DIR>Inputs
--tau2 <DIR>(required): Path to a τ²-bench checkout.--domain <NAME>(repeatable; defaultretail,airline): Domains to analyze.--no-baselines: Do not train on τ²-bench's published baselines in the checkout (then pass --source).--source <[LABEL=]PATH>(repeatable): Extra τ²-bench results to train on:pathorlabel=path(repeatable; files for other domains are skipped).--target <[LABEL=]PATH>(repeatable): τ²-bench results for an agent model the habit never trains on, measured as a transfer target:pathorlabel=path(repeatable; files for other domains are skipped).--train-fraction <SHARE>(default1; not with--train-tasks): Train on this share of the training tasks: a fixed sample by task id, each smaller share part of every larger one. To see how flows do with fewer traces.--train-tasks <IDS>(repeatable; not with--train-fraction): Train on these training tasks only (comma-separated), in place of--train-fraction: a sample named exactly.
The habit
--order <N>(default2): Context length for coverage and transfer.--alpha-samples <N>(default600): MH draws for the posterior over α (0 to use --alpha).--alpha <ALPHA>(default1): α to use when --alpha-samples is 0.--min-evidence <N>(default5): Training observations a context needs before the habit may act on it.--seed <N>(default7): Seed for the MH chain.--no-features: Skip learning code features from tool outputs.--no-intent: Do not let flows know the episode's goal: the habit is not conditioned on it and System-One questions do not name it (for live flows nobody names).
Phase 0b: the System-One model
--oracle <ORACLE>(one ofjev,replay,mock): Phase 0b: who answers the System-One questions.jevneeds TYPESAFE_API_KEY and pays once per distinct question;replayreads the cache only;mockchecks the pipeline for free. The other options under this heading need it.--questions <QUESTIONS>(one ofv1,v2; defaultv1): Which Phase 0b questions to ask:v1(one question over every tool, for flows that may call any tool) orv2(RFC-001 §3.5: read-only flows, a site's own lookups as options, a state slice, the stop decision asked on its own, and the answers combined with the habit).--dataflow-hints: v2: also describe each lookup by what its results supply, learned from argument dataflow in training.--predicates <FILE>: v2: a JSON file of yes/no predicates about the state (seedata/predicates-v2.json) to ask with every next-step question and weigh in the arbiter.--no-predicate-features: v2: ask the predicates but leave them out of the arbiter, to measure what they add.--candidates <FILE>: v2,phase0: candidate predicates (same format as --predicates), each asked alone at every asked next-step decision with the same state, so the other answers keep their cache keys. Their answers go to --oracle-log, forstretto refine.--weigh <ID>(repeatable): v2,phase0: also weigh this candidate in the arbiter (repeatable).--manifest-options: v2: offer every read-only tool at every site, not only the lookups seen there in training, for the System-One model to choose from. A lookup training never made is bound by argument name.--pooled-arbiter: v2,compile: give the flow one arbiter, fitted on every held-out decision, in place of one per fold, so thatexport-arbitercan ship it. Replayed on the same test tasks it has seen other agents' decisions on them, so compare flows without it.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--oracle-concurrency <N>(default8): Oracle requests in flight at once.--oracle-limit <N>: Ask at most this many distinct questions (a stable sample), for a pilot run.--oracle-budget <DOLLARS>(default5): Refuse to start if uncached questions could cost more than this many dollars.--oracle-model <MODEL>: Model id to request (default: TYPESAFE_DEFAULT_MODEL, else jev-latest).--oracle-dump <FILE>: Write every distinct oracle request to this file (JSON lines; the domain is added to the file name).--oracle-log <FILE>: Write every decision, with the agent's option and the oracle's pick (with v2, also the features the arbiter weighs), to this file (JSON lines; the domain is added to the file name).
Output
--out <FILE>: Write the Markdown report here (default: stdout).--json <FILE>: Also write the full report as JSON here.
stretto flow-serve
Serve a live read-only flow (RFC-001 §3.13): compile it from the same data and questions as phase0 --questions v2, goal free, then answer one query per connection on a local TCP port. A query is a JSON line {"task_id", "messages"} (τ²-bench messages so far, ending with a tool result); the answer is a JSON line with "action": "lookup" (and tool, arguments) or "action": "hand_back" (and reason).
Usage: stretto flow-serve [OPTIONS] --tau2 <DIR>Inputs
--tau2 <DIR>(required): Path to a τ²-bench checkout.--domain <NAME>(repeatable; defaultretail,airline): Domains to analyze.--no-baselines: Do not train on τ²-bench's published baselines in the checkout (then pass --source).--source <[LABEL=]PATH>(repeatable): Extra τ²-bench results to train on:pathorlabel=path(repeatable; files for other domains are skipped).--target <[LABEL=]PATH>(repeatable): τ²-bench results for an agent model the habit never trains on, measured as a transfer target:pathorlabel=path(repeatable; files for other domains are skipped).--train-fraction <SHARE>(default1; not with--train-tasks): Train on this share of the training tasks: a fixed sample by task id, each smaller share part of every larger one. To see how flows do with fewer traces.--train-tasks <IDS>(repeatable; not with--train-fraction): Train on these training tasks only (comma-separated), in place of--train-fraction: a sample named exactly.
The habit
--order <N>(default2): Context length for coverage and transfer.--alpha-samples <N>(default600): MH draws for the posterior over α (0 to use --alpha).--alpha <ALPHA>(default1): α to use when --alpha-samples is 0.--min-evidence <N>(default5): Training observations a context needs before the habit may act on it.--seed <N>(default7): Seed for the MH chain.--no-features: Skip learning code features from tool outputs.--no-intent: Do not let flows know the episode's goal: the habit is not conditioned on it and System-One questions do not name it (for live flows nobody names).
Phase 0b: the System-One model
--oracle <ORACLE>(one ofjev,replay,mock): Phase 0b: who answers the System-One questions.jevneeds TYPESAFE_API_KEY and pays once per distinct question;replayreads the cache only;mockchecks the pipeline for free. The other options under this heading need it.--questions <QUESTIONS>(one ofv1,v2; defaultv1): Which Phase 0b questions to ask:v1(one question over every tool, for flows that may call any tool) orv2(RFC-001 §3.5: read-only flows, a site's own lookups as options, a state slice, the stop decision asked on its own, and the answers combined with the habit).--dataflow-hints: v2: also describe each lookup by what its results supply, learned from argument dataflow in training.--predicates <FILE>: v2: a JSON file of yes/no predicates about the state (seedata/predicates-v2.json) to ask with every next-step question and weigh in the arbiter.--no-predicate-features: v2: ask the predicates but leave them out of the arbiter, to measure what they add.--candidates <FILE>: v2,phase0: candidate predicates (same format as --predicates), each asked alone at every asked next-step decision with the same state, so the other answers keep their cache keys. Their answers go to --oracle-log, forstretto refine.--weigh <ID>(repeatable): v2,phase0: also weigh this candidate in the arbiter (repeatable).--manifest-options: v2: offer every read-only tool at every site, not only the lookups seen there in training, for the System-One model to choose from. A lookup training never made is bound by argument name.--pooled-arbiter: v2,compile: give the flow one arbiter, fitted on every held-out decision, in place of one per fold, so thatexport-arbitercan ship it. Replayed on the same test tasks it has seen other agents' decisions on them, so compare flows without it.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--oracle-concurrency <N>(default8): Oracle requests in flight at once.--oracle-limit <N>: Ask at most this many distinct questions (a stable sample), for a pilot run.--oracle-budget <DOLLARS>(default5): Refuse to start if uncached questions could cost more than this many dollars.--oracle-model <MODEL>: Model id to request (default: TYPESAFE_DEFAULT_MODEL, else jev-latest).--oracle-dump <FILE>: Write every distinct oracle request to this file (JSON lines; the domain is added to the file name).--oracle-log <FILE>: Write every decision, with the agent's option and the oracle's pick (with v2, also the features the arbiter weighs), to this file (JSON lines; the domain is added to the file name).
Serving
--listen <ADDR>(default127.0.0.1:0): Address to listen on (port 0: any free port; the ready line on stderr names it).--threshold <P>(default0.3): Take the most likely lookup when its probability, times the chance that its arguments are the agent's, is at least this.--decider <DECIDER>(one ofarbiter,habit,reach; defaultarbiter): Where each option's probability comes from:arbiter(the habit, the System-One model and the predicates, combined: arm D0),habit(the habit alone, never asking the System-One model: a flow compiled from traces only, arm C at a high threshold), orreach(the habit's counts for whether the agent makes the lookup before its next write, now or later; a flow learned with them).--max-questions <N>(default300): Stop asking the System-One model after this many live questions.--log <FILE>: Append every query's answer here (JSON lines).--explore <EPSILON>: Explore: with this probability, take a lookup other than the rule's choice, drawn by the decider's probabilities among those that bind. Each answer then carries itspolicy: every option and the chance that the flow took what it took, forevaluate. 0 explores nothing but still logs it.--explore-seed <N>(default0): Seed for the exploration draws.
stretto compile
Compile a live read-only flow and write its IR (JSON) to a file: the same data and questions as phase0 --questions v2, goal free, from cached System-One answers only. serve and stretto-proxy --flow load the file.
Usage: stretto compile [OPTIONS] --tau2 <DIR> --out <FILE>Inputs
--tau2 <DIR>(required): Path to a τ²-bench checkout.--domain <NAME>(repeatable; defaultretail,airline): Domains to analyze.--no-baselines: Do not train on τ²-bench's published baselines in the checkout (then pass --source).--source <[LABEL=]PATH>(repeatable): Extra τ²-bench results to train on:pathorlabel=path(repeatable; files for other domains are skipped).--target <[LABEL=]PATH>(repeatable): τ²-bench results for an agent model the habit never trains on, measured as a transfer target:pathorlabel=path(repeatable; files for other domains are skipped).--train-fraction <SHARE>(default1; not with--train-tasks): Train on this share of the training tasks: a fixed sample by task id, each smaller share part of every larger one. To see how flows do with fewer traces.--train-tasks <IDS>(repeatable; not with--train-fraction): Train on these training tasks only (comma-separated), in place of--train-fraction: a sample named exactly.
The habit
--order <N>(default2): Context length for coverage and transfer.--alpha-samples <N>(default600): MH draws for the posterior over α (0 to use --alpha).--alpha <ALPHA>(default1): α to use when --alpha-samples is 0.--min-evidence <N>(default5): Training observations a context needs before the habit may act on it.--seed <N>(default7): Seed for the MH chain.--no-features: Skip learning code features from tool outputs.--no-intent: Do not let flows know the episode's goal: the habit is not conditioned on it and System-One questions do not name it (for live flows nobody names).
Phase 0b: the System-One model
--oracle <ORACLE>(one ofjev,replay,mock): Phase 0b: who answers the System-One questions.jevneeds TYPESAFE_API_KEY and pays once per distinct question;replayreads the cache only;mockchecks the pipeline for free. The other options under this heading need it.--questions <QUESTIONS>(one ofv1,v2; defaultv1): Which Phase 0b questions to ask:v1(one question over every tool, for flows that may call any tool) orv2(RFC-001 §3.5: read-only flows, a site's own lookups as options, a state slice, the stop decision asked on its own, and the answers combined with the habit).--dataflow-hints: v2: also describe each lookup by what its results supply, learned from argument dataflow in training.--predicates <FILE>: v2: a JSON file of yes/no predicates about the state (seedata/predicates-v2.json) to ask with every next-step question and weigh in the arbiter.--no-predicate-features: v2: ask the predicates but leave them out of the arbiter, to measure what they add.--candidates <FILE>: v2,phase0: candidate predicates (same format as --predicates), each asked alone at every asked next-step decision with the same state, so the other answers keep their cache keys. Their answers go to --oracle-log, forstretto refine.--weigh <ID>(repeatable): v2,phase0: also weigh this candidate in the arbiter (repeatable).--manifest-options: v2: offer every read-only tool at every site, not only the lookups seen there in training, for the System-One model to choose from. A lookup training never made is bound by argument name.--pooled-arbiter: v2,compile: give the flow one arbiter, fitted on every held-out decision, in place of one per fold, so thatexport-arbitercan ship it. Replayed on the same test tasks it has seen other agents' decisions on them, so compare flows without it.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--oracle-concurrency <N>(default8): Oracle requests in flight at once.--oracle-limit <N>: Ask at most this many distinct questions (a stable sample), for a pilot run.--oracle-budget <DOLLARS>(default5): Refuse to start if uncached questions could cost more than this many dollars.--oracle-model <MODEL>: Model id to request (default: TYPESAFE_DEFAULT_MODEL, else jev-latest).--oracle-dump <FILE>: Write every distinct oracle request to this file (JSON lines; the domain is added to the file name).--oracle-log <FILE>: Write every decision, with the agent's option and the oracle's pick (with v2, also the features the arbiter weighs), to this file (JSON lines; the domain is added to the file name).
Output
--out <FILE>(required): Where to write the flow.
stretto learn
Learn a live flow from sessions recorded by stretto-proxy (JSONL logs in a directory) and write its IR, as compile does from τ²-bench results. The tools come from the sessions' tools/list responses (their readOnlyHint annotations), or from --manifest. With --results, the sessions are τ²-bench episodes on a checkout's training tasks instead, as if a deployment had recorded them.
Usage: stretto learn [OPTIONS] --domain <NAME> --out <FILE>Inputs
--sessions <DIR>(not with--results): Directory of session logs (*.jsonl); needed unless--resultsis given.--results <FILE>(repeatable; not with--sessions,--manifest,--rewards): τ²-bench results to learn from in place of--sessions(repeatable): their episodes on the training tasks of the--tau2checkout's split, with their rewards. The tools come from the checkout.--tau2 <DIR>: The τ²-bench checkout that--resultsbelong to.--train-fraction <SHARE>(default1; not with--train-tasks): With--results, learn from this share of the training tasks: the samplecompile --train-fractiontakes.--train-tasks <IDS>(repeatable; not with--train-fraction): With--results, learn from these training tasks only (comma-separated), in place of--train-fraction.--trials <N>...(repeatable): With--results, only these trials of each task (default: all).--domain <NAME>(required): The domain to name the flow for.--manifest <FILE>(not with--results): A tool manifest (JSON, as stretto-trace writes it) instead of the sessions' owntools/list.--rewards <FILE>(not with--results): Rewards by session id (JSON object). Sessions without one count as successful.
The arbiter
--habit-only(not with--refit-habit,--arbiter-from,--manifest-options,--oracle,--oracle-cache,--oracle-budget,--predicates): Ask no System-One model: every session trains the habit, and the flow has no arbiter (serve it with--flow-decider reach, asstretto initdoes).--refit-habit(not with--habit-only,--arbiter-from): Once the arbiter is fitted on the held-out sessions, learn the habit, the sites and the bindings again from every session.--arbiter-from <FILE>(not with--habit-only,--refit-habit,--manifest-options,--oracle,--oracle-cache,--oracle-budget,--predicates): Ask no System-One model while learning: every session trains the habit, and the flow serves this arbiter instead. It is an arbiter file (export-arbiter;data/arbiters/ships two), or a flow whose arbiter to take, such as onecompilefitted on other agents' traces.--manifest-options(not with--habit-only,--arbiter-from): Offer every read-only tool at every site (seecompile). Only the arbiter a flow fits here weighs such options.--oracle <ORACLE>(one ofjev,replay,mock; defaultjev; not with--habit-only,--arbiter-from): Who answers the held-out questions the arbiter is fitted on.--habit-onlyand--arbiter-fromask nothing, so they take none of the oracle's options.--oracle-cache <DIR>(default.oracle-cache; not with--habit-only,--arbiter-from): Replay cache for oracle answers.--oracle-budget <DOLLARS>(default1; not with--habit-only,--arbiter-from): Refuse to start if uncached questions could cost more than this many dollars.--predicates <FILE>(not with--habit-only,--arbiter-from): Yes/no predicates to ask and weigh (seedata/predicates-v2.json). An arbiter from--arbiter-frombrings its own.
The bindings
--constants: Also pass, as the agent did, each argument it passed with one value in every call of a lookup, at least five, and in at least half of them, such as a page size, so that the flow's lookups are the agent's own calls.
Output
--out <FILE>(required): Where to write the flow.
stretto serve
Serve a compiled flow (from compile), as flow-serve does.
Usage: stretto serve [OPTIONS] --flow <FILE>Inputs
--flow <FILE>(required): The flow IR to load.
The System-One model
--oracle <ORACLE>(one ofjev,replay,mock; defaultjev): Who answers live questions:jev(needs TYPESAFE_API_KEY),replay(the cache only) ormock.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.
Serving
--listen <ADDR>(default127.0.0.1:0): Address to listen on (port 0: any free port; the ready line on stderr names it).--threshold <P>(default0.3): Take the most likely lookup when its probability, times the chance that its arguments are the agent's, is at least this.--decider <DECIDER>(one ofarbiter,habit,reach; defaultarbiter): Where each option's probability comes from:arbiter(the habit, the System-One model and the predicates, combined: arm D0),habit(the habit alone, never asking the System-One model: a flow compiled from traces only, arm C at a high threshold), orreach(the habit's counts for whether the agent makes the lookup before its next write, now or later; a flow learned with them).--max-questions <N>(default300): Stop asking the System-One model after this many live questions.--log <FILE>: Append every query's answer here (JSON lines).--explore <EPSILON>: Explore: with this probability, take a lookup other than the rule's choice, drawn by the decider's probabilities among those that bind. Each answer then carries itspolicy: every option and the chance that the flow took what it took, forevaluate. 0 explores nothing but still logs it.--explore-seed <N>(default0): Seed for the exploration draws.
stretto guards
Test the policy guards (typed checks a proxy runs before a write) against recorded τ²-bench trajectories: every write is checked against what came before it, as stretto-proxy --guards checks it. A rule that fails the writes of successful episodes is too strict, or wrong.
Usage: stretto guards [OPTIONS] --tau2 <DIR>Inputs
--tau2 <DIR>(required): Path to a τ²-bench checkout (its published baselines are read).--domain <NAME>(repeatable; defaultretail,airline): Domains to audit.--source <[LABEL=]PATH>(repeatable): Extra τ²-bench results to audit:pathorlabel=path(repeatable; files for other domains are skipped).
Output
--out <FILE>: Write the Markdown report here (default: stdout).--json <FILE>: Also write the audit as JSON here.
stretto confirm
Judge the customer's confirmation before each write with a System-One model, next to the guards' word list, on τ²-bench's published trajectories: one yes/no question per write, sorted by whether the tool accepted it and the episode passed.
Usage: stretto confirm [OPTIONS] --tau2 <DIR>Inputs
--tau2 <DIR>(required): Path to a τ²-bench checkout (its published baselines are read).--domain <NAME>(repeatable; defaultretail,airline): Domains to judge.--source <[LABEL=]PATH>(repeatable): Extra τ²-bench results to judge:pathorlabel=path(repeatable; files for other domains are skipped).
The judge
--oracle <ORACLE>(one ofjev,replay,mock; defaultreplay): Who judges:jev(needs TYPESAFE_API_KEY; pays once per distinct question),replay(the cache only) ormock.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--oracle-budget <DOLLARS>(default2): Refuse to start if uncached questions could cost more than this many dollars.--threshold <P>(default0.5): The judge fails a write below this probability of an explicit yes.--second-question <SECOND_QUESTION>(one ofdescribed,proposed): Also ask a second question about each write: whether the agent hadproposedthis change before the customer's reply, or (the first wording, too literal) whether its messagedescribedit. The judge then fails a write unless both answers are yes.
Output
--examples <N>(default8): Disagreements to show, of each kind, per domain.--oracle-dump <FILE>: Write every distinct question to this file (JSON lines; the domain is added to the file name).--out <FILE>: Write the Markdown report here (default: stdout).--json <FILE>: Also write every judged write, and the audits, as JSON here.
stretto match
Match descriptions to records (RFC-001): at each write in τ²-bench's published trajectories that picks records out of earlier results (items of an order, a new variant, a payment method, a reservation), ask the System-One model which one the customer means, without the agent's pick, and score both against the task's expected actions.
Usage: stretto match [OPTIONS] --tau2 <DIR>Inputs
--tau2 <DIR>(required): Path to a τ²-bench checkout.--domain <NAME>(repeatable; defaultretail,airline): Domains to judge.--source <[LABEL=]PATH>(repeatable): Extra τ²-bench results to judge, besides the published baselines (repeatable).
The System-One model
--oracle <ORACLE>(one ofjev,replay,mock; defaultjev): Who answers:jev(needs TYPESAFE_API_KEY),replay(the cache only) ormock.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--oracle-budget <DOLLARS>(default2): Refuse to start if uncached questions could cost more than this many dollars.
Output
--examples <N>(default8): Disagreements to show, of each kind, per domain.--oracle-dump <FILE>: Write every distinct question to this file (JSON lines; the domain is added to the file name).--out <FILE>: Write the Markdown report here (default: stdout).--json <FILE>: Also write every choice, and the audits, as JSON here.
stretto audit
Audit a flow against recorded episodes: score the agent's own steps under the flow's decisions, run as a fugue program, for agreement, calibration and surprise per site and per episode. Run it on new sessions before trusting a flow compiled from older ones.
Usage: stretto audit [OPTIONS] --flow <FILE>Inputs
--flow <FILE>(required): The flow IR to audit.--sessions <DIR>: Sessions recorded by stretto-proxy (a directory of*.jsonl).--results <FILE>(repeatable): τ²-bench results files (repeatable); files for other domains are skipped.--tau2 <DIR>: With --results: keep only the test split of this τ²-bench checkout, the tasks a flow compiled from it never trained on.
The System-One model
--oracle <ORACLE>(one ofjev,replay,mock; defaultreplay): Who answers the flow's questions:replay(the cache only; decisions it cannot answer are left out),jev(needs TYPESAFE_API_KEY; about $0.0001 per decision) ormock.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--decider <DECIDER>(one ofarbiter,habit,reach): How the flow decides:arbiter(weighing the System-One model's answers),habit(the habit alone, asking nothing) orreach(the chance of each lookup before the agent's next write, asking nothing). Default: the arbiter, or the habit for a flow without one (learn --habit-only).
Output
--out <FILE>: Write the Markdown report here (default: stdout).--json <FILE>: Also write the audit as JSON here.
stretto flow-show
Show a flow as a reviewer reads it (Markdown): the tools it may call, the lookups it may make after each call and where their arguments come from, what it does there with the habit alone, and how its arbiter weighs the System-One model's answers.
Usage: stretto flow-show [OPTIONS] <FILE>Arguments
<FILE>(required): The flow IR.
Options
--threshold <P>(default0.3): The threshold the flow will be served with (stretto-proxy --flow-threshold).--out <FILE>: Write the Markdown here (default: stdout).
stretto flow-diff
What changed from one flow to another, as a change list for a pull request (Markdown). Exits with 1 when a change needs review (the flow may call a tool, make a lookup, bind an argument from a source, or ask a model or a question it did not before), 2 on an error, and 0 otherwise.
Usage: stretto flow-diff [OPTIONS] <OLD> <NEW>Arguments
<OLD>(required): The flow before.<NEW>(required): The flow after.
Options
--tolerance <X>(default0.05): Leave out shares, chances and weights that moved by less than this.--threshold <P>(default0.3): The threshold the flow will be served with (stretto-proxy --flow-threshold).--out <FILE>: Write the Markdown here (default: stdout).
stretto evaluate
Estimate what another rule would have done on a flow's logged decisions (RFC-001 §3.7). The decisions are JSON lines, each a flow answer logged with its policy (serve --explore, stretto-proxy --flow-explore) and a labels list: each option's outcome (used, detour, turn), as pilot/check_flow.py --explore writes them. For each target, per site and in total: the lookups, used lookups, detours and turns spared, estimated directly from every option's label, and by IPS, self-normalized IPS and doubly robust estimates from the taken option's label alone.
Usage: stretto evaluate [OPTIONS] --decisions <FILE> --target <RULE>Inputs
--decisions <FILE>(required; repeatable): Labelled decisions (JSON lines).
Targets
--target <RULE>(required; repeatable): A rule to evaluate:NAME=DECIDER@THRESHOLD, the deciderarbiter(the logged decider's probabilities) orhabit.--min-ess <N>(default10): Refuse weighted estimates, a site's or the total's, below this effective sample size.
Output
--out <FILE>: Write the Markdown here (default: stdout).--json <FILE>: Also write every estimate as JSON here.
stretto refine
Refine the arbiter's predicates (RFC-001 §3.4). From one domain's decision log (phase0 --questions v2 --oracle-log, with the candidates asked by --candidates), fit the arbiter by cross-validation over tasks with each candidate added, and keep candidates greedily while one raises the held-out log-likelihood of the agents' steps by more than a penalty. With --examples, first write examples from the sites where the arbiter is weakest, for whoever proposes the candidates: a person or a model.
Usage: stretto refine [OPTIONS] --log <FILE>Inputs
--log <FILE>(required): One domain's decision log (phase0 --oracle-log).--candidates <FILE>: The candidates to weigh (thephase0 --candidatesfile the log was written with).
Search
--penalty <NATS>: Keep a candidate only when it raises the held-out log-likelihood by more than this many nats (default: half the log of the decisions scored, what BIC charges a parameter).
Examples
--examples <FILE>: Write examples from the sites where the arbiter is weakest here (Markdown), for a proposer.--dump <FILE>: The requestsphase0 --oracle-dumpwrote in the same run, for each example's state.--sites <N>(default4): Sites to take examples from.--per-site <N>(default6): Examples per site, one per task.
Output
--domain <NAME>(defaultairline): The domain, for the report's title.--out <FILE>: Write the Markdown here (default: stdout).--json <FILE>: Also write the search as JSON here.
stretto search
Search a flow's settings (RFC-001 §3.10): NSGA-II, from fugue-evo, over each site's threshold (0.10 to 0.95, or off) and the decider, to save the most turns for the fewest detours. Each setting is scored by the replay command after --, run with --flow FILE --flow-decider NAME --out DIR added; it must print a CHECK {…} line of totals, as pilot/check_flow.py does. A decision the replay cannot answer, such as a question missing from a replay cache, hands back and makes a setting look safer than it is. The report counts them (unanswered), so let the replay ask what its cache lacks. The hand-set flows start the search: every site at 0.3 with the arbiter (D0), and with the habit alone at 0.3 and at 0.9 (arm C). With --rescore, replay an earlier search's hand-set settings and front on the command's episodes instead.
Usage: stretto search [OPTIONS] --flow <FILE> --dir <DIR> -- <COMMAND>...Inputs
--flow <FILE>(required): The flow whose settings to search.--site <SITE>(repeatable): Search this site only (repeatable; default: every site where the flow may look something up).
Search
--population <N>(default16): Settings in each generation.--generations <N>(default10): Generations to breed.--seed <N>(default25): Seed for the search.--rescore <FILE>: Replay the hand-set settings and front of this earlier search (its --json) instead of searching.
Output
--dir <DIR>(required): Where each setting's flow and replay go.--out <FILE>: Write the Markdown here (default: stdout).--json <FILE>: Also write every setting replayed, and the front, as JSON here.
Arguments
<COMMAND>...(required): The replay command and its arguments, after--.
stretto promote
Promote a flow's sites (RFC-001 §3.7). Wherever the flow would decide in recorded sessions or τ²-bench results, score the lookup it would make: used if the agent made it later in the session, a detour if it never did. The promoted flow acts only after the calls whose record meets the bar, and hands back after the rest. Sessions recorded with stretto-proxy --flow-shadow have the flow's questions answered in the proxy's cache: pass it as --oracle-cache.
Usage: stretto promote [OPTIONS] --flow <FILE> --out <FILE>Inputs
--flow <FILE>(required): The flow to promote.--sessions <DIR>: Sessions recorded by stretto-proxy (a directory of*.jsonl); each counts as its own task.--results <FILE>(repeatable): τ²-bench results files (repeatable); files for other domains are skipped.--tau2 <DIR>: With --results: keep only the test split of this τ²-bench checkout, the tasks a flow compiled from it never trained on.--task-ids <IDS>(repeatable): With --results: keep only these tasks.
The System-One model
--oracle <ORACLE>(one ofjev,replay,mock; defaultreplay): Who answers the flow's questions:replay(the cache only; decisions it cannot answer are left out),jev(needs TYPESAFE_API_KEY) ormock.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--decider <DECIDER>(one ofarbiter,habit,reach): How the flow decides:arbiter,habitorreach, as it will be served (stretto-proxy --flow-decider). Default: as the proxy serves it by default, with its arbiter, elsereach, else, for a flow learned before flows held reach's counts,habit.
The bar
--threshold <P>(default0.3): The threshold the flow will be served with (stretto-proxy --flow-threshold).--min-used <X>(default0.7): The least share of the flow's lookups at a site that the agent made later in the session.--min-lower <X>(default0.5): The least lower bound on that share (Wilson, 90% two-sided).--min-tasks <N>(default3): The fewest distinct tasks (or sessions) the lookups came from.
Output
--out <FILE>(required): Write the promoted flow here.--report <FILE>: Write each site's record here, as Markdown (default: stdout).
stretto redact
Pseudonymize recorded sessions. A value that fewer than --keep-shared sessions contain becomes a salted hash, the same wherever it appears; a value more sessions share is kept, unless its field is named with --hash-field. A flow learned from the redacted logs matches one learned from the originals (docs/privacy.md).
Usage: stretto redact [OPTIONS] --sessions <DIR> --out <DIR>Options
--sessions <DIR>(required): Sessions recorded by stretto-proxy (a directory of*.jsonl).--out <DIR>(required): Write the redacted logs here, one per session.--salt-env <VAR>(defaultSTRETTO_REDACT_SALT): The environment variable holding the salt. Keep the salt secret, and the same for logs whose hashes should match.--keep-shared <N>(default3): Keep a value that at least this many sessions share.--hash-field <NAME>(repeatable): Hash every value of this field (a JSON key, at any depth of the arguments and results), and each word of it, wherever it appears, however many sessions share it: for the ids and names that a returning customer shares among their own sessions. Repeatable, or comma-separated.
stretto jev-check
Check that Jev is reachable with TYPESAFE_API_KEY: ask one small question (uncached) and print the answer, model version and latency.
Usage: stretto jev-checkstretto export-arbiter
Write a flow's arbiter to its own file, to ship: the predicates it weighs, the model it asks, and one fit of its weights. learn --arbiter-from serves it with a habit learned from new sessions. The flow's folds must share one fit (compile --pooled-arbiter, or a flow from learn).
Usage: stretto export-arbiter --flow <FILE> --out <FILE>Options
--flow <FILE>(required): The flow whose arbiter to write.--out <FILE>(required): Where to write it.
stretto fit-arbiter
Fit one arbiter, to ship, on the held-out decisions of one or more compiles: each --log is a decision log from compile --oracle-log asked with the same predicates. With logs of several domains, the arbiter is fitted on all of them at once.
Usage: stretto fit-arbiter [OPTIONS] --log <FILE> --predicates <FILE> --domain <NAME> --out <FILE>Options
--log <FILE>(required; repeatable): A decision log (repeatable).--predicates <FILE>(required): The predicates the logs' questions weighed (seedata/predicates-v2.json).--domain <NAME>(required): The name to give its domain, such asretail+airline.--model <MODEL>(defaultjev-latest): The System-One model id it asks.--out <FILE>(required): Where to write it.
stretto export-answers
Write every cached oracle answer to stdout as JSON lines ({"key", "response"}). Answers carry no benchmark text, so the bundle can be shared to replay Phase 0b without a key.
Usage: stretto export-answers [OPTIONS]Options
--oracle-cache <DIR>(default.oracle-cache): Replay cache to read.
stretto ask
Ask a System-One model questions of your own: JSON lines of requests (bare, or {"key", "request"} as --oracle-dump writes them), answered through the replay cache and written as {"key", "response"} lines, as export-answers writes them. For experiments that change the questions, such as text injected into their state.
Usage: stretto ask [OPTIONS] --requests <FILE> --out <FILE>Options
--requests <FILE>(required): The requests (JSON lines).--oracle <ORACLE>(one ofjev,replay,mock; defaultreplay): Who answers:replay(the cache only),jev(needs TYPESAFE_API_KEY; pays once per distinct question) ormock.--oracle-cache <DIR>(default.oracle-cache): Replay cache for oracle answers.--oracle-budget <DOLLARS>(default1): Refuse to start if uncached questions could cost more than this many dollars.--oracle-concurrency <N>(default8): Requests in flight at once.--out <FILE>(required): Where to write the answers (JSON lines; a request that failed gets"error"in place of"response").
stretto import-answers
Read JSON lines from export-answers on stdin into a replay cache.
Usage: stretto import-answers [OPTIONS]Options
--oracle-cache <DIR>(default.oracle-cache): Replay cache to fill.
stretto init
Print the configuration that runs an MCP server behind stretto-proxy in an MCP host, recording its sessions, and the steps from there to a served flow: for Claude Code a claude mcp add command, for the other hosts their JSON. The configuration goes to stdout and the steps to stderr. Nothing is written without --write.
Usage: stretto init [OPTIONS] --host <HOST> [-- <SERVER_COMMAND>...]Options
--host <HOST>(required; one ofclaude-code,claude-desktop,cursor,vscode): The MCP host to configure.--domain <NAME>: The server's name in the host, which is also the domain of its sessions and flows (default, with --flow: the flow's).--flow <FILE>: Run this flow (fromstretto learnorstretto promote) after the agent's calls. A flow without an arbiter is served on the chance of each lookup before the agent's next write (--flow-decider reach), or on its habit alone (habit) if it was learned before flows held those counts.--shadow: Run the flow in shadow: it decides and logs, but looks nothing up, forstretto promote.--record <DIR>: Where the proxy records sessions (default: ~/.stretto/logs/NAME, or ~/.stretto/shadow/NAME with --shadow).--write <PATH>: Write the host's configuration file here instead of printing it (for Claude Code, a project's .mcp.json). An existing file is left as it is, unless --force is given.--force: With --write, replace an existing file, with any other servers in it.--upstream <URL>: A Streamable HTTP server, such ashttps://example.com/mcp, in place of a server command: the proxy connects to it (stretto-proxy --upstream).--upstream-header <NAME=VAR>(repeatable): With --upstream: send header NAME with the value of environment variable VAR, which the host gives the proxy in itsenv(stretto-proxy --upstream-header). The value is not written.
Arguments
<SERVER_COMMAND>...(not with--upstream): The MCP server's command and its arguments, after--.
stretto doctor
Check the installation: the versions of stretto-proxy, stretto-procedure and stretto-mcp-demo on PATH, whether the data directory (~/.stretto) is writable, whether TYPESAFE_API_KEY is set (never its value), and the flows and recorded sessions in the data directory. Exits with 1 when something needs fixing.
Usage: stretto doctor [OPTIONS]Options
--network: Also ask Jev one question (uncached) if a key is set, asjev-checkdoes. Without it, doctor makes no network request.--data <DIR>: The data directory to check, instead of ~/.stretto, asstretto-console --dataserves it.
stretto completions
Print the completion script for stretto in bash, zsh, fish, powershell or elvish to stdout. docs/install.md says where each shell reads it.
Usage: stretto completions <SHELL>Arguments
<SHELL>(required; one ofbash,elvish,fish,powershell,zsh): The shell.
stretto-proxy
Forward a stdio MCP server's traffic, record it for stretto, and optionally run a flow, policy guards and a commit tool on it. Put this in an MCP host's configuration in place of the server's command, with the real command after --. stdout carries only the protocol; the proxy's own messages go to stderr. The exit status is the server's, or 125 if the proxy itself fails. Without --flow, --guards, --commit, --context or --confirm-judge, every line is forwarded byte for byte and nothing is parsed.
Usage: stretto-proxy [OPTIONS] [-- <SERVER_COMMAND>...]Recording
--record <DIR>: Write a session log (JSONL) into this directory, created if missing. A leading~is expanded, since hosts start servers without a shell.--domain <NAME>: Domain for the log header, e.g.retail; --guards uses its rules.--agent-model <MODEL>: Model that drives the agent, for the log header.--retain-days <DAYS>: On start, delete what is older than this many days: the session logs in --record with the flow and confirmation logs beside them, and the answers in --oracle-cache.
Flows
--flow <FILE>: Run this flow (fromstretto compileorstretto learn) after each of the agent's calls, and append its lookups to the result.--flow-threshold <P>(default0.3): Take a lookup when the tool's probability times its arguments' agreement is at least this.--flow-decider <FLOW_DECIDER>(one ofarbiter,habit,reach): Where the tool's probability comes from:arbiter(the habit, the System-One model's answers and the predicates, combined),habit(the habit alone, which never asks a System-One model and needs no key), orreach(the habit's counts for whether the agent makes the lookup before its next write; no key either). Default: the flow's arbiter, elsereach, else, for a flow learned before flows held reach's counts,habit.--flow-per-call <N>(default8): Lookups appended to one result, at most.--flow-per-session <N>(default40): Lookups per session, at most.--flow-questions <N>(default300): Questions to the System-One model per session, at most.--flow-log <FILE>: Append the flow's decisions here (default: next to the session log).--flow-tools <TOOLS>(repeatable): The only tools the flow may call on its own (comma-separated, or the option repeated). A server'sreadOnlyHintsays a call changes nothing, not that it is free, unlogged or fine to make unasked: a read can be metered, rate-limited, or recorded as an access. Without this, the flow may call every tool it reads as a lookup.--flow-shadow: Shadow mode: the flow decides after each call and logs what it would look up ("shadow": true), but makes no lookups, so the agent gets the server's results unchanged.stretto promote --sessionsthen makes the same decisions again from the answers cached in --oracle-cache, and scores them against what the agent did.--flow-explore <EPSILON>: Explore: with this probability, take a lookup other than the rule's choice, drawn by the decider's probabilities among those that bind. Each decision in the flow log then carries itspolicy: every option and the chance that the flow took what it took, forstretto evaluate. 0 explores nothing but still logs it.--flow-explore-seed <N>(default0): Seed for the exploration draws.--task-id <ID>: Task id, which picks the flow's fold (default: the session).
The System-One model
--oracle <ORACLE>(one ofjev,replay,mock; defaultjev): Who answers the flow's and the confirmation judge's questions:jev(needs TYPESAFE_API_KEY),replay(the cache only) ormock.--oracle-cache <DIR>(default~/.stretto/oracle-cache): Replay cache for the System-One model's answers.
Writes
--guards: Check each of the agent's calls against the policy guards of --domain (retailorairline), and refuse the ones they fail.--confirm-judge <MODE>(one oflog,enforce): Put each write the guards check for a confirmation to the System-One model too, with the questionsstretto confirmasks:logrecords each judgment;enforcealso refuses a write the judge fails. Needs --guards and --context. A judge that cannot answer refuses nothing.--confirm-second <QUESTION>(one ofproposed,described): Also ask the second question (proposed: had the agent proposed the change?); a write then fails unless both answers are yes.--confirm-second-shadow: Ask the second question but only log its answer (shadow mode): a write then fails on the first answer alone.--confirm-threshold <P>(default0.5): A write fails when an answer's probability of a yes is below this.--confirm-questions <N>(default100): The judge's questions per session, at most.--confirm-log <FILE>: Append the judgments here (default: next to the session log).--commit: Addstretto_commit, which makes several calls in one, in order, each checked by the guards.
The conversation
--context <FILE>: Read the conversation from this file, which the host appends to as JSON lines:{"role": "user" | "assistant", "content": text}.
The server
--upstream <URL>: A Streamable HTTP server to proxy for, such ashttps://example.com/mcp, in place of a server command. The host still runs the proxy as a stdio server.--upstream-header <NAME=VAR>(repeatable): With --upstream: send header NAME with the value of environment variable VAR, such asAuthorization=GITHUB_AUTHfor a variable that holdsBearer …. Values are never logged.
Arguments
<SERVER_COMMAND>...(not with--upstream): The MCP server to run, and its arguments.
stretto-procedure
Run a compiled procedure on a ticket against an MCP server (stdio), with no model. The run goes to stdout as JSON: its calls, its check of the outcome the ticket states, and its verdict (resolved, transferred, or hand_back when the check failed and the ticket should go to a model).
Usage: stretto-procedure [OPTIONS] --procedure <FILE> -- <SERVER>...Options
--procedure <FILE>(required): The procedure (JSON), asscripts/telecom_workflow.py --exportwrites it.--ticket <TEXT>(not with--ticket-file): The ticket, as text.--ticket-file <FILE>(not with--ticket): The ticket, from a file.--out <FILE>: Write the run here instead of to stdout.
Arguments
<SERVER>...(required): The MCP server to run against, and its arguments (after--).