File formats
Two kinds of stretto file are meant to be kept, reviewed and shared:
- a flow (
*.flow.json), written bystretto compileandstretto learnand served bystretto serve,stretto flow-serveandstretto-proxy --flow; - an arbiter (
data/arbiters/*.json), written bystretto export-arbiterand read bystretto learn --arbiter-from.
Both are JSON. This page names every field, says what sets it, and says what a reviewer should check. The proxy's session logs are documented in its README. The Rust types are Flow, Arbiter and Fitted in crates/stretto-report/src/flow.rs and arbitrate.rs; a test (crates/stretto-report/tests/formats.rs) fails when a serialized field is missing from this page.
Flows are written on one line. jq . some.flow.json prints one for reading. Maps whose keys are not strings are written as lists of [key, value] pairs sorted by key, so two flows compiled from the same data differ only in compiled_unix_ms.
Reviewing a flow
A flow makes read-only calls on the agent's behalf, so a review asks what it may call, when, and with what. stretto flow-show renders these fields for a reviewer, and stretto flow-diff lists what changed between two flows (reviewing flows). In the file:
manifest.tools. The flow calls only tools markedread. A tool that writes but is markedreadis the one mistake that matters here. Kinds come from the server'sreadOnlyHintannotations or a--manifestfile, so check them against what the tools do.sites.next. After each call, these are the lookups the flow may make next, with how often the agent made each in training. A lookup seen once or twice is a thin basis.bindings.sources. These say where each lookup's arguments come from. A required argument with no source, such as an email only the customer knows, means the flow never makes that lookup itself.provenance. Check which sources the flow learned from, how many successful episodes (habit_episodes) and how many held-out decisions (arbiter_cases). The cold start found an arbiter fitted on a handful of decisions worse than none.map.ids. Holds values copied from training outputs, such as user ids. Treat a flow with features like the data it was learned from before sharing it. The rest of the file holds tool names, argument names, JSON paths, counts, the tools' documentation, and the questions' text.
The flow IR (stretto_flow: 1, or 2 with per-site thresholds)
| Field | What it holds |
|---|---|
stretto_flow | The format version, 1 |
provenance | Where the flow came from |
vocab | The actions the habit predicts, in id order |
manifest | The domain's tools, their kinds and their documentation |
map | Code features: tool output fields that sharpen the habit |
group | The habit's condition for a live episode |
habit | The habit: counts of what the agent did next after each short history |
sites | The lookups the flow may make after each call |
predicates, weighed | The yes/no questions asked with each next-step question, and those the arbiter weighs |
folds | The arbiter's fitted weights and Jev's record, one per fold of tasks |
bindings | Where each lookup's arguments came from in training, and how often binding them that way matched the agent |
model | The System-One model the flow asks |
program | The flow's run after each call, as a fugue program |
promoted | Where the flow may act, once promoted; absent until stretto promote writes it |
thresholds | Per-site thresholds, in place of the served one at those sites; absent unless a search set them |
contracts | Each tool's input contract when the flow was learned from recorded sessions: {tool: "order_id:string!, reason:string"}, each argument with its JSON type and ! when required. stretto-proxy makes no lookup of a tool whose server lists another. Absent for a flow compiled from τ²-bench results |
reach | Over the habit's histories, how often each action came before the agent's next write, for --decider reach; absent in a flow learned before flows counted it |
provenance
stretto: the version of stretto that wrote the flow.sources: the training sources, by label. Forcompilethese are τ²-bench's agent models. Forlearnit is the agent model the session logs name: the proxy's--agent-model, or else the name the host gave ininitialize. A flow serving an arbiter from elsewhere also lists that arbiter's sources, asarbiter (<domain>): <source>when the arbiter came from a file andarbiter: <source>when it came from another flow.habit_episodes: the successful training episodes the habit learned from.arbiter_cases: the held-out decisions the arbiter was fitted on; 0 for a flow with no arbiter. A flow serving a shipped arbiter carries that arbiter's count.compiled_unix_ms: when the flow was written, in milliseconds since the Unix epoch.
vocab
The actions the habit predicts: "Respond" (a message to the customer with no tool call) first, then {"Tool": name} for each tool, sorted. An action's id is its position. One more id, the list's length, stands for any action not in the list. The habit's counts use these ids.
manifest
domain: the domain's name, from--domain.tools: each tool's kind:read(reads without changing anything),write(changes state) orgeneric(neither, or unknown).compiletakes kinds from τ²-bench's tool types.learntakes them from the servers'readOnlyHintannotations in the recordedtools/list, or from--manifest. A tool that gave no hint isgeneric. Onlyreadtools are ever looked up by the flow.docs: each tool's documentation, where the source had any:summary, one paragraph, andargs, each argument's description. The System-One questions show it.
map
Code features (RFC-001 §3.3): values in tool outputs that change what the agent does next, such as whether an order has one item or several. compile and learn choose them from training outputs. A flow learned from a few sessions often has none; compile --no-features skips them.
fields: for each tool, the output fields read:{"Scalar": path}, a scalar at a JSON path such asstatusoraddress.state, or{"Len": path}, whether an array at a path (""for the whole output) is empty, has one element or more ("0","1","2+").ids:[[tool, values], id]: each combination of values seen in training, per tool, and its feature id (1 to 63). Combinations never seen, and those past 63, share id 0, "no feature".
Check: ids holds training values verbatim.
group
The habit's condition for a live episode. The habit can be conditioned on an episode's goal, but a flow serves sessions nobody has named, so flows are compiled goal free: every training episode has group 1, and so does every live one.
habit
The habit: a hierarchical Dirichlet back-off model of the agent's next action, given its last order steps (RFC-001 §3.3; stretto_model::world). With c_j the last j steps,
P_j(a | c_j) = (n(c_j, a) + α · P_{j-1}(a | c_{j-1})) / (n(c_j) + α)and P_{-1} uniform over the vocabulary, so a history never seen in training falls back to a shorter one.
base.order: the steps of history, 2.base.alpha: the concentration α, shared by every level. It is the posterior median given the successful training episodes, from--alpha-samplesMetropolis–Hastings draws (with fugue), or--alphawhen that is 0.base.vocab_size: the number of action ids, the vocabulary's length plus one.base.levels: one list per history length, 0 toorder. Each entry is[history, counts], wherecountsis{"by_action": [[action id, count], …], "total": count}. A history is a list of step symbols; a symbol is(action id × 4 + outcome) × 4096 + feature id, where the outcome is 0 for a tool call that succeeded, 1 for one that failed, 2 for a message the customer answered and 3 for one that ended the episode. So8192is a message to the customer that they answered.beta: the concentration of the layer conditioned ongroup, equal to α.top: that layer's counts,[[group, history], counts], over histories ofordersteps. In a goal-free flow they repeat the longest level's.
reach
The habit's histories counted for another event: not which action came next, but whether an action came at all before the agent's next write (stretto_model::world::BackoffModel::fit_reach). It has the fields of habit.base, and each history's total counts the steps that followed it, where by_action counts, for each action, the steps from which that action came before the next write. Each action is then its own Beta–Bernoulli, backed off as the habit is:
R_j(a | c_j) = (m(c_j, a) + α · R_{j-1}(a | c_{j-1})) / (n(c_j) + α)A read's result stays current until the next write, so R is the chance that a lookup made now answers a call the agent will make, and --decider reach takes it in place of the habit's next-step probability. The chances need not sum to one. A build that does not know the field serves the flow as before, so it keeps the format version.
sites
reads: the read-only tools, from the manifest.next:[[tool, failed], {lookup: count}]: after a call totoolthat failed (true) or not, the lookups the agent made next in training, and how often. They and handing back are the options at that site. At a site with no lookups the flow hands back without asking.feeds: withcompile --dataflow-hints, for each lookup,[[write, argument], count]: how many training episodes passed a value from its results to that write argument. The System-One questions describe each lookup by the two arguments it supplied most often ("Its results supplyorder_idforcancel_pending_order"). Empty otherwise.every_read: written only when true, with--manifest-options: every read-only tool is offered at every site, not only the lookups seen there.
predicates and weighed
The yes/no questions asked with each next-step question (RFC-001 §3.4; data/predicates-v2.json), and those of them the arbiter weighs. Each has an id (asked as pred_<id>), favors (the options its answer bears on: same_lookup, the lookup just made, made again; any_lookup, every lookup; hand_back; or {"lookup": TOOL}, one named lookup wherever a site offers it), the question, and what yes and no mean. With --no-predicate-features they are asked but not weighed, and weighed is empty. A flow with no arbiter has neither.
folds
The arbiter (RFC-001 §3.6): for each option at a site, P(option) ∝ exp(Σ weights[i] · x[i]), times the site's coverage. The flow makes the most likely lookup when that probability, times its binding's chance (see bindings), reaches the serving threshold; otherwise it hands back. A compiled flow has five arbiters, one per fold of tasks, each fitted on the other folds' held-out decisions, and a task is judged by its fold's. A flow compiled with --pooled-arbiter, learned from sessions, or serving a shipped arbiter has five copies of one fit. A flow with no arbiter (learn --habit-only) has none, and decides with the habit alone.
weights: the conditional logit's weights, in this order:- the log of the habit's probability of the option (every step the flow cannot take counts as handing back);
- the log of the System-One model's probability for it in the next-step question, a choice among the options;
- the log of its probability from the split questions: whether to go on at all, then which lookup;
- 1 for handing back;
- one for each weighed predicate: the log-odds of its yes, on the options it favors;
- last, the reliability weight: the log-odds of Jev's agreement with the agent at this site, on Jev's pick.
With the three predicates of
data/predicates-v2.json, that is eight weights.rates.sites:[[site, [decisions, agreed, covered]], …]: for each site (a tool, with(error)after its name when the call failed), the held-out decisions there, how many of them Jev's pick matched the agent's step, and how many offered the agent's step among the options.rates.agreeandrates.cover: the same shares over every site. A site's reliability is(agreed + 10 · agree) / (decisions + 10), and its coverage(covered + 10 · cover) / (decisions + 10). At a site the arbiter never saw, such as any site of another domain, they areagreeandcover.
bindings
How the flow fills a lookup's arguments, learned from the agent's own lookups in training.
args: for each lookup,[calls, {argument: calls that passed it}]. An argument passed in at least 90% of the calls is required. The flow passes only required arguments.sources:[[lookup, argument], {"values": n, "found": [[[tool, path], count], …]}]: of thenstring values the argument took, how many were found in an earlier successful output oftoolat the JSON pathpath($is the whole output,[*]any element). An output that is not JSON is read as one string, or, when it has several lines, as the list of its lines, so$[*]is one line of it. The flow binds each required argument to the first value at one of its sources that it has not passed already, taking the sources in the ordersite_sourcesgives at the site, and otherwise starting with the most recent output. It prefers a value the customer mentioned, or one whose record they mentioned by another of its fields. It uses only sources found at least twice that account for at least 10% of the values. A lookup with no required arguments is made once per session.agreed: for each lookup,[[agreed, calls], [agreed, calls]]: at the agent's own lookups in training, how often the binding picked the agent's arguments, first when the customer had not mentioned the values picked, then when they had. The binding's chance is(agreed + 1) / (calls + 2).named_other(written only when training counted any): for each lookup,[used, picks]: where the customer had mentioned a value at the lookup's sources and the binding would pass another, such as a second reservation after the customer asked about one, how many of the distinct values it would pass there the agent went on to pass itself. Its chance there is(used + 2c) / (picks + 2), wherecis the unmentioned chance. Scored at the agent's own lookups, asagreedis, such picks look right, since an agent that reads a second record picks it as the binding does; what that misses is that the agent mostly reads only the record the customer named. A flow without it gives such picks the unmentioned chance.described_read(written only when training counted any): for each lookup,[used, picks], asnamed_othercounts them, where the record the customer described had been read: the binding would pass a value of a list other values of which the lookup was already called with, and what one of those calls returned held a value the customer gave that another's did not. A value the customer gave has at least 4 characters and a digit, and either the customer wrote it or the agent passed it before any result held it, as a ticket's phone number. After the agent has read the line with the customer's number, the next of their lines is such a pick. Its chance is as withnamed_other, which comes first where both apply. A flow without it gives such picks the unmentioned chance.lists(written only when a lookup took lists of strings):[[lookup, argument], {"values": n, "found": [[[tool, path], count], …]}], assourcescounts them, for an argument that took lists: of thenlists, how many had every value atpathof one earlier successful output oftool(the most recent that held them all). The flow passes every value at the path of the most recent output of such a source, in order, unless the lookup was already given that list: after a city's hotel listing, the names in it, to a lookup of their prices. It uses the sourcessourceswould. A flow without it never makes a lookup whose required argument is a list.bare(written only for lookups with no required argument, none the agent passed in 90% of its calls, that the agent called with some argument in training): how many of the agent's calls passed none. Such a lookup is made with no argument, and is the agent's own call only when the agent's is bare too: a search the agent always narrows by one of its optional arguments never is. Its chance is(bare + 1) / (calls + 2). Any other lookup without arguments, such as a read that takes none, has the chance 1.constants(written only bylearn --constants):[[lookup, argument], value]pairs, each an argument the agent passed with one value in every call of the lookup that passed it, at least five calls and at least half of the lookup's, and that was never found in an earlier output, such as a page size; a string a user wrote with a digit or an @, such as an id or an email, is never one, however many sessions shared it. The binding passes the value as the agent did, so that the lookup is the agent's own call; a required argument with a constant needs no source.site_sources(written only for arguments with more than one source):[{"tool", "arg", "site", "source": [tool, path], "values"}, …]: how many of the argument's values the agent took from each source at each site, the tool whose result came last before the call. Each read of a batch counts at the site of the read before it, since a flow makes a batch's reads one after another. Where a site has at least 5 such values, the binding takes the sources in that order, the most recent output first within each, in place of recency alone: after reading a phone line, an agent working through the customer's lines takes the next line id, not the plan id of the line it just read. A flow without it orders every source by recency.
model
The System-One model id the flow requests, jev-latest by default. It is part of each question's cache key. Empty in a flow with no arbiter.
program
The flow's run after each of the agent's calls, as a program in fugue's serializable format (fugue::program; RFC-001 §3.5). It is a JSON object: the format's version, fugue_program (1), then the program's body and ret, whose nodes fugue documents. stretto compile and stretto learn write the standard run, which stretto flow-show prints as text:
let prev = call;
let failed = call_failed;
for i in 0..max_lookups {
let d <- sample(addr!("decide", i), Decide(prev, failed));
if d == 0 {
break;
}
let ok <- sample(addr!("outcome", i), Outcome(d));
prev = d;
failed = !ok;
}
pure(prev)- Tools by number. 0 is handing back (
respond), then comesites.readsin order, then the manifest's other tools. A decision's value is the number of the lookup it makes. - Data.
callis the number of the tool whose call just returned, andcall_failedwhether it failed.max_lookupsis the proxy's--flow-per-call. - Distributions.
Decide(prev, failed)is the decision after a call to toolprev. Handing back and the lookupssites.nextoffers there get a flat Dirichlet's posterior predictive, given what the agent did next in training; the other lookups get 0.Outcome(d)is whether a call to tooldsucceeds, a flat Beta's posterior predictive given how often its calls did. These two are the only distributions a flow's program can use. - Checked on load. A program that uses another name, or gives
DecideorOutcomethe wrong number of arguments, does not load. A flow written withoutprogramruns the standard one.
The proxy decides at each decide#i with the decider it serves the flow with (--flow-decider: the arbiter, the habit alone, or reach), and takes each outcome#i from the server.
Check: any program other than the standard run is code the proxy runs. stretto flow-diff lists a change to it as needing review.
promoted
Written by stretto promote (RFC-001 §3.7), and absent until then. A promoted flow acts only after the calls whose record met the bar. After the rest it hands back, with the reason the site is not promoted.
bar:threshold, the one the flow was scored at, as it will be served;min_used, the least share of its lookups at a site that the agent made in a later LLM turn;min_lower, the least lower bound on that share (Wilson, 90% two-sided); andmin_tasks, the fewest distinct tasks the lookups came from. Each recorded session counts as its own task.sites: for each site scored, by name (the tool, and(error)after a failed call):decisions, the times the flow decided there;lookups, the lookups it would have made;used, the ones the agent made in a later LLM turn (the rest are detours);tasks;lower; andpromoted. A site never scored is not promoted.
Check: a site newly promoted lets the flow act where it handed back, and stretto flow-diff lists it as needing review.
thresholds
A map from site to threshold, written when a flow's settings come from stretto search (RFC-001 §3.10), and absent otherwise. A flow with thresholds is format version 2 (Versions). At a listed site the flow takes a lookup when it clears that site's threshold instead of the one it is served with (--threshold). Above 1, the flow never acts at the site: it hands back with the reason the site is switched off, before asking anything.
A flow, walked through
The live cold start's flow was learned from five retail sessions GLM-5.3 ran through the proxy, and it served the three tasks of the cold start's live run.
provenance:sourcesis["glm-5.3"]. The habit learned from 3 successful sessions; the arbiter was fitted on the other sessions' 15 decisions.manifestlists retail's 16 tools: 7read, 7writeand 2generic(calculate,transfer_to_human_agents).mapis empty: three sessions support no code feature.sites.nexthas five sites. Afterfind_user_id_by_emailthe agent calledget_user_details(once). Afterget_user_detailsit calledget_order_details(3 times). Afterget_order_detailsit calledget_order_details7 times andget_product_detailsonce. Those are the only lookups the flow can make.bindings.sources:order_idforget_order_detailscame fromget_user_detailsat$.orders[*], all 10 times;product_idcame fromget_order_detailsat$.items[*].product_id, 4 of 4;user_idcame fromfind_user_id_by_name_zipat$twice and fromfind_user_id_by_emailonce. A source found once is not used, so afterfind_user_id_by_emailthe flow cannot bindget_user_detailsand hands back.
The email, name and zip code were never found in an output, since they come from the customer, so the flow never finds the user itself.
bindings.agreed: the binding picked the agent's own order id at all 10 of itsget_order_detailscalls, a chance of 11/12. Forget_product_detailsit matched 2 of 4 times, all where the customer had mentioned the product, a chance of 3/6.foldsholds five copies of one fit. Jev's pick matched the agent at 9 of the 15 held-out decisions (agree0.6), and the agent's step was always among the options (cover1.0). Atget_order_detailsJev matched 3 of 9, so the arbiter trusts it less there than elsewhere.
The procedure IR (stretto_procedure: 1)
A whole workflow compiled once from agents' traces, writes included, for where no user speaks: scripts/telecom_workflow.py --export writes it from τ²-bench telecom's solo runs, and stretto-procedure runs it against an MCP server with no model (results). The published one is the workflow of those results, with its identifiers held as where they came from.
| Field | What it holds |
|---|---|
stretto_procedure | The format version, 1 |
domain | The domain it was compiled for |
provenance | sources, the runs it was compiled from, and training_tasks, how many tasks they cover |
symbolic | Whether its identifiers are held as where they came from and bound again at run time, or are constants of the traces |
max_calls | The most calls a run makes |
stop | The action that ends a run |
handoff, handoff_arguments | The tool that hands the customer to a person, which also ends a run, and what it is called with |
reads | The tools that only read: the guard skips a read made since the last write, and a write already made |
vocab | The ticket's words and word pairs the trees may ask about |
trees | Per site, the tool that just returned (! after a failed call) or start, its tree |
fallback | The node for a site training never saw |
checks | Each outcome a ticket may state: its phrase, the read that checks it (probe), and what that read's result, lowercased, contains and lacks when the outcome holds |
A tree's node is a question, {"feature": F, "yes": NODE, "no": NODE}, which goes to yes when the run's state has feature F, or a leaf, {"calls": [[ACTION, COUNT], …]}: the actions taken there in training, by count, in the order the run tries them. It takes the first that binds and is not a repeat. An action is a tool with its arguments as JSON (toggle_roaming, enable_roaming{"customer_id": …}), or stop. With symbolic, an argument that is an identifier is held as where it came from: @TOOL$.path[*], the most recent result of TOOL with a value at that path, the first in a list not yet passed as that argument, with ?{…} the fields that set its record apart; or @ticket and a pattern, the ticket's first text of that shape (\d{N} for N digits, \ before a literal character).
A run's state is a set of features: ticket: WORD for each word of the ticket in vocab; called TOOL and made ACTION for each tool called and write made; and, from each tool's last result, TOOL:error or TOOL:error=False, and for a record TOOL:FIELD=value for each short field (ids and free text left out), TOOL:len(FIELD)=0, 1 or 2+ for its lists, and TOOL:FIELD.*.status=value for a map of records; for a list TOOL:len($); and for a phone's check, TOOL:KEY=value for each part of each Key: value line and TOOL:says LINE for a short line of its own, parts with digits left out. After a record that holds a value the ticket gives (the line with the ticket's number), another record of its kind does not replace it in the state.
The arbiter file (stretto_arbiter: 1)
A flow's arbiter on its own, to serve with a habit learned elsewhere (data/arbiters). stretto export-arbiter writes it from a flow whose folds share one fit, and refuses a flow whose five folds were fitted apart, since none of those is the arbiter for new tasks. It is pretty-printed.
| Field | What it holds |
|---|---|
stretto_arbiter | The format version, 1 |
domain | The domain whose decisions it was fitted on |
provenance | The flow it came from, as in a flow |
predicates, weighed | The questions it asks and weighs, as in a flow |
fitted | One fit: weights and rates, as in one of a flow's folds |
model | The System-One model it asks |
learn --arbiter-from gives the learned flow the arbiter's predicates, weighed and model, five copies of fitted, its arbiter_cases, and its sources. The file holds no answer or conversation text, only counts, weights and the questions.
Versions
Every reader checks the version field first and refuses any other version, with a message naming both. Before the first release, fields added, such as every_read, kept version 1. Since it:
- Flow version 2 adds
thresholds,bindings.named_other,bindings.described_read,bindings.site_sources,bindings.lists,bindings.bareandbindings.constants, which a build that reads only version 1 would ignore and must not: it would give a record the customer did not ask about the chance of any other, bind from another source, make fewer lookups than the flow was learned to, make searches the agent never makes, or make lookups without the arguments the agent always passes. A flow is written as version 2 only when it has any of them, and this build reads versions 1 and 2.
The rules:
- A new version comes with any change after which one build would read another's file wrongly. That covers a field removed, renamed, or given a new meaning or encoding (the symbols, the vocabulary's ids, the order of the weights). It also covers a new field an older build would ignore but must not, as it would ignore
every_readand offer fewer lookups. - The same version holds for a new field that an older build can ignore without acting differently, read with a default that keeps today's behavior.
- Old files. Before 1.0, a build reads only its own version and those it extends, and the release notes say which version each release reads. A flow is cheap to learn again from the recorded sessions or results it came from, so a new version may not convert old files. A shipped arbiter is rebuilt from the published answer bundles (data/arbiters).