Flows
A flow is what stretto serves behind an agent: a JSON file that says, for one server's tools, which lookups the agent made after each call, how often, and where each lookup's arguments came from. It holds counts, tool and argument names and JSON paths, and no model weights, so a person can read it, and a change to it can be reviewed like code.
Learn one
stretto learn --sessions ~/.stretto/logs/orders --domain orders \
--habit-only --out ~/.stretto/orders.flow.json--sessions DIRreads the session logsstretto-proxy --recordwrote.--domain NAMEnames the flow's domain.--habit-onlylearns from counts alone and asks no model. Without it,learnalso fits an arbiter on held-out sessions, which asks a System-One model and needsTYPESAFE_API_KEY(deciders).--manifest FILEgives tool kinds when the server'stools/listhas noreadOnlyHintannotations (tool kinds).--constantsalso learns arguments the agent always passes the same way, such as a page size (bindings).
Every option is in the CLI reference. stretto compile writes a flow from τ²-bench's published results instead, for research.
The console runs learn as a job, and its page for a flow draws the flow's graph of calls and lookups. Its threshold slider shows which lookups would act at another threshold.
What is in it
| Part | What it holds |
|---|---|
manifest | The server's tools and their kinds: read, write or generic. A flow calls only read tools. |
habit | Counts of what the agent did next after each short history of steps. |
reach | The same histories, counted for whether each tool was called before the agent's next write. |
sites | After each tool, the lookups the agent made next in training, and how often. They are the only lookups the flow can make there. |
bindings | For each lookup's arguments, the earlier results and JSON paths their values came from, and how often binding them that way matched the agent. |
contracts | Each tool's input contract as the server listed it. The proxy makes no lookup of a tool whose server now lists another. |
program | The flow's run after each call, as a fugue program. |
provenance | What the flow learned from: the agent model the logs name, and how many successful sessions. |
A flow may also hold an arbiter (folds, predicates, model), code features (map), a promotion record (promoted) and per-site thresholds (thresholds). File formats names every field and says what a reviewer should check. Flows are written on one line; jq . orders.flow.json prints one for reading, and stretto flow-show renders one for review.
The counts
Each step of a session is abstracted to its tool, whether it returned or failed, and sometimes a feature of its result. The habit is a hierarchical Dirichlet back-off model of the next step given the last two, so that a history never seen in training falls back to a shorter one. With
The reach counts ask a different question of the same histories: not which action came next, but whether an action came at all before the agent's next write,
The chances need not sum to one: after finding a customer, an agent reads their record and then an order before it writes anything. Deciders says how a flow uses each.
Learning is counting
Every estimate in a flow is a posterior predictive of conjugate counts, so learning from another session adds its counts. There is no gradient and no training run to schedule, and learning is quick. That makes the loop cheap:
- Serve a flow, and keep recording.
- Learn again from the sessions recorded since, or from all of them.
- Compare the new flow with the one being served, with
stretto flow-diff. It exits with 1 when the new flow can do something the old one could not, such as a new lookup or a new source for an argument.
Other agents' sessions can teach a flow too. Where agents act alike, they teach as much as the agent's own; where they do not, only the agent's own sessions teach its habits. Replayed in τ²-bench's telecom domain, a hundred of an agent's own sessions saved 15.2% of its turns, against 12.6% from all 1,184 of four other agents' (the paper, §4.3).
The run after each call
A flow's run after each of the agent's calls is a probabilistic program, in fugue's serializable format. stretto flow-show prints it:
let prev = call;
let failed = call_failed;
for i in 0..max_lookups {
let d <- sample(addr!("decide", i), Decide(prev, failed));
if d == 0 {
break;
}
let ok <- sample(addr!("outcome", i), Outcome(d));
prev = d;
failed = !ok;
}
pure(prev)Decide is a decision between handing back (0) and the lookups offered after the call just made; Outcome is whether a lookup succeeds. The proxy decides each decide#i with the flow's decider and takes each outcome#i from the server. Because the run is a fugue program, the same flow can be executed against a server, simulated, and used to score a recorded session: that is how stretto audit works.
Related
- Lookups and detours: what serving a flow does
- Bindings: where arguments come from
- File formats: every field
- The console: a flow's graph at any threshold