How stretto works
Where stretto sits
stretto-proxy wraps any MCP server. The MCP host runs it in place of the server's command, and it forwards every message both ways, unchanged. The agent keeps its prompt and its tools. Until you give the proxy a flow, it only forwards and records.
stretto init prints the host's configuration, for Claude Code, Claude Desktop, Cursor or VS Code. One proxy per server: the server's command after --, or a Streamable HTTP server with --upstream.
- →find_user_id_by_email {"email":"c7@example.com"}
- ←user_7
- →get_user_details {"user_id":"user_7"}
- ←{"orders":["#W7a","#W7b"],…}
- →get_order_details {"order_id":"#W7a"}
- ←{"order_id":"#W7a","status":"pending",…}
- →cancel_pending_order {"order_id":"#W7a",…} write
Each log holds every message the proxy forwarded, with timestamps: the calls, their results, and the conversation if the host shares it.
Record sessions
With the configuration stretto init prints, the proxy writes each session to a log in ~/.stretto/logs/, under the server's name. Use the agent as usual. The logs are all a flow learns from, so record the kinds of request the agent will see.
Ten of an agent's own sessions give 96% (retail) and 93% (airline) of what all of them do. Replay on τ²-bench, three agents, three orders each.
→ get_user_details {"user_id":"user_7"} ← {"orders":["#W7a","#W7b"],…} → get_order_details {"order_id":"#W7a"}
Learn a flow
stretto learn counts, across the sessions, which reads followed which calls, and where each read's arguments came from: here, the order ids in the user's record. The result is a flow, a file you review like code with stretto flow-show and stretto flow-diff.
A flow only calls tools the server does not mark readOnlyHint: false; --flow-tools narrows that to a list.
→ get_user_details {"user_id":"user_7"} {"email":"c7@example.com","orders":["#W7a","#W7b"],…} --- Also looked up automatically (current results; no need to repeat these calls) --- get_order_details {"order_id":"#W7a"}: {"order_id":"#W7a","status":"pending",…} get_order_details {"order_id":"#W7b"}: {"order_id":"#W7b","status":"pending",…}
Serve it
Served, the flow runs after each of the agent's calls. It makes the lookups it expects the agent to need, with arguments bound from the results so far, and they ride in the same tool result. The agent sees them before it would ask, so it spends fewer LLM turns. There is no new tool and no change to the prompt.
Each block is an LLM turn; the outlined ones are the turns the lookups save. Run a flow in shadow first (stretto init --flow FILE --shadow): it decides and logs, but looks nothing up. Then stretto promote keeps the sites whose lookups were the agent's own.
used: it answered the agent's next call
a detour: never used
It changes no state: a flow only reads, and reads leave the state as it was.
What a detour costs
A detour is a lookup the agent did not use. Its result stays in the context, so every later LLM turn reads it again: that is its cost, in tokens. It cannot change the state, since a flow only reads: it reaches an episode only through what the agent reads, never through the tools.
On the live paired run, a detour carried about 2,530 input tokens over the rest of its episode, and a saved turn saved about 6,000. GLM-5.3 on τ²-bench retail and airline.
When to make a lookup
The reach decider asks no model and needs no key. It makes a lookup when the chance that the agent uses it before its next write clears a threshold set by costs, δ / (β + δ). That chance, not the chance the read comes next, is the one that pays: the agent may reply first, and still read it before it writes.
Replay, τ²-bench retail, two agents: after a user's details, the habit puts an order next at 0.64, and the agents read one before their next write 94% of the time; after an order, a product: 0.08, and 38%. Live costs set the threshold: 2,530 / (6,000 + 2,530) ≈ 0.3.
The numbers as a table
| Agent | Trials | Without | With stretto | Change (95% CI) |
|---|---|---|---|---|
| Claude Sonnet 5 | 3 | 9.92 | 7.88 | −20.5% (−24.4, −16.5) |
| Claude Haiku 4.5 | 3 | 8.81 | 6.83 | −22.4% (−29.4, −15.8) |
| GLM-5.3 | 1 | 11.09 | 8.00 | −27.9% (−35.9, −19.1) |
Results, with scope
Live, each agent ran in Claude Code on τ²-bench's own prompts and tools, with the flow in stretto-proxy between the agent and the tools. The Claude models ran three trials of every task in each arm, with Claude Haiku 4.5 as the customer; GLM-5.3 ran one, as agent and customer, against its recorded baseline. The flows were learned from τ²-bench's published runs, not from these agents. Every other result is a replay.
Arrow keys step through; space plays or pauses; M turns the narration on or off. On a touch screen, swipe.