Definition
Using agent session data for model routing means building and evaluating a routing policy from a team's own captured sessions rather than only from public benchmarks or generic preference data. Captured requests, responses, and outcomes reveal which requests in the team's real workload needed the strong model, which is the signal a router needs to route that workload well.
A team adopts an off-the-shelf model router: a small piece of software that reads each request and decides whether a cheap model or an expensive one should answer it. The router’s published numbers are strong, quality retained at 95% on a public benchmark. In production it feels worse. Support tickets tick up, and nobody can say why.
The answer is sitting in the team’s own records. Replaying the router against a month of the team’s captured agent sessions shows the problem: their traffic is heavy on a class of multi-step reasoning the benchmark barely contains, and the router keeps sending that class to the weak model. Retuning the router’s cutoff on those same sessions restores quality on the class that mattered, at a small cost in savings. The benchmark was not wrong. It measured someone else’s workload.
That is the practice this page teaches: fit and check the routing policy against the traffic the team actually serves, using the sessions it has already captured, instead of trusting a public stand-in.
Why your own traffic beats a public benchmark for model routing
A router’s whole job is telling apart the requests that need the strong model from the ones a cheaper model can handle. To draw that line well, it needs evidence about the requests it will actually see.
A benchmark like MT Bench or MMLU is a fixed, public collection of test prompts chosen to be broadly representative. Your production traffic is a specific distribution shaped by your product, your users, and your agents. A router tuned to the benchmark inherits the benchmark’s distribution, so it draws the cut between weak and strong models where the benchmark says to draw it rather than where your traffic does.
A router’s job is to tell requests that need the strong model apart from those that do not. The best evidence for that is your traffic rather than a benchmark that stands in for it.
Captured sessions are that evidence. When a team’s agent traffic flows through a gateway (a company-controlled service that requests pass through on their way to the model provider), the gateway can record each exchange: what was sent, what came back, and how the session went. Those records carry two things a benchmark structurally cannot: the workload context each request sat in (the input workload-aware routing needs), and the outcome of the request, which is the ground truth a router should learn from and be evaluated against.
What session capture records, and what it does not
One boundary matters before building anything on this data. Session capture records provider traffic: the model requests, the responses, and the tool calls and errors that pass through that traffic. Work that stays on the local machine and never reaches the provider is not in the record. An agent’s message saying it will run a command crosses the provider connection and gets captured; the command’s actual execution and its side effects do not. A routing policy fit to captured sessions inherits exactly this boundary: it learns from what the provider connection carried.
What agent session data adds to model routing
| Real distribution | The actual mix of easy and hard requests you serve, rather than a public average. |
|---|---|
| Workload context | The task, tools, and session phase each request belonged to. |
| Outcome signals | Session status and the recorded exchange, the raw material a judge turns into a quality label. |
| A replay corpus | Stored requests to replay a candidate router against, offline and repeatably. |
| Drift signal | Fresh sessions that show when the workload has moved away from the router's assumptions. |
The judge in that third row is a model that reads a recorded exchange and scores whether the answer worked. Capture supplies the sessions rather than the labels; turning captured traffic into a fitted policy or a quality score still means running a candidate router and a judge over it.
The mechanics reuse the rest of this cluster. Session data is the corpus that routing evals replay against (replay means running a candidate router over stored requests offline, without touching production), and it is the source of the contextual signals workload-aware routing reads. Session-data routing is less a separate technique than the practice of grounding the whole routing loop in your own captured traffic.
captured sessions (requests, responses, session status, context)
│
├──► fit / tune the routing policy on real distribution
│
├──► replay candidate routers offline, score quality kept
│
└──► re-check on fresh sessions → detect drift → re-tune
│
▼
route production traffic, capture moreThe last arrow closes the loop. Workloads move, and a policy fit to last quarter’s sessions can quietly stop matching this quarter’s traffic; that slow mismatch is called drift. Fresh captures are how you notice it, and re-tuning on them is how you correct it.
What session data is not
- Not a model. Session data does not improve the models. It improves the decision about which model to use, and the confidence that the decision is right for your traffic.
- Not a substitute for evaluation. Learning a policy from your data and evaluating it are different steps. The data feeds both, but a policy still has to clear routing evals on held-out sessions (sessions set aside during fitting so the check is honest).
- Not automatically complete. Capture has a boundary. What was not captured is not in the signal, and the router inherits that gap.
Concepts related to session-data routing
- Model routing: the parent concept.
- Workload-aware model routing: the strategy this data most directly feeds.
- Model routing evals: the evaluation session data makes real rather than generic.
- AI session capture: how the session data is produced.
- Agent session replay: stepping through the captured sessions the router learns from.
Failure modes of session-data routing
- Too little data. A policy fit to a handful of sessions overfits to noise. Representativeness needs volume.
- No outcome signal. Without knowing whether an answer worked, the data shows what was asked. It cannot show which requests needed the strong model, and that is the signal a router learns from.
- Capture bias. If capture missed a class of traffic, the router is blind to it. The router is only as representative as the capture behind it.
- Learning once. A policy fit to last quarter’s sessions and never refreshed drifts as the workload moves, the same drift evals exist to catch.
When to route from session data
Reach for it when a generic router is underperforming on your traffic and you have the capture to do better:
- You have enough captured sessions to be representative of your real workload.
- You have a usable outcome signal, so the data teaches what actually needed the strong model.
- Your workload differs from public benchmarks enough that a generic router misroutes a class you care about.
If you have no capture yet, start by capturing sessions, so a router has something honest to learn from.
Research on data-driven model routing
The public routing frameworks make the case for data-driven routing without having your data. RouteLLM (LMSYS and Anyscale, 2024; published at ICLR 2025 as arXiv 2406.18665) trains routers on preference data and notes that its routers generalize across model pairs, which is useful precisely because it lets you swap in your own models. The step this concept adds is swapping in your own distribution and outcomes as well, which a public framework cannot supply and only your capture can.