Definition
Model routing evals are the measurements that establish whether a router meets its objective without breaking quality: how much of the strong model's quality it retains, how much cost it saves, what fraction of requests still reach the strong model, and how often it misroutes. They are run against a baseline and, ideally, against the team's own traffic.
A team turns on a model router and the bill drops 40% in a week. The dashboard is green. Everyone is pleased. A month later a user complains about a wrong answer, and nobody connects it to the router, because a misrouted request produces a worse answer without producing an error.
That is the problem routing evals exist to solve. A router’s most visible output is a cost saving, and its most important output is the quality it kept. Cost savings alone are trivial to produce: route everything to the cheapest model and you save the most. The number that makes a saving meaningful is the quality number beside it, and the only way to get that number is to measure it. Routing evals are that measurement. They are what separates a router you can trust from one that is quietly shipping worse answers.
Why routing without evals is a trap
The failure is asymmetric. When a router saves money, the saving shows up immediately on the bill. When a router degrades an answer, nothing errors, the cost still drops, and the damage surfaces later as a user complaint or a wrong action nobody traced back to the router. Cost is loud and quality loss is silent, so a router optimized on cost alone drifts toward worse answers with positive-looking metrics.
Cost savings are loud. Quality loss is silent. Evals are how you hear the second one.
How to evaluate a model router
Start with the question the eval answers: compared to sending everything to the strong model, how much quality did the router give up, and how much cost did it save? That comparison point (always use the strong model) is the ceiling. The floor is a comparison against something dumb, such as routing requests at random at the same cost. Any comparison point like these is called a baseline, and a router only means something relative to one.
Four numbers do most of the work:
| Quality retained | How close the routed system gets to always using the strong model, on your quality measure. |
|---|---|
| Strong-model call fraction | The share of requests still sent to the expensive model. The cost proxy. |
| Misroute rate | How often a request that needed the strong model went to the weak one. |
| Cost per correct outcome | Total cost to a right answer, including retries and escalations, rather than cost per call. |
The right way to report a router is as a point on a cost-quality curve against a baseline. RouteLLM is explicitly a framework for both serving and evaluating routers, and it reports its routers this way: quality retained at a given fraction of strong-model calls, compared against a random-routing baseline at the same cost. Their matrix-factorization router reached 95% of GPT-4 quality using 26% of GPT-4 calls on MT Bench, which is a paired claim (a quality number and a cost number together) rather than a bare saving.
quality retained
100% ┤ ● always strong (baseline ceiling)
│ ● good router
│ ●
│ ● random routing
50% ┤ ●
└────────────────────────────► strong-model call fraction
low cost high cost
A good router sits above the random line: more quality
at the same cost, or the same quality for less.Offline replay versus online evaluation
There are two ways to run these measurements, and teams use both. Offline replay takes requests you already recorded, runs them through the candidate router, and compares the routed answers to what the strong model would have produced. Nothing touches a user. Comparing answers needs a quality signal: either an outcome you stored with the original request, or a model that scores answers, called a judge. Online evaluation routes a small slice of live traffic and measures real outcomes.
| Offline replay | Online evaluation | |
|---|---|---|
| What it does | Replays captured requests through the router | Routes a slice of live traffic and measures outcomes |
| Strength | Fast, cheap, repeatable, no user risk | Real conditions, real outcomes |
| Weakness | Needs a stored quality signal or a judge | Slower, and a bad router touches real users |
| Best for | Tuning the knob and catching regressions | Confirming the offline result holds in production |
Offline replay is where a router is tuned, because it can be run repeatedly against the same captured requests with no cost to users. Online evaluation confirms the offline result under real conditions. A router should clear the offline bar before it sees a slice of live traffic.
Example: catching a regression a new model introduced
A team runs a router validated at 95% quality retained on their own replayed traffic. A provider ships a new version of the weak model. The team re-runs the offline eval on the same captured requests with the new weak model in place, and quality retained drops to 88%. Nothing broke, no error fired, and the bill actually improved because the new weak model is cheaper. Only the eval caught that the router was now trading more quality than the team had agreed to. They re-tune the knob and restore the floor.
What routing evals are not
- Not a one-time gate. A router validated once is validated for one moment. Models, prices, and traffic move, so the eval has to be re-run to stay meaningful.
- Not a public benchmark score. A benchmark is a sanity check. The decision to trust a router in production is made on your own traffic.
- Not cost measurement alone. Cost without a paired quality number is the exact trap evals exist to close.
Concepts related to routing evals
- Model routing: the parent concept and the savings evidence.
- Cost-aware model routing: the objective the eval holds to a floor.
- Using agent session data for model routing: the captured traffic offline replay runs on.
- Workload-aware model routing: a richer router whose extra signal has to earn its place in the eval.
Failure modes of routing evaluation
- Measuring cost without quality. The original trap: a saving with no paired quality number is not an evaluation.
- Evaluating on the wrong distribution. A benchmark or a stale sample that does not match current traffic produces a number that does not predict production.
- No quality signal to replay against. Offline replay needs a stored outcome or a judge. Without one, there is nothing to compare the routed answer to.
- Never re-running. A router evaluated once and trusted forever is the definition of undetected drift.
When you need routing evals
You need them the moment you route real traffic. Specifically:
- Before trusting any router in production, to know the quality you are keeping.
- On a schedule and on every model or price change, to catch drift.
- Whenever you add a signal (moving from cost-aware to workload-aware), to prove the added complexity earns its place.
A router with no eval is not a cost optimization. It is an unmeasured change to the quality of every answer it touches.
Evidence and research on routing evaluation
RouteLLM (LMSYS and Anyscale, ICLR 2025, arXiv 2406.18665) is as much an evaluation framework as a routing one: it defines routers as things to be measured against a baseline on paired cost and quality, and releases the datasets to do it. Its practice of reporting quality retained at a given strong-model call fraction is the template for evaluating any router.