Skip to content

Paper Compute Concept

Model Routing Evals: Proving a Router Before You Trust It

A model router that saves money but quietly ships worse answers is a liability, and the damage is invisible without evaluation. Model routing evals are how you prove a router holds quality before you route real traffic through it, and how you catch it when it stops.

Published September 9, 2026· Updated September 16, 2026
Model RoutingEvalsAI GatewayMeasurement

Definition

Model routing evals are the measurements that establish whether a router meets its objective without breaking quality: how much of the strong model's quality it retains, how much cost it saves, what fraction of requests still reach the strong model, and how often it misroutes. They are run against a baseline and, ideally, against the team's own traffic.

A team turns on a model router and the bill drops 40% in a week. The dashboard is green. Everyone is pleased. A month later a user complains about a wrong answer, and nobody connects it to the router, because a misrouted request produces a worse answer without producing an error.

That is the problem routing evals exist to solve. A router’s most visible output is a cost saving, and its most important output is the quality it kept. Cost savings alone are trivial to produce: route everything to the cheapest model and you save the most. The number that makes a saving meaningful is the quality number beside it, and the only way to get that number is to measure it. Routing evals are that measurement. They are what separates a router you can trust from one that is quietly shipping worse answers.

Why routing without evals is a trap

The failure is asymmetric. When a router saves money, the saving shows up immediately on the bill. When a router degrades an answer, nothing errors, the cost still drops, and the damage surfaces later as a user complaint or a wrong action nobody traced back to the router. Cost is loud and quality loss is silent, so a router optimized on cost alone drifts toward worse answers with positive-looking metrics.

Cost savings are loud. Quality loss is silent. Evals are how you hear the second one.

How to evaluate a model router

Start with the question the eval answers: compared to sending everything to the strong model, how much quality did the router give up, and how much cost did it save? That comparison point (always use the strong model) is the ceiling. The floor is a comparison against something dumb, such as routing requests at random at the same cost. Any comparison point like these is called a baseline, and a router only means something relative to one.

Four numbers do most of the work:

The core routing metrics
Quality retainedHow close the routed system gets to always using the strong model, on your quality measure.
Strong-model call fractionThe share of requests still sent to the expensive model. The cost proxy.
Misroute rateHow often a request that needed the strong model went to the weak one.
Cost per correct outcomeTotal cost to a right answer, including retries and escalations, rather than cost per call.

The right way to report a router is as a point on a cost-quality curve against a baseline. RouteLLM is explicitly a framework for both serving and evaluating routers, and it reports its routers this way: quality retained at a given fraction of strong-model calls, compared against a random-routing baseline at the same cost. Their matrix-factorization router reached 95% of GPT-4 quality using 26% of GPT-4 calls on MT Bench, which is a paired claim (a quality number and a cost number together) rather than a bare saving.

Reading a router as a curve
quality retained
 100% ┤                      ● always strong (baseline ceiling)
      │                 ●  good router
      │            ●
      │       ●  random routing
  50% ┤  ●
      └────────────────────────────► strong-model call fraction
        low cost                 high cost

A good router sits above the random line: more quality
at the same cost, or the same quality for less.

Offline replay versus online evaluation

There are two ways to run these measurements, and teams use both. Offline replay takes requests you already recorded, runs them through the candidate router, and compares the routed answers to what the strong model would have produced. Nothing touches a user. Comparing answers needs a quality signal: either an outcome you stored with the original request, or a model that scores answers, called a judge. Online evaluation routes a small slice of live traffic and measures real outcomes.

Two ways to evaluate a router
Offline replayOnline evaluation
What it doesReplays captured requests through the routerRoutes a slice of live traffic and measures outcomes
StrengthFast, cheap, repeatable, no user riskReal conditions, real outcomes
WeaknessNeeds a stored quality signal or a judgeSlower, and a bad router touches real users
Best forTuning the knob and catching regressionsConfirming the offline result holds in production

Offline replay is where a router is tuned, because it can be run repeatedly against the same captured requests with no cost to users. Online evaluation confirms the offline result under real conditions. A router should clear the offline bar before it sees a slice of live traffic.

Example: catching a regression a new model introduced

A team runs a router validated at 95% quality retained on their own replayed traffic. A provider ships a new version of the weak model. The team re-runs the offline eval on the same captured requests with the new weak model in place, and quality retained drops to 88%. Nothing broke, no error fired, and the bill actually improved because the new weak model is cheaper. Only the eval caught that the router was now trading more quality than the team had agreed to. They re-tune the knob and restore the floor.

What routing evals are not

  • Not a one-time gate. A router validated once is validated for one moment. Models, prices, and traffic move, so the eval has to be re-run to stay meaningful.
  • Not a public benchmark score. A benchmark is a sanity check. The decision to trust a router in production is made on your own traffic.
  • Not cost measurement alone. Cost without a paired quality number is the exact trap evals exist to close.

Failure modes of routing evaluation

  • Measuring cost without quality. The original trap: a saving with no paired quality number is not an evaluation.
  • Evaluating on the wrong distribution. A benchmark or a stale sample that does not match current traffic produces a number that does not predict production.
  • No quality signal to replay against. Offline replay needs a stored outcome or a judge. Without one, there is nothing to compare the routed answer to.
  • Never re-running. A router evaluated once and trusted forever is the definition of undetected drift.

When you need routing evals

You need them the moment you route real traffic. Specifically:

  • Before trusting any router in production, to know the quality you are keeping.
  • On a schedule and on every model or price change, to catch drift.
  • Whenever you add a signal (moving from cost-aware to workload-aware), to prove the added complexity earns its place.

A router with no eval is not a cost optimization. It is an unmeasured change to the quality of every answer it touches.

Evidence and research on routing evaluation

RouteLLM (LMSYS and Anyscale, ICLR 2025, arXiv 2406.18665) is as much an evaluation framework as a routing one: it defines routers as things to be measured against a baseline on paired cost and quality, and releases the datasets to do it. Its practice of reporting quality retained at a given strong-model call fraction is the template for evaluating any router.

Routing evaluation resources

Frequently asked questions

What are model routing evals?+
They are the tests that tell you whether a router is doing its job: keeping quality high while sending less traffic to the expensive model. A routing eval reports how much of the strong model's quality the router retained, how much it cut cost or strong-model calls, and how often it sent a request to the wrong model. Without them, a router's cost savings are just a number with no idea what they cost in quality.
What metrics matter for a router?+
Quality retained versus always using the strong model, the fraction of requests routed to the strong model (the cost proxy), the misroute rate on requests that needed the strong model, and cost per correct outcome rather than cost per call. A good router is reported as a point on a curve: this much quality at this much cost, against a baseline like random routing at the same cost.
Can I evaluate a router on a public benchmark?+
You can, and it is a useful sanity check, but a public benchmark is not your workload. A router that retains 95% of quality on MT Bench may route your production traffic badly because your traffic is different. The evaluation that decides whether to trust a router in production runs on your own captured requests with your own quality measure.
How do routing evals catch drift?+
By being re-run. A router validated once is validated against the models, prices, and traffic of that moment. When a new model version ships or the workload shifts, the same eval re-run on fresh traffic shows the quality or savings moving, which is the signal to re-tune. Evals are the instrument that turns silent routing drift into a visible number.

Where to go next