Skip to content

Pillar Concept

Model Routing: Choosing the Right Model for Each Request

Model routing sends each request to the model that fits it: a cheaper model for the easy ones, a stronger model for the hard ones. Done well, it cuts cost and latency while holding quality. Done blindly, it degrades answers in ways that are hard to see.

Published September 9, 2026· Updated September 16, 2026
PillarModel RoutingAI GatewayCost OptimizationPillar

Definition

Model routing is the practice of directing each request to a specific model chosen to meet a target such as cost, quality, or latency. A router decides, per request or per class of request, whether a cheaper or stronger model should serve it, and forwards the request accordingly.

A coding agent works through an afternoon. Most of its requests are small: rename this variable, write this docstring, fix this import. A few are genuinely hard: untangle a race condition, redesign an interface. Every one of those requests goes to the same frontier model, and every one costs frontier prices.

Model routing puts a decision in front of that. Before a request reaches a model, a router looks at it and asks one question: does this need the strong model, or would a cheaper one answer it just as well? Easy requests go to a small, cheap model. Hard ones go to the frontier model. The agent’s code does not change. The path its requests take does.

Teams have always picked models by hand for different jobs. Agentic workloads broke that habit, because an agent fires hundreds of requests an hour across a wide spread of difficulty, and one hand-picked model for all of them is wrong in one of two expensive directions. Send everything to a frontier model and you overpay. Send everything to a cheap one and you underperform. Routing is the attempt to be right per request instead of on average.

Why model routing matters

The economics moved model routing from a nice idea to a line item. Frontier models cost multiples of small ones per token, and a large fraction of real requests are easy enough that a small model answers them correctly. Paying frontier prices for those requests is pure waste, and at agent volumes the waste compounds fast.

Sending every request to a frontier model overpays. Sending every request to a cheap one underperforms. Routing is the attempt to be right per request.

The catch is that both the savings and the damage are invisible by default. Route well and the bill drops with no one noticing. Route badly and a hard request lands on a weak model: the answer degrades, and nothing errors. That asymmetry is why routing is worth doing carefully and worth evaluating.

Example: one workload, two policies

A support-triage agent handles a stream of tickets. Half are “how do I reset my password” and half are “reconcile these conflicting refund records.”

With no routing, every ticket hits a frontier model. The password resets cost the same as the refund reconciliations, and most of the bill is spent on requests a small model would have answered correctly.

With routing, a decision step reads each ticket first. The resets go to a small model. The reconciliations go to the frontier model. The bill drops sharply because the easy half moved to a cheap model, and quality holds because the hard half still gets the strong one. The entire gain came from telling the two halves apart. That is the router’s whole job.

How model routing works

A router sits on the request path and makes its decision before the expensive model runs. To decide, it needs a fact about the request it can act on: how much the request is allowed to cost, how difficult it looks, what it is about, or what similar requests needed in the past. That fact is the signal. A rule then maps the signal to a model. That rule is the policy, and it usually carries a knob that trades cost against quality. The models the policy chooses between are the destinations, and every router needs a plan for errors and uncertainty: the fallback.

The parts of a model router
SignalWhat the decision is based on: a cost threshold, a learned score, request meaning, workload shape, or past sessions.
PolicyThe rule that maps a signal to a destination model, with a tunable cost/quality knob.
DestinationsThe model pool: at minimum a strong and a weak model, often more.
FallbackWhat happens on error or low confidence: retry, escalate to the stronger model, or fail over.

The cleanest published result routes between just two models. RouteLLM, the open-source framework from LMSYS and Anyscale (ICLR 2025, arXiv 2406.18665), puts a small learned judge in front of a strong model and a weak one. The judge is trained on human preference data, so it predicts when the weak model’s answer would have been judged good enough. Its reported cost reductions vary by benchmark, which is exactly what you would expect: easy benchmarks route more traffic to the cheap model and save more.

What model routing saved
Cost reduction at roughly 95% of the strong model's quality
over 85%
lower cost than always using GPT-4
on MT Bench
85% cost saved15% of baseline cost
The matrix-factorization router reached 95% of GPT-4 quality using 26% of GPT-4 calls. RouteLLM's reported figures; savings depend on the workload and the strong/weak model pair.
SOURCE LMSYS + Anyscale, RouteLLM (2024)LIVE — switch the benchmark; the savings and bar follow

Routing also splits into static and dynamic. Static routing maps a request to a model by fixed rules written down ahead of time: this model name, this team, this fallback list. Dynamic routing decides per request, from a signal about the request itself. The specialized strategies in this cluster are all forms of dynamic routing that differ in the signal they read.

A request through a model router
incoming request
    │
    ▼
router: read the signal, apply the policy
    │
    ├── easy / low-risk    → weak model (cheap, fast)
    ├── hard / high-stakes → strong model (costly, capable)
    └── uncertain          → escalate to the strong model
    │
    ▼
forward, then record the decision for evaluation

Model routing strategies in this cluster

Every strategy below is a different answer to the same question: what signal should the router read, and what target should it optimize?

What model routing is not

  • Not load balancing. Load balancing spreads traffic across replicas of the same model. Routing chooses between different models. See the FAQ above for the full distinction.
  • Not the same as semantic routing. Model routing names the target (a model). Semantic routing names one way to decide (by meaning). A router can pick a model from a cost threshold with no semantics at all.
  • Not automatic quality. A router does not improve a model. At best it preserves the strong model’s quality on the requests that need it while saving cost on the rest.
  • Not a one-time setup. The right routing policy drifts as models, prices, and workloads change. See model routing evals and the drift they catch.

Failure modes of model routing

  • Silent quality loss. A misrouted hard request produces a worse answer with no error. Without evals, you find out from users.
  • A router that costs more than it saves. If the routing decision itself is expensive (a large model judging every request), the overhead eats the savings. The decision has to be cheap relative to what it routes.
  • Overfitting to a benchmark. A router tuned to MT Bench may route your production traffic badly, because your traffic is not MT Bench. Routing policy should be validated on your workload.
  • Set-and-forget drift. New model versions and price changes move the right cutline. A router left alone slowly routes against a world that no longer exists.

When you need model routing

Reach for model routing when the workload is mixed and the bill is dominated by requests that did not need the model they got:

  • Your spend is high and a meaningful share of requests are plausibly answerable by a cheaper model.
  • Your workload mixes trivial and hard requests through one path, so a single model choice is wrong much of the time.
  • You can measure quality on your own traffic, so you can route without flying blind.

If your requests are uniform in difficulty, or the cheapest model already meets your quality bar, routing adds complexity for little gain. The value comes from variance in the workload.

Evidence and research on model routing

RouteLLM is the most-cited open reference: a framework for both serving and evaluating routers, with reported cost reductions of over 85% on MT Bench, 45% on MMLU, and 35% on GSM8K while retaining roughly 95% of GPT-4’s performance, and a matrix-factorization router reaching 95% of GPT-4 quality using 26% of GPT-4 calls on MT Bench. The vLLM Semantic Router shows the routing decision moving onto serving infrastructure, classifying intent to pick a fast or reasoning path. Read together, they mark the two signals a router can use: a learned preference score and a semantic classifier.

Model routing implementation resources

Frequently asked questions

What is model routing?+
Model routing is deciding which model should serve a given request instead of sending everything to one model. In practice a router looks at a request, judges whether a cheaper or weaker model can handle it, and forwards it to the strong model only when it needs to. The point is to stop paying frontier prices for requests a small model would answer just as well.
How is model routing different from load balancing?+
Load balancing spreads requests across interchangeable replicas of the same model to manage throughput. Model routing sends requests to different models chosen for their cost or capability. Load balancing treats the destinations as equivalent; model routing treats them as different tools for different requests.
Does model routing hurt quality?+
It can, and that is the risk to manage. A router that sends a hard request to a weak model degrades the answer silently, with no error to catch it. Good routing is conservative about uncertainty (route to the stronger model when unsure) and is validated with evals that measure quality retained rather than only cost saved. Reported frameworks hold roughly 95% of the strong model's quality at a large cost reduction, but that is a measured result on a benchmark rather than a default you get for free.
Where does model routing run?+
It usually runs at the AI gateway, on the request path, so every caller benefits without changing application code. It can also run inside an application or an SDK, but that scatters the routing policy across services. Running it at the gateway keeps one routing policy the platform team owns and can evaluate against real traffic.
What signals can a router use?+
A router can decide from a cost threshold, a learned score trained on preference data, the meaning of the request, the shape of the workload, or a team's own captured sessions. Those signals are the difference between the specialized strategies: cost-aware routing optimizes spend, workload-aware routing reads the shape of the work, and session-data routing learns from what actually needed the strong model in your traffic.

Where to go next