Definition
Model routing is the practice of directing each request to a specific model chosen to meet a target such as cost, quality, or latency. A router decides, per request or per class of request, whether a cheaper or stronger model should serve it, and forwards the request accordingly.
A coding agent works through an afternoon. Most of its requests are small: rename this variable, write this docstring, fix this import. A few are genuinely hard: untangle a race condition, redesign an interface. Every one of those requests goes to the same frontier model, and every one costs frontier prices.
Model routing puts a decision in front of that. Before a request reaches a model, a router looks at it and asks one question: does this need the strong model, or would a cheaper one answer it just as well? Easy requests go to a small, cheap model. Hard ones go to the frontier model. The agent’s code does not change. The path its requests take does.
Teams have always picked models by hand for different jobs. Agentic workloads broke that habit, because an agent fires hundreds of requests an hour across a wide spread of difficulty, and one hand-picked model for all of them is wrong in one of two expensive directions. Send everything to a frontier model and you overpay. Send everything to a cheap one and you underperform. Routing is the attempt to be right per request instead of on average.
Why model routing matters
The economics moved model routing from a nice idea to a line item. Frontier models cost multiples of small ones per token, and a large fraction of real requests are easy enough that a small model answers them correctly. Paying frontier prices for those requests is pure waste, and at agent volumes the waste compounds fast.
Sending every request to a frontier model overpays. Sending every request to a cheap one underperforms. Routing is the attempt to be right per request.
The catch is that both the savings and the damage are invisible by default. Route well and the bill drops with no one noticing. Route badly and a hard request lands on a weak model: the answer degrades, and nothing errors. That asymmetry is why routing is worth doing carefully and worth evaluating.
Example: one workload, two policies
A support-triage agent handles a stream of tickets. Half are “how do I reset my password” and half are “reconcile these conflicting refund records.”
With no routing, every ticket hits a frontier model. The password resets cost the same as the refund reconciliations, and most of the bill is spent on requests a small model would have answered correctly.
With routing, a decision step reads each ticket first. The resets go to a small model. The reconciliations go to the frontier model. The bill drops sharply because the easy half moved to a cheap model, and quality holds because the hard half still gets the strong one. The entire gain came from telling the two halves apart. That is the router’s whole job.
How model routing works
A router sits on the request path and makes its decision before the expensive model runs. To decide, it needs a fact about the request it can act on: how much the request is allowed to cost, how difficult it looks, what it is about, or what similar requests needed in the past. That fact is the signal. A rule then maps the signal to a model. That rule is the policy, and it usually carries a knob that trades cost against quality. The models the policy chooses between are the destinations, and every router needs a plan for errors and uncertainty: the fallback.
| Signal | What the decision is based on: a cost threshold, a learned score, request meaning, workload shape, or past sessions. |
|---|---|
| Policy | The rule that maps a signal to a destination model, with a tunable cost/quality knob. |
| Destinations | The model pool: at minimum a strong and a weak model, often more. |
| Fallback | What happens on error or low confidence: retry, escalate to the stronger model, or fail over. |
The cleanest published result routes between just two models. RouteLLM, the open-source framework from LMSYS and Anyscale (ICLR 2025, arXiv 2406.18665), puts a small learned judge in front of a strong model and a weak one. The judge is trained on human preference data, so it predicts when the weak model’s answer would have been judged good enough. Its reported cost reductions vary by benchmark, which is exactly what you would expect: easy benchmarks route more traffic to the cheap model and save more.
Routing also splits into static and dynamic. Static routing maps a request to a model by fixed rules written down ahead of time: this model name, this team, this fallback list. Dynamic routing decides per request, from a signal about the request itself. The specialized strategies in this cluster are all forms of dynamic routing that differ in the signal they read.
incoming request
│
▼
router: read the signal, apply the policy
│
├── easy / low-risk → weak model (cheap, fast)
├── hard / high-stakes → strong model (costly, capable)
└── uncertain → escalate to the strong model
│
▼
forward, then record the decision for evaluationModel routing strategies in this cluster
Every strategy below is a different answer to the same question: what signal should the router read, and what target should it optimize?
- Cost-aware model routing: route to a spend target.
- Workload-aware model routing: route by the shape of the workload rather than only per-prompt difficulty.
- Frontier vs open-weight model routing: route between hosted frontier models and self-hostable open-weight ones for cost, privacy, and latency.
- Using agent session data for model routing: learn the routing policy from your own captured traffic.
- Model routing evals: prove a router holds quality before you trust it.
- Semantic routing vs model routing: the strategy-versus-target distinction underneath all of these.
What model routing is not
- Not load balancing. Load balancing spreads traffic across replicas of the same model. Routing chooses between different models. See the FAQ above for the full distinction.
- Not the same as semantic routing. Model routing names the target (a model). Semantic routing names one way to decide (by meaning). A router can pick a model from a cost threshold with no semantics at all.
- Not automatic quality. A router does not improve a model. At best it preserves the strong model’s quality on the requests that need it while saving cost on the rest.
- Not a one-time setup. The right routing policy drifts as models, prices, and workloads change. See model routing evals and the drift they catch.
Failure modes of model routing
- Silent quality loss. A misrouted hard request produces a worse answer with no error. Without evals, you find out from users.
- A router that costs more than it saves. If the routing decision itself is expensive (a large model judging every request), the overhead eats the savings. The decision has to be cheap relative to what it routes.
- Overfitting to a benchmark. A router tuned to MT Bench may route your production traffic badly, because your traffic is not MT Bench. Routing policy should be validated on your workload.
- Set-and-forget drift. New model versions and price changes move the right cutline. A router left alone slowly routes against a world that no longer exists.
When you need model routing
Reach for model routing when the workload is mixed and the bill is dominated by requests that did not need the model they got:
- Your spend is high and a meaningful share of requests are plausibly answerable by a cheaper model.
- Your workload mixes trivial and hard requests through one path, so a single model choice is wrong much of the time.
- You can measure quality on your own traffic, so you can route without flying blind.
If your requests are uniform in difficulty, or the cheapest model already meets your quality bar, routing adds complexity for little gain. The value comes from variance in the workload.
Evidence and research on model routing
RouteLLM is the most-cited open reference: a framework for both serving and evaluating routers, with reported cost reductions of over 85% on MT Bench, 45% on MMLU, and 35% on GSM8K while retaining roughly 95% of GPT-4’s performance, and a matrix-factorization router reaching 95% of GPT-4 quality using 26% of GPT-4 calls on MT Bench. The vLLM Semantic Router shows the routing decision moving onto serving infrastructure, classifying intent to pick a fast or reasoning path. Read together, they mark the two signals a router can use: a learned preference score and a semantic classifier.