Skip to content

Paper Compute Concept

Model Routing Drift

A model router is tuned against the models, prices, and traffic of one moment. All three move. Model routing drift is what happens when the policy stays fixed while the world it was tuned for changes, degrading quality or savings with no error to announce it.

Published October 6, 2026
Model RoutingDriftEvalsAI Gateway

Definition

Model routing drift is the degradation of a routing policy over time as the models, prices, providers, and workload it was tuned against change. The policy keeps applying its original rules to a world that has moved, so its quality or its savings slip silently, with no error to signal the decline.

A team tunes a model router in March. It holds 95% of the strong model’s quality at 40% of the cost, the eval says so, and everyone moves on. In June a provider ships a new version of the cheap model. In July the bill is lower than ever. In August the support queue has a new kind of complaint, and nobody thinks to look at the router, because the router has not changed.

That is model routing drift: a routing policy that keeps applying its original rules to a world that has moved. A router is validated at a moment: these models, these prices, this traffic. All of those move, and when they do, the policy’s quality or its savings slip. Because nothing errors, the slip is silent.

The concept is the routing sibling of skill drift, where a skill’s procedure falls behind the system it describes. Same shape, different artifact: a decision encoded against one version of reality, still running after reality moved. Naming it is what makes it something you monitor rather than something you discover from a complaint.

Why routing drifts

A router is a tuned tradeoff, and every input to that tradeoff moves over time. The strong and weak models get new versions with different capability. Providers change token prices, which moves the cost-optimal cutline even if nothing else changes. Endpoints degrade and fallbacks fire more often. And the workload itself shifts as your product and users change, altering the share of requests a cheap model can handle. The router was correct for the inputs it saw. It cannot be correct for inputs it never saw and was never re-tuned against.

A router is a tuned tradeoff, and every input to that tradeoff moves. Drift is the gap between the world it was tuned for and the world it now runs in.

The four kinds of routing drift

What moves underneath a routing policy
Model driftA new version of the strong or weak model changes its capability, so the same request routes to a different-behaving model.
Price driftA provider changes token prices, moving the cost-optimal cutline even when capability is unchanged.
Provider driftAn endpoint degrades or a fallback fires more often, changing latency and reliability the policy assumed.
Workload driftThe mix of easy and hard requests shifts, changing the share a cheaper model can safely absorb.

Why drift is invisible without evals

The reason drift is dangerous is the same asymmetry that makes routing evals necessary in the first place: cost is loud and quality is silent. A drifted router still produces a bill, and the bill often looks fine. The quality it traded away produces no error, no alert, no log line. Without evals, the first signal is usually a slow rise in bad answers that surfaces as user complaints, weeks after the model version or price change that caused it, with nothing obvious connecting the two.

How drift hides
a change lands (new model version / price / traffic shift)
    │
    ▼
router keeps applying the old policy
    │
    ├── cost: still visible, often looks fine or better
    └── quality: slips silently, no error emitted
    │
    ▼
weeks later: user complaints, cause not obvious
    │
    ▼
only a re-run eval on fresh traffic connects them

Example: a price change that inverted the cutline

A team runs a cost-aware router tuned so that the cost-optimal cutline sends 60% of traffic to the weak model. A provider cuts the price of its strong model by half. Nothing in the router changes, and the bill drops, so it looks like good news. But the cheaper strong model has also shifted the cost-optimal point: now more traffic should go to the strong model, because it is nearly as cheap and clearly better. The router keeps sending 60% to the weak model out of habit, trading quality it no longer needs to trade. Only a re-run eval, comparing quality and cost at the new prices, reveals that the router is now leaving quality on the table for savings that no longer matter.

What routing drift is not

  • Not a bug. The router is working as configured. The configuration is stale. That distinction matters because the fix is re-tuning rather than debugging.
  • Not always a cost problem. Drift can improve cost while quietly degrading quality, which is why watching cost alone misses it.
  • Not self-correcting. A fixed policy does not adapt on its own. Catching drift requires re-measurement; fixing it requires re-tuning.

Failure modes around routing drift

  • No re-evaluation cadence. A router validated once and trusted forever is drift waiting to happen. The absence of a re-run schedule is the root failure.
  • Watching cost without quality. Monitoring only the bill misses the axis drift degrades silently.
  • Re-tuning on stale data. Re-tuning a drifted router against last year’s captured traffic re-introduces the drift it was meant to fix. Re-tuning needs recent sessions.
  • Treating a model upgrade as risk-free. Adopting a new model version without re-running the routing eval assumes the router still holds, which is exactly the assumption drift breaks.

When to watch for routing drift

Watch for it whenever a router is in production and any of its inputs can move, which is always:

  • On a regular cadence, so slow workload drift is caught before it compounds.
  • On every model version change from either the strong or weak side.
  • On every provider price change, which silently moves the cost-optimal cutline.
  • After any noticeable shift in what your users or agents are asking for.

A router with no drift watch is an unmeasured decision. It was correct once and is now running on trust.

What's next

Frequently asked questions

What is model routing drift?+
It is a routing policy quietly getting worse because the things it was tuned against changed. A router set up to hold 95% quality at a given cost was tuned for specific models, prices, and traffic. When a provider ships a new model version, changes a price, or your traffic mix shifts, the same rules now produce different results. Nothing errors; the router just drifts away from the tradeoff you signed off on.
What causes routing to drift?+
Four things, usually. Model drift: a new version of the strong or weak model changes its capability. Price drift: providers change token prices, moving the cost-optimal cutline. Provider drift: an endpoint degrades or a fallback fires more often. And workload drift: the mix of easy and hard requests you serve changes, so the share a cheap model can handle changes with it. Any of these moves the right routing decision while the policy stays put.
Why is routing drift hard to notice?+
Because it is silent on the axis you watch. Cost stays visible and often looks fine or better, while quality (the thing that actually degraded) produces no error when it slips. A drifted router ships slightly worse answers at a slightly different cost, and unless you re-measure quality on current traffic, the first signal is usually a slow rise in user complaints no one connects to the router.
How do you catch and fix routing drift?+
By re-running the routing evals on fresh traffic, on a schedule and after any model or price change. The eval that validated the router originally, re-run on recent captured requests, shows the quality or savings moving. The fix is to re-tune the cost knob or re-learn the policy from recent session data, which is the same loop used to build the router, run again.

Where to go next