Definition
Model routing drift is the degradation of a routing policy over time as the models, prices, providers, and workload it was tuned against change. The policy keeps applying its original rules to a world that has moved, so its quality or its savings slip silently, with no error to signal the decline.
A team tunes a model router in March. It holds 95% of the strong model’s quality at 40% of the cost, the eval says so, and everyone moves on. In June a provider ships a new version of the cheap model. In July the bill is lower than ever. In August the support queue has a new kind of complaint, and nobody thinks to look at the router, because the router has not changed.
That is model routing drift: a routing policy that keeps applying its original rules to a world that has moved. A router is validated at a moment: these models, these prices, this traffic. All of those move, and when they do, the policy’s quality or its savings slip. Because nothing errors, the slip is silent.
The concept is the routing sibling of skill drift, where a skill’s procedure falls behind the system it describes. Same shape, different artifact: a decision encoded against one version of reality, still running after reality moved. Naming it is what makes it something you monitor rather than something you discover from a complaint.
Why routing drifts
A router is a tuned tradeoff, and every input to that tradeoff moves over time. The strong and weak models get new versions with different capability. Providers change token prices, which moves the cost-optimal cutline even if nothing else changes. Endpoints degrade and fallbacks fire more often. And the workload itself shifts as your product and users change, altering the share of requests a cheap model can handle. The router was correct for the inputs it saw. It cannot be correct for inputs it never saw and was never re-tuned against.
A router is a tuned tradeoff, and every input to that tradeoff moves. Drift is the gap between the world it was tuned for and the world it now runs in.
The four kinds of routing drift
| Model drift | A new version of the strong or weak model changes its capability, so the same request routes to a different-behaving model. |
|---|---|
| Price drift | A provider changes token prices, moving the cost-optimal cutline even when capability is unchanged. |
| Provider drift | An endpoint degrades or a fallback fires more often, changing latency and reliability the policy assumed. |
| Workload drift | The mix of easy and hard requests shifts, changing the share a cheaper model can safely absorb. |
Why drift is invisible without evals
The reason drift is dangerous is the same asymmetry that makes routing evals necessary in the first place: cost is loud and quality is silent. A drifted router still produces a bill, and the bill often looks fine. The quality it traded away produces no error, no alert, no log line. Without evals, the first signal is usually a slow rise in bad answers that surfaces as user complaints, weeks after the model version or price change that caused it, with nothing obvious connecting the two.
a change lands (new model version / price / traffic shift)
│
▼
router keeps applying the old policy
│
├── cost: still visible, often looks fine or better
└── quality: slips silently, no error emitted
│
▼
weeks later: user complaints, cause not obvious
│
▼
only a re-run eval on fresh traffic connects themExample: a price change that inverted the cutline
A team runs a cost-aware router tuned so that the cost-optimal cutline sends 60% of traffic to the weak model. A provider cuts the price of its strong model by half. Nothing in the router changes, and the bill drops, so it looks like good news. But the cheaper strong model has also shifted the cost-optimal point: now more traffic should go to the strong model, because it is nearly as cheap and clearly better. The router keeps sending 60% to the weak model out of habit, trading quality it no longer needs to trade. Only a re-run eval, comparing quality and cost at the new prices, reveals that the router is now leaving quality on the table for savings that no longer matter.
What routing drift is not
- Not a bug. The router is working as configured. The configuration is stale. That distinction matters because the fix is re-tuning rather than debugging.
- Not always a cost problem. Drift can improve cost while quietly degrading quality, which is why watching cost alone misses it.
- Not self-correcting. A fixed policy does not adapt on its own. Catching drift requires re-measurement; fixing it requires re-tuning.
Concepts related to model routing drift
- Model routing evals: the instrument that turns silent drift into a visible number.
- Model routing: the parent concept that drifts.
- Using agent session data for model routing: the fresh traffic drift is re-measured and re-tuned against.
- Skill drift: the same phenomenon for skills, and the closest conceptual sibling.
Failure modes around routing drift
- No re-evaluation cadence. A router validated once and trusted forever is drift waiting to happen. The absence of a re-run schedule is the root failure.
- Watching cost without quality. Monitoring only the bill misses the axis drift degrades silently.
- Re-tuning on stale data. Re-tuning a drifted router against last year’s captured traffic re-introduces the drift it was meant to fix. Re-tuning needs recent sessions.
- Treating a model upgrade as risk-free. Adopting a new model version without re-running the routing eval assumes the router still holds, which is exactly the assumption drift breaks.
When to watch for routing drift
Watch for it whenever a router is in production and any of its inputs can move, which is always:
- On a regular cadence, so slow workload drift is caught before it compounds.
- On every model version change from either the strong or weak side.
- On every provider price change, which silently moves the cost-optimal cutline.
- After any noticeable shift in what your users or agents are asking for.
A router with no drift watch is an unmeasured decision. It was correct once and is now running on trust.