Definition
The AI gateway maturity model is a five-level framework for how a team's control over AI traffic evolves: Level 0 ungoverned, Level 1 captured, Level 2 governed, Level 3 routed, and Level 4 adaptive. Each level adds a capability the previous one lacked, and each is reached when a specific trigger makes the previous level insufficient.
A platform lead reads about semantic routing on Monday and cannot report last month’s AI spend by team on Tuesday. Both feel urgent. Neither tells them what to build next.
The AI gateway maturity model is a way to answer that. Few teams adopt a full AI gateway in one step. They climb a ladder: from ungoverned key sprawl, to capturing traffic, to governing it, to routing it intelligently, to an adaptive loop that learns from its own sessions. Each level adds a capability the previous one lacked, and each is reached when a specific trigger makes the previous level insufficient.
The model is useful for two things: locating where you are honestly, and naming the single next thing worth building. It is a map rather than a mandate. The right level for a team is the one its workload, spend, and constraints justify, which for many teams is somewhere short of the top.
Why a maturity model helps here
The AI gateway space is noisy, and it is easy to feel behind because someone is talking about semantic routing while you still cannot report spend by team. A maturity model cuts through that by ordering the capabilities: you cannot govern what you have not captured, and you cannot route well what you cannot evaluate. Knowing the order tells you that the team without cost attribution should build that before it worries about an adaptive routing loop. Progress is sequential, and the model makes the sequence explicit.
You cannot govern what you have not captured, and you cannot route well what you cannot evaluate. The maturity model is that ordering made explicit.
The five levels
The levels build on each other. Select one in the figure to see what a team at that level has, what it is still missing, and the trigger that pushes it up.
Pick a level above to see what a team at that level has, what it is missing, and the trigger that moves it up.
| Level | What defines it | The trigger to advance |
|---|---|---|
| 0 · Ungoverned | Personal keys, per-tool dashboards, no central record | The first spend or security question no one can answer |
| 1 · Captured | A proxy records traffic to a durable archive | A second team onboards and the flat log is not enough |
| 2 · Governed | Auth, per-team policy, model allowlists, cost attribution | The bill is dominated by requests a cheaper model would serve |
| 3 · Routed | Model routing to cost and quality targets, with evals | A mixed workload that static or per-prompt routing mis-serves |
| 4 · Adaptive | Routing learned from session data, drift caught by evals | Maintained as models, prices, and workloads move |
What each level adds
Level 0, Ungoverned. AI works and no one can see it. Engineers use personal keys and vendor dashboards; there is no central record, policy, or cost view. The level ends when finance or security asks a question the org cannot answer from one place.
Level 1, Captured. A capture proxy records the provider traffic it routes into a durable archive. For the first time there is a record of what was sent. What is missing is everything built on that record: per-team policy, a surface non-engineers can read, budget enforcement, routing. Level 1 ends when a second team wants in and the flat log stops scaling.
Level 2, Governed. A gateway adds authentication, per-team policy, control over which providers and models are reachable, and cost attribution by team and model. Much of the governance value lives here, and a team can reasonably stop here. What is still missing is intelligent routing: every request hits a statically chosen model. Level 2 ends when the frontier bill is visibly dominated by requests a cheaper model would answer.
Level 3, Routed. The gateway adds model routing to cost and quality targets, validated with evals on real traffic. The aim is spend that drops without quality slipping, and the evals are what show whether it did. What is missing is adaptivity: the routing is blind to request meaning and workload shape, and it is set once. Level 3 ends when a mixed or agentic workload keeps getting mis-served by static and per-prompt routing.
Level 4, Adaptive. Routing becomes semantic and workload-aware, learned from the team’s own session data, with drift caught by re-run evals. This is a maintained state rather than a finish line: the loop is kept current as models, prices, and workloads move.
L0 Ungoverned ─► L1 Captured ─► L2 Governed ─► L3 Routed ─► L4 Adaptive key sprawl a record policy+cost routing+evals learned loop each rung is reached when the rung below stops being enough
Example: placing a team on the ladder
A platform lead reads the model and places their org honestly. They capture Claude Code sessions (past Level 0) and can report rough spend, but policy is per-tool and there is no central control over which providers a tool can reach, so cost attribution is partial. They are between Level 1 and Level 2, having built capture but not governance. The model tells them the next thing to build is per-team policy and proper cost attribution rather than the semantic routing they had been reading about. Naming the level turned a vague sense of being behind into one concrete next project.
Concepts related to the maturity model
- AI gateway: the capability the model matures.
- AI gateway build vs buy: how to obtain the next level, sized by the level you target.
- Enterprise AI gateway: the governance depth of Levels 2 and up at organizational scale.
- Model routing and model routing evals: the capabilities that define Levels 3 and 4.
Failure modes of using a maturity model
- Skipping levels. Reaching for Level 3 routing before Level 1 capture means routing against traffic you cannot see or evaluate. The order is load-bearing.
- Treating the top as the goal. Building an adaptive loop for a uniform, low-stakes workload is over-engineering. The right level is the one the workload justifies.
- Mistaking tools for levels. Buying a routing product does not make you Level 3 if you cannot evaluate the routing. A level is a capability you actually have rather than a feature you purchased.
- Reading a single level as a static state. Expect to be between two levels, and Level 4 is maintained rather than reached. The model describes motion rather than a resting place.
When to use the maturity model
Reach for it when you need to locate yourself and plan the next step:
- You are unsure whether the next investment should be governance, routing, or something earlier.
- You feel behind on AI infrastructure and want an honest, ordered picture instead of a feature race.
- You are making a build-vs-buy decision and need to size it to a target level.
The model is most useful as a planning instrument: assess the current level, identify the single trigger you are hitting, and build the one capability that resolves it.