Skip to content

Paper Compute Concept

LLM Proxy: What It Is and How It Works

An LLM proxy sits between an AI tool and a model provider. The tool sends its requests to the proxy; the proxy records them, applies rules, and forwards them on. This page covers the mechanism in ordinary terms first: where the proxy sits, what changes when you use one, what it can record and control, and what it cannot see.

Published September 15, 2026
LLM ProxyAI GatewayInfrastructureAI PlatformTelemetry

Definition

An LLM proxy is a service that sits between an AI tool and a model provider. Instead of sending a request directly to the provider, the tool sends it to the proxy; the proxy can record the request, apply rules, and forward it on, and the response comes back through the same path. In the AI gateway architecture used throughout this cluster, this request-handling layer is called the data plane.

An LLM proxy sits between an AI tool and a model provider. Instead of sending a request directly to OpenAI, Anthropic, or another provider, the tool sends it to the proxy. The proxy can inspect or record the request, apply rules, and then forward it to the provider. The response comes back through the same path.

The simplest view of an LLM proxy
without a proxy:   AI tool ─────────────────► model provider

with a proxy:      AI tool ──► LLM PROXY ──► model provider

Here is what that looks like in practice. Install a local proxy and it sets one environment variable, so your coding agent sends its Anthropic-compatible requests to localhost instead of to Anthropic. The proxy writes each request to an archive on your machine, then forwards it. The response streams back through, untouched. From the agent’s perspective, the API behaves the same way. You now have a searchable record of every model exchange that passed through the proxy.

In the AI gateway architecture used throughout this cluster, that request-handling layer is called the data plane. The rest of this page covers what the proxy does and what its data enables, how tools send traffic through it, what it can and cannot see, and what it takes to run one in production.

What an LLM proxy does on each request

What an LLM proxy does on each request
InterceptReceive the request the tool sent, in the provider's own API format.
RecordWrite prompt metadata, model, timestamp, and configured request fields to a durable archive.
PolicyApply allowed-models, redaction, rate-limit, routing, or cache-safety rules.
ForwardSend the request on to the real provider, keeping streaming intact.
Record responseCapture configured response fields, reassemble streams when needed, and return output to the tool.

What a proxy does directly, and what its data makes possible

Keep two lists separate: what the proxy itself does on the request path, and what other systems can later build from the data it captured.

The proxy itself, on the request path:

  • Captures requests and responses as they pass through.
  • Enforces policy in one place: allowed models, redaction, rate limits.
  • Routes requests to models and providers, with fallback and retries.
  • Holds provider credentials so individual tools do not have to.
  • Measures tokens and cost per request, when the provider reports them.
  • Feeds cross-tool telemetry from one source.

Built later, by reading the captured archive:

  • Session replay and debugging of past agent runs.
  • Historical analysis and recurring-pattern discovery.
  • Skill extraction and workflow improvement.
  • Evaluations and organizational learning.
  • Cache analysis: finding repeated prompt shapes and deciding where caching is safe. Read why AI workflows need new telemetry primitives for what cache topology reveals beyond a simple hit rate.

The proxy does not do skills or evaluations itself. It produces the record those systems read, which is why the quality of the archive matters more than any single feature on the first list.

How is an LLM proxy different from an AI gateway?

A proxy is one part of a gateway. The proxy handles model requests as they pass through. The gateway adds the systems around it: routing configuration, policies, credentials, budgets, administration, and observability. In the architecture described in this cluster, the proxy is the request-handling core of the gateway, and a bare proxy that only forwards requests is not yet a gateway.

The proxy and the gateway, side by side
LLM proxyAI gateway
Core jobIntercept, record, forwardRoute, govern, observe, control cost, secure
ScopeOne path, often one providerMany callers, many providers, one control point
RoutingUsually forwards to one destinationRoutes across providers and models
GovernanceCan apply simple rulesPer-team policy, egress, audit
Operated byA developer, per machineA platform team, for the org

The rows are a spectrum rather than a wall. A proxy can grow a rule, and in this cluster’s architecture a gateway still forwards each request through a proxy at its core. What moves a deployment from the left column to the right is the shift from a process on one machine to a system the organization operates.

From proxy to gateway
LLM PROXY
intercept ─► record ─► forward
      │
      │  add: multi-provider routing, per-team policy,
      │       central archive, cost accounting, a UI
      ▼
AI GATEWAY
authenticate ─► route ─► govern ─► forward ─► capture & meter

The arc is common enough to plan for. A single capturing proxy on one laptop proves the value in a week. Then a second team joins, and the flat log starts meeting questions it cannot answer: which team spent what, can we block one team from an expensive model, what happens when the provider goes down, how does a non-engineer read this. Each of those is a responsibility the bare proxy skipped. A bare proxy is enough while capture is the only requirement and one person reads the output. Move toward the gateway when a second team, a second provider, or a question from outside engineering shows up. The enterprise AI gateway page covers that transition as a build-vs-buy and rollout decision.

How does a tool send traffic through a proxy?

The mechanics depend on the tool, but the patterns are small in number.

Three ways a tool ends up talking to the proxy
MechanismHow it worksWhere it's used
Environment variableTool reads ANTHROPIC_API_URL or equivalent; proxy sets it to localhostAny tool that reads a base URL env var (e.g., Claude Code, Cursor, custom agents)
Config fileTool reads a config that points at the proxyIDE extensions, internal services
HTTP CONNECTOS-level proxy setting that the HTTP client honorsBrowsers, generic clients

When the proxy preserves streaming, error behavior, and the provider’s request format, the tool’s behavior does not change. It sends a request, gets a response, and streams output if it’s a streaming endpoint. The only difference is that the proxy was in the middle.

LLM proxy architecture diagram showing AI tools sending traffic through a proxy that captures requests, applies policy, and forwards to model providers (Anthropic, OpenAI, Bedrock, Vertex AI, self-hosted models)

How does a proxy avoid breaking the tool?

A correctly written LLM proxy preserves three properties:

  • Streaming. Forward chunks as they arrive. Latency should stay dominated by the provider’s generation time, with only a small amount of proxy overhead.
  • Faithful forwarding. By default, the provider receives the same request the tool sent. When the proxy is configured to change something (redaction, routing, header changes) those changes should be explicit and traceable.
  • Failure passthrough. By default, provider errors should reach the tool in the shape it expects. If the proxy introduces its own errors (policy blocks, rate limits, upstream failures) those should be easy to distinguish from provider failures.

Lose any of those and the proxy starts breaking the tools that depend on it. The discipline of a proxy is “be neutral by default, opinionated only when configured to be.”

What an LLM proxy can and cannot see

A proxy sees exactly what passes through it, and nothing else.

It can see:

  • Requests sent through it, and the responses that come back.
  • Which model and provider served each request.
  • Token usage and cost information, when the provider reports it.
  • Request metadata the tool sends: headers, model settings, tool-call messages.

It cannot automatically see:

  • Shell commands an agent executes locally.
  • Files written on the machine, or any side effects between model calls.
  • Browser activity that does not pass through the proxy.
  • Direct provider calls that bypass it.
  • Personal AI accounts outside the managed path.

This is why enterprise deployments pair the proxy with policy, device management, network rules, and an approved-tool rollout, and why the gateway reference architecture draws the visibility boundary explicitly.

How a capture-oriented LLM proxy works in practice

There are two deployment models for a capture-oriented proxy, and most teams end up using both.

Local proxy

A local proxy runs as a background process on the developer’s machine, listening on a localhost port. On install, it sets the relevant environment variables so the tool’s requests go to localhost first. The proxy writes each request and response to a structured archive and forwards the call to the configured provider.

Request lifecycle through a local capture proxy
tool                      proxy                        provider
|                            |                             |
| POST /v1/messages          |                             |
|--------------------------->|                             |
|                            | record request              |
|                            | apply policy (if any)       |
|                            | forward request             |
|                            |---------------------------->|
|                            |                             |
|                            |    streaming response       |
|                            |<----------------------------|
|                            | record response             |
|    streamed back            |                             |
|<---------------------------|                             |
|                            |                             |

The whole loop should add only a small amount of overhead. The session archive grows by one record per request.

Cloud proxy

A cloud proxy is a managed gateway that multiple developers or services share. Instead of each machine running its own proxy and storing records locally, the team provisions a centralized gateway (typically backed by a proxy like Envoy and a durable store like Postgres) and points their tools at it.

The request lifecycle is the same: intercept, record, apply policy, forward. The difference is where the proxy runs and where the records live. A cloud proxy gives the platform team a single capture archive for the whole org, model allowlisting at the edge, and centralized auth, without requiring anything on each developer’s machine beyond a pointer to the gateway URL.

Most teams start with one or the other and add the second when the need arrives. A local proxy gives individual developers capture and replay immediately. A cloud proxy gives the platform team visibility, policy, and cost attribution across teams. The two compose: a local daemon can authenticate to a cloud gateway, so the developer gets a localhost endpoint while the org gets a central archive.

How proxy capture compares with SDK and OpenTelemetry instrumentation

You don’t have to choose between a proxy, SDK instrumentation, and OpenTelemetry. They solve different parts of the problem, and they can work together.

Proxy interception vs SDK instrumentation vs OpenTelemetry tracing
LLM proxySDK instrumentationOpenTelemetry tracing
Where it livesOn the network pathInside each tool or appInside app code, exported to a collector
CoverageAny compatible tool pointed at the endpoint, with no code changesOnly tools whose code you can modifyOnly services you instrument
Payload fidelityFull request and response payloads: prompts, responses, streaming, tool-call messagesWhatever the SDK surfaces, at the SDK version you shipSpans and attributes; payloads commonly truncated, sampled, or elided
EnforcementCentralized: shared rules on one common request pathPossible, but implemented and maintained inside each toolObservation only
App contextSees only what crosses the wireRich: user ids, feature flags, app stateRich, plus correlation with the rest of the distributed trace
Maintenance surfaceProvider API formatsEvery tool and SDK version in the fleetInstrumentation libraries; GenAI semantic conventions are newer and still stabilizing
  • A proxy gives you broad coverage and centralized enforcement. One compatible proxy can cover many tools without adding instrumentation to each one. If a new tool can point at the same proxy endpoint and uses a supported protocol, it inherits the same capture and policy controls. And because the proxy sits on the shared request path, a platform team can enforce rules once, for every tool on that path at the same time.
  • SDK and OpenTelemetry instrumentation give you deeper context. This instrumentation lives inside your application, so it knows things the proxy can’t: which user clicked what, which feature triggered the model call, and where that call sits inside a larger application trace. SDK hooks can also block or modify requests, but that behavior has to be built and maintained inside each application, which is exactly the per-tool work a proxy avoids.
  • Use both when you need both. The proxy captures and governs the AI requests. OpenTelemetry captures what the application was doing around those requests. A shared request id connects the two records.
Joining proxy records and application traces on a request id
PROXY ARCHIVE                        APPLICATION TRACE (OpenTelemetry)
request id: req_8f3a                 request id: req_8f3a
prompt, response, model,             user id, feature flag,
tokens, tool-call messages           checkout-flow span, timings
      └────────────── join on req_8f3a ──────────────┘
 what the model exchange was    +    what the app was doing

A single captured proxy session can contain hundreds of messages across multiple model tiers. Capturing model switches and fallback events requires no SDK; the proxy sees the model field on every routed request.

Why provider compatibility matters

The tool thinks it is talking to the provider, so the proxy has to speak the provider’s language exactly. That language is the wire format: the exact shape of requests and responses on the network. Get it wrong and tools break in confusing ways; get it right and the proxy goes unnoticed. Five things define how compatible a proxy is:

  • Provider dialects. OpenAI-style chat-completions endpoints, Anthropic’s Messages API, Gemini, and the cloud-hosted variants (Bedrock, Vertex) each have their own request shape. A proxy either speaks each dialect it sits in front of, or says clearly which ones it supports.
  • The OpenAI-compatible ecosystem. Many self-hosted servers (vLLM, Ollama, and most gateway-friendly runtimes) expose OpenAI-compatible endpoints, so supporting that one dialect covers a long tail of backends.
  • Streaming semantics. Responses stream back in chunks, and the chunk format differs by provider. The proxy has to pass chunks through untouched while also reassembling them into one complete response for the archive.
  • Tool-call message shapes. Each provider encodes tool calls and tool results differently, and those are exactly the fields replay and audit care about. A proxy that drops or flattens them loses most of the archive’s value.
  • Auth translation and churn. The proxy usually gives callers its own credentials and keeps the real provider keys on its side; cloud providers add request signing on top. Providers also add new fields all the time, so the proxy has to pass through fields it does not recognize instead of rejecting them.

What an LLM proxy is not

A few common confusions worth resolving.

  • Not a model router by default. A proxy can route, but the simplest configuration forwards to one provider. Routing is a capability layered on top; the routing page covers the decision signals.
  • Not a model wrapper. The proxy doesn’t replace the model or modify the response content (unless explicitly configured to redact). It records and forwards.
  • Not a security boundary. The proxy sits on the path requests take out of the company; it can enforce policy but it doesn’t isolate tools from each other. Sandboxing is a separate problem (see stereOS for the sandboxing primitive).
  • Not the whole gateway. The comparison section above draws that line.

LLM proxy implementation resources

Frequently asked questions

What is the difference between an LLM proxy and an AI gateway?+
A proxy is one part of a gateway. The proxy handles model requests as they pass through: it records them, applies rules, and forwards them. The gateway adds the systems around it: routing configuration, policies, credentials, budgets, administration, and observability. In the architecture described in this cluster, the proxy is the request-handling core of the gateway; a bare proxy that only forwards requests is not yet a gateway.
Does an LLM proxy slow down inference?+
Not meaningfully. The dominant latency in any inference call is the model's generation time, which is unchanged. A local proxy adds a localhost-to-localhost hop and the time to write a record. A cloud proxy adds a network hop to the gateway, but the gateway forwards to the provider from a well-connected data center, so the total round trip is often comparable. Streaming responses pass through chunk-for-chunk in both models, so a well-built proxy is hard to notice from the tool's side.
Can an LLM proxy work with self-hosted models?+
Yes. The proxy cares about the request format rather than who serves it. A self-hosted Ollama or vLLM server speaking an OpenAI-compatible endpoint is captured the same way as the OpenAI cloud endpoint. The destination URL is configurable; the capture logic is unchanged.
How does an LLM proxy handle TLS?+
How the proxy handles TLS depends on deployment. In a local setup, the tool may talk to a localhost endpoint and the proxy opens the TLS connection to the provider. In managed or network-level deployments, the proxy may terminate TLS and establish a separate TLS session to the provider. The important point is that full request-body capture requires the proxy to see plaintext at some point in the request path; a pure pass-through TCP proxy can only log metadata.
Is a proxy required for every machine?+
Not necessarily. With a cloud proxy, a single managed gateway can serve the whole team: each developer points their tools at the gateway URL and the proxy captures centrally. With a local proxy, each machine runs its own instance. Many teams combine both: a local daemon authenticates to a cloud gateway, giving the developer a localhost endpoint while the org gets a central archive. The right model depends on the coverage and policy requirements.

Where to go next