Skip to content
← All posts
ProductMarch 4, 2026

Agents Need Black Box Recorders

Every commercial aircraft carries a black box. Not because crashes are common — modern aviation is extraordinarily safe — but because when something goes wrong at 35,000 feet, you need to know exactly what happened. No guessing. No reconstructing from memory. The data is there, or it isn’t.

Agent systems don’t have black boxes. They should.

A Dead Session

I was an hour deep into a Claude Code session, designing a Postgres integration for an agent memory system. Tool calls, schema decisions, trade-off discussions — all living in the context window. Then the session crashed. The context window closed and the entire conversation evaporated.

No write-ahead log. No checkpoint. No recovery path. An hour of design work, gone. Not because the model was bad, but because there was no durable record of what happened.

I got lucky. I’d been running tapes as a proxy between Claude and the API, and every request and response had been written to a local SQLite database. I pointed a fresh Claude session at the tapes DB, had it query the conversation history, and reconstructed the full context — design decisions, tool calls, exactly where I’d left off. The integration shipped to production that night. The full story is here.

But this isn’t a recovery anecdote. It’s a systems design problem that distributed computing solved thirty years ago.

Distributed Systems Already Solved This

Every production database writes a Write-Ahead Log (WAL) before committing a transaction. Every message queue checkpoints consumer offsets. Every event-sourced system can replay its entire history from the log. These aren’t optional features — they’re foundational primitives that exist because distributed systems assume failure.

Agent systems don’t assume failure. They assume the session will complete, the context window will hold, and the provider will stay up. When any of those assumptions break, you lose everything.

The primitives map directly:

Primitive What It Means for Agents
WAL Every tool call, prompt, and response logged before execution. If the session dies, the log survives.
Checkpointing Periodic snapshots of agent state — context, decisions made, progress markers. Resume from the last good checkpoint instead of starting over.
Replay Reconstruct a session from its log, exactly as it happened. Debug failures by replaying the exact sequence of events.
Event Sourcing The log is the source of truth. Derive any view of the session from the immutable event stream.

None of this is novel computer science. It’s infrastructure engineering that agent systems haven’t adopted yet.

What the Black Box Records

A useful agent telemetry system needs to capture every interaction, verbatim, in a record that outlives the session that produced it. In tapes, that record is an append-only log of raw turns: for every request/response pair that crosses the proxy, tapes stores the provider’s response byte-for-byte alongside the reduced, structured turn — role, content, tool calls, token usage, model, timing. Sessions, traces, and spans are derived views, rebuilt from the raw log rather than edited in place.

Three properties of that log matter:

  • 1.Fidelity — the verbatim bytes are the record. tapes can re-reduce the stored bytes and check the result against the stored turn (tapes raw equivalence), so the structured view is provably faithful to what the provider actually sent.
  • 2.Append-only durability — raw turns are never edited in place. Every derived view — sessions, traces, spans — rebuilds from the log, so re-deriving can’t rewrite history.
  • 3.Structure — every session breaks down into traces (turns) and spans (LLM calls, tool activity, subagent work), with links that preserve causality: which tool call spawned which subagent, and what context each step saw.

Two capture lanes feed this log. The proxy records the wire — every request/response pair between agent and provider. For harnesses that keep a local transcript, like Claude Code, tapes also ingests the transcript, which carries subagent structure the wire alone can’t show. Either lane works on its own; together they make the record complete.

From Recovery to Self-Healing

Recording telemetry is the starting point. The real value compounds as you move up the stack:

  • 1.Recovery — replay a crashed session from the log. Reconstruct context, resume where you left off. This is the baseline.
  • 2.Diagnosis — query across sessions to find failure patterns. Which tool calls fail most often? Where do agents get stuck in loops? What prompts produce hallucinations?
  • 3.Prevention — detect anomalies in real-time. Token consumption spikes, repeated tool call failures, context window degradation. Flag these before the session fails.
  • 4.Self-healing — agents that checkpoint themselves, detect when they’re drifting, and recover automatically. Not a human replaying a log — the agent doing it for itself.

This is the trajectory. Black box recorders don’t just explain crashes — they prevent them.

And the same telemetry that enables self-healing also closes the accountability gap. When your security team asks what the agent did, when compliance needs an audit trail, when your provider changes terms and you need to prove your session history is yours — the data is there. I wrote about this problem in detail in why agent visibility matters: organizations are banning agent tools not because the tools are bad, but because there’s no visibility into what they do. Telemetry solves both problems with the same infrastructure.

The Architecture

The simplest architecture that works: a transparent proxy between your agent and the model provider.

tapes proxy architecture — Agent sends requests through the tapes proxy to the Model Provider, with all interactions recorded to a durable Postgres session store for search, replay, and export

No code changes to your agent. No SDK integration. Point your agent at the proxy instead of the provider, and every interaction gets recorded into a durable session store. You get semantic search across your entire history and a complete, replayable record of every run — one that survives crashes, closed context windows, and provider outages.

This is what tapes does. It’s open source, runs locally, and works with Anthropic-, OpenAI-, and Ollama-compatible providers.

The Invisible Gap

Teams are starting to collect data on how their developers use AI — acceptance rates, model usage, friction points. That’s a good start. But it doesn’t go far enough.

You also need to collect data on how your agents build software. Every tool call, every decision branch, every failure and recovery. Not for dashboards. For durability.

Once you have that durable record, patterns inside it tell you things the bill never can — like which sessions wrote a pile of cache that never paid back, or which “stable” prompts are stable because nobody has cleaned them up. We took that exact lens to nineteen days of Claude Code usage in Prompt Caching Is Subsidizing Bad AI Architecture.

The gap between “agent that works” and “agent you can trust in production” is telemetry. Black box recorders are infrastructure. The teams that build for durability from the start are the ones that’ll still be running agents when everyone else is dealing with bans, outages, and audit failures.

The black-box pattern has two parts: writing the record and reading it back. The proxy that writes it is AI session capture. The interface that reads it is agent session replay.

Give your agents a black box. Start with tapes.

Found this useful? Share it.
ShareY

Start with paper

Turn every session into knowledge at team scale.

Get started