I built a pipeline to fine-tune a coding model on labeled agent sessions exported from Paper Compute. I used Databricks to store the conversations alongside their labels so I could inspect what would go into my datasets. The pipeline turns selected sessions into training examples and uses corrections to build tests. I track the runs in MLflow to compare the original model with the tuned version.
My source data is the conversation from a coding task: what the engineer asked for, what the agent did, and any corrections that followed. I use labels to mark good examples I’ve reviewed and places where the engineer corrected the agent. Those labels tell the pipeline which sessions to select for training and which to save for evaluation.
My Pokémon experiments led me to the argument in Every Agent Session Should Be a Heist: a model earns its role through measured results. This pipeline builds on that idea. Once I have the record of what worked and what failed, I can use it to train a model for the work and check it against those failures.
This is the path I built with labeled coding sessions: export the evidence, select the examples, keep training and evaluation separate, then train and compare. The labels carry the selection decisions through each step.
Start with what the label means
A label is a named annotation attached to a session or a specific part of its trace. It gives me a way to select records without searching the text again. A session-level label might say the work produced a good result. A label on an individual exchange can identify where the engineer corrected the agent.
I use golden for manually selected good sessions and regression for failures I want to test again. pushback marks an engineer correcting the agent, and apology marks the agent backtracking. I use apologies to find exchanges worth reviewing and manually select the sessions marked golden.
Labels can come from a person or an automated labeler. In Tracing. Where Jev Fits, I compared Jev with other models for finding candidate labels in session turns. That experiment covers how I find the labels; this pipeline uses them to select training examples and evaluation cases. Each label stays linked to the session’s messages. For a proposed pushback label, I can read the agent’s response and the engineer’s reply to see what needed correcting.
I also use no-outcome to mark sessions that produced no result. That helps filter the collection before training. Marking a session golden is a separate decision: I’ve reviewed it and want the model to learn from that approach.
Export the evidence with the labels
The export carries each label with its evidence. For a correction, that means the engineer’s instruction, the agent’s response, and the context before it. For a training example, it means the conversation that produced the outcome.
My pipeline exports the records from Paper Compute to local files, then loads sessions, turns, labels, and training_input into Unity Catalog, Databricks’ catalog for managing data and permissions. Session and turn identifiers connect the labels to the exchanges they describe. That connection lets me inspect why an example was selected.
The tables also let me check the input before training anything. I can group correction labels by model, project, or week, then read the engineer’s words behind a count. To compare rates, I include eligible non-empty sessions alongside labeled sessions. Ten corrections out of twenty sessions means something different from ten out of a thousand. Those rates describe the sampled work; differences in tasks still matter.
I added checks to keep dataset updates complete: skip empty sessions, reject incomplete records, and stop fetching if the source service becomes unavailable. Before updating the tables in Databricks, the sync checks that the export is complete and contains data. That keeps an interrupted export from replacing the dataset I already have.
Give training and evaluation different jobs
Training examples show the model behavior I want it to learn. Evaluation cases check whether it can produce useful behavior on inputs held back from training. Putting the same session in both makes the result harder to trust: the model may already have seen the answer.
For training, I want conversations that show behavior worth repeating. I select them in two ways. The automatic selection excludes sessions marked as producing no outcome, needing correction, or containing a failure. I also manually select good sessions with the golden label.
From those golden sessions, I reserve roughly one fifth for evaluation and use the rest for training. This gives me good examples to test against that the model hasn’t seen during training. Each session keeps its assignment when I rebuild the dataset.
Sessions with correction or regression labels stay out of training, including ones also marked golden. I use them to test whether a model repeats the same mistake. For example, if an agent changed a function’s return type against the engineer’s instructions, I can give another model the original request and check whether its response preserves that type. The correction becomes the guideline for judging the new response.
Play the split below to follow ten example sessions from the export into separate datasets. The same session never appears in both.
into training and evaluation datasets
Each record carries its conversation, labels, and outcome.
The split happens at the session level. Dividing individual turns at random could put the beginning of a conversation in training and its continuation in evaluation. Keeping the session together avoids that overlap. It doesn’t, by itself, catch copied content or repeated tasks across different sessions; those need a separate check.
Turn a correction into a test
To build that test, I stop the recorded conversation just before the agent’s mistaken response. The model receives the earlier messages and writes a new response. I keep the engineer’s later correction separately as the judging guideline.
Take the return-type example. Imagine the original request was to add validation while keeping the function’s return type. The agent changed the type anyway, and the engineer replied, “Keep the existing return type. The callers depend on it.” The test gives the model the original request and earlier messages, then checks whether its new response follows that instruction.
The resulting record has two parts: the earlier conversation as input, and “preserves the existing return type” as the expectation. The model never sees the later correction. This works because the requirement was already in the original request. A correction that introduces a new requirement would need a different test.
I use the same source of feedback in a private harness that reviews my Go and Rust code. It learns from engineers’ code reviews and corrections made during their sessions in Paper Compute. Pairing a correction with the code that prompted it gives the reviewer an example of when to flag a change. That project uses corrections to train the reviewer; the Databricks pipeline here uses them as tests.
Track the experiment in MLflow
MLflow is an open-source tool for tracking machine-learning experiments and evaluating models. An experiment groups related runs. A run records a particular attempt, including its settings, outputs, and measured results. In this pipeline, I use it to keep the evaluation dataset and the base-versus-tuned comparison together.
The sync loads the session tables and selected training examples into Databricks. It also saves the evaluation inputs and their judging guidelines as an MLflow evaluation dataset. Each comparison run uses that dataset, so I can trace the scores back to the cases that produced them.
I give the base model and the tuned model the same inputs held back from training. A separate LLM, hosted on Databricks, judges each response through MLflow’s ExpectationsGuidelines scorer. This is an LLM-as-judge evaluation: it reads the response and checks it against the guideline for that case. In the example above, it checks whether the response preserves the function’s return type. It evaluates the generated response; it doesn’t execute the code.
The results land in separate MLflow runs. I use the same judge and guidelines for both models, compare their scores, then inspect individual responses to understand what changed.
using recorded corrections
Use the earlier conversation. Leave out the original answer and the later correction.
Fine-tune an existing coding model
I used Qwen3-4B as the starting model. Fine-tuning means continuing training on selected examples so the model can adapt to the behavior in them. In this pipeline, each training example is a sequence of user and assistant messages from an eligible session.
I use LoRA, which trains a small set of adapter weights instead of updating every weight in the model. This training step can run on Databricks GPU compute, and I plan to move it there so training runs alongside the datasets and MLflow experiment records.
That gives me a model I can run through the same evaluation as the base model. For each correction case, both get the same earlier conversation. The judge checks both responses against the same guideline. I want to see which failures the tuned model avoids, which it still repeats, and whether it loses behavior the base model got right.
Reuse the work in the next run
A reviewed good session can supply a training example. A correction can supply an evaluation case. Keeping their labels and session identifiers attached gives me a dataset I can rebuild as more work arrives.
The data boundary matters too. Session exports can contain code, prompts, and credentials. Review and redact the selected records before moving them to a training environment. Sending excerpts to an external labeler or judge is a separate decision from processing them locally.
I plan to move training into Databricks next. The selection rules and evaluation cases carry forward, so the next model run starts from the work I have already recorded and reviewed.
Related reading
Start with paper
Turn every session into knowledge at team scale.
