Skip to content

Paper Compute Concept

Agent Skills vs Memory, RAG, and Fine-Tuning

Fine-tuning changes the model. RAG retrieves content into context. Memory recalls previous interactions. Skills package procedural knowledge an agent can activate. The real question is what kind of improvement you need to make durable, rather than which one to choose.

Published June 26, 2026· Updated September 4, 2026
SkillsComparisonRAGFine-Tuning

Definition

Agent skills, memory, RAG, and fine-tuning are four complementary ways to improve how an agent performs. Skills package together reusable task instructions, workflows, scripts, references, and error-handling guides that an agent can activate when it needs them; fine-tuning changes the model's weights; RAG retrieves relevant content into context; memory preserves information and experience from previous interactions for later recall. These approaches differ primarily in what kind of improvement they make durable and how that improvement reaches the agent.

Skills, memory, RAG, and fine-tuning are four ways to make an agent better at your problems, and they are often discussed as if you have to pick one. You do not. Skills package together reusable task instructions, workflows, scripts, references, and error-handling guides that an agent can activate when it needs them. Fine-tuning changes the model’s weights, RAG retrieves relevant content into context, and memory preserves information and experience from previous interactions for later recall. These approaches differ primarily in what kind of improvement they make durable and how that improvement reaches the agent.

Why skills, memory, RAG, and fine-tuning get confused

All four are sold with the same vocabulary. Each one promises that the agent will “learn,” “know your codebase,” or “get better over time,” so from the outside they look like four vendors solving one problem. They are not. Each mechanism makes a different kind of improvement durable (trained weights, a retrieval index, a memory store, or a packaged skill), and the promises only come true when the mechanism matches the kind of improvement your problem needs.

The distinction matters because the failure is silent. An agent with the wrong mechanism does not error out; it just underperforms in a way that looks like a model limitation. A team that stuffed its deployment runbook into a retrieval index will watch the runbook arrive in context without becoming behavior, conclude that “RAG doesn’t work,” and never notice that what was missing was packaging rather than retrieval: an artifact with an activation boundary and defined instructions. Naming the distinctions is what makes the debugging possible.

Fine-tuning, RAG, memory, and skills each fill a slot the others leave empty, and they can be combined rather than treated as exclusive choices. This page will help you to identify what kind of improvement your problem needs, how it should reach the agent, and what each mechanism contributes when you stack them.

How skills, memory, RAG, and fine-tuning each work

Each approach answers a different question about how to improve an agent.

Four mechanisms, four jobs
Fine-tuningChanges the model's weights. Makes desired behavior (style, output structure, domain conventions) more intrinsic to the model.
RAGRetrieves relevant content into the prompt. Supplies information.
MemoryPreserves information and experience from previous interactions for later recall. Supplies continuity.
SkillsPackage instructions, scripts, and references the agent activates for a task. Supply a course of action.

Fine-tuning adjusts the model’s weights on curated examples. It is the right tool when you want a behavior to be intrinsic to the model rather than re-instructed on every run: task behavior, output structure, style, classification conventions, specialized domain behavior. What it does not do well is encode a specific, inspectable procedure: weights are not readable, iteration takes training cycles, and a fine-tune is tied to the model it was trained on.

RAG retrieves relevant content and places it in context. It is the right tool when the agent needs information it does not carry: current documentation, internal facts, reference material. And retrieval is not limited to facts: RAG can retrieve a runbook or a how-to as readily as a reference table. What it does not inherently do is turn that content into an activatable, governed procedure. RAG answers what information should enter this context; a skill answers what reusable capability or workflow should activate for this task. It is the difference between finding a shell script pasted in a wiki and having it installed on your PATH. It’s the same content, but only one is packaged, versioned, and runnable on demand.

RAG can retrieve a procedure. A skill packages one for reuse.

Memory covers more ground than any one sentence. The cognitive taxonomy runs from working and sensory memory to episodic, semantic, and even procedural memory, and agent frameworks borrow different pieces of it. In many agent systems, memory primarily preserves information or experience from previous interactions so it can be recalled later: what was said earlier in a conversation, what happened in a previous session, what a user prefers. Unlike a skill, that recalled experience is not necessarily packaged as an explicit task-specific artifact with an activation boundary and defined instructions. Recall brings the past back into context; it does not by itself turn the past into a governed capability.

Skills package together reusable task instructions, workflows, scripts, references, and error-handling guides that an agent can activate when a task calls for them. Activation is the mechanism: when a skill fires, its contents load into the agent’s context. A skill is not a separate architectural layer beside the prompt. It is a packaged, activatable form of procedural knowledge delivered through it. What distinguishes a skill from the other three is the packaging: a named, bounded, reviewed, versioned artifact that says when it applies, what to do, and what to do when it goes wrong.

RAG can retrieve a runbook. Memory can remember that you used it. Fine-tuning can make the model better at reasoning about it. A skill turns the workflow into an explicit capability the agent can activate again. The gap between recallable knowledge and packaged capability is measurable: when we read 100 of our own team’s recorded sessions closely, most contained knowledge worth reusing, and almost none of the archive had been packaged.

In a 766-session team archive, 70 of the 100 sessions read closely carried reusable knowledge. Only 29 of the 766 had been packaged as skills.

What each approach changes, and at what cost
ApproachWhat it changesIterationInspectable?Model-bound?
Fine-tuningModel weights (default behavior)Training cyclesNo (weights are opaque)Yes (tied to the model)
RAGPrompt context (information)Update the indexPartly (sources are visible)No
MemoryRecall of previous interactionsUpdates as interactions occurPartly (entries are readable)No
SkillsPackaged procedure (steps, decisions, fixes)Edit and republish the packageYes (reviewed and versioned)Largely model-agnostic

Because each makes a different kind of improvement durable, the four stack rather than compete. A capable system might fine-tune so that domain conventions are the model’s default, retrieve current documentation with RAG, carry conversation context in memory, and activate a skill for the procedure that solved the task last time.

Which mechanism fits the problem
What kind of improvement does the agent need?

Better default behavior               →  Fine-tuning
Missing or current information        →  RAG
Relevant prior context or experience  →  Memory
A repeatable way to perform a task    →  Skill

Example: rolling back a bad deploy with fine-tuning, RAG, memory, and a skill

Take a concrete recurring task: an agent asked to roll back a bad deploy of an internal service. Walk it through all four mechanisms and each contributes something different.

  • Fine-tuning can make certain behavior more native to the model: the organization’s incident-classification conventions, the structured output schema its tooling expects, or specialized troubleshooting behavior without supplying the facts of this particular incident.
  • RAG retrieves what this task needs to know: the service’s current configuration reference, the on-call doc that names the rollback window, the changelog for the release being reverted. Facts that change too often to live in weights.
  • Memory carries the session’s continuity: which environment the operator said matters, what the agent already tried three turns ago, the preference stated last week for staged rollbacks over instant ones.
  • A skill activates the packaged course of action: confirm the failing health check first, snapshot the current state, revert to the previous tagged release, verify the health check clears, and if the revert itself fails with a migration conflict, run the schema rollback before retrying.

Remove any one mechanism and a specific thing degrades. Without RAG the agent works from stale facts; without memory it re-asks what it was told; without the skill it rediscovers the procedure and the migration-conflict fix from scratch. None of the four substitutes for another, which is the whole argument of this page in one scenario.

Adjacent concepts for agent skills

The skills side of this comparison connects to a cluster of related terms:

  • AI agent skills: the procedural artifact itself (trigger, steps, decisions, troubleshooting).
  • Skill extraction: how a recorded session becomes a skill.
  • Skill library: where reviewed, versioned skills live and how a team governs them.
  • Skill drift: what happens when a skill’s procedure falls behind the system it describes.
  • Continuous agent improvement: the capture → extract → apply loop all four mechanisms feed.

Failure modes: the wrong mechanism for the job

Most failures in this space are category errors, meaning the knowledge was real, but it was carried by the wrong mechanism.

  • Expecting retrieval to provide workflow guarantees. A runbook can live in a retrieval system, and the agent may retrieve exactly the right document. Retrieval solves discovery and context selection (which content should enter this context). It does not inherently provide invocation rules, versioning, validation, error handling, or execution guarantees, so retrieving the correct document is not the same as activating and following a tested procedure.
  • Team-shared procedures kept in one agent’s memory. What one person’s agent recalls is not packaged as an artifact a teammate’s agent can activate. It has no review, no version, and no activation boundary, so it never reaches the rest of the team.
  • Fine-tuning on volatile knowledge. Facts that change weekly (config values, API surfaces, org structure) baked into weights are stale by the time the training run finishes, and the only fix is another training run.
  • Skills used as a fact store. A skill that is mostly reference material is a worse retrieval index. If nothing in it is a step, a decision, or a fix, it belongs in the retrieval index.
  • Skills adopted and never revisited. A skill encodes the procedure that worked when it was written or extracted. When the system underneath changes and the skill does not, you get skill drift: the procedural version of the stale index.

When you need a skill instead of memory, RAG, or fine-tuning

The decision signals fall out of what each mechanism makes durable:

  • Reach for fine-tuning when you want a behavior to be the model’s default (a style, an output structure, a classification convention) instead of something you re-instruct on every run.
  • Reach for RAG when the agent lacks information: the facts exist in documents, they change too often for weights, and the agent needs them in context to write correctly.
  • Reach for memory when the problem is continuity: the agent re-asks what it was told, forgets preferences, or loses the thread across sessions.
  • Reach for a skill when the same class of task recurs, the solution is an ordered procedure with decisions and known fixes, and more than one person (or agent) needs to follow it. If you find yourself pasting the same instructions into prompts, or an agent keeps rediscovering a fix your team already found, that is the skill signal.

A system can use all four because they solve overlapping but distinct problems. The signal to watch is which kind of improvement your failures are asking for.

Additional resources

Frequently asked questions

What is the difference between agent skills, memory, RAG, and fine-tuning?+
The practical difference shows up in what each is for. Reach for fine-tuning when a behavior should become the model's default rather than something you re-instruct every run. Reach for RAG when the agent is missing information that lives in documents. Reach for memory when the problem is continuity: the agent forgets what previous interactions established. Reach for a skill when a repeatable task procedure should activate on demand, with a version and known fixes. None replaces the others, and picking the wrong mechanism fails silently: the agent underperforms instead of erroring.
Are agent skills a replacement for fine-tuning?+
No. Fine-tuning changes model behavior at the weight level, which is the right tool for making desired behavior (style, output structure, domain conventions) more intrinsic to the model. Skills change the procedure the model follows on a specific class of task without touching weights. A team can fine-tune for default behavior and use skills for procedure at the same time.
When should you use RAG instead of a skill?+
Use RAG when the agent needs relevant content pulled into context: reference material, documentation, facts. RAG can retrieve a procedure too; what it does not inherently do is turn that content into an activatable, governed capability. Reach for a skill when a reusable workflow should activate for the task: packaged instructions with an activation boundary, a version, and known fixes. Many tasks want both: retrieve the relevant facts and activate the known procedure.
How is a skill different from agent memory?+
In many agent systems, memory primarily preserves information or experience from previous interactions so it can be recalled later. A skill differs in packaging as much as content: it is an explicit task-specific artifact with an activation boundary and defined instructions, reviewed and versioned. Recalled experience can be an input a skill is built from, but recall on its own does not produce an activatable artifact.
Can you use all four together?+
Yes. A system can use all four because they solve overlapping but distinct problems. Fine-tuning makes desired behavior more intrinsic to the model, RAG supplies current information, memory carries what previous interactions established, and skills supply activatable procedures that worked before. Each makes a different kind of improvement durable, so combining them tends to compound rather than conflict.

Where to go next