A skill evaluation checks a skill against explicit criteria and real session evidence. It tells you whether the skill holds up. When it fails, it can draft the fix. Run one when you create a skill, when you change one, or when you suspect one has gone stale.
Where to find it
Open a skill in paper console. The Skill evals panel is on the skill detail page, with three tabs:
- Evaluation — run an eval and read the results.
- Improvements — review the rewrites the evaluator proposed.
- Eval spec — the criteria evals run against.
Evaluations run against a saved revision of the skill. A brand-new skill asks you to save a revision first.
The eval spec
An eval spec defines what good looks like for one skill. It is a list of criteria. Each criterion is one rule the evaluator checks, with four parts:
- Rule ID — a short name for the rule.
- Kind — which part of the skill the rule checks:
structure(how the skill is organized),content(what its instructions say), oroutput-property(what the results of using it must satisfy). - Description — the rule itself, written so a judge can decide pass or fail.
- Weight — how much the rule matters: 1 is useful, 2 is important, 3 is essential.
You don’t have to write the spec yourself. Choose Generate full spec and the console drafts criteria from the skill and its session evidence. To add one rule, describe it in plain words — “Require every recommendation to cite supporting evidence” — and choose Generate criterion.
Review what was generated. Edit descriptions, change weights, remove rules that don’t matter. Once you edit a spec it is yours: the console keeps your version and won’t regenerate over it. Spec changes apply to future runs across every version of the skill; completed runs keep the rubric they used.
For how to write criteria worth trusting, and how to pick weights, see Write a good skill eval.
Run an evaluation
From the Evaluation tab, choose Run evaluation. If the skill has no eval spec yet, the console asks you to draft one first. The draft takes seconds and lands on the Eval spec tab for review; the run is the next step.
The evaluator judges the exact revision you’re viewing against the spec’s criteria, using your captured sessions as evidence. Results appear in the panel when the run finishes. You can re-run at any time.
Read the result
Every run produces a verdict:
- Pass — the skill met its criteria.
- Needs revision — something has to change. A failed criterion with weight 2 or 3 requires revision on its own. Findings from session evidence can also require revision.
- Not enough evidence — the evaluator had nothing to score against. Add criteria or capture related sessions, then evaluate again.
The score is the criteria score: passed weight divided by total weight. A skill that passes 6 of 7 weight points scores 86%.
Below the score, each criterion shows whether it passed, with the judge’s reasoning. Findings from session evidence are listed with their severity. The details record how many sessions the run considered and which judge model scored it.
Runs accumulate as history on each revision. Each result also shows how the score moved against the previous run, and against the version this revision was based on, so you can see whether a change actually helped.
Improve a failing skill
When a run fails criteria, the panel offers Generate improvement. The evaluator drafts a rewrite of the skill grounded in the evaluation: what failed, and what the evidence says it should say instead.
The draft opens in the Improvements tab as a proposal with a diff and a rationale. From there:
- Accept — appends the rewrite as a new revision, and the console takes you to it. Publishing a new version for your team is still a separate step.
- Reject — records the decision and leaves the skill as it was.
After accepting, run the evaluation on the new revision. The result compares its score against the revision it replaced.
Next steps
- Write a good skill eval — criteria and weights that produce verdicts you can trust.
- Skills overview — the full lifecycle from captured session to published skill.
- Stop trusting skills you haven’t measured — why every skill needs an eval.