Skip to content

Evaluate a skill

Run a skill evaluation in paper console — check a skill against weighted criteria and real session evidence, read the verdict, and accept an improvement when it fails.

A skill evaluation checks a skill against explicit criteria and real session evidence. It tells you whether the skill holds up. When it fails, it can draft the fix. Run one when you create a skill, when you change one, or when you suspect one has gone stale.

Where to find it

Open a skill in paper console. The Skill evals panel is on the skill detail page, with three tabs:

  • Evaluation — run an eval and read the results.
  • Improvements — review the rewrites the evaluator proposed.
  • Eval spec — the criteria evals run against.

Evaluations run against a saved revision of the skill. A brand-new skill asks you to save a revision first.

The eval spec

An eval spec defines what good looks like for one skill. It is a list of criteria. Each criterion is one rule the evaluator checks, with four parts:

  • Rule ID — a short name for the rule.
  • Kind — which part of the skill the rule checks: structure (how the skill is organized), content (what its instructions say), or output-property (what the results of using it must satisfy).
  • Description — the rule itself, written so a judge can decide pass or fail.
  • Weight — how much the rule matters: 1 is useful, 2 is important, 3 is essential.

You don’t have to write the spec yourself. Choose Generate full spec and the console drafts criteria from the skill and its session evidence. To add one rule, describe it in plain words — “Require every recommendation to cite supporting evidence” — and choose Generate criterion.

Review what was generated. Edit descriptions, change weights, remove rules that don’t matter. Once you edit a spec it is yours: the console keeps your version and won’t regenerate over it. Spec changes apply to future runs across every version of the skill; completed runs keep the rubric they used.

For how to write criteria worth trusting, and how to pick weights, see Write a good skill eval.

Run an evaluation

From the Evaluation tab, choose Run evaluation. If the skill has no eval spec yet, the console asks you to draft one first. The draft takes seconds and lands on the Eval spec tab for review; the run is the next step.

The evaluator judges the exact revision you’re viewing against the spec’s criteria, using your captured sessions as evidence. Results appear in the panel when the run finishes. You can re-run at any time.

Read the result

Every run produces a verdict:

  • Pass — the skill met its criteria.
  • Needs revision — something has to change. A failed criterion with weight 2 or 3 requires revision on its own. Findings from session evidence can also require revision.
  • Not enough evidence — the evaluator had nothing to score against. Add criteria or capture related sessions, then evaluate again.

The score is the criteria score: passed weight divided by total weight. A skill that passes 6 of 7 weight points scores 86%.

Below the score, each criterion shows whether it passed, with the judge’s reasoning. Findings from session evidence are listed with their severity. The details record how many sessions the run considered and which judge model scored it.

Runs accumulate as history on each revision. Each result also shows how the score moved against the previous run, and against the version this revision was based on, so you can see whether a change actually helped.

Improve a failing skill

When a run fails criteria, the panel offers Generate improvement. The evaluator drafts a rewrite of the skill grounded in the evaluation: what failed, and what the evidence says it should say instead.

The draft opens in the Improvements tab as a proposal with a diff and a rationale. From there:

  • Accept — appends the rewrite as a new revision, and the console takes you to it. Publishing a new version for your team is still a separate step.
  • Reject — records the decision and leaves the skill as it was.

After accepting, run the evaluation on the new revision. The result compares its score against the revision it replaced.

Next steps

Frequently asked questions

Do I have to write the eval spec myself?+
No. Choose Generate full spec and the console drafts criteria from the skill and its session evidence. You can also describe one rule in plain words and the console writes the criterion. Everything stays editable, and once you edit a spec the console keeps your version.
What does the weight on a criterion do?+
Weight sets how much a criterion counts. The score is passed weight divided by total weight. A failed criterion with weight 2 or 3 sends the skill to Needs revision on its own. Use 1 for useful, 2 for important, 3 for essential.
What happens when I accept an improvement?+
The proposed rewrite is appended as a new revision of the skill, and the console takes you to it. Publishing it as a new version for your team is still a separate step. Rejecting a proposal records the decision and leaves the skill unchanged.
Why does my evaluation say "Not enough evidence"?+
The evaluator had nothing to score against — no usable criteria and no related captured sessions. Add criteria to the eval spec or capture sessions that use the skill, then evaluate again.
Copied to clipboard