Skip to content

Write a good skill eval

How to write eval criteria a judge can actually check, and how to choose weights so a pass means the skill works and a failure tells you what to fix.

An eval is only as good as its spec. This guide shows you how to write criteria the judge can check, and how to choose weights so a pass means the skill works and a failure tells you what to fix. It assumes you know how skill evals run.

Start from the generated spec

Choose Generate full spec before writing anything by hand. The console drafts criteria from the skill and the sessions it came from. A drafted spec is easier to sharpen than a blank page.

Your job is review. Keep the rules that describe what you actually care about. Cut the rest. Adjust the weights. Once you edit a spec it is yours: the console keeps your edits and won’t regenerate over them.

Make each criterion one checkable rule

The judge reads each criterion and decides: passed or failed. Write every criterion so that decision is easy.

  • One rule per criterion. “Steps are numbered and each step names a command” is two rules. Split it. When it fails, you want to know which half failed.
  • Describe something observable. The judge can check what the skill says and what its output looks like. It can’t check intent. “The skill encourages care with credentials” is a feeling. “The skill tells the agent to read secrets from the environment and never print them” is checkable.
  • Be concrete. Vague: “The skill should be clear.” Checkable: “Every step names the exact command to run.”

A useful test: could a teammate look at a result and agree with the pass or fail call without asking what you meant? If yes, the judge can too.

Pick the kind

Kind says which part of the skill the rule checks:

Kind It checks Example
structure How the skill is organized “The skill lists prerequisites before the first step.”
content What the instructions say “The rollback step names the exact command, not just ‘undo the change’.”
output-property What the results of using the skill must satisfy “Every recommendation cites supporting evidence.”

Kind doesn’t change the scoring. It keeps the spec organized and helps the judge read a rule the way you meant it.

Choose weights

Weights are the part most people get stuck on. A weight does two things:

  1. It sets the score. The score is passed weight divided by total weight. A weight-3 criterion moves the score three times as much as a weight-1.
  2. It decides what forces a revision. A failed criterion with weight 2 or 3 means the skill needs revision, on its own, whatever the rest scored. A failed weight-1 criterion only lowers the score.

So the question to ask when picking a weight is: what should happen when only this rule fails?

  • The skill has failed at its job. Weight 3 (essential). The skill exists to get this right.
  • A teammate would send the work back. Weight 2 (important). Wrong enough to fix before anyone relies on it.
  • You’d accept the work and mention it. Weight 1 (useful). Worth checking, never worth blocking.

A few habits keep weights honest:

  • Start everything at 1 and promote with a reason. “This one is essential because…” is a sentence you should be able to finish.
  • Most specs need one or two 3s. A skill usually has one job. If every criterion is essential, a failure can’t tell you what matters most.
  • Reserve 2 for rules you’d act on. If a failed criterion wouldn’t change the skill, it’s a 1. Or it doesn’t belong in the spec.

How the math plays out

Say a deploy-runbook skill has four criteria:

Criterion Weight
Names the exact deploy command for each environment 3
Includes a rollback step 2
Steps are numbered 1
Links the on-call dashboard 1

Total weight is 7.

  • Fail “Links the on-call dashboard”: score 6 ÷ 7 = 86%. The run can still pass.
  • Fail “Includes a rollback step”: score 5 ÷ 7 = 71%, and the skill needs revision no matter what else passed. A weight-2 rule failed.
  • Fail “Names the exact deploy command”: score 4 ÷ 7 = 57%, needs revision. The skill missed the thing it exists for.

The score tells you how much of the spec held. The weights decide which failures block the skill.

Keep the spec small

A handful of criteria beats a long checklist. Every criterion is one more judgment call per run, and long specs collect rules nobody would act on. If you can’t say what you’d do when a criterion fails, cut it.

Let the runs teach you

Read the judge’s reasoning on each criterion after a run.

  • When the judge misread a rule, rewrite the description. The reasoning shows you how it was read.
  • When a criterion passes every run and never changes a decision, demote it to weight 1 or remove it.
  • When a real problem slips through, add a criterion for it. Describe it in plain words and choose Generate criterion if you want the console to write it.

Edits apply to future runs. Completed runs keep the rubric they used, so old scores stay comparable to the spec they were scored against.

Next steps

Frequently asked questions

How do I choose a weight for a criterion?+
Ask what should happen when only this rule fails. If the skill has failed at its job, use 3. If a teammate would send the work back, use 2. If you would accept the work and mention it, use 1. A failed weight-2 or weight-3 criterion forces Needs revision on its own; a failed weight-1 only lowers the score.
How many criteria should an eval spec have?+
A handful. Most skills have one job, so most specs need one or two weight-3 criteria and a few supporting rules. If you cannot say what you would do when a criterion fails, cut it.
If I edit the spec, do old evaluation results change?+
No. Spec changes apply to future runs. Completed runs keep the rubric they used, so old scores stay comparable to the spec they were scored against.
Copied to clipboard