An eval is only as good as its spec. This guide shows you how to write criteria the judge can check, and how to choose weights so a pass means the skill works and a failure tells you what to fix. It assumes you know how skill evals run.
Start from the generated spec
Choose Generate full spec before writing anything by hand. The console drafts criteria from the skill and the sessions it came from. A drafted spec is easier to sharpen than a blank page.
Your job is review. Keep the rules that describe what you actually care about. Cut the rest. Adjust the weights. Once you edit a spec it is yours: the console keeps your edits and won’t regenerate over them.
Make each criterion one checkable rule
The judge reads each criterion and decides: passed or failed. Write every criterion so that decision is easy.
- One rule per criterion. “Steps are numbered and each step names a command” is two rules. Split it. When it fails, you want to know which half failed.
- Describe something observable. The judge can check what the skill says and what its output looks like. It can’t check intent. “The skill encourages care with credentials” is a feeling. “The skill tells the agent to read secrets from the environment and never print them” is checkable.
- Be concrete. Vague: “The skill should be clear.” Checkable: “Every step names the exact command to run.”
A useful test: could a teammate look at a result and agree with the pass or fail call without asking what you meant? If yes, the judge can too.
Pick the kind
Kind says which part of the skill the rule checks:
| Kind | It checks | Example |
|---|---|---|
structure |
How the skill is organized | “The skill lists prerequisites before the first step.” |
content |
What the instructions say | “The rollback step names the exact command, not just ‘undo the change’.” |
output-property |
What the results of using the skill must satisfy | “Every recommendation cites supporting evidence.” |
Kind doesn’t change the scoring. It keeps the spec organized and helps the judge read a rule the way you meant it.
Choose weights
Weights are the part most people get stuck on. A weight does two things:
- It sets the score. The score is passed weight divided by total weight. A weight-3 criterion moves the score three times as much as a weight-1.
- It decides what forces a revision. A failed criterion with weight 2 or 3 means the skill needs revision, on its own, whatever the rest scored. A failed weight-1 criterion only lowers the score.
So the question to ask when picking a weight is: what should happen when only this rule fails?
- The skill has failed at its job. Weight 3 (essential). The skill exists to get this right.
- A teammate would send the work back. Weight 2 (important). Wrong enough to fix before anyone relies on it.
- You’d accept the work and mention it. Weight 1 (useful). Worth checking, never worth blocking.
A few habits keep weights honest:
- Start everything at 1 and promote with a reason. “This one is essential because…” is a sentence you should be able to finish.
- Most specs need one or two 3s. A skill usually has one job. If every criterion is essential, a failure can’t tell you what matters most.
- Reserve 2 for rules you’d act on. If a failed criterion wouldn’t change the skill, it’s a 1. Or it doesn’t belong in the spec.
How the math plays out
Say a deploy-runbook skill has four criteria:
| Criterion | Weight |
|---|---|
| Names the exact deploy command for each environment | 3 |
| Includes a rollback step | 2 |
| Steps are numbered | 1 |
| Links the on-call dashboard | 1 |
Total weight is 7.
- Fail “Links the on-call dashboard”: score 6 ÷ 7 = 86%. The run can still pass.
- Fail “Includes a rollback step”: score 5 ÷ 7 = 71%, and the skill needs revision no matter what else passed. A weight-2 rule failed.
- Fail “Names the exact deploy command”: score 4 ÷ 7 = 57%, needs revision. The skill missed the thing it exists for.
The score tells you how much of the spec held. The weights decide which failures block the skill.
Keep the spec small
A handful of criteria beats a long checklist. Every criterion is one more judgment call per run, and long specs collect rules nobody would act on. If you can’t say what you’d do when a criterion fails, cut it.
Let the runs teach you
Read the judge’s reasoning on each criterion after a run.
- When the judge misread a rule, rewrite the description. The reasoning shows you how it was read.
- When a criterion passes every run and never changes a decision, demote it to weight 1 or remove it.
- When a real problem slips through, add a criterion for it. Describe it in plain words and choose Generate criterion if you want the console to write it.
Edits apply to future runs. Completed runs keep the rubric they used, so old scores stay comparable to the spec they were scored against.
Next steps
- Evaluate a skill — the full evaluation flow: spec, run, verdict, improvement.
- Skills overview — where evals fit in the skill lifecycle.
- Stop trusting skills you haven’t measured — why unmeasured skills quietly go wrong.