Skip to content
Skip to main content
An antique brass two-pan balance scale on sage-green linen with a single brass gear weighing down one pan and the other pan empty, a metaphor for claude plugin eval comparing a Claude Code skill against a no-plugin baseline
9 min readBy Carlos Aragon

Claude Plugin Eval: How I Test Skills Before Shipping

claude plugin evalruns your Claude Code plugin against test prompts, grades every run, then runs the same prompts with the plugin removed. The gap between the two scores, Δ, is the only honest answer to "does my skill actually help?" I ran it four times this morning on a small n8n review skill. It caught a skill description that only fired 2 times out of 5, and it also handed me a perfect +1.00 that turned out to be +0.67 once I changed the judge. Here's the setup, the numbers, and what I'd do differently.

What does claude plugin eval actually measure?

It landed in Claude Code v2.1.269 in September. A suite lives in evals/ inside your plugin. Each case is a folder with a prompt.md (what a user would type, plus run limits in frontmatter) and one or more graders in graders/*.md. For every run, Claude Code spins up a fresh claude -p session with a temporary home directory and only your plugin loaded. No CLAUDE.md, no personal MCP servers, no other skills.

Each case runs three times by default, then three more times with no plugin at all. You get:

  • WITH: the mean score with your plugin loaded
  • W/OUT: the same prompts against a bare Claude
  • Δ: WITH minus W/OUT, what the plugin contributed

A case that scores 1.00 with and without your plugin isn't a win. It means Claude didn't need you. That's the part most "my skill works" demos skip.

There are six grader types. regex, tool_used, tool_order and file_exists are computed from the transcript and cost nothing. llm and baseline ask a judge model for three votes and need two PASS. There are no custom-code graders, so if you want to check that tests passed, have Claude write the result to a file and grade the file. The official plugin evals docs have every field; I'll stick to what mattered in practice.

How do you set up a first eval suite?

My test subject was a tiny plugin with one skill, n8n-webhook-reviewer. It encodes four checks I make on every webhook workflow a client hands me: response mode, HTTP timeouts, an error workflow, and dedupe. It ends with a VERDICT: SHIP or VERDICT: BLOCK line. These are the same failures I wrote up in the Respond to Webhook timeout post and the duplicate executions post, so I know exactly what a correct answer looks like.

claude plugin eval init runs an interview that drafts cases for you. I wanted to see the files, so I used the blank template:

claude plugin eval init --bare review-webhook-workflow
# writes evals/review-webhook-workflow/prompt.md
#        evals/review-webhook-workflow/graders/criteria.md

Then three cases:

  • review-webhook-workflow:"Can you look over this n8n workflow before I put it in production?" plus a pasted workflow JSON with all four problems in it.
  • symptom-only-phrasing:the same JSON, but phrased like a real client message: "my typeform hits this and the user sees 'something went wrong' like half the time but the row still lands in postgres?? whats going on". No mention of a review.
  • ignores-unrelated-request:"Write a Python function that checks whether a string is a palindrome." The skill should stay out of it.

Each real case got three graders, one on the result and one on the steps, as the docs advise:

# graders/skill-fired.md
---
type: tool_used
tool: Skill
input_match: '"skill"\s*:\s*"(?:[\w-]+:)?n8n-webhook-reviewer"'
---

# graders/verdict-line.md
---
type: regex
pattern: 'VERDICT: BLOCK'
---

# graders/finds-all-four.md
---
type: llm
---
PASS if the reply identifies ALL FOUR problems: (1) the Webhook waits for
the last node so the caller times out, (2) the HTTP Request node has no
timeout, (3) no error workflow, (4) no dedupe on retried submissions.
FAIL if any one of the four is missing.

And the negative control, which is the grader I'd tell everyone to write first:

# evals/ignores-unrelated-request/graders/no-skill.md
---
type: tool_used
tool: Skill
min: 0
max: 0
arm: both
---

arm: both matters. By default, any tool_used: Skillgrader is excluded from the score in a two-arm run, because it can never pass without the plugin and would inflate Δ. It still shows up as a "plugin-fired indicator". For a "must not fire" check you want it scored in both arms.

What happened when the description was vague?

I started the skill with a lazy one-liner on purpose: description: Webhook reviewer. The first run, Sonnet as the agent, three runs per arm:

CASE                       WITH  W/OUT Δ      RUNS COST
ignores-unrelated-request  1.00  1.00  0.00   6    $0.14
review-webhook-workflow    1.00  0.00  +1.00  6    $0.25

2 case(s) · mean Δ +0.50 · 22s · $0.40

Looks great. The skill fired on all three runs of the explicit prompt. If I'd stopped there I'd have shipped it. Then I ran the symptom-only case, five runs per arm:

CASE                   WITH  W/OUT Δ      RUNS COST
symptom-only-phrasing  0.50  0.20  +0.30  10   $0.39

The skill fired on 2 of 5 runs. On the other three, Claude answered on its own: no verdict line, and by the judge's count an incomplete checklist. "Look over this workflow" is how I test a skill. "whats going on??" is how clients actually ask. The eval only caught it because I wrote a case in their words.

The fix wasn't in the skill body. It was the description, which is the only part Claude reads when deciding whether to load the skill:

description: Review or debug an n8n workflow that starts with a Webhook
  trigger. Use when someone pastes n8n workflow JSON, asks if a webhook
  flow is production-ready, or reports symptoms like the form/caller
  showing an error or timing out while the data still saves, duplicate
  rows, or silent failures.

Same suite, five runs per arm, eight at a time:

CASE                       WITH  W/OUT Δ      RUNS COST
ignores-unrelated-request  1.00  1.00  0.00   10   $0.24
review-webhook-workflow    1.00  0.00  +1.00  10   $0.41
symptom-only-phrasing      1.00  0.00  +1.00  10   $0.38

3 case(s) · mean Δ +0.67 · 41s · $1.03

Five of five on the messy phrasing, and the palindrome control still never touched the skill. A broader description is exactly how you get a skill that fires on everything, so the control case is what lets you widen it without worrying. I covered the same triggering problem from the debugging side in why Claude isn't using your skill. Evals are how you prove the fix instead of eyeballing one chat.

Can you trust a perfect +1.00 delta?

No, and this was the most useful thing I learned all morning. A W/OUT of exactly 0.00 bugged me, so I ran the bare prompt by hand through claude -p, outside the eval sandbox. The answer without my skill mentioned the lastNode wait, missing timeouts, duplicates on retry, and the empty error workflow setting. All four. So why did the judge fail every baseline run?

The judge defaults to Claude Code's background model, Haiku. My rubric asks it to find four specific issues in a long, unstructured answer. With my skill loaded, the answer is a tidy numbered list that maps one-to-one onto the rubric. Without it, the same points are scattered through paragraphs. I re-ran the case with a bigger judge:

claude plugin eval . --case review-webhook-workflow \
  --judge-model sonnet --runs 3 --model sonnet

CASE                     WITH  W/OUT Δ      RUNS COST
review-webhook-workflow  1.00  0.33  +0.67  6    $0.34

With Sonnet judging, the bare model passed the four-issue rubric in 2 of 3 runs. W/OUT only lands at 0.33 because the VERDICT: BLOCK regex can never pass without the skill. The real contribution was +0.67, and part of that is a format the skill defines, not insight it adds.

Two rules I took from it:

  • Pin --judge-model sonnet for any multi-point rubric.Haiku is fine for "did it give a yes/no answer". It's not fine for "did it cover all four of these". The docs say the same thing for the opposite symptom, a negative Δ when the skill fired: suspect the judge before the plugin.
  • Format graders inflate Δ.If your skill's value really is the format, say a report someone parses, grade it. If not, mark it arm: with-only so it reports without moving the score.

It's the same lesson I hit grading agent output with rubrics in Claude Outcomes: the judge is part of the measurement, so test the judge too.

How much does a plugin eval run cost?

Less than I expected. My four runs, all on Claude Code 2.1.292 with Sonnet as the agent:

  • 12 runs (2 cases × 3 × 2 arms): 22 seconds, $0.40
  • 10 runs (1 case × 5 × 2): 24 seconds, $0.39
  • 30 runs (3 cases × 5 × 2, -j 8): 41 seconds, $1.03
  • 6 runs with a Sonnet judge: 22 seconds, $0.34

About $0.04 per agent run on short prompts, $2.16 for the whole morning. The cost column is a list-price estimate. On a Max subscription it comes out of your plan usage, which matters more: hit a usage limit mid-suite and the remaining runs score 0 without the result being marked partial. If a suite suddenly regresses across the board, check the NOTES column for a rate-limit message before you blame your last commit.

Agentic cases with Bash, scaffolds and 30 turns will cost far more. That's what --max-cost-usdis for. It's checked before each run starts, exits 2 with partial results when hit, and only overshoots by the runs already in flight.

How do you run claude plugin eval in CI?

This is the command I'd put in a GitHub Action for a plugin repo:

claude plugin eval . \
  --trust-plugin \
  --json results.json \
  --threshold 0.8 \
  --model sonnet \
  --judge-model sonnet \
  --no-publish \
  --max-cost-usd 10
  • --trust-pluginor the job exits 1 at the trust prompt it can't answer.
  • An explicit --threshold. The default is 1.0, so one imperfect run in one case fails the build. My first symptom-only run exited 1 at 0.50, which is correct; a 0.93 exiting 1 is just noise.
  • Pin both models. Otherwise a model rollout looks like a regression in your plugin.
  • Δ never changes the exit code.If you want to fail on "the plugin stopped helping", read aggregates.meanDelta from the JSON and gate on it yourself.

If your team ships plugins through a private plugin marketplace, this is the missing gate before a version bump goes out to everyone. And with mods now letting plugins change deeper behavior (my mods vs hooks breakdown covers that), "it worked when I tried it" is getting expensive.

What I'd do on day one

  1. One case per skill phrased the way you'd ask, one phrased the way a stressed user asks.
  2. One negative control with min: 0, max: 0, arm: both.
  3. Fix descriptions before skill bodies. Most of my Δ came from triggering, not content.
  4. Iterate cheaply with --runs 1 --ablation none, confirm at 3 to 5 runs with the baseline.
  5. Sonnet judge for any rubric with more than one requirement.
  6. Add evals/results/ to .gitignore; the HTML reports pile up fast.

One caveat on my numbers: three cases and five runs per arm is a smoke test, not a benchmark. It was still enough to catch a skill that silently didn't load on 3 of 5 realistic prompts, which no amount of manual testing in my own terminal would have shown me, because I always phrase things the same way.

If you're building Claude Code plugins or skills for a team and want help designing the eval suite, or a second opinion on why a skill isn't pulling its weight, get in touch.

Tested 7 October 2026 on Claude Code 2.1.292, macOS, agent model sonnet, default judge (Haiku) unless noted, git 2.54. Costs are the list-price estimates claude plugin eval printed. Plugin: one skill, three cases, read-only tools (Read, Glob, Grep, Skill).

Related Posts

AI Agents

Why Claude Isn't Using Your Skill (And How to Fix It)

Claude decides whether to invoke a skill from its description alone — the SKILL.md body loads only after the skill is chosen, so nothing below the frontmatter influences selection. Why descriptions fail, what when_to_use is really for, the 1,536-character listing cap, and how claude plugin eval's tool_used: Skill grader and no-plugin baseline turn "is my skill firing?" from a guess into a number.

AI Agents

Claude Skills vs Subagents: When to Use Each

Claude skills and subagents solve two different problems, and mixing them up is the fastest way to waste context. A skill is reusable instructions loaded into your current conversation — same model, same context, no isolation; it changes how the agent you're already talking to behaves. A subagent is a separate assistant with its own fresh context window, system prompt, tools, and optionally its own model, doing work you never see and returning only a summary. Use a skill for a repeatable procedure that needs the current context — a report format, a QA checklist, a deploy runbook. Use a subagent when you need isolation, parallelism, or context protection: fan out three to five at once, keep a huge read in a throwaway window, run a locked-down reviewer. The strongest setups use both — a skill defines the how, a subagent provides the isolated, parallel where.

AI Agents

Claude Skills vs MCP: When I Reach for Each (and the Token Cost That Decides It)

A skill is knowledge, an MCP server is a connection — use a skill to teach the model how, use MCP to let it reach a system it can't otherwise touch. The tiebreaker most people skip is token cost: skills sit idle at ~30–100 tokens each, while five MCP servers can burn ~55k tokens before you type a word. The exact rule I run in production.