Lance Martin (claude.dev) explains principles for production-like evals with held-out sets, then shows how the claude-api skill’s build-eval and hillclimb commands automate design and overfitting-aware improvement—including cost and performance case studies.
Automating eval design and hillclimbing with Claude
Principles for designing evals and hillclimbing against them without fooling yourself, and how the claude-api skill's build-eval and hillclimb commands put them to work.
Author: Lance Martin · Published: Sep 28, 2026
Evaluations provide a signal on how your app or skill is performing on specific tasks. But designing evaluations, and improving performance on them without fooling yourself, is hard. We've added guidance for both to the claude-api skill.
With the skill, you can run /claude-api build-eval to build an evaluation inside your codebase, and run /claude-api hillclimb to improve your application against it, one change at a time, with a held-out set of examples to catch overfitting.
Eval design
Well designed evaluations have a few common elements:
Eval tasks mirror production. Sample tasks that you care about in the setting where the capability will be used.
Performance improves with stronger models and more thinking. If they don't, ambiguous tasks or a miscalibrated grader often are hobbling performance.
There is "passable" headroom at the frontier. The most capable model at the highest effort should be well below 100%.
Low run-to-run variance. High variance is often due to poorly designed, ambiguous tasks or a grader that produces different verdicts on identical output.
Adversarial sampling
Model capability is jagged. If you pick cases because today's model fails them, you are sampling the valleys of one model's capability surface. Pick hard cases because a human judged them hard.
/claude-api build-eval
When you run /claude-api build-eval in Claude Code, Claude interviews you, builds the eval inside your codebase, and pauses for approval at specific points. It samples inputs from production transcripts, bug reports, hand-written cases, and synthetic cases anchored in real examples. It proposes the cheapest grader that fits (programmatic verification or LLM-as-judge), validates the grader with you, runs a baseline, and prints a score with a confidence interval.
Hillclimbing
Hillclimbing is effective for tuning parameters like effort or prompts. Prefer surfaces that are cheap to iterate, attributable to the score, and well-scoped. Overfitting is common: split train/test, never paste failures into the prompt, and keep answers structurally out of the model's reach.
/claude-api hillclimb
Claude iterates with one patch per round. If the train set improves but the test set is flat, it suspects overfitting and reverts. When the score stalls, it buckets remaining failures by cause. Cost-focused and performance-focused case studies in the original post show large gains (e.g., cutting support-ticket token cost roughly fivefold while improving held-out accuracy; lifting a docs/API skill from 66% toward ~88%).
Getting started
Run /claude-api build-eval to generate an evaluation set. Run /claude-api hillclimb when you have an evaluation and a goal (better performance, or lower cost while performance holds).