{"article":{"slug":"automating-eval-design-and-hillclimbing-with-claude","title":"Automating eval design and hillclimbing with Claude","subtitle":null,"summary":"Lance Martin (claude.dev) explains principles for production-like evals with held-out sets, then shows how the claude-api skill’s build-eval and hillclimb commands automate design and overfitting-aware improvement—including cost and performance case studies.","content_type":"tutorial","language":"en","canonical_url":"https://claude.dev/blog/automating-eval-design-and-hillclimbing/","author":{"name":"Lance Martin","url":"https://claude.dev/blog/automating-eval-design-and-hillclimbing/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Anthropic","url":"https://www.anthropic.com/","listing_slug":"anthropic","listing":{"slug":"anthropic","name":"Anthropic","listing_type":"company","url":"https://listedstartups.com/companies/anthropic"}},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"}],"about_listings":[{"slug":"claude","name":"Claude","listing_type":"product","url":"https://listedstartups.com/products/claude"}],"cover_image_url":null,"license":"all-rights-reserved","word_count":462,"reading_minutes":2,"published_at":"2026-09-28T12:00:00.000Z","added_at":"2026-09-30T00:16:07.307Z","updated_at":"2026-09-30T00:16:07.307Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/automating-eval-design-and-hillclimbing-with-claude","markdown_url":"https://listedarticles.com/articles/automating-eval-design-and-hillclimbing-with-claude.md","example":false,"citation":"Lance Martin, Anthropic. \"Automating eval design and hillclimbing with Claude.\" 28 Sept 2026. https://claude.dev/blog/automating-eval-design-and-hillclimbing/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://claude.dev/blog/automating-eval-design-and-hillclimbing/"},"body_markdown":"# Automating eval design and hillclimbing with Claude\n\nPrinciples for designing evals and hillclimbing against them without fooling yourself, and how the claude-api skill's `build-eval` and `hillclimb` commands put them to work.\n\n**Author:** Lance Martin · **Published:** Sep 28, 2026\n\nEvaluations provide a signal on how your app or skill is performing on specific tasks. But designing evaluations, and improving performance on them without fooling yourself, is hard. We've added guidance for both to the claude-api skill.\n\nWith the skill, you can run `/claude-api build-eval` to build an evaluation inside your codebase, and run `/claude-api hillclimb` to improve your application against it, one change at a time, with a held-out set of examples to catch overfitting.\n\n## Eval design\n\nWell designed evaluations have a few common elements:\n\n1. **Eval tasks mirror production.** Sample tasks that you care about in the setting where the capability will be used.\n2. **Performance improves with stronger models and more thinking.** If they don't, ambiguous tasks or a miscalibrated grader often are hobbling performance.\n3. **There is \"passable\" headroom at the frontier.** The most capable model at the highest effort should be well below 100%.\n4. **Low run-to-run variance.** High variance is often due to poorly designed, ambiguous tasks or a grader that produces different verdicts on identical output.\n\n### Adversarial sampling\n\nModel capability is jagged. If you pick cases because today's model fails them, you are sampling the valleys of one model's capability surface. Pick hard cases because a human judged them hard.\n\n## `/claude-api build-eval`\n\nWhen you run `/claude-api build-eval` in Claude Code, Claude interviews you, builds the eval inside your codebase, and pauses for approval at specific points. It samples inputs from production transcripts, bug reports, hand-written cases, and synthetic cases anchored in real examples. It proposes the cheapest grader that fits (programmatic verification or LLM-as-judge), validates the grader with you, runs a baseline, and prints a score with a confidence interval.\n\n## Hillclimbing\n\nHillclimbing is effective for tuning parameters like effort or prompts. Prefer surfaces that are cheap to iterate, attributable to the score, and well-scoped. Overfitting is common: split train/test, never paste failures into the prompt, and keep answers structurally out of the model's reach.\n\n## `/claude-api hillclimb`\n\nClaude iterates with one patch per round. If the train set improves but the test set is flat, it suspects overfitting and reverts. When the score stalls, it buckets remaining failures by cause. Cost-focused and performance-focused case studies in the original post show large gains (e.g., cutting support-ticket token cost roughly fivefold while improving held-out accuracy; lifting a docs/API skill from 66% toward ~88%).\n\n## Getting started\n\nRun `/claude-api build-eval` to generate an evaluation set. Run `/claude-api hillclimb` when you have an evaluation and a goal (better performance, or lower cost while performance holds).\n\n*Full article with figures and case studies: [claude.dev/blog/automating-eval-design-and-hillclimbing](https://claude.dev/blog/automating-eval-design-and-hillclimbing/).*","body_html":"<h1 id=\"automating-eval-design-and-hillclimbing-with-claude\">Automating eval design and hillclimbing with Claude</h1>\n<p>Principles for designing evals and hillclimbing against them without fooling yourself, and how the claude-api skill&#39;s <code>build-eval</code> and <code>hillclimb</code> commands put them to work.</p>\n<p><strong>Author:</strong> Lance Martin · <strong>Published:</strong> Sep 28, 2026</p>\n<p>Evaluations provide a signal on how your app or skill is performing on specific tasks. But designing evaluations, and improving performance on them without fooling yourself, is hard. We&#39;ve added guidance for both to the claude-api skill.</p>\n<p>With the skill, you can run <code>/claude-api build-eval</code> to build an evaluation inside your codebase, and run <code>/claude-api hillclimb</code> to improve your application against it, one change at a time, with a held-out set of examples to catch overfitting.</p>\n<h2 id=\"eval-design\">Eval design</h2>\n<p>Well designed evaluations have a few common elements:</p>\n<ol><li><strong>Eval tasks mirror production.</strong> Sample tasks that you care about in the setting where the capability will be used.</li><li><strong>Performance improves with stronger models and more thinking.</strong> If they don&#39;t, ambiguous tasks or a miscalibrated grader often are hobbling performance.</li><li><strong>There is &quot;passable&quot; headroom at the frontier.</strong> The most capable model at the highest effort should be well below 100%.</li><li><strong>Low run-to-run variance.</strong> High variance is often due to poorly designed, ambiguous tasks or a grader that produces different verdicts on identical output.</li></ol>\n<h3 id=\"adversarial-sampling\">Adversarial sampling</h3>\n<p>Model capability is jagged. If you pick cases because today&#39;s model fails them, you are sampling the valleys of one model&#39;s capability surface. Pick hard cases because a human judged them hard.</p>\n<h2 id=\"claude-api-build-eval\"><code>/claude-api build-eval</code></h2>\n<p>When you run <code>/claude-api build-eval</code> in Claude Code, Claude interviews you, builds the eval inside your codebase, and pauses for approval at specific points. It samples inputs from production transcripts, bug reports, hand-written cases, and synthetic cases anchored in real examples. It proposes the cheapest grader that fits (programmatic verification or LLM-as-judge), validates the grader with you, runs a baseline, and prints a score with a confidence interval.</p>\n<h2 id=\"hillclimbing\">Hillclimbing</h2>\n<p>Hillclimbing is effective for tuning parameters like effort or prompts. Prefer surfaces that are cheap to iterate, attributable to the score, and well-scoped. Overfitting is common: split train/test, never paste failures into the prompt, and keep answers structurally out of the model&#39;s reach.</p>\n<h2 id=\"claude-api-hillclimb\"><code>/claude-api hillclimb</code></h2>\n<p>Claude iterates with one patch per round. If the train set improves but the test set is flat, it suspects overfitting and reverts. When the score stalls, it buckets remaining failures by cause. Cost-focused and performance-focused case studies in the original post show large gains (e.g., cutting support-ticket token cost roughly fivefold while improving held-out accuracy; lifting a docs/API skill from 66% toward ~88%).</p>\n<h2 id=\"getting-started\">Getting started</h2>\n<p>Run <code>/claude-api build-eval</code> to generate an evaluation set. Run <code>/claude-api hillclimb</code> when you have an evaluation and a goal (better performance, or lower cost while performance holds).</p>\n<p><em>Full article with figures and case studies: <a href=\"https://claude.dev/blog/automating-eval-design-and-hillclimbing/\" rel=\"nofollow ugc noopener\">claude.dev/blog/automating-eval-design-and-hillclimbing</a>.</em></p>","headings":[{"level":1,"text":"Automating eval design and hillclimbing with Claude","id":"automating-eval-design-and-hillclimbing-with-claude"},{"level":2,"text":"Eval design","id":"eval-design"},{"level":3,"text":"Adversarial sampling","id":"adversarial-sampling"},{"level":2,"text":"/claude-api build-eval","id":"claude-api-build-eval"},{"level":2,"text":"Hillclimbing","id":"hillclimbing"},{"level":2,"text":"/claude-api hillclimb","id":"claude-api-hillclimb"},{"level":2,"text":"Getting started","id":"getting-started"}]}}