{"article":{"slug":"evals-skills-for-coding-agents","title":"Evals Skills for Coding Agents","subtitle":null,"summary":"Hamel Husain publishes evals-skills—agent skills for AI product evaluation covering audit, error analysis, synthetic data, judge prompts, evaluator validation, and RAG evals, distilled from work with dozens of companies.","content_type":"blog_post","language":"en","canonical_url":"https://hamelhusain.substack.com/p/evals-skills-for-coding-agents","author":{"name":"Hamel Husain","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Hamel Husain","url":"https://hamelhusain.substack.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":585,"reading_minutes":3,"published_at":"2026-09-22T21:21:27.264Z","added_at":"2026-09-22T21:21:27.264Z","updated_at":"2026-09-22T21:21:27.264Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/evals-skills-for-coding-agents","markdown_url":"https://listedarticles.com/articles/evals-skills-for-coding-agents.md","example":false,"citation":"Hamel Husain, Hamel Husain. \"Evals Skills for Coding Agents.\" 22 Sept 2026. https://hamelhusain.substack.com/p/evals-skills-for-coding-agents (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://hamelhusain.substack.com/p/evals-skills-for-coding-agents"},"body_markdown":"Today, I’m publishing [evals-skills](https://github.com/hamelsmu/evals-skills), a set of skills that for AI product evals <sup>[1]</sup>. They distill what I’ve learned helping 50+ companies and teaching 4,000+ students build evaluation systems.\n\n## **Why Skills for Evals**\n\nCoding agents now instrument applications, run experiments, analyze data, and build interfaces. I’ve been pointing them at evals.\n\nOpenAI’s Harness Engineering [article](https://openai.com/index/harness-engineering/) makes the case well: they built a product entirely with Codex agents (~1 million lines of code, 1,500 PRs, three engineers, five months) and found that **improving the infrastructure around the agent** yielded better returns than improving the model. Their agents queried distributed traces to verify their own work against runtime evidence. Documentation tells the agent what to do. Telemetry tells it whether it worked. Evals apply the same principle to AI output quality.\n\nAll major eval vendors now ship an MCP server <sup>[2]</sup>. The tedious parts: instrumenting your app, orchestrating experiments, building annotation tools, now belong to coding agents.\n\nBut an agent with an eval platform still needs to know what to do with it. Say a support bot tells a customer “your plan includes free returns” when it doesn’t. Another says “I’ve canceled your order” when nobody asked. Both are hallucinations, but one gets a fact wrong and the other makes up a user action. If you lump them together in a generic “hallucination score” real problems will likely go undetected.\n\nThese skills fill in the gaps. They complement the vendor MCP servers: those give your agent access to traces and experiments, these teach it what to do with them.\n\n## **The Skills**\n\nIf you’re new to evals or inheriting an existing eval pipeline, start with **eval-audit**. It inspects your current setup (or lack of one), runs diagnostic checks across six areas (error analysis, evaluator design, judge validation, human review, labeled data, pipeline hygiene), and produces a prioritized list of problems with next steps. Install the skills and give your agent this prompt:\n\nInstall the eval skills plugin from https://github.com/hamelsmu/evals-skills, then run /evals-skills:eval-audit on my eval pipeline. Investigate each diagnostic area using a separate subagent in parallel, then synthesize the findings into a single report. Use other skills in the plugin as recommended by the audit.\n\n\nIf you’re experienced with evals, you can skip the audit and pick the skill you need:\n\n- **error-analysis** : Read traces, categorize failures, build a vocabulary of what’s broken\n- **generate-synthetic-data** : Create diverse test inputs when real data is sparse\n- **write-judge-prompt** : Design binary Pass/Fail LLM-as-Judge evaluators\n- **validate-evaluator** : Calibrate judges against human labels using TPR/TNR and bias correction\n- **evaluate-rag** : Evaluate retrieval and generation quality separately\n- **build-review-interface** : Generate annotation interfaces for human trace review\n\nThese skills are a starting point. They only cover parts of evals that generalize across projects. Skills grounded in your stack, your domain, and your data will outperform them. Start here, then write your own. Our [AI Evals course](https://maven.com/parlance-labs/evals) teaches the end to end workflow.\n\n👉 The repo is here: [github.com/hamelsmu/evals-skills](https://github.com/hamelsmu/evals-skills) 👈\n\n*P.S. Our next* [AI Evals course](https://maven.com/parlance-labs/evals?promoCode=newsletter-25) *cohort starts March 16th. The one after that won’t run until fall. You’ll get lifetime access to all materials and learn with students from OpenAI, Meta, Google, Walmart, Airbnb and more. Use code* [](https://maven.com/parlance-labs/evals?promoCode=newsletter-25)[newsletter-25](https://maven.com/parlance-labs/evals?promoCode=newsletter-25)[](https://maven.com/parlance-labs/evals?promoCode=newsletter-25) *for a discount.*\n\n##### **Footnotes**\n\n1. Not foundation model benchmarks like MMLU or HELM that measure general LLM capabilities. Product evals measure whether *your* pipeline works on*your* task with*your* data. If you aren’t familiar with product-specific AI evals, check out my[AI Evals FAQ](https://hamel.dev/blog/posts/evals-faq/) .\n2. [Braintrust](https://www.braintrust.dev/docs/reference/mcp) ,[LangSmith](https://github.com/langchain-ai/langsmith-mcp-server) ,[Phoenix](https://github.com/Arize-ai/phoenix/tree/main/js/packages/phoenix-mcp) ,[Truesight](https://truesight.goodeyelabs.com/docs/mcp-integration) , and others.","body_html":"<p>Today, I’m publishing <a href=\"https://github.com/hamelsmu/evals-skills\" rel=\"nofollow ugc noopener\">evals-skills</a>, a set of skills that for AI product evals &lt;sup&gt;[1]&lt;/sup&gt;. They distill what I’ve learned helping 50+ companies and teaching 4,000+ students build evaluation systems.</p>\n<h2 id=\"why-skills-for-evals\"><strong>Why Skills for Evals</strong></h2>\n<p>Coding agents now instrument applications, run experiments, analyze data, and build interfaces. I’ve been pointing them at evals.</p>\n<p>OpenAI’s Harness Engineering <a href=\"https://openai.com/index/harness-engineering/\" rel=\"nofollow ugc noopener\">article</a> makes the case well: they built a product entirely with Codex agents (~1 million lines of code, 1,500 PRs, three engineers, five months) and found that <strong>improving the infrastructure around the agent</strong> yielded better returns than improving the model. Their agents queried distributed traces to verify their own work against runtime evidence. Documentation tells the agent what to do. Telemetry tells it whether it worked. Evals apply the same principle to AI output quality.</p>\n<p>All major eval vendors now ship an MCP server &lt;sup&gt;[2]&lt;/sup&gt;. The tedious parts: instrumenting your app, orchestrating experiments, building annotation tools, now belong to coding agents.</p>\n<p>But an agent with an eval platform still needs to know what to do with it. Say a support bot tells a customer “your plan includes free returns” when it doesn’t. Another says “I’ve canceled your order” when nobody asked. Both are hallucinations, but one gets a fact wrong and the other makes up a user action. If you lump them together in a generic “hallucination score” real problems will likely go undetected.</p>\n<p>These skills fill in the gaps. They complement the vendor MCP servers: those give your agent access to traces and experiments, these teach it what to do with them.</p>\n<h2 id=\"the-skills\"><strong>The Skills</strong></h2>\n<p>If you’re new to evals or inheriting an existing eval pipeline, start with <strong>eval-audit</strong>. It inspects your current setup (or lack of one), runs diagnostic checks across six areas (error analysis, evaluator design, judge validation, human review, labeled data, pipeline hygiene), and produces a prioritized list of problems with next steps. Install the skills and give your agent this prompt:</p>\n<p>Install the eval skills plugin from <a href=\"https://github.com/hamelsmu/evals-skills\" rel=\"nofollow ugc noopener\">https://github.com/hamelsmu/evals-skills</a>, then run /evals-skills:eval-audit on my eval pipeline. Investigate each diagnostic area using a separate subagent in parallel, then synthesize the findings into a single report. Use other skills in the plugin as recommended by the audit.</p>\n<p>If you’re experienced with evals, you can skip the audit and pick the skill you need:</p>\n<ul><li><strong>error-analysis</strong> : Read traces, categorize failures, build a vocabulary of what’s broken</li><li><strong>generate-synthetic-data</strong> : Create diverse test inputs when real data is sparse</li><li><strong>write-judge-prompt</strong> : Design binary Pass/Fail LLM-as-Judge evaluators</li><li><strong>validate-evaluator</strong> : Calibrate judges against human labels using TPR/TNR and bias correction</li><li><strong>evaluate-rag</strong> : Evaluate retrieval and generation quality separately</li><li><strong>build-review-interface</strong> : Generate annotation interfaces for human trace review</li></ul>\n<p>These skills are a starting point. They only cover parts of evals that generalize across projects. Skills grounded in your stack, your domain, and your data will outperform them. Start here, then write your own. Our <a href=\"https://maven.com/parlance-labs/evals\" rel=\"nofollow ugc noopener\">AI Evals course</a> teaches the end to end workflow.</p>\n<p>👉 The repo is here: <a href=\"https://github.com/hamelsmu/evals-skills\" rel=\"nofollow ugc noopener\">github.com/hamelsmu/evals-skills</a> 👈</p>\n<p><em>P.S. Our next</em> <a href=\"https://maven.com/parlance-labs/evals?promoCode=newsletter-25\" rel=\"nofollow ugc noopener\">AI Evals course</a> <em>cohort starts March 16th. The one after that won’t run until fall. You’ll get lifetime access to all materials and learn with students from OpenAI, Meta, Google, Walmart, Airbnb and more. Use code</em> <a href=\"https://maven.com/parlance-labs/evals?promoCode=newsletter-25\" rel=\"nofollow ugc noopener\"></a><a href=\"https://maven.com/parlance-labs/evals?promoCode=newsletter-25\" rel=\"nofollow ugc noopener\">newsletter-25</a><a href=\"https://maven.com/parlance-labs/evals?promoCode=newsletter-25\" rel=\"nofollow ugc noopener\"></a> <em>for a discount.</em></p>\n<h5 id=\"footnotes\"><strong>Footnotes</strong></h5>\n<ol><li>Not foundation model benchmarks like MMLU or HELM that measure general LLM capabilities. Product evals measure whether <em>your</em> pipeline works on<em>your</em> task with<em>your</em> data. If you aren’t familiar with product-specific AI evals, check out my<a href=\"https://hamel.dev/blog/posts/evals-faq/\" rel=\"nofollow ugc noopener\">AI Evals FAQ</a> .</li><li><a href=\"https://www.braintrust.dev/docs/reference/mcp\" rel=\"nofollow ugc noopener\">Braintrust</a> ,<a href=\"https://github.com/langchain-ai/langsmith-mcp-server\" rel=\"nofollow ugc noopener\">LangSmith</a> ,<a href=\"https://github.com/Arize-ai/phoenix/tree/main/js/packages/phoenix-mcp\" rel=\"nofollow ugc noopener\">Phoenix</a> ,<a href=\"https://truesight.goodeyelabs.com/docs/mcp-integration\" rel=\"nofollow ugc noopener\">Truesight</a> , and others.</li></ol>","headings":[{"level":2,"text":"**Why Skills for Evals**","id":"why-skills-for-evals"},{"level":2,"text":"**The Skills**","id":"the-skills"}]}}