Today, I’m publishing evals-skills, a set of skills that for AI product evals <sup>[1]</sup>. They distill what I’ve learned helping 50+ companies and teaching 4,000+ students build evaluation systems.

Why Skills for Evals

Coding agents now instrument applications, run experiments, analyze data, and build interfaces. I’ve been pointing them at evals.

OpenAI’s Harness Engineering article makes the case well: they built a product entirely with Codex agents (~1 million lines of code, 1,500 PRs, three engineers, five months) and found that improving the infrastructure around the agent yielded better returns than improving the model. Their agents queried distributed traces to verify their own work against runtime evidence. Documentation tells the agent what to do. Telemetry tells it whether it worked. Evals apply the same principle to AI output quality.