This document curates the most common questions Shreya and I received while teaching 5,000+ engineers and PMs AI Evals. Warning: These are sharp opinions about what works in most cases. They are not universal truths. Use your judgment.

How to use this FAQ

Browse the questions that interest you, or choose a guide below for a curated reading path through the FAQs and related articles.

Where are you?This sounds like me
I’m new to evalsI’ve heard the term, but I’m not sure what evals involve or whether I need them.
I don’t know what to testI’m building an AI product, but I haven’t figured out which failures to measure or what good performance looks like.
I don’t trust my eval scoresWe have evals, but the scores don’t match our judgment of the outputs, or tests pass while users still encounter problems.
My product feels too hard to evaluateOur outputs are subjective, long, or involve many steps. Even a knowledgeable person has trouble deciding whether they’re right.
Evals take too much time or moneyWe’re spending too much effort reviewing outputs, maintaining tests, or running evaluators.