By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres argues that AI evals should be a core discovery habit for product teams, not just an engineering concern. Because LLM outputs are probabilistic and semantic, traditional pass/fail unit tests don't work; evals instead measure how often an AI behaves correctly. She walks through a three-step process: run error analysis to see what mistakes occur, choose a way to count each error (golden datasets, code assertions, LLM-as-a-judge, or customer feedback), and then use baselines and experiments to improve the product. The piece matters because it gives PMs a concrete, hands-on method to build confidence in AI features before they reach customers.
01Key takeaways
- Evals measure how often an AI output is correct, rather than checking a single pass or fail as unit tests do.
- Define correctness yourself, because what counts as good depends on your product's context, not a vendor's default.
- Run error analysis first and categorize mistakes before deciding which errors deserve an eval.
- Prefer cheap code assertions where possible, reserving LLM-as-a-judge for errors that truly require semantic judgment.
- Collect a baseline score on fixed inputs before changing anything, then compare each experiment variant against it.
02Key sections
- Why product teams need evals
- Evals help teams verify AI outputs for personal workflows and customer-facing products alike. Torres shows how she used them to check AI-written interview summaries and to decide whether her Interview Coach was ready for students.
- Why LLMs break traditional testing
- Deterministic code gives the same result every time, so unit tests suffice, but LLMs vary across runs and tackle tasks with many acceptable answers. Evals therefore measure a rate of correctness rather than a single pass or fail.
- Defining what good looks like
- Teams must define correctness in their own context, since vendor default definitions rarely fit. Error analysis, done by reviewing outputs and categorizing mistakes, reveals which errors matter and which need evals.
- Choosing the right eval type
- Golden datasets suit small inputs with one right answer, code assertions are cheap and fast for structural checks, and LLM-as-a-judge handles semantic judgments at higher cost. Customer feedback is the ultimate signal but is hard to make actionable.
- Using evals to improve products
- Establish a baseline by running fixed inputs through the product and scoring them, then test variants one hypothesis at a time and compare results across all error categories. Torres ties this loop back to continuous discovery.
03From the post
“AI evals tell you whether your AI product or workflow is any good. A hands-on guide to error analysis, the four types of evals, and your first experiment.”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.