Product Talk · Free post · Building AI products · Metrics, data & experimentation

AI Evals: A Hands-On Guide for Product Teams

Teresa TorresSep 2, 202626 min
SourceProduct Talk
KindFree post
PublishedSep 2, 2026
Originalproducttalk.org ↗
N:

Teresa Torres argues that AI evals should be a core discovery habit for product teams, not just an engineering concern. Because LLM outputs are probabilistic and semantic, traditional pass/fail unit tests don't work; evals instead measure how often an AI behaves correctly. She walks through a three-step process: run error analysis to see what mistakes occur, choose a way to count each error (golden datasets, code assertions, LLM-as-a-judge, or customer feedback), and then use baselines and experiments to improve the product. The piece matters because it gives PMs a concrete, hands-on method to build confidence in AI features before they reach customers.

01Key takeaways

  • Evals measure how often an AI output is correct, rather than checking a single pass or fail as unit tests do.
  • Define correctness yourself, because what counts as good depends on your product's context, not a vendor's default.
  • Run error analysis first and categorize mistakes before deciding which errors deserve an eval.
  • Prefer cheap code assertions where possible, reserving LLM-as-a-judge for errors that truly require semantic judgment.
  • Collect a baseline score on fixed inputs before changing anything, then compare each experiment variant against it.

02Key sections

Why product teams need evals
Evals help teams verify AI outputs for personal workflows and customer-facing products alike. Torres shows how she used them to check AI-written interview summaries and to decide whether her Interview Coach was ready for students.
Why LLMs break traditional testing
Deterministic code gives the same result every time, so unit tests suffice, but LLMs vary across runs and tackle tasks with many acceptable answers. Evals therefore measure a rate of correctness rather than a single pass or fail.
Defining what good looks like
Teams must define correctness in their own context, since vendor default definitions rarely fit. Error analysis, done by reviewing outputs and categorizing mistakes, reveals which errors matter and which need evals.
Choosing the right eval type
Golden datasets suit small inputs with one right answer, code assertions are cheap and fast for structural checks, and LLM-as-a-judge handles semantic judgments at higher cost. Customer feedback is the ultimate signal but is hard to make actionable.
Using evals to improve products
Establish a baseline by running fixed inputs through the product and scoring them, then test variants one hypothesis at a time and compare results across all error categories. Torres ties this loop back to continuous discovery.

03From the post

“AI evals tell you whether your AI product or workflow is any good. A hands-on guide to error analysis, the four types of evals, and your first experiment.”

“Correctness is context dependent. Don't let a vendor define correctness for your product. This is the product team's job.”Teresa Torres · Product Talk
“The only way to know if our AI products and workflows are any good is with evals.”Teresa Torres · Product Talk
“Evals allow us to measure the rate at which the LLM returns a correct response.”Teresa Torres · Product Talk

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.