Q · Building AI products · answered from 5 notes

Q:How should product teams use AI evals?

Notes cited5
People3
UpdatedOct 8, 2026
A:

Teresa Torres argues that AI evals should be a core discovery habit for product teams, measuring how often an AI behaves correctly rather than running pass/fail checks1. Start with error analysis on real traces, then pick the simplest eval that works17.

01Start from observed failures

  • Torres says to run error analysis first and categorize mistakes before deciding which errors deserve an eval1.
  • She warns that off-the-shelf eval tools measure generic things, so your product's specific failures need to surface first7.
  • Torres and Petra Wille suggest turning recurring error modes seen in real traces into specific evals4.

02Choose the cheapest eval that works

  • Torres prefers code assertions where possible because they are fast and cheap, using golden datasets for smaller tasks and LLM-as-Judge only when semantic judgment is required1.
  • Aman Khan recommends choosing an approach per step: human feedback for ground truth, code checks for objective rules, and LLM judges for scalable subjective grading8.
  • Start with 10 to 100 human-labeled examples and aim for roughly 90% agreement before trusting an LLM judge8.

03Use evals as an ongoing loop

  • Torres describes evals as a fast loop: analyze traces, measure failure modes, improve, and repeat, including A/B testing prompt and model changes7.
  • Evals need upkeep because criteria drift over time4.
  • Marty Cagan and Marily Nika note that PMs should define acceptable error rates before launch3.
“Evaluating AI systems is less like traditional software testing and more like giving someone a driving test.”Aman Khan · Lenny’s Newsletter

Written by PM Atlas from the cited notes only, drawing on Teresa Torres, Marty Cagan, Lenny Rachitsky. Quotes are short excerpts; read the originals for the full argument.

Q:
Answers quote and cite the source notes