Teresa Torres argues that AI evals should be a core discovery habit for product teams, measuring how often an AI behaves correctly rather than running pass/fail checks1. Start with error analysis on real traces, then pick the simplest eval that works17.
01Start from observed failures
- Torres says to run error analysis first and categorize mistakes before deciding which errors deserve an eval1.
- She warns that off-the-shelf eval tools measure generic things, so your product's specific failures need to surface first7.
- Torres and Petra Wille suggest turning recurring error modes seen in real traces into specific evals4.
02Choose the cheapest eval that works
- Torres prefers code assertions where possible because they are fast and cheap, using golden datasets for smaller tasks and LLM-as-Judge only when semantic judgment is required1.
- Aman Khan recommends choosing an approach per step: human feedback for ground truth, code checks for objective rules, and LLM judges for scalable subjective grading8.
- Start with 10 to 100 human-labeled examples and aim for roughly 90% agreement before trusting an LLM judge8.
03Use evals as an ongoing loop
- Torres describes evals as a fast loop: analyze traces, measure failure modes, improve, and repeat, including A/B testing prompt and model changes7.
- Evals need upkeep because criteria drift over time4.
- Marty Cagan and Marily Nika note that PMs should define acceptable error rates before launch3.
“Evaluating AI systems is less like traditional software testing and more like giving someone a driving test.”Aman Khan · Lenny’s Newsletter
Written by PM Atlas from the cited notes only, drawing on Teresa Torres, Marty Cagan, Lenny Rachitsky. Quotes are short excerpts; read the originals for the full argument.