Product Talk · Free post · Building AI products · Metrics, data & experimentation

AI Evals & Discovery - All Things Product Podcast with Teresa Torres & Petra Wille

Teresa TorresSep 23, 2025
SourceProduct Talk
KindFree post
PublishedSep 23, 2025
Originalproducttalk.org ↗
N:

Teresa Torres and Petra Wille explore AI evals: what they mean, why they matter beyond quality assurance, and how teams can build them. Drawing on Teresa's experience building the Interview Coach tool, the conversation covers golden datasets, synthetic data, real-world traces, and turning observed error modes into evals. It contrasts code-based checks with LLM-as-judge methods, shows how discovery practice feeds evaluation, and stresses that evals need ongoing maintenance because criteria drift over time. The episode frames evals as a continuing loop of measurement, guardrails, and human oversight rather than a one-time test.

01Key takeaways

  • Treat evals as a measurement system for whether your AI product works, not just a QA checklist.
  • Start from real traces and observed failures, then turn recurring error modes into specific evals.
  • Use a golden dataset for stable benchmarking and synthetic data to expand coverage, knowing each has limits.
  • Use code-based checks for objective, rule-like criteria and LLM-as-judge for subjective quality dimensions.
  • Ground evaluation criteria in discovery work so you measure what customers actually care about.
  • Revisit and maintain evals regularly, because criteria drift as users, models, and expectations change.

02Key sections

What evals are
The episode defines evals in the AI/ML sense and separates them from traditional QA. Evals measure whether an AI product's outputs are actually good.
Building the Interview Coach
Teresa walks through the hard lessons from building her Interview Coach tool and how they shaped her evaluation approach.
Datasets and error modes
Golden datasets, synthetic data, and real-world traces are compared, along with how to find recurring error modes and convert them into evals.
Code-based vs. LLM-as-judge evals
The discussion explains when deterministic code checks are the right tool and when an LLM should judge output quality.
Maintenance and criteria drift
Evals need continuous upkeep, since what counts as good output shifts over time, and human oversight remains part of the system.

03From the post

“Listen to this episode on: Spotify | Apple Podcasts Building AI products isn’t just about clever prompts and orchestration—it’s about knowing if what you’ve built actually works. In this episode, Teresa Torres and Petra Wille dive deep into AI evals: how they’re defined, why they’re”

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.