By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres and Hamel Husain discuss how to evaluate and debug AI products with a scientist's mindset. Hamel draws on his machine learning background at Airbnb and GitHub, and his consulting work with startups like Nurture Boss, to show why error analysis should come before building elaborate tooling. The conversation covers data leakage, using synthetic data to probe failure modes, choosing between code-based assertions and LLM-as-judge evals, and keeping broken cases in regression test sets. For product teams, the core message is that reliable AI features come from systematically finding and prioritizing failures rather than from intuition about model quality.
01Key takeaways
- Treat AI debugging as a scientific process: form hypotheses about failures and test them against evidence.
- Watch for data leakage, which can make model performance look better than it really is.
- Use synthetic data to deliberately probe failure modes that production traffic rarely surfaces.
- Prefer code-based assertions for objectively checkable behavior and reserve LLM-as-judge for subjective judgments.
- Keep known broken cases in your CI/CD eval set so regressions and fixes stay visible.
- Prioritize failure modes by impact rather than trying to fix every error at once.
02Key sections
- Debugging as scientific thinking
- Hamel frames AI debugging as a scientific practice rooted in forming hypotheses and testing them against data. He argues that teams should understand their models and failure patterns before reaching for tooling.
- Lessons from machine learning past
- Hamel recounts forecasting guest lifetime value at Airbnb to illustrate how data leakage can quietly corrupt model results. The example shows why careful validation matters.
- Case study with Nurture Boss
- The discussion turns to an AI-native assistant for apartment complexes, showing how real product conversations expose failure modes that must be found and categorized.
- Synthetic data and error analysis
- Synthetic inputs can stress-test failure modes that rarely appear in production logs. Error analysis then turns observed failures into a prioritized, documented list.
- Evals: assertions and LLM-as-judge
- Deterministic code-based assertions suit checkable behavior, while LLM-as-judge evals handle subjective qualities. Both should feed a continuous improvement loop with broken cases kept in the test set.
03From the post
“Listen to this episode on: Spotify | Apple Podcasts How do you know if your AI product is actually any good? Hamel Husain has been answering that question for over 25 years. As a former machine learning engineer and data scientist at Airbnb and GitHub (where he worked on research that”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.