Lenny’s Podcast · Podcast episode · Metrics, data & experimentation · Building AI products

Why AI evals are the hottest new skill for product builders

Lenny Rachitskywith Hamel Husain & Shreya ShankarSep 25, 2025131 min♥ 151
SourceLenny’s Podcast
KindPodcast episode
PublishedSep 25, 2025
Readers♥ 151
Originallennysnewsletter.com ↗
N:

By Lenny Rachitsky · with Hamel Husain & Shreya Shankar · lennysnewsletter.com · @lennysan on X · YouTube · LinkedIn

Hamel Husain and Shreya Shankar, creators of a leading online course on AI evaluations, explain what evals are and walk through a live error-analysis process on a real estate AI assistant. They cover common misconceptions, the role of the product person, and how to turn observed failures into automated checks and LLM-as-judge evaluators.

01Key learnings

  • Evals are systematic ways to measure and improve an AI application, essentially data analytics on LLM behavior, replacing guesswork and vibe checks as the product grows.
  • Start with error analysis by manually reviewing around 100 sampled traces and writing short notes on the first upstream failure you see, rather than jumping straight to writing tests.
  • Product people with domain expertise should lead open coding; appoint one 'benevolent dictator' whose judgment you trust instead of running a committee.
  • Use an LLM to cluster your open-coded notes into failure-mode categories, but review and refine those categories yourself, and avoid vague labels like 'janky'.
  • Count failure categories with a simple pivot table to prioritize which problems matter most, and fix obvious prompt or engineering bugs before building any eval.
  • Build LLM-as-judge evaluators for subjective failure modes using binary pass/fail decisions, and validate them against your human labels with a confusion matrix; be wary of raw agreement percentages.
  • Reuse evaluators in unit tests and production monitoring, and build simple tools to make reviewing data frictionless; the whole process is high-ROI and takes about a week up front.
  • Do not rely on an AI tool to evaluate your product out of the box; domain context is needed to spot issues like hallucinated features.
“The goal is not to do evals perfectly, it's to actionably improve your product.”Shreya Shankar · Lenny’s Podcast · 00:00:18
“You can appoint one person whose taste that you trust. It should be the person with domain expertise.”Hamel Husain · Lenny’s Podcast · 00:01:00
“It's the highest ROI activity you can engage in.”Hamel Husain · Lenny’s Podcast · 00:30:00

02Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.