Lenny’s Newsletter · Subscriber post · Building AI products · Metrics, data & experimentation

Building eval systems that improve your AI product

A practical guide to moving beyond generic scores and measuring what matters

Lenny RachitskySep 9, 202520 min
SourceLenny’s Newsletter
KindSubscriber post
PublishedSep 9, 2025
Originallennysnewsletter.com ↗
N:

Hamel Husain and Shreya Shankar outline a practical playbook for building AI evaluation systems that improve a product rather than just producing dashboards. The core argument is to start with error analysis by a domain expert, build validated LLM-as-a-judge and code-based evaluators measured by TPR/TNR, and operationalize them in CI and production as an improvement flywheel. The post also includes a promotional course offer.

Subscriber post — summary only

01Key takeaways

  • Start with error analysis on real user interactions rather than off-the-shelf metrics like hallucination or toxicity scores.
  • Designate one principal domain expert as the quality arbiter, using binary pass/fail judgments with detailed critiques.
  • Use code-based evaluators for objective rules and validated LLM judges for subjective failures, split data into train, dev, and test sets.
  • Measure judges with true positive and true negative rates rather than accuracy, since imbalanced datasets can mislead.
  • Run a golden dataset in CI as a regression safety net and use asynchronous production evals to discover new failure modes.
“Instead of reporting a hallucination or toxicity score on a dashboard, calculate the scores on your traces and sort them by high/low score.”Hamel Husain and Shreya Shankar · Lenny’s Newsletter
“Binary decisions force clarity. An output either meets the quality bar or it does not.”Hamel Husain and Shreya Shankar · Lenny’s Newsletter

02Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.