Read the original at Lenny’s Newsletter ↗lennysnewsletter.com · subscriber post
By Lenny Rachitsky · lennysnewsletter.com · @lennysan on X · YouTube · LinkedIn
Hamel Husain and Shreya Shankar outline a practical playbook for building AI evaluation systems that improve a product rather than just producing dashboards. The core argument is to start with error analysis by a domain expert, build validated LLM-as-a-judge and code-based evaluators measured by TPR/TNR, and operationalize them in CI and production as an improvement flywheel. The post also includes a promotional course offer.
Subscriber post — summary only01Key takeaways
- Start with error analysis on real user interactions rather than off-the-shelf metrics like hallucination or toxicity scores.
- Designate one principal domain expert as the quality arbiter, using binary pass/fail judgments with detailed critiques.
- Use code-based evaluators for objective rules and validated LLM judges for subjective failures, split data into train, dev, and test sets.
- Measure judges with true positive and true negative rates rather than accuracy, since imbalanced datasets can mislead.
- Run a golden dataset in CI as a regression safety net and use asynchronous production evals to discover new failure modes.
“Instead of reporting a hallucination or toxicity score on a dashboard, calculate the scores on your traces and sort them by high/low score.”Hamel Husain and Shreya Shankar · Lenny’s Newsletter
“Binary decisions force clarity. An output either meets the quality bar or it does not.”Hamel Husain and Shreya Shankar · Lenny’s Newsletter
02Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.