Lenny’s Newsletter · Subscriber post · Building AI products · Metrics, data & experimentation

Beyond vibe checks: A PM’s complete guide to evals

How to master the emerging skill that can make or break an AI product

Lenny RachitskyApr 8, 202515 min♥ 579
SourceLenny’s Newsletter
KindSubscriber post
PublishedApr 8, 2025
Readers♥ 579
Originallennysnewsletter.com ↗
N:

Aman Khan argues that writing evaluations is becoming a defining skill for AI product managers, since evals measure how each part of a system affects quality. The post explains what evals are, compares human, code-based and LLM-judge approaches, outlines a four-part eval structure, and walks through an iterative workflow from data collection to production monitoring.

Subscriber post — summary only

01Key takeaways

  • Evals measure AI quality against defined criteria, acting like regression tests rather than simple pass/fail software checks.
  • Choose an eval approach per step: human feedback for ground truth, code checks for objective rules, LLM judges for scalable subjective grading.
  • A strong LLM-judge eval sets a role, supplies context, states the goal, and defines key terms and labels precisely.
  • Start with 10 to 100 human-labeled examples, aim for roughly 90% agreement, then iterate prompts and expand edge cases.
  • Keep early evals simple and validate them against real user feedback to avoid noisy signals that erode team trust.
“Evals are how you measure the quality and effectiveness of your AI system.”Aman Khan · Lenny’s Newsletter
“Writing good evals forces you into the shoes of your user—they are how you catch "bad" scenarios and know what to improve on.”Aman Khan · Lenny’s Newsletter

02Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.