Product Talk · Free post · Building AI products · Metrics, data & experimentation

How I Designed & Implemented Evals for Product Talk’s Interview Coach

Teresa TorresSep 3, 202533 min
SourceProduct Talk
KindFree post
PublishedSep 3, 2025
Originalproducttalk.org ↗
N:

Teresa Torres describes building Product Talk's Interview Coach, an AI tool that gives students feedback on their customer interviews, and how she designed evals to measure its quality. She explains the core idea of evals as a fast feedback loop: analyze traces, measure common failure modes, improve, and repeat. The key lesson is that error analysis must come first, since off-the-shelf eval tools measure generic things rather than the failures specific to your product. She shares how she moved from human labeling and simple code assertions to LLM-as-Judge evals, and how the process even sharpened her own teaching rubric. The write-up is useful because it shows a PM-led, cross-functional, and honest approach to shipping AI products that stay improving over time.

01Key takeaways

  • Do error analysis on real traces before choosing what to measure, because generic eval tools rarely catch your product's specific failures.
  • Start with the simplest eval that might work, such as keyword matching, before building complex LLM-as-Judge systems.
  • Compare eval outputs to human labels so your automated judges stay aligned with real judgment.
  • Use evals as a fast feedback loop to A/B test prompt, model, and temperature changes before releasing.
  • Expect writing evals to sharpen your own domain rubric, since deterministic checks can expose vague human intuitions.
  • Treat evals as continuous work, since AI products drift and need ongoing investment to keep quality.

02Key sections

Why build the Interview Coach
Torres wanted students to get better feedback on practice interviews, since deliberate practice depends on expert feedback. The Coach was an attempt to model what good feedback looks like across four rubric dimensions.
What evals are and how they work
She frames evals as the way to judge whether an AI product is good, analogous to unit and integration tests. Three approaches are covered: datasets with expected outputs, code-based assertions, and LLM-as-Judge.
Error analysis comes first
Reviewing traces and tagging recurring failure modes tells you what to measure. Her two most serious failures, the Coach suggesting leading or general questions, became her first evals.
Building and scoring evals
She wrote simple evals, started with the dumbest one that might work, and compared eval outputs against human labels. Writing evals also revealed bugs and clarified her own rubric.
Where the process stands now
Evals now run on dev and test sets for A/B comparison of prompt, model, and temperature changes, with custom annotation tools and cross-functional collaboration.

03From the post

“In this webinar, I dive deep on how I built code-based and LLM-as-judge evals for Product Talk’s Interview Coach.”

“You have to do the error analysis to figure out what to measure in the first place.”Teresa Torres · Product Talk
“Start small and build on what you know.”Teresa Torres · Product Talk
“Don’t forget: This is continuous work.”Teresa Torres · Product Talk

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.