Product Talk · Free post · Building AI products · Metrics, data & experimentation

Behind the Scenes: Building the Product Talk Interview Coach

Teresa TorresAug 6, 202547 min
SourceProduct Talk
KindFree post
PublishedAug 6, 2025
Originalproducttalk.org ↗
N:

Teresa Torres recounts building an AI-powered 'interview coach' that gives students feedback on practice customer interviews, grounded in her Continuous Interviewing course rubric. She describes moving from a ChatGPT/Claude prototype to Replit, then to LMS plus Zapier, and eventually AWS Step Functions, while handling data privacy and prompt-complexity problems. The core lesson is that splitting prompts and building error-analysis-driven evals gave her a fast feedback loop to measure whether changes actually improve quality. The piece is useful to PMs considering how to build, test, and responsibly deploy LLM products in learning or discovery contexts.

01Key takeaways

  • Use real traces and error analysis to discover failure categories before writing any evals.
  • Validate LLM-as-judge evaluators against human grades so you can trust the judges measuring your product.
  • Splitting a long prompt into per-dimension calls can improve each output but may create cross-dimension inconsistencies to watch for.
  • Be transparent with users about data handling and collect explicit consent before storing transcripts, especially without compliance certifications.
  • Prototype tools can outgrow their platforms quickly; plan a migration path to production infrastructure when usage and data needs grow.
  • Version-control experiments in notebooks so every prompt, model, or parameter change can be compared and rolled back.

02Key sections

Motivation and the coaching concept
Torres uses generative AI as a thought partner and coach rather than a replacement for human interviewing. Her goal is deliberate practice, with AI modeling what good feedback looks like.
Rapid prototyping and deployment hurdles
An initial custom GPT/Claude prototype worked surprisingly well, but access control, rate limiting, and LMS embedding pushed her toward Replit, which she later outgrew.
Launch learnings and prompt redesign
Real students exposed rate-limit problems, unexpected customer data uploads, and rubric confusion from an overlong prompt. Splitting each dimension into its own LLM call improved precision but introduced inconsistency across dimensions.
Evals and error analysis
Torres learned that evals are essentially a product in themselves, starting with error analysis on annotated traces and building code-based and LLM-as-judge evaluators validated against human grading.
Scaling, partnership, and data handling
To avoid SOC 2 burdens, she partners with Vistaly for real customer transcripts, runs a controlled beta, and orchestrates calls through Step Functions with cost-conscious eval cascades.

03From the post

“I’ve been diving deep into how generative AI can help us build new skills. I’m firmly in the “AI should augment (not replace) humans at work” camp. I’m concerned about the short-term impact on entry-level employees. I’m concerned about brain rot in the rest”

“It turns out all the work that we do to teach humans is very applicable to teaching an LLM.”Teresa Torres · Product Talk
“I'm a discovery coach. What is discovery about? It's about fast feedback loops.”Teresa Torres · Product Talk
“The biggest thing that's hard about evals is when you do code-based evals and LLM as judge evals, your evals are almost a product in…”Teresa Torres · Product Talk

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.