By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres recounts building an AI-powered 'interview coach' that gives students feedback on practice customer interviews, grounded in her Continuous Interviewing course rubric. She describes moving from a ChatGPT/Claude prototype to Replit, then to LMS plus Zapier, and eventually AWS Step Functions, while handling data privacy and prompt-complexity problems. The core lesson is that splitting prompts and building error-analysis-driven evals gave her a fast feedback loop to measure whether changes actually improve quality. The piece is useful to PMs considering how to build, test, and responsibly deploy LLM products in learning or discovery contexts.
01Key takeaways
- Use real traces and error analysis to discover failure categories before writing any evals.
- Validate LLM-as-judge evaluators against human grades so you can trust the judges measuring your product.
- Splitting a long prompt into per-dimension calls can improve each output but may create cross-dimension inconsistencies to watch for.
- Be transparent with users about data handling and collect explicit consent before storing transcripts, especially without compliance certifications.
- Prototype tools can outgrow their platforms quickly; plan a migration path to production infrastructure when usage and data needs grow.
- Version-control experiments in notebooks so every prompt, model, or parameter change can be compared and rolled back.
02Key sections
- Motivation and the coaching concept
- Torres uses generative AI as a thought partner and coach rather than a replacement for human interviewing. Her goal is deliberate practice, with AI modeling what good feedback looks like.
- Rapid prototyping and deployment hurdles
- An initial custom GPT/Claude prototype worked surprisingly well, but access control, rate limiting, and LMS embedding pushed her toward Replit, which she later outgrew.
- Launch learnings and prompt redesign
- Real students exposed rate-limit problems, unexpected customer data uploads, and rubric confusion from an overlong prompt. Splitting each dimension into its own LLM call improved precision but introduced inconsistency across dimensions.
- Evals and error analysis
- Torres learned that evals are essentially a product in themselves, starting with error analysis on annotated traces and building code-based and LLM-as-judge evaluators validated against human grading.
- Scaling, partnership, and data handling
- To avoid SOC 2 burdens, she partners with Vistaly for real customer transcripts, runs a controlled beta, and orchestrates calls through Step Functions with cost-conscious eval cascades.
03From the post
“I’ve been diving deep into how generative AI can help us build new skills. I’m firmly in the “AI should augment (not replace) humans at work” camp. I’m concerned about the short-term impact on entry-level employees. I’m concerned about brain rot in the rest”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.