By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres shares answers from the Q&A of her webinar on building AI evals for the Interview Coach, a teaching tool that gives students feedback on their customer interviews. She explains her lean, hands-on tooling (plain code in VS Code, Jupyter notebooks, serverless Lambdas), how she manages per-transcript cost, and how guardrails run in production. She covers dev and holdout test sets, why evals should not hit 100%, and how she checks that her LLM judges stay aligned with human labels. The post matters because it shows a pragmatic, incremental way for product people to build and maintain AI features, treating quality as a continuous investment rather than a one-time launch.
01Key takeaways
- Begin with tools you already know, and add complexity only when your current setup actually blocks you.
- Use the smallest model that matches the best model's performance on each narrow eval to keep costs low.
- Keep separate dev and holdout test sets, and make sure both reflect the real diversity of production traffic.
- Treat evals that reach 100% as saturated, since they no longer detect errors worth measuring.
- Regularly compare LLM judge results to human labels, because evals can be wrong and can drift over time.
- Expect AI products to need continuous investment, since fixing one error often surfaces a new one.
02Key sections
- Start with the tools you already have
- Torres avoids dedicated eval platforms to avoid being overwhelmed and to keep outside tools from shaping her thinking. She writes evals in plain code and uses a notebook for analysis, choosing tools based on the immediate goal.
- Cost and model selection
- She cuts costs by caching long transcripts and using the smallest model that still performs as well as the best one for each judge. Simple, narrow eval questions are very cheap, while complex ones cost more.
- Guardrails, production checks, and privacy
- Some evals run as guardrails before responses reach users, retrying or repairing output when they fail. She limits data retention to 90 days, keeps only LLM responses for real customer interviews, and checks third-party vendors' data policies.
- Dev sets, test sets, and saturation
- A small dev set guides iteration while a larger holdout test set validates releases. Evals that reach 100% are no longer measuring anything useful, so the goal is a judgment-based comparison against production rather than a perfect score.
- Checking the evals themselves
- She compares judge outputs to human labels over time to confirm alignment, because evals can be wrong or drift. Ongoing human labeling and error analysis reveal new failure modes after each fix.
- Learning and working with AI assistants
- Torres treats LLMs as thought partners in pair programming, not vibe-coding black boxes, and insists on understanding every line. She advises beginners to ask simple questions and start small.
03From the post
“I am falling deeper into the AI rabbit hole every day. Now, that’s not to say that I believe all the hype about AI. What I mean is that I’m excited to see what’s just now possible with AI (some may even say I’ve become a”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.