Product Talk · Free post · Building AI products · Metrics, data & experimentation

Q&A Session: Building AI Evals for the Interview Coach

Teresa TorresSep 10, 202522 min
SourceProduct Talk
KindFree post
PublishedSep 10, 2025
Originalproducttalk.org ↗
N:

Teresa Torres shares answers from the Q&A of her webinar on building AI evals for the Interview Coach, a teaching tool that gives students feedback on their customer interviews. She explains her lean, hands-on tooling (plain code in VS Code, Jupyter notebooks, serverless Lambdas), how she manages per-transcript cost, and how guardrails run in production. She covers dev and holdout test sets, why evals should not hit 100%, and how she checks that her LLM judges stay aligned with human labels. The post matters because it shows a pragmatic, incremental way for product people to build and maintain AI features, treating quality as a continuous investment rather than a one-time launch.

01Key takeaways

  • Begin with tools you already know, and add complexity only when your current setup actually blocks you.
  • Use the smallest model that matches the best model's performance on each narrow eval to keep costs low.
  • Keep separate dev and holdout test sets, and make sure both reflect the real diversity of production traffic.
  • Treat evals that reach 100% as saturated, since they no longer detect errors worth measuring.
  • Regularly compare LLM judge results to human labels, because evals can be wrong and can drift over time.
  • Expect AI products to need continuous investment, since fixing one error often surfaces a new one.

02Key sections

Start with the tools you already have
Torres avoids dedicated eval platforms to avoid being overwhelmed and to keep outside tools from shaping her thinking. She writes evals in plain code and uses a notebook for analysis, choosing tools based on the immediate goal.
Cost and model selection
She cuts costs by caching long transcripts and using the smallest model that still performs as well as the best one for each judge. Simple, narrow eval questions are very cheap, while complex ones cost more.
Guardrails, production checks, and privacy
Some evals run as guardrails before responses reach users, retrying or repairing output when they fail. She limits data retention to 90 days, keeps only LLM responses for real customer interviews, and checks third-party vendors' data policies.
Dev sets, test sets, and saturation
A small dev set guides iteration while a larger holdout test set validates releases. Evals that reach 100% are no longer measuring anything useful, so the goal is a judgment-based comparison against production rather than a perfect score.
Checking the evals themselves
She compares judge outputs to human labels over time to confirm alignment, because evals can be wrong or drift. Ongoing human labeling and error analysis reveal new failure modes after each fix.
Learning and working with AI assistants
Torres treats LLMs as thought partners in pair programming, not vibe-coding black boxes, and insists on understanding every line. She advises beginners to ask simple questions and start small.

03From the post

“I am falling deeper into the AI rabbit hole every day. Now, that’s not to say that I believe all the hype about AI. What I mean is that I’m excited to see what’s just now possible with AI (some may even say I’ve become a”

“I'm not vibe coding this. I am relying a lot on Claude Code, but not in a vibe coding sense”Teresa Torres · Product Talk
“It's really important to do this on a percentage of traces over time because your evals can drift with time.”Teresa Torres · Product Talk
“I think when we build AI products, we have to commit to evolving them continuously over time because they will drift.”Teresa Torres · Product Talk

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.