By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres recounts a 90-day, self-directed journey building her first AI product, an Interview Coach that gives feedback on customer interviews, which she built while recovering from a broken ankle. She argues that AI products should start from real customer opportunities where AI can genuinely help, and that a product's value depends on experimentation and rigorous evaluation rather than a one-time build. She walks through prototyping with off-the-shelf LLM tools, testing against real transcripts, moving from a chat interface to a workflow architecture, and building evals based on error analysis. The essay closes with a sober look at data ethics and transparency once real customer data entered the picture, making it a practical guide for PMs building AI features.
01Key takeaways
- Start from real customer opportunities that were previously hard to solve, and where you hold unique domain expertise, rather than adding AI features for their own sake.
- Test a prototype against your own judgment on real examples to see where the model misses things and where it adds value.
- Choose a workflow over an agent when the process is deterministic, and break complex evaluation tasks into smaller, single-purpose steps.
- Build evals grounded in error analysis of real traces before scaling, since they are how you know whether the product is any good.
- Be transparent about data collection up front, get explicit consent, delete data you do not need, or partner with compliant infrastructure.
02Key sections
- Recognizing AI-shaped problems
- Torres advises starting from customer opportunities that were hard to solve before, and where her own domain expertise gives an edge, rather than chasing AI for its own sake.
- Prototyping and choosing test users
- She prototyped in web-based LLM projects, compared Claude and ChatGPT against her own opportunity synthesis, and tested with a small, opt-in alumni group.
- Architecture and user experience
- She moved away from chat toward a homework-triggered workflow with separate prompts per rubric dimension, then to a Step Function for reliable error handling, noting agents and RAG as likely next steps.
- Consistency through evals
- Splitting prompts created contradictory feedback, which led her to build code-based and LLM-as-Judge evals grounded in error analysis of traces.
- Continuous investment and experimentation
- Weekly trace review, eval updates, and small A/B experiments drove a dramatic drop in one error rate, showing AI product work never truly finishes.
- Data responsibility
- Students submitted real customer transcripts with sensitive data, prompting consent gates, 90-day deletion, and ultimately a partnership with a SOC 2 compliant platform.
03From the post
“From broken ankle to breakthrough: How I built my first production AI app in 90 days. Real lessons on architecture, testing, and continuous improvement.”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.