Product Talk · Free post · Building AI products · Discovery & customer research

Building My First AI Product: 6 Lessons from My 90-Day Deep Dive

Teresa TorresAug 20, 202526 min
SourceProduct Talk
KindFree post
PublishedAug 20, 2025
Originalproducttalk.org ↗
N:

Teresa Torres recounts a 90-day, self-directed journey building her first AI product, an Interview Coach that gives feedback on customer interviews, which she built while recovering from a broken ankle. She argues that AI products should start from real customer opportunities where AI can genuinely help, and that a product's value depends on experimentation and rigorous evaluation rather than a one-time build. She walks through prototyping with off-the-shelf LLM tools, testing against real transcripts, moving from a chat interface to a workflow architecture, and building evals based on error analysis. The essay closes with a sober look at data ethics and transparency once real customer data entered the picture, making it a practical guide for PMs building AI features.

01Key takeaways

  • Start from real customer opportunities that were previously hard to solve, and where you hold unique domain expertise, rather than adding AI features for their own sake.
  • Test a prototype against your own judgment on real examples to see where the model misses things and where it adds value.
  • Choose a workflow over an agent when the process is deterministic, and break complex evaluation tasks into smaller, single-purpose steps.
  • Build evals grounded in error analysis of real traces before scaling, since they are how you know whether the product is any good.
  • Be transparent about data collection up front, get explicit consent, delete data you do not need, or partner with compliant infrastructure.

02Key sections

Recognizing AI-shaped problems
Torres advises starting from customer opportunities that were hard to solve before, and where her own domain expertise gives an edge, rather than chasing AI for its own sake.
Prototyping and choosing test users
She prototyped in web-based LLM projects, compared Claude and ChatGPT against her own opportunity synthesis, and tested with a small, opt-in alumni group.
Architecture and user experience
She moved away from chat toward a homework-triggered workflow with separate prompts per rubric dimension, then to a Step Function for reliable error handling, noting agents and RAG as likely next steps.
Consistency through evals
Splitting prompts created contradictory feedback, which led her to build code-based and LLM-as-Judge evals grounded in error analysis of traces.
Continuous investment and experimentation
Weekly trace review, eval updates, and small A/B experiments drove a dramatic drop in one error rate, showing AI product work never truly finishes.
Data responsibility
Students submitted real customer transcripts with sensitive data, prompting consent gates, 90-day deletion, and ultimately a partnership with a SOC 2 compliant platform.

03From the post

“From broken ankle to breakthrough: How I built my first production AI app in 90 days. Real lessons on architecture, testing, and continuous improvement.”

“Claude is not ready to replace me when it comes to synthesizing interviews.”Teresa Torres · Product Talk
“Observability (being able to see your traces) is everything. But to do this ethically, we have to be transparent with our customers about this.”Teresa Torres · Product Talk
“Assume users will use your tool in unexpected ways.”Teresa Torres · Product Talk

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.