💭 "I wish I could click on this card and be like, 'Clean this up.'"
When a customer pointed out a flat, unstructured branch in her AI-generated opportunity solution tree, it sparked a three-week journey to fix the problem at its source. Here's what you'll learn from this article:
How one customer complaint led to building 4 new AI evaluation metrics and testing 16 different experiment variations to get to the root cause—and why a quick fix wasn't the answer.
Key takeaways:
🔍 Building reliable AI products requires more than just prompt engineering—you need evals, guardrails, and orchestration working together
📊 Calibrating your AI judges against production data is harder than it sounds; your judges can fail in ways you don't expect
⚖️ Fixing upstream errors (like poorly framed parents) is essential, but it can accidentally make downstream errors worse
🔄 Sometimes the best solution isn't tweaking prompts endlessly—it's building an agent that audits its own work and fixes mistakes
🎯 Sweating the details matters because the real benefit of an opportunity solution tree is knowing exactly what to work on next
Read the article:
❓ What's one quality issue you've noticed in AI outputs that seemed simple to fix but turned out to be more complicated than expected? Share your thoughts in the comments below.Teresa Torres on X · Sep 16, 2026
Case study on debugging an AI-generated opportunity solution tree, building evals and an audit agent after a customer flagged poor structure.
01Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.