By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres interviews SallyAnn DeLucia and Jack Zhou of Arize AI about building Alyx, the AI agent that helps teams debug, optimize, and evaluate AI applications. The team describes scrappy beginnings in Jupyter notebooks and hacked-together apps, refined through weekly dogfooding sessions with customer success. They explain how evals started messy and were layered across tool calls, sessions, and system-level decisions. The conversation argues that cross-functional, boundary-spanning teams and customer insight matter as much as technical skill when shipping AI products. It is most useful for PMs trying to move from ad hoc prototypes and vibe checks to systematic AI product improvement.
01Key takeaways
- Start prototyping in messy tools like notebooks; speed of learning matters more than polish early on.
- Use your own product daily (dogfooding) to find gaps that scripted testing will never reveal.
- Involve customer-facing staff, since they spot repeatable workflows customers actually need.
- Begin evals imperfectly and iterate, layering them from tool calls up to whole-system decisions.
- Analyze failure cases systematically before designing evals, rather than relying on vibe checks.
- Favor cross-functional, boundary-spanning team members who can move between engineering, product, and customers.
02Key sections
- Platform and core concepts
- The guests frame tracing, observability, and evals as the foundation for understanding GenAI application behavior. These concepts set up why the team could use its own product to build an agent.
- Genesis of Alyx
- Alyx began as experiments in notebooks and local web apps rather than a planned product. Its evolution was shaped by what the team learned in practice.
- Dogfooding with customer success
- Customer success engineers helped surface repeatable workflows, and weekly internal testing built confidence before launch. Using the product daily exposed gaps that synthetic testing missed.
- Evals that start messy
- The team began evaluating without polished criteria and gradually layered evals across individual tool calls, whole sessions, and system-level decisions. Structured analysis of failure data guided which evals to build.
- Looking ahead
- Alyx is moving from fixed, on-rails workflows toward more autonomous planning loops. The next phase depends on the same evaluation discipline that made the first version work.
03From the post
“Listen to this episode on: Spotify | Apple Podcasts What does it really take to build an AI agent inside an AI platform—especially when you’re using that same platform to build the agent? In this episode of Just Now Possible, Teresa Torres talks with SallyAnn DeLucia (Director of Product”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.