Product Talk · Free post · Building AI products · Metrics, data & experimentation

Building Alyx: How Arize AI Dogfooded Its Way to an Agentic Future

Teresa TorresOct 9, 2025
SourceProduct Talk
KindFree post
PublishedOct 9, 2025
Originalproducttalk.org ↗
N:

Teresa Torres interviews SallyAnn DeLucia and Jack Zhou of Arize AI about building Alyx, the AI agent that helps teams debug, optimize, and evaluate AI applications. The team describes scrappy beginnings in Jupyter notebooks and hacked-together apps, refined through weekly dogfooding sessions with customer success. They explain how evals started messy and were layered across tool calls, sessions, and system-level decisions. The conversation argues that cross-functional, boundary-spanning teams and customer insight matter as much as technical skill when shipping AI products. It is most useful for PMs trying to move from ad hoc prototypes and vibe checks to systematic AI product improvement.

01Key takeaways

  • Start prototyping in messy tools like notebooks; speed of learning matters more than polish early on.
  • Use your own product daily (dogfooding) to find gaps that scripted testing will never reveal.
  • Involve customer-facing staff, since they spot repeatable workflows customers actually need.
  • Begin evals imperfectly and iterate, layering them from tool calls up to whole-system decisions.
  • Analyze failure cases systematically before designing evals, rather than relying on vibe checks.
  • Favor cross-functional, boundary-spanning team members who can move between engineering, product, and customers.

02Key sections

Platform and core concepts
The guests frame tracing, observability, and evals as the foundation for understanding GenAI application behavior. These concepts set up why the team could use its own product to build an agent.
Genesis of Alyx
Alyx began as experiments in notebooks and local web apps rather than a planned product. Its evolution was shaped by what the team learned in practice.
Dogfooding with customer success
Customer success engineers helped surface repeatable workflows, and weekly internal testing built confidence before launch. Using the product daily exposed gaps that synthetic testing missed.
Evals that start messy
The team began evaluating without polished criteria and gradually layered evals across individual tool calls, whole sessions, and system-level decisions. Structured analysis of failure data guided which evals to build.
Looking ahead
Alyx is moving from fixed, on-rails workflows toward more autonomous planning loops. The next phase depends on the same evaluation discipline that made the first version work.

03From the post

“Listen to this episode on: Spotify | Apple Podcasts What does it really take to build an AI agent inside an AI platform—especially when you’re using that same platform to build the agent? In this episode of Just Now Possible, Teresa Torres talks with SallyAnn DeLucia (Director of Product”

“What does it really take to build an AI agent inside an AI platform—especially when you’re using that same platform to build the agent?”Teresa Torres · Product Talk · 00:00
“If you’ve ever wondered how to move from vibe checks and one-off prototypes to systematic improvement in your AI product, this episode is for you.”Teresa Torres · Product Talk · 00:00

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.