Product Talk · Free post · Building AI products · Metrics, data & experimentation

Building Trainline’s AI Travel Assistant: How a 25-Year-Old Company Went Agentic

Teresa TorresOct 30, 2025
SourceProduct Talk
KindFree post
PublishedOct 30, 2025
Originalproducttalk.org ↗
N:

This episode of Teresa Torres's podcast walks through how Trainline, a 25-year-old rail and coach ticketing platform, built Travel Assistant, an AI-powered companion for travelers facing disruptions and needing real-time answers. The team describes starting from underserved needs beyond ticketing and building a fully agentic system with orchestration, tools, and reasoning loops from the outset. They explain layered guardrails for safety, grounding, and human handoff, and a major scale-up of curated retrieval content from 450 to 700,000 pages. The conversation also covers LLM-as-judge evaluations and a custom user context simulator used to measure quality. It matters as a behind-the-scenes case of an established company adopting new AI architectures at scale.

01Key takeaways

  • Pair scalable reasoning with deep domain context, since an assistant is only useful when it knows the specific domain well.
  • Treat tool design and guardrails as seriously as prompt design when building agent systems.
  • Use LLM-as-judge evals to measure open-ended outputs without the expense of large human labeling efforts.
  • Build a simulator of user context so quality can be tested across varied scenarios before and after release.
  • Legacy companies can move fast by pairing PM and engineering closely and embracing experimentation.
  • Plan for a clear human handoff path so customers are never stuck when the AI cannot resolve their issue.

02Key sections

Mission and unmet traveler needs
The team frames Trainline's mission and explains how they looked beyond ticket purchase to find where travelers still struggled. This grounded the assistant's scope in real journey pain points.
Building an agentic system from day one
Rather than bolting AI onto existing flows, they designed the assistant around orchestration, tool use, and reasoning loops. Iteration happened quickly within that architecture.
UX, guardrails and human handoff
Layered guardrails were designed for safety and grounding, with clear paths to a human when the assistant cannot help. Trust on the go depended on balancing latency and reliability.
Scaling retrieval content
The knowledge base grew from 450 to 700,000 curated pages, creating data retrieval and processing challenges. Managing this breadth of information was a central technical problem.
Evaluating quality in open-ended systems
The team used LLM-as-judge evals and a custom user context simulator to measure performance in near real time without costly manual labeling.

03From the post

“Listen to this episode on: Spotify | Apple Podcasts Trainline—the world’s leading rail and coach platform—helps millions of travelers get from point A to point B. Now, they’re using AI to make every step of the journey smoother. In this episode, Teresa Torres talks with David Eason”

“AI assistants need both scalable reasoning and deep domain context to be useful.”Teresa Torres · Product Talk

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.