By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
This podcast episode follows Perk, a company that eliminates 'shadow work' like travel booking, as its team built a voice AI agent that calls hotels to verify virtual credit card payments. The project grew from a Make.com hackathon prototype into a production system handling over 10,000 calls weekly across several languages. The discussion covers how the team learned prompt engineering for voice, broke one monolithic prompt into staged conversation flows, and built two evaluation systems. It matters because it shows how real-time, high-stakes voice AI can be shipped and improved through close listening to actual calls rather than theorizing alone.
01Key takeaways
- Tie an AI prototype to a real, recurring operational problem so its value is obvious from the start.
- Break one monolithic prompt into structured stages to make a complex multi-step task more reliable.
- Voice prompts need explicit handling of numbers, booking references, and text-to-speech formatting.
- Combine automated success classification with LLM-as-judge evals to measure both outcomes and conversational quality.
- Keep listening to real calls manually, since automated metrics can miss what actually happens.
- Write prompts in the target language when expanding to new markets to get better results.
02Key sections
- Origin of the use case
- The team connected earlier experimentation with a concrete operational pain point in travel booking. A hotel payment verification problem gave the prototype a clear, valuable job.
- Prototyping without backend code
- The first version was built in a no-code automation tool, and it was pushed toward production without writing new backend infrastructure. This shortened the path from idea to live calls.
- Structuring the conversation
- A single large prompt was split into discrete stages such as IVR handling, booking confirmation, and payment request. Separating the task made the agent markedly more reliable.
- Evaluating quality
- The team ran two evaluation systems: one classifying call success and one using an LLM as judge for conversational behavior. Manual call listening remained part of the process alongside the automated metrics.
- Scaling and language expansion
- Volume grew from about five calls a day to tens of thousands per week while quality was maintained. Writing prompts in the native language, as with German, improved results.
03From the post
“Listen to this episode on: Spotify | Apple Podcasts What happens when you combine a real customer problem, a no-code prototype, and a team willing to listen to every single call? In this episode of Just Now Possible, Teresa Torres talks with Steven Payne (Product Manager), Gabriel Stock (Senior Engineering”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.