Product Talk · Free post · Building AI products · Metrics, data & experimentation

From Prototype to Production: How Perk Built a Voice AI Agent That Makes 10,000 Calls a Week

Teresa TorresDec 4, 2025
SourceProduct Talk
KindFree post
PublishedDec 4, 2025
Originalproducttalk.org ↗
N:

This podcast episode follows Perk, a company that eliminates 'shadow work' like travel booking, as its team built a voice AI agent that calls hotels to verify virtual credit card payments. The project grew from a Make.com hackathon prototype into a production system handling over 10,000 calls weekly across several languages. The discussion covers how the team learned prompt engineering for voice, broke one monolithic prompt into staged conversation flows, and built two evaluation systems. It matters because it shows how real-time, high-stakes voice AI can be shipped and improved through close listening to actual calls rather than theorizing alone.

01Key takeaways

  • Tie an AI prototype to a real, recurring operational problem so its value is obvious from the start.
  • Break one monolithic prompt into structured stages to make a complex multi-step task more reliable.
  • Voice prompts need explicit handling of numbers, booking references, and text-to-speech formatting.
  • Combine automated success classification with LLM-as-judge evals to measure both outcomes and conversational quality.
  • Keep listening to real calls manually, since automated metrics can miss what actually happens.
  • Write prompts in the target language when expanding to new markets to get better results.

02Key sections

Origin of the use case
The team connected earlier experimentation with a concrete operational pain point in travel booking. A hotel payment verification problem gave the prototype a clear, valuable job.
Prototyping without backend code
The first version was built in a no-code automation tool, and it was pushed toward production without writing new backend infrastructure. This shortened the path from idea to live calls.
Structuring the conversation
A single large prompt was split into discrete stages such as IVR handling, booking confirmation, and payment request. Separating the task made the agent markedly more reliable.
Evaluating quality
The team ran two evaluation systems: one classifying call success and one using an LLM as judge for conversational behavior. Manual call listening remained part of the process alongside the automated metrics.
Scaling and language expansion
Volume grew from about five calls a day to tens of thousands per week while quality was maintained. Writing prompts in the native language, as with German, improved results.

03From the post

“Listen to this episode on: Spotify | Apple Podcasts What happens when you combine a real customer problem, a no-code prototype, and a team willing to listen to every single call? In this episode of Just Now Possible, Teresa Torres talks with Steven Payne (Product Manager), Gabriel Stock (Senior Engineering”

“What happens when you combine a real customer problem, a no-code prototype, and a team willing to listen to every single call?”Teresa Torres · Product Talk · 00:00

04Frameworks mentioned

Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.