By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
This episode of Just Now Possible, hosted by Teresa Torres, is an in-depth look at how Doist built Ramble, a voice-to-task feature in Todoist that turns spoken brain dumps into structured tasks in real time. Ramble was Todoist's first pure AI feature, emerging from a two-to-three month AI exploration phase, and it uses a Gemini live audio model to make tool calls while the user is still speaking, with no text output. The guests explain how user research revealed people brainstorming tasks with pen and paper or voice chat before committing them, and how that shaped the design. The conversation shows why keeping AI features constrained, simple, and easy to correct can be more useful than chasing perfect first-time accuracy.
01Key takeaways
- Start AI exploration broadly, then pick the use case that shows the clearest user pull, as Doist did with voice-to-task.
- Study how users already behave outside your product, such as brainstorming on paper or with voice assistants, to find the real job.
- Skipping a transcription step can simplify a live audio pipeline and reduce latency for voice features.
- Constrain the model to a small set of tools and literal capture rules so it doesn't over-interpret or act on the user's intent.
- Injecting a bounded context list into the prompt can beat building a RAG pipeline when the data is small enough.
- Build evals from real, diverse user recordings across languages to catch prompt regressions before shipping.
02Key sections
- Origins of Ramble
- Doist's AI exploration phase surfaced voice-to-task as the strongest candidate. The team describes how it became their first pure AI feature.
- Brain dump research insight
- User research showed people using paper or voice assistants to brainstorm before adding tasks, which motivated a capture-first design.
- Live audio architecture
- Ramble skips transcription and feeds raw audio to a Gemini live model that calls add, edit, and delete tools in real time, with visual cards and sound cues for driving use.
- Context and date handling
- The team injected the full project and label list into the prompt instead of building RAG, and worked through the complexity of dates in a live audio pipeline.
- Evals and what's next
- Multi-language evals built from employee recordings catch regressions, and the roadmap extends to multimodal capture and integrations.
03From the post
“Listen to this episode on: Spotify | Apple Podcasts How do you turn a rambling stream of consciousness into a clean task list — while the person is still talking? That's the core challenge Doist solved with Ramble, a voice-to-task feature inside Todoist that uses live audio AI to”
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.