By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
This episode of Just Now Possible features the eSpark team discussing how they built Teacher Assistant, an AI feature that helps K–5 teachers align eSpark's supplemental lessons with district-mandated core curricula. The conversation traces how post-COVID pressures shaped the problem, why a chatbot interface failed in testing and was replaced by a more structured workflow, and the practical difficulties of building a first retrieval-augmented generation system and wrangling embeddings. The team also explains how their background in education led to a rigorous evaluation process before "evals" became a common term. The discussion matters because it offers a grounded account of iterating on an AI product with real users in a high-stakes domain, rather than abstract advice about AI.
01Key takeaways
- Test the interaction model with real users early; a familiar chatbot pattern may not fit the workflow.
- Domain expertise, such as knowing how teachers plan, can ground rigorous evaluation criteria.
- Semantic and keyword search fail differently, so retrieval quality needs direct inspection.
- Human-in-the-loop review of rubric-scored outputs catches quality issues automation misses.
- Metadata and data enrichment often determine recommendation quality as much as the model does.
- Ship, observe thousands of real sessions, and iterate rather than optimizing in a lab alone.
02Key sections
- Origins and classroom context
- The team describes eSpark's decade of adaptive learning work for K–5 classrooms and how post-COVID shifts created new pressure on teachers and administrators. Administrator mandates and teacher needs did not always line up, which set up the problem Teacher Assistant addresses.
- Abandoning the chatbot interface
- The first instinct was a chatbot, but testing showed teachers struggled with it. The team moved to a more structured workflow that better matched how educators actually plan lessons.
- Building the RAG architecture
- Building the first retrieval-augmented generation system meant learning to manage embeddings, and the team found semantic search and keyword search behaved differently in practice. Metadata and data enrichment became important for useful recommendations.
- Evaluation approach
- Drawing on their education background, the team built rubric-based evals with a human-in-the-loop process, using tooling such as Braintrust. This rigor came before the term "evals" became widespread.
- Learnings from real usage and next steps
- Thousands of teachers used the product during the school year, generating feedback that shapes the roadmap. Future plans include more contextual recommendations drawing on student data.
03From the post
“Listen to this episode on: Spotify | Apple Podcasts How do you build an AI-powered assistant that teachers will actually use? In this episode of Just Now Possible, Teresa Torres talks with Thom van der Doef (Principal Product Designer), Mary Gurley (Director of Learning Design & Product Manager), and Ray Lyons”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.