By Lenny Rachitsky · lennysnewsletter.com · @lennysan on X · YouTube · LinkedIn
Lenny Rachitsky and prompt engineer Mike Taylor ran a blind test to see how close AI comes to doing core product manager work. They argue that most headlines saying AI 'can't do X' rely on weak models and basic prompts, so they used the newest model with prompt engineering on three hard PM tasks drawn from reader suggestions. Public polls compared AI and human answers without revealing which was which. The AI answer won two of three tasks, and most voters guessed the AI answer correctly yet often still preferred it. The authors caution this was a small, non-scientific test, but argue it suggests AI is further along than many people assume.
01Key takeaways
- Judge AI using the latest model with strong prompting, not default chat output, before dismissing it.
- Give AI a role, planning steps, clear instructions, and at least one example to improve output reliability.
- Blind comparisons reveal preferences that brand or AI-suspicion bias can hide.
- AI tends to excel at comprehensive, structured answers but can read as a feature list rather than true strategy.
- Human PMs should lean into specific context, niche references, and strategic judgment that AI lacks.
02Key sections
- Why most AI skepticism is outdated
- Claims that AI cannot perform a task often come from older models or unrefined prompts. Prompt engineering can substantially raise output quality.
- Test design and selected tasks
- Three real PM tasks were chosen: product strategy, defining KPIs, and estimating feature ROI. Human answers came from Exponent's interview question database, and voting was done blind on social media.
- Results by task
- AI won the strategy and KPI tasks, while the human answer won the ROI estimation task narrowly. Many voters preferred AI even when they guessed it was AI, and ties were counted as AI wins.
- Prompting method and limitations
- The prompts used a role, chain-of-thought planning, instructions, and a single example to shape the output. The authors note limits including data contamination, poll methodology, and subjective judgment.
- Mapping AI capability to the PM role
- The authors propose benchmarking across a broader framework of PM skills to measure how much of the role is automatable, and expect AI to keep improving.
03From the post
“1. Counterintuitive advice for building AI products 2. General management, functional, and hybrid models: Which org design works best for top companies? 3. When and how to run a billboard campaign Subscribe to get access to these posts, and every post. For more: Best of Lenny’s Newsletter | Hire your next product leader | Podcast | Lennybot | Swag In my quest to develop a comprehensive benchmark to measure progress toward AI replacing PMs, I teamed up with full-time prompt engineer (and past collaborator) Mike Taylor on a piece that will surely blow your mind. I’d love to hear your reactions in the comments. Leave a comment The AI industry moves fast, which leads to lots of confusion about what AI is actually good at. OpenAI, Anthropic, and all the other AI companies are constantly testing their latest models’ math, language, and coding abilities. However, these abstract benchmarks don’t tell us how much of your job AI is able to potentially replace, now or in the near future—which is really what we care about. To make matters more challenging, expert…”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.