By Teresa Torres · producttalk.org · @ttorres on X · LinkedIn
Teresa Torres argues that A/B testing should follow the scientific method rather than being used for trivial tweaks like button colors or wording. She recommends starting from a real insight about why recipients ignore an email, turning it into a testable hypothesis that can be refuted, and designing a controlled experiment that changes only one variable. The piece walks through measuring unique recipients and actions, checking statistical significance, and recognizing that small samples only reveal big wins. It closes by urging honest conclusions that avoid overgeneralizing, since each test should build a knowledge base over time.
01Key takeaways
- Start A/B tests from a strategic insight about user behavior, not from a long list of random variations.
- Write hypotheses that can be measured and refuted, such as predicting an increase in open rates rather than an abstract feeling.
- Change only one variable at a time and use a true control group, or you cannot attribute results to the change.
- Decide the test duration beforehand and do not act on results until it ends, to avoid false significance.
- Only trust results that reach roughly 95% statistical significance, and aim for big wins when sample sizes are small.
- Draw conclusions narrowly and build a body of evidence across tests before generalizing about what works.
02Key sections
- Start with an insight
- Rather than randomly testing many variations, begin by reasoning from the recipient's perspective about why they would not open or click. This surfaces strategic levers instead of tactical tweaks.
- Formulate a testable hypothesis
- A good hypothesis must be measurable and capable of being disproven by an experiment. Vague claims about credibility or urgency should be restated as outcomes like open rates that can actually be tested.
- Design a controlled experiment
- Split the audience into a variable group and a control group, changing only one thing at a time. Comparing time periods introduces confounding factors that make results unreliable.
- Run the experiment and measure results
- Set the test duration in advance and ignore interim results to avoid misleading significance. Count unique people taking action rather than total actions.
- Evaluate results and draw conclusions
- Check statistical significance first, recognizing that small samples need large effects while bigger samples can detect small lifts. Interpret results carefully and avoid overreaching beyond what the test actually showed.
03From the post
“Now that you are measuring open, click, and conversion rates for your emails, it's time to look at how to improve them. Let's talk about A/B testing, sometimes called split testing. Far too many people hear A/B testing and think button colors and small wording changes.”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.