By Lenny Rachitsky · lennysnewsletter.com · @lennysan on X · YouTube · LinkedIn
Guest author Olga Berezovsky explains the difference between correlation and linear regression analysis for product teams. Correlation measures how strongly two variables move together on a scale from -1 to 1, and is a fast first check for which user behaviors relate to engagement or retention. Linear regression goes further by estimating how much one variable changes another and whether one can predict the other. The piece walks through a MyFitnessPal example of food logging and retention, shows how to run these analyses in Amplitude, Mixpanel, Excel, and Google Sheets, and warns that neither method proves causation and that outliers can skew results. It matters because these are foundational tools for finding predictable user actions and forecasting outcomes.
01Key takeaways
- Run a correlation analysis first to confirm whether two metrics are related and in which direction before attempting regression.
- Use correlation to measure how strongly a behavior relates to retention or engagement, and regression to estimate how much change in one drives change in the other.
- Remember that neither correlation nor regression proves causation; treat results as signals to test further.
- Product analytics tools like Amplitude's Compass and Mixpanel's Signal can quickly surface behaviors correlated with retention.
- Inspect your data for outliers before trusting a linear regression, since extreme values can skew the trend line.
- Check whether a behavior's frequency matters, not just whether it happens, to find the threshold that drives retention.
02Key sections
- Correlation analysis
- Correlation scores how closely two variables are related, from -1.0 to 1.0, and is the quickest way to spot which behaviors go with higher or lower engagement. The author recommends running it first when exploring new data.
- Linear regression
- Regression estimates how much one variable affects another and whether its pattern can predict the other, useful for forecasting and sizing the effect of product changes.
- Difference between the two
- Correlation confirms that a relationship exists, while regression explains how it works and enables prediction. A simple test is whether swapping the variables changes the answer.
- Worked example and tools
- Using a food-logging retention example, the author shows how to run correlation in Amplitude's Compass and Mixpanel's Signal reports, then use Excel or Google Sheets for regression.
- Handling outliers
- Outliers can distort a regression trend line, so analysts must examine the full data distribution and decide whether to keep or remove them based on their position and the overall variance.
03From the post
“1. How to use ChatGPT in your PM work 2. Discussion: How and where are you finding the best job opportunities? 3. What jury duty taught me about product management Subscribe to get access to these posts, and every post. Q: You’ve mentioned regression analysis a few times in your posts . What exactly is a regression analysis, and how do I run one? I’ll be honest. Though it’s come up in the newsletter a few times, I’ve also never truly understood what a regression analysis is. I know it helps you understand how two metrics are connected, but when exactly to run one, how to run one, and how regression differs from metrics being correlated, I’ve never deeply understood. I imagine many of you feel the same way. Considering how often these two methods come up, and (as you’ll see below) how powerful these tools can be, it’s important we all get smarter about this. To help us out, making her return appearance, I’ve pulled in Olga Berezovsky, author of the wonderful Data Analysis Journal newsletter, to explain what…”
04Frameworks mentioned
Summary and takeaways written by PM Atlas; quotes are short excerpts. © the original author.