Drawer 11 · 316 notes · 58 people · 2005–2026

Metrics, data & experimentation

North-star metrics, analytics, A/B tests, and data-informed decisions

Q:
Answers from Metrics, data & experimentation notes, cited
@ttorres · Teresa Torres on X♥ 97
"The only way to know if our AI products and workflows are any good is with evals." 💡 If you're using AI to write PRDs, analyze customer feedback, or build customer-facing AI products, you need to understand evals. This guide breaks down what evals actually are and why product teams should be…
@ttorres · Teresa Torres on X♥ 9
💭 "I wish I could click on this card and be like, 'Clean this up.'" When a customer pointed out a flat, unstructured branch in her AI-generated opportunity solution tree, it sparked a three-week journey to fix the problem at its source. Here's what you'll learn from this article: How one customer…
@lissijean · Melissa Perri on X♥ 1
This is what separates teams that accelerate with AI from those that get stuck. You don't find perfect data lying around. You create it using the deep product intuition you already have. /end Check out the whole episode with @vlaurenlee here:
@lissijean · Melissa Perri on X♥ 2
They're taking everything they know about commerce, customer behavior, and product strategy, and using that knowledge to systematically generate the training data they need. It transforms the data bottleneck from a passive waiting game into an active engineering challenge. /3
@ttorres · Teresa Torres on X♥ 2
🎙️Quality of Evidence You've got behavioral analytics, support tickets, sales call notes, and a feedback inbox that never empties. With all that data coming in, do product teams actually still need to go out and talk to users? In this episode, Petra Wille and Teresa Torres dig into why not all…
@hnshah · Hiten Shah on X♥ 10
If you use AI for competitor research, steal this rule. Every claim gets a date. Last week Claude answered a simple question about Linear’s competitors using an old source as if it were recent. It also stated a guess as fact. The skill forced the answer to show the dates,
@ttorres · Teresa Torres on X♥ 3
I'm seeing a huge difference in performance (in a bad direction) across my evals moving from Sonnet 4.6 to Sonnet 5. Previously, my service ran at temperature 0. I'm wondering what prompt strategies people are using to constrain Sonnet 5 now that temperature is no longer an option.