◀ Course contents Part 3 · Module 3-06

Metrics

Pick the number that actually tells you the truth

Not all metrics are equal. Some tell you what's about to happen; some only tell you what already happened. Some can be gamed into looking good while the product quietly gets worse. This guide covers how to choose metrics you can actually trust and act on.

Ready?

1

What Makes a Metric Good

A good metric has four properties. It is specific — everyone means the same thing by it. It is comparable — this week against last week, without an argument about definitions. It is actionable — a team can do something on Monday that plausibly moves it. And it is hard to game — there is no cheap way to improve the number without improving the underlying reality.

That last property is where most metrics fail, and the reason is structural rather than moral. People respond to what they are measured on; that is the entire purpose of measuring. So the question is never “would anyone game this?” but “what is the cheapest way to move this number without doing the real work, and how bad would that be?”

It is worth noticing that a bad metric is usually not a vague one. It is frequently a very sharp, very precise number that measures the wrong thing — which is far more dangerous, because precision reads as rigour.

If you cannot immediately think of the cheap path, you have not looked hard enough. The team being measured will find it within a month, and they will not be doing anything wrong when they do.

A1

Actionable

A specific decision can actually move it, it's not just a number that sits there.

A2

Accessible

Easy to get and understand, not buried in a report only one analyst can pull.

A3

Auditable

You can trace it back to source data and trust it's counted consistently.

The four properties of a good metric A checklist of what makes a metric usable against what makes one misleading: specific versus ambiguous, comparable versus drifting, actionable versus untouchable, and hard to game versus trivially gameable. Good metric Bad metric Specific — everyone means the same thing Comparable week to week Actionable — a team can move it Monday Hard to game without doing real work Ambiguous — two people, two readings Definition drifts between reports Nobody can influence it directly A cheap shortcut moves the number
Read the right-hand column as four ways a number stays precise while becoming useless. None of them makes the metric look broken — a definition that drifts still reports cleanly, and a gamed number goes up — which is why a metric is worth auditing while it is still behaving.

Everyday example

A bathroom scale is a good metric: comparable day to day, hard to game, and it moves when the thing you care about moves. "How healthy do you feel?" is a bad one — it drifts with mood, and two people mean different things by it.

The metric that got gamed

A support team was measured on tickets closed per hour. Closures rose sharply. Agents had started closing tickets before the problem was resolved, customers re-opened them, and total load went up. The metric was precise, easy to collect, and pointed at the opposite of the thing it was meant to protect.

Quick check

A team tracks "total page views, all-time" as their main dashboard metric. Which of the 3 A's does it fail hardest?

2

Leading vs. Lagging, Input vs. Output

Two distinctions, crossed together, explain most of what is wrong with the average weekly metrics review.

The first is timing. A lagging metric reports what already happened — quarterly revenue, last month’s churn. A leading metric moves early and hints at what is coming — trials started today, demos booked this week.

The second is control. An input is something a team does directly: calls made, experiments run, pages published. An output is the result of those inputs meeting the world: sign-ups, revenue, retention.

Cross them and you get four quadrants, and only one of them is genuinely useful for steering. Leading inputs are both early and yours — they tell you something in time to act, and acting is within your power. Lagging outputs are the opposite on both counts: by the time quarterly revenue has moved, the decisions that moved it are months old and the people who made them have moved on.

Yet lagging outputs are what almost every weekly review opens with, because they are the numbers the business cares about most. The confusion is between what you report and what you steer by. Leadership can reasonably ask for revenue. A team trying to change something this week needs a number that will have moved by Friday.

One team switched their Monday review from quarterly revenue to demos booked and trials started. Nothing about the business changed that week. What changed was how early they could tell something was going wrong — which is the only thing a review is for.

Leading indicator

Moves first, predicts a future result, e.g. trial signups predicting future revenue.

Lagging indicator

Confirms what already happened, e.g. quarterly revenue.

Input metric

Something the team directly controls, e.g. features shipped.

Output / outcome metric

The result inputs are meant to produce, e.g. activation, retention.

Leading and lagging, input and output A two-by-two grid. The horizontal axis runs from lagging to leading. The vertical axis runs from output, which you observe, to input, which you control. The leading-input quadrant is highlighted as the only one a team can act on this week. Lagging input last quarter’s headcount Leading input demos booked this week Lagging output quarterly revenue Leading output trials started today lagging → leading output → input
Only the top-right is both early and yours. Teams report from the bottom-left and are then surprised they cannot move it — by the time a lagging output has changed, the decisions that changed it are months old.

Everyday example

A smoke alarm and a fire report both tell you about a fire. One of them tells you in time to do something. Choosing which to review each week is the whole distinction, and most teams review the report.

Reporting from the wrong quadrant

A team reviewed quarterly revenue every Monday. By the time the number moved, the decisions behind it were three months old and the people who made them had moved on. Switching the weekly review to demos booked and trials started changed nothing about the business and everything about how early the team could react.

Quick check

Quarterly revenue confirms how the last three months went, but by the time you see it, the quarter is over. What kind of metric is this?

3

Choosing a North Star Metric

A north star metric is the single number a whole company agrees to steer by. Not the only number anyone looks at — the one that wins when two teams disagree about priorities.

The reason to have one is unglamorous: without it, every prioritisation argument restarts from first principles and is settled by whoever is most senior or most persistent. With one, the argument becomes “which option moves the north star more?”, which is a question evidence can answer.

A good one has three qualities. It reflects customer value — it rises when customers get something they want, not when you extract more from them. It predicts revenue without being revenue, so it moves early enough to act on. And it is influenceable by the teams being asked to move it.

Revenue itself fails the third test in most companies, and the second by definition. It is the thing the business needs; it is rarely the thing a team can steer by, because it moves too late to react to.

The common failure is choosing something easy to move and disconnected from value — registered users, page views, features shipped. All of these rise obligingly and tell you nothing.

The test is whether the number could climb for a full quarter while customers were getting steadily worse service. If it could, it is not a north star; it is an activity counter with an inspiring name.

Choosing a north star A checklist for a north star metric: it should reflect customer value, predict revenue without being revenue, and be influenceable by the teams asked to move it. North star Vanity metric Rises when customers get value Predicts revenue, is not revenue Teams can actually influence it Would fall if service got worse Rises when you extract more Is revenue, so moves too late Nobody owns it in practice Could climb all quarter regardless
The fourth pair is the one that settles it, and it is a question about direction rather than size. A north star has to be able to fall — if nothing you could do to customers over three months would bend it downward, the number is not watching them, and it will keep rising while they leave.

Everyday example

A ship needs one north star, not forty. Navigating by forty stars means arguing about which one to steer by every night, which is how a team ends up with a dashboard everyone reads and nobody uses.

Not revenue directly, but what leads to it

Spotify tracks time spent listening. Airbnb tracks nights booked. Neither is revenue itself, both are the moment a customer actually gets the core value the product promises. When that number is healthy, revenue reliably follows.

A North Star isn't your only metric, and a shallow proxy like "app opens", easy to measure but disconnected from real value, is a common trap.

A north star that was not revenue

Spotify is widely described as steering by time spent listening rather than by subscription revenue directly. The logic is that listening precedes renewal, so a team can act on it in a week instead of waiting a quarter — the north star is chosen for how early it moves, not for how close it sits to the bank account.

Quick check

A team picks "number of app opens" as their North Star, even though users often open the app, get frustrated, and leave without doing anything useful. What's the problem?

4

The Right Metric for the Decision

So the sequence matters, and it is the reverse of what most teams do. Start with the decision, then choose the metric. What will you do differently depending on the answer? If nothing, you do not need the number. If something, which number would actually change your mind — and what value of it would push you one way rather than the other?

That last part is the one worth forcing yourself to write down in advance. “We will roll this out if activation reaches 34%” is a decision rule. “Let’s see how activation looks” is an invitation to interpret whatever happens as support for what you already wanted to do.

A related habit: pair every metric you are trying to move with a guardrail — a second number that must not get worse. Optimising conversion alone will eventually produce a checkout that tricks people; conversion paired with refund rate will not. Almost every metric can be improved by damaging something the metric does not measure, and naming the guardrail up front is how you find out before your customers do.

Start from the decision A chooser: the decision you face determines the metric. Judging whether a feature worked needs an adoption measure among its intended users; forecasting capacity needs a volume measure; judging business health needs retention and revenue. What decision are you making? feature worked? Adoption among the users it was built for, over its natural cycle capacity? Volume and peak load, blended across all users business health? Retention and revenue, with a guardrail
One team used weekly active users for the first two at once. It was reasonable for capacity and useless for the feature, which was designed to be used monthly and so looked like a failure every week. The number was fine; it was being asked two questions and could only answer one.

Everyday example

A thermometer is the right instrument for a fever and useless for a broken arm. There is no such thing as the best metric — only the right one for the decision in front of you.

One number, two decisions, one wrong answer

Weekly active users was used both to judge whether a new feature worked and to forecast infrastructure spend. It was reasonable for the second and useless for the first, because a feature used monthly by design looked like a failure every week. The number was fine; it was being asked two questions and could only answer one.

Quick check

A team wants to judge whether a new onboarding flow works, and defaults to checking "total company revenue" a week later. What's wrong with this metric choice?

Drill what you learned
Scenario 1 easy

A team's dashboard headlines "total signups since launch," which has risen every month for three years, even during a period when the product was clearly declining.

What kind of metric is this, and what's the risk?

Scenario 2 easy

A team wants a metric to catch problems early enough to react, not just confirm results after the fact.

What type of metric should they prioritize?

Scenario 3 easy

A support team is evaluated purely on "tickets closed per hour," and closures soar while customer satisfaction quietly drops.

Which of the 3 A's does this metric fail?

Scenario 4 easy

A PM proposes "employee happiness with the roadmap" as the company's North Star metric.

What's the issue?

Scenario 5 easy

A team wants to evaluate a pricing-page redesign and defaults to checking "monthly active users" a week later.

What's the concern with this metric choice?

Scenario 6 medium

A metric is described as: "the number our finance team can trace back to raw transaction records at any time, with a documented calculation method."

Which of the 3 A's does this describe?

Scenario 7 medium

A growth team reports "trial signups are up 30% this week" and predicts a strong revenue quarter.

What kind of metric are they using to make that prediction?

Scenario 8 medium

Airbnb famously tracks "nights booked" rather than just "signups" as closer to its North Star.

Why is "nights booked" the better choice?

Scenario 9 medium

A team's "onboarding completion rate" has been flat for months, and they're debating whether it's a leading or lagging indicator of retention.

What's the right way to think about it?

Scenario 10 medium

A stakeholder asks for "one metric that tells us everything about how the product is doing."

What's the realistic response?

Scenario 11 hard

A metric named "engagement score" turns out to be calculated differently by two different teams, with no shared documentation.

Which of the 3 A's is this failing?

Scenario 12 hard

A team is deciding between raising prices or improving retention, and someone suggests checking "daily app opens" to guide the choice.

Is this the right metric for this decision?

Scenario 13 hard

A team notices their North Star metric has been rising for six months, but customer support tickets about frustration have also been rising just as fast.

What should the team do?

Scenario 14 hard

A junior PM says "let's just track everything so we don't miss anything important."

What's the issue with this instinct?

Scenario 15 hard

A team wants to A/B test a checkout redesign and picks "checkout completion rate" as the metric to judge success.

Is this a good match?

Scenario 16 medium Select all
Activation dashboard, four weeks
WeekSignupsActivatedActivation rate
w/c 5 May4,1201,03025.0%
w/c 12 May4,4801,12025.0%
w/c 19 May6,9001,24018.0%
w/c 26 May7,2401,30018.0%

A paid campaign started on 19 May.

This is the dashboard your team reviews every Monday. Nobody has changed the product in these four weeks.

Which of these statements does the table actually support?

Select all that apply — there are 3 to find.

Scenario 17 medium Put in order

A PM has been asked to "pick a metric" for a checkout redesign shipping next sprint.

Put the steps in the order that produces a metric you can defend.

Drag the rows, or use the arrows, then check.

  1. Name the decision the number has to inform First, always. Every later step depends on knowing what the number is for; a metric chosen before the decision is a metric looking for a use.
  2. Write down what result would change that decision Second. If no possible value would change what you do, the metric is decoration and you can stop here.
  3. Pick the metric closest to the change being made Third. Only now is there a standard to choose against — closeness to the change is what makes the reading attributable.
  4. Check it against the 3 A's Fourth. This screens a candidate you already have; it does not generate one.
  5. Add the guardrail that would catch the damage Last. A guardrail is defined against the metric you have chosen, so it cannot be picked before it.
Scenario 18 hard Case
Slack#product-leadershipToday
  1. DanaVP Growth09:14 MAU is up 12% this month — brilliant work. Let's make MAU the company North Star and get it on the exec dashboard before the board meeting.
  2. PriyaSupport Lead09:31 Congrats! Unrelated, but refund requests are up about 20% over the same period. Probably noise.

You are the PM on the account. Dana wants an answer today.

Work the thread.

  1. Step 1 of 3

    Before agreeing or disagreeing, what is the first thing to establish?

    You check: MAU counts any session longer than three seconds. It cannot tell a user who got what they came for from one who bounced.

  2. Step 2 of 3

    Priya's refund signal is rising over the same period. What does that now suggest?

    You now have a specific, testable worry: retries inflating the count. That is something you can check, not just an objection.

  3. Step 3 of 3

    What do you take back to Dana?

    Dana still gets a number for the board. It is just a number that would go down if the product got worse.

Scenario 19 hard Written

Your team ships a redesigned onboarding flow next sprint. In the planning review, the director asks how you will know whether it worked.

Answer them in four or five sentences: the metric you would judge it on, the guardrail you would watch alongside it, and why company revenue is the wrong choice here.

One strong answer

I would judge it on activation rate — the share of new signups who reach the first real value moment within seven days. It is the metric closest to what the flow actually changes, so a movement in it is attributable to us rather than to everything else happening that month.

Alongside it I would watch seven-day retention of the users who activated. A flow can push more people over the activation line by rushing them, and retention is what would show that we bought the number rather than earned it.

Revenue is the wrong judge here. It is influenced by pricing, sales, seasonality and churn from cohorts that never saw this flow, so by the time it moves we cannot say what moved it — and it moves far too slowly to inform the decision of whether to keep the change.

Tick every point your own answer actually made.

  • "Activation rate", "seven-day activation" — something a query could return. "Engagement" or "user happiness" is not yet a metric.

  • The reason it is the right metric has to be that this flow is what moves it. Attribution is the argument.

  • A guardrail with no stated failure mode is a second headline metric. Say what going wrong would look like.

  • Too many other inputs feed it, so a movement cannot be traced back to the onboarding change.

  • Even if it were attributable, it arrives after the point where you would have had to decide.

Drilled it. Now apply it to a real situation.

Put it to work
From lesson 1

A metric that passed every check and still broke the bank

Wells Fargo

Through the 2010s Wells Fargo ran its retail bank on cross-sell ratio — the average number of products per customer household. The public target was eight. The number was reported to investors quarterly, branch managers saw it daily, and every banker knew their own.

Run it through the 3 A's before you read on. A specific decision moved it: sell another product. It was trivially accessible — on a dashboard, refreshed nightly. And it was auditable to the account, because every product counted was a real row in a real system.

In September 2016 the bank was fined $185 million. Around 5,300 employees had been dismissed for opening accounts customers had never asked for — on the order of two million of them, later revised upward.

The 3 A's did not catch this. Which of these statements about the cross-sell metric are true?

Select all that apply — there are 3 to find.

From lesson 2

The number that told Slack a team would stay

Slack

Slack's problem was one every team product has. The number that really mattered — is this team still using us in six months, and will they pay — could only be known in six months. By the time it arrived, whatever had gone wrong in week one was long past fixing.

So the growth team went looking for something that moved early and predicted that outcome. What they landed on, and have discussed publicly ever since, was a threshold: a team that had exchanged around 2,000 messages was very likely to keep going.

Work through how that number gets found.

  1. Step 1 of 3

    Why can "still active in six months" not be the number the team steers on week to week?

    So the search is on for something that happens in week one and predicts what week twenty-four will say.

  2. Step 2 of 3

    Several things can be counted in a team's first week. Which is the strongest candidate for a leading indicator?

    Note what makes it a good candidate: it is the value moment itself, counted, not an activity that happens near the value moment.

  3. Step 3 of 3

    You now have a threshold that predicts retention. What is it actually for?

    A leading indicator you do not act on is trivia. One you build onboarding around is a strategy.

From lesson 3

Medium threw away pageviews on purpose

Medium

In 2014 Medium did something publishing companies did not do: it said publicly that it was not going to run on pageviews or unique visitors, and moved to Total Time Reading — the aggregate time people actually spent reading, rather than the number of times a page was opened.

The reasoning was blunt. Pageviews reward a headline that gets clicked, whether or not anything behind it gets read. A publisher steering on them is being pulled, one small decision at a time, toward writing for the click.

Make Medium's case in four or five sentences. Why does Total Time Reading capture value where pageviews do not, and what would you watch alongside it to catch the ways it could go wrong?

One strong answer

Total Time Reading counts the thing the product exists to cause. Medium is not in the business of getting pages opened; it is in the business of getting things read, and time spent reading is the closest countable stand-in for that happening.

Pageviews cannot tell a piece someone read to the end from one they closed in two seconds, so the two register identically. Worse, the cheapest way to raise pageviews is a more misleading headline — which means steering on them actively rewards making the product worse.

Alongside it I would watch completion rate, or reading time relative to a piece's length. Total Time Reading is an aggregate, so it can be held up by a small number of very long pieces, or inflated by people struggling through something badly written. The guardrail is what separates "read attentively" from "took a long time".

Tick every point your own answer actually made.

  • Reading is the thing the product is for. A page being opened is the step before the thing the product is for.

  • They cannot distinguish a piece read to the end from one abandoned instantly — both count once.

  • The cheapest lever on them is a more clickable headline. A metric whose easiest win damages the thing it measures is not a neutral one.

  • Completion rate, or time relative to length — because an aggregate of time can be held up by a few long pieces, or by prose that is simply hard going.

  • The argument has to stand on value delivered to readers. If it only works because TTR happens to monetise better, it is not a North Star argument.

From lesson 4

Four metrics, one decision

A DIY marketplace

Your team rebuilt the checkout page: fewer form fields, and delivery options moved above the fold. It has been live to 50% of traffic for eleven days. On Thursday the director asks a single question: do we keep it?

Four numbers are available to you. Only one of them can answer that question on its own; the others are somewhere between useful and useless here.

Order these from best to worst at answering "do we keep the new checkout?"

Drag the rows, or use the arrows, then check.

  1. Checkout completion rate, new page vs. old, over the eleven days First. It is the closest metric to the thing that changed, it is measured on both variants at once, and eleven days is long enough for the rate to be readable. It answers the question directly.
  2. Average order value, new page vs. old Second, as a guardrail rather than a verdict. Moving delivery options above the fold could plausibly push people to cheaper delivery and smaller baskets — a completion win that costs more than it earns. Worth checking; not the answer on its own.
  3. Support tickets mentioning checkout, this week vs. last Third. Directionally useful and it catches breakage the funnel would not show, but ticket volume is noisy, slow, and only a small fraction of affected people ever write in.
  4. Total company revenue for the month so far Last, and not close. Half the traffic never saw the change, the month is not finished, and revenue moves for a dozen reasons that have nothing to do with a checkout page. It cannot isolate this decision at all.

Notification