All 29 modules are open from the start — nothing here is
locked, and nothing costs anything. Sign in so your progress, titles and
credentials stay with you, on every device you use.
Free forever, with your Google account. No password, no payment.
Not all metrics are equal. Some tell you what's about to happen; some only tell you what already happened. Some can be gamed into looking good while the product quietly gets worse. This guide covers how to choose metrics you can actually trust and act on.
Ready?
1
What Makes a Metric Good
A good metric has four properties. It is specific — everyone means the same thing by it. It is comparable — this week against last week, without an argument about definitions. It is actionable — a team can do something on Monday that plausibly moves it. And it is hard to game — there is no cheap way to improve the number without improving the underlying reality.
That last property is where most metrics fail, and the reason is structural rather than moral. People respond to what they are measured on; that is the entire purpose of measuring. So the question is never “would anyone game this?” but “what is the cheapest way to move this number without doing the real work, and how bad would that be?”
It is worth noticing that a bad metric is usually not a vague one. It is frequently a very sharp, very precise number that measures the wrong thing — which is far more dangerous, because precision reads as rigour.
If you cannot immediately think of the cheap path, you have not looked hard enough. The team being measured will find it within a month, and they will not be doing anything wrong when they do.
A1
Actionable
A specific decision can actually move it, it's not just a number that sits there.
A2
Accessible
Easy to get and understand, not buried in a report only one analyst can pull.
A3
Auditable
You can trace it back to source data and trust it's counted consistently.
Read the right-hand column as four ways a number stays precise while becoming useless. None of them makes the metric look broken — a definition that drifts still reports cleanly, and a gamed number goes up — which is why a metric is worth auditing while it is still behaving.
Everyday example
A bathroom scale is a good metric: comparable day to day, hard to game, and it moves when the thing you care about moves. "How healthy do you feel?" is a bad one — it drifts with mood, and two people mean different things by it.
The metric that got gamed
A support team was measured on tickets closed per hour. Closures rose sharply. Agents had started closing tickets before the problem was resolved, customers re-opened them, and total load went up. The metric was precise, easy to collect, and pointed at the opposite of the thing it was meant to protect.
Quick check
A team tracks "total page views, all-time" as their main dashboard metric. Which of the 3 A's does it fail hardest?
2
Leading vs. Lagging, Input vs. Output
Two distinctions, crossed together, explain most of what is wrong with the average weekly metrics review.
The first is timing. A lagging metric reports what already happened — quarterly revenue, last month’s churn. A leading metric moves early and hints at what is coming — trials started today, demos booked this week.
The second is control. An input is something a team does directly: calls made, experiments run, pages published. An output is the result of those inputs meeting the world: sign-ups, revenue, retention.
Cross them and you get four quadrants, and only one of them is genuinely useful for steering. Leading inputs are both early and yours — they tell you something in time to act, and acting is within your power. Lagging outputs are the opposite on both counts: by the time quarterly revenue has moved, the decisions that moved it are months old and the people who made them have moved on.
Yet lagging outputs are what almost every weekly review opens with, because they are the numbers the business cares about most. The confusion is between what you report and what you steer by. Leadership can reasonably ask for revenue. A team trying to change something this week needs a number that will have moved by Friday.
One team switched their Monday review from quarterly revenue to demos booked and trials started. Nothing about the business changed that week. What changed was how early they could tell something was going wrong — which is the only thing a review is for.
Leading indicator
Moves first, predicts a future result, e.g. trial signups predicting future revenue.
Lagging indicator
Confirms what already happened, e.g. quarterly revenue.
Input metric
Something the team directly controls, e.g. features shipped.
Output / outcome metric
The result inputs are meant to produce, e.g. activation, retention.
Only the top-right is both early and yours. Teams report from the bottom-left and are then surprised they cannot move it — by the time a lagging output has changed, the decisions that changed it are months old.
Everyday example
A smoke alarm and a fire report both tell you about a fire. One of them tells you in time to do something. Choosing which to review each week is the whole distinction, and most teams review the report.
Reporting from the wrong quadrant
A team reviewed quarterly revenue every Monday. By the time the number moved, the decisions behind it were three months old and the people who made them had moved on. Switching the weekly review to demos booked and trials started changed nothing about the business and everything about how early the team could react.
Quick check
Quarterly revenue confirms how the last three months went, but by the time you see it, the quarter is over. What kind of metric is this?
3
Choosing a North Star Metric
A north star metric is the single number a whole company agrees to steer by. Not the only number anyone looks at — the one that wins when two teams disagree about priorities.
The reason to have one is unglamorous: without it, every prioritisation argument restarts from first principles and is settled by whoever is most senior or most persistent. With one, the argument becomes “which option moves the north star more?”, which is a question evidence can answer.
A good one has three qualities. It reflects customer value — it rises when customers get something they want, not when you extract more from them. It predicts revenue without being revenue, so it moves early enough to act on. And it is influenceable by the teams being asked to move it.
Revenue itself fails the third test in most companies, and the second by definition. It is the thing the business needs; it is rarely the thing a team can steer by, because it moves too late to react to.
The common failure is choosing something easy to move and disconnected from value — registered users, page views, features shipped. All of these rise obligingly and tell you nothing.
The test is whether the number could climb for a full quarter while customers were getting steadily worse service. If it could, it is not a north star; it is an activity counter with an inspiring name.
The fourth pair is the one that settles it, and it is a question about direction rather than size. A north star has to be able to fall — if nothing you could do to customers over three months would bend it downward, the number is not watching them, and it will keep rising while they leave.
Everyday example
A ship needs one north star, not forty. Navigating by forty stars means arguing about which one to steer by every night, which is how a team ends up with a dashboard everyone reads and nobody uses.
Not revenue directly, but what leads to it
Spotify tracks time spent listening. Airbnb tracks nights booked. Neither is revenue
itself, both are the moment a customer actually gets the core value the product
promises. When that number is healthy, revenue reliably follows.
A North Star isn't your only metric, and a shallow proxy like "app opens", easy to
measure but disconnected from real value, is a common trap.
A north star that was not revenue
Spotify is widely described as steering by time spent listening rather than by subscription revenue directly. The logic is that listening precedes renewal, so a team can act on it in a week instead of waiting a quarter — the north star is chosen for how early it moves, not for how close it sits to the bank account.
Quick check
A team picks "number of app opens" as their North Star, even though users often open the app, get frustrated, and leave without doing anything useful. What's the problem?
4
The Right Metric for the Decision
So the sequence matters, and it is the reverse of what most teams do. Start with the decision, then choose the metric. What will you do differently depending on the answer? If nothing, you do not need the number. If something, which number would actually change your mind — and what value of it would push you one way rather than the other?
That last part is the one worth forcing yourself to write down in advance. “We will roll this out if activation reaches 34%” is a decision rule. “Let’s see how activation looks” is an invitation to interpret whatever happens as support for what you already wanted to do.
A related habit: pair every metric you are trying to move with a guardrail — a second number that must not get worse. Optimising conversion alone will eventually produce a checkout that tricks people; conversion paired with refund rate will not. Almost every metric can be improved by damaging something the metric does not measure, and naming the guardrail up front is how you find out before your customers do.
One team used weekly active users for the first two at once. It was reasonable for capacity and useless for the feature, which was designed to be used monthly and so looked like a failure every week. The number was fine; it was being asked two questions and could only answer one.
Everyday example
A thermometer is the right instrument for a fever and useless for a broken arm. There is no such thing as the best metric — only the right one for the decision in front of you.
One number, two decisions, one wrong answer
Weekly active users was used both to judge whether a new feature worked and to forecast infrastructure spend. It was reasonable for the second and useless for the first, because a feature used monthly by design looked like a failure every week. The number was fine; it was being asked two questions and could only answer one.
Quick check
A team wants to judge whether a new onboarding flow works, and defaults to checking "total company revenue" a week later. What's wrong with this metric choice?
Drill what you learned
Scenario 1
easy
A team's dashboard headlines "total signups since launch," which has risen every month for three years, even during a period when the product was clearly declining.
What kind of metric is this, and what's the risk?
Scenario 2
easy
A team wants a metric to catch problems early enough to react, not just confirm results after the fact.
What type of metric should they prioritize?
Scenario 3
easy
A support team is evaluated purely on "tickets closed per hour," and closures soar while customer satisfaction quietly drops.
Which of the 3 A's does this metric fail?
Scenario 4
easy
A PM proposes "employee happiness with the roadmap" as the company's North Star metric.
What's the issue?
Scenario 5
easy
A team wants to evaluate a pricing-page redesign and defaults to checking "monthly active users" a week later.
What's the concern with this metric choice?
Scenario 6
medium
A metric is described as: "the number our finance team can trace back to raw transaction records at any time, with a documented calculation method."
Which of the 3 A's does this describe?
Scenario 7
medium
A growth team reports "trial signups are up 30% this week" and predicts a strong revenue quarter.
What kind of metric are they using to make that prediction?
Scenario 8
medium
Airbnb famously tracks "nights booked" rather than just "signups" as closer to its North Star.
Why is "nights booked" the better choice?
Scenario 9
medium
A team's "onboarding completion rate" has been flat for months, and they're debating whether it's a leading or lagging indicator of retention.
What's the right way to think about it?
Scenario 10
medium
A stakeholder asks for "one metric that tells us everything about how the product is doing."
What's the realistic response?
Scenario 11
hard
A metric named "engagement score" turns out to be calculated differently by two different teams, with no shared documentation.
Which of the 3 A's is this failing?
Scenario 12
hard
A team is deciding between raising prices or improving retention, and someone suggests checking "daily app opens" to guide the choice.
Is this the right metric for this decision?
Scenario 13
hard
A team notices their North Star metric has been rising for six months, but customer support tickets about frustration have also been rising just as fast.
What should the team do?
Scenario 14
hard
A junior PM says "let's just track everything so we don't miss anything important."
What's the issue with this instinct?
Scenario 15
hard
A team wants to A/B test a checkout redesign and picks "checkout completion rate" as the metric to judge success.
Is this a good match?
Scenario 16
mediumSelect all
Activation dashboard, four weeks
Week
Signups
Activated
Activation rate
w/c 5 May
4,120
1,030
25.0%
w/c 12 May
4,480
1,120
25.0%
w/c 19 May
6,900
1,240
18.0%
w/c 26 May
7,240
1,300
18.0%
A paid campaign started on 19 May.
This is the dashboard your team reviews every Monday. Nobody has changed the product in these four weeks.
Which of these statements does the table actually support?
Select all that apply — there are 3 to find.
Scenario 17
mediumPut in order
A PM has been asked to "pick a metric" for a checkout redesign shipping next sprint.
Put the steps in the order that produces a metric you can defend.
Drag the rows, or use the arrows, then check.
Name the decision the number has to informFirst, always. Every later step depends on knowing what the number is for; a metric chosen before the decision is a metric looking for a use.
Write down what result would change that decisionSecond. If no possible value would change what you do, the metric is decoration and you can stop here.
Pick the metric closest to the change being madeThird. Only now is there a standard to choose against — closeness to the change is what makes the reading attributable.
Check it against the 3 A'sFourth. This screens a candidate you already have; it does not generate one.
Add the guardrail that would catch the damageLast. A guardrail is defined against the metric you have chosen, so it cannot be picked before it.
DanaVP Growth09:14MAU is up 12% this month — brilliant work. Let's make MAU the company North Star and get it on the exec dashboard before the board meeting.
PriyaSupport Lead09:31Congrats! Unrelated, but refund requests are up about 20% over the same period. Probably noise.
You are the PM on the account. Dana wants an answer today.
Work the thread.
Step 1 of 3
Before agreeing or disagreeing, what is the first thing to establish?
You check: MAU counts any session longer than three seconds. It cannot tell a user who got what they came for from one who bounced.
Step 2 of 3
Priya's refund signal is rising over the same period. What does that now suggest?
You now have a specific, testable worry: retries inflating the count. That is something you can check, not just an objection.
Step 3 of 3
What do you take back to Dana?
Dana still gets a number for the board. It is just a number that would go down if the product got worse.
Scenario 19
hardWritten
Your team ships a redesigned onboarding flow next sprint. In the planning review, the director asks how you will know whether it worked.
Answer them in four or five sentences: the metric you would judge it on, the guardrail you would watch alongside it, and why company revenue is the wrong choice here.
One strong answer
I would judge it on activation rate — the share of new signups who reach
the first real value moment within seven days. It is the metric closest
to what the flow actually changes, so a movement in it is attributable
to us rather than to everything else happening that month.
Alongside it I would watch seven-day retention of the users who
activated. A flow can push more people over the activation line by
rushing them, and retention is what would show that we bought the number
rather than earned it.
Revenue is the wrong judge here. It is influenced by pricing, sales,
seasonality and churn from cohorts that never saw this flow, so by the
time it moves we cannot say what moved it — and it moves far too slowly
to inform the decision of whether to keep the change.
Tick every point your own answer actually made.
"Activation rate", "seven-day activation" — something a query could return. "Engagement" or "user happiness" is not yet a metric.
The reason it is the right metric has to be that this flow is what moves it. Attribution is the argument.
A guardrail with no stated failure mode is a second headline metric. Say what going wrong would look like.
Too many other inputs feed it, so a movement cannot be traced back to the onboarding change.
Even if it were attributable, it arrives after the point where you would have had to decide.
Drilled it. Now apply it to a real situation.
Put it to work
From lesson 1
A metric that passed every check and still broke the bank
Wells Fargo
Through the 2010s Wells Fargo ran its retail bank on
cross-sell ratio — the average number of products per
customer household. The public target was eight. The number was reported
to investors quarterly, branch managers saw it daily, and every banker
knew their own.
Run it through the 3 A's before you read on. A specific decision moved
it: sell another product. It was trivially accessible — on a dashboard,
refreshed nightly. And it was auditable to the account, because every
product counted was a real row in a real system.
In September 2016 the bank was fined $185 million. Around 5,300
employees had been dismissed for opening accounts customers had never
asked for — on the order of two million of them, later revised upward.
The 3 A's did not catch this. Which of these statements about the cross-sell metric are true?
Select all that apply — there are 3 to find.
What actually happened
Wells Fargo retired the cross-sell ratio as a public metric in 2017 and
replaced its branch incentives with measures built on customer
experience and retention — numbers that do not move when an account is
opened and never used.
The lesson worth carrying: the 3 A's screen a candidate you already
have. They do not ask whether moving the number and doing the job are
still the same thing. Add that question yourself, and add it again the
moment anyone proposes paying people against it.
From lesson 2
The number that told Slack a team would stay
Slack
Slack's problem was one every team product has. The number that really
mattered — is this team still using us in six months, and will they pay
— could only be known in six months. By the time it arrived, whatever
had gone wrong in week one was long past fixing.
So the growth team went looking for something that moved early and
predicted that outcome. What they landed on, and have discussed publicly
ever since, was a threshold: a team that had exchanged around
2,000 messages was very likely to keep going.
Work through how that number gets found.
Step 1 of 3
Why can "still active in six months" not be the number the team steers on week to week?
So the search is on for something that happens in week one and predicts what week twenty-four will say.
Step 2 of 3
Several things can be counted in a team's first week. Which is the strongest candidate for a leading indicator?
Note what makes it a good candidate: it is the value moment itself, counted, not an activity that happens near the value moment.
Step 3 of 3
You now have a threshold that predicts retention. What is it actually for?
A leading indicator you do not act on is trivia. One you build onboarding around is a strategy.
What actually happened
Slack organised its onboarding around getting a team talking to each
other quickly rather than around setup steps, and the 2,000-message mark
has since become one of the most cited activation benchmarks in
software.
The transferable part is not the number — 2,000 means nothing for your
product. It is the method: find the lagging outcome you actually care
about, then hunt for the earliest countable behaviour that reliably
precedes it, and point the team at that.
Lagging indicators tell you how it went. Leading indicators are the only
ones that give you time to change how it goes.
From lesson 3
Medium threw away pageviews on purpose
Medium
In 2014 Medium did something publishing companies did not do: it said
publicly that it was not going to run on pageviews or unique visitors,
and moved to Total Time Reading — the aggregate time
people actually spent reading, rather than the number of times a page
was opened.
The reasoning was blunt. Pageviews reward a headline that gets clicked,
whether or not anything behind it gets read. A publisher steering on
them is being pulled, one small decision at a time, toward writing for
the click.
Make Medium's case in four or five sentences. Why does Total Time Reading capture value where pageviews do not, and what would you watch alongside it to catch the ways it could go wrong?
One strong answer
Total Time Reading counts the thing the product exists to cause. Medium
is not in the business of getting pages opened; it is in the business
of getting things read, and time spent reading is the closest countable
stand-in for that happening.
Pageviews cannot tell a piece someone read to the end from one they
closed in two seconds, so the two register identically. Worse, the
cheapest way to raise pageviews is a more misleading headline — which
means steering on them actively rewards making the product worse.
Alongside it I would watch completion rate, or reading time relative to
a piece's length. Total Time Reading is an aggregate, so it can be held
up by a small number of very long pieces, or inflated by people
struggling through something badly written. The guardrail is what
separates "read attentively" from "took a long time".
Tick every point your own answer actually made.
Reading is the thing the product is for. A page being opened is the step before the thing the product is for.
They cannot distinguish a piece read to the end from one abandoned instantly — both count once.
The cheapest lever on them is a more clickable headline. A metric whose easiest win damages the thing it measures is not a neutral one.
Completion rate, or time relative to length — because an aggregate of time can be held up by a few long pieces, or by prose that is simply hard going.
The argument has to stand on value delivered to readers. If it only works because TTR happens to monetise better, it is not a North Star argument.
What actually happened
Medium built its recommendation and ranking systems around reading time
rather than clicks, and published the reasoning so writers on the
platform knew what they were being measured on. The idea spread: reading
or watch time is now the default engagement measure across most
publishing and video products.
It is the same move as Airbnb counting nights booked and Spotify
counting time spent listening. None of the three is revenue. All three
are the moment the customer receives what the product promised — and in
every case revenue followed the value, not the other way round.
From lesson 4
Four metrics, one decision
A DIY marketplace
Your team rebuilt the checkout page: fewer form fields, and delivery
options moved above the fold. It has been live to 50% of traffic for
eleven days. On Thursday the director asks a single question:
do we keep it?
Four numbers are available to you. Only one of them can answer that
question on its own; the others are somewhere between useful and
useless here.
Order these from best to worst at answering "do we keep the new checkout?"
Drag the rows, or use the arrows, then check.
Checkout completion rate, new page vs. old, over the eleven daysFirst. It is the closest metric to the thing that changed, it is measured on both variants at once, and eleven days is long enough for the rate to be readable. It answers the question directly.
Average order value, new page vs. oldSecond, as a guardrail rather than a verdict. Moving delivery options above the fold could plausibly push people to cheaper delivery and smaller baskets — a completion win that costs more than it earns. Worth checking; not the answer on its own.
Support tickets mentioning checkout, this week vs. lastThird. Directionally useful and it catches breakage the funnel would not show, but ticket volume is noisy, slow, and only a small fraction of affected people ever write in.
Total company revenue for the month so farLast, and not close. Half the traffic never saw the change, the month is not finished, and revenue moves for a dozen reasons that have nothing to do with a checkout page. It cannot isolate this decision at all.
What actually happened
Completion rose from 61.4% to 66.9%. Average order value fell about
£1.10 — enough to notice, not enough to outweigh the extra completed
orders. They kept it.
Note what did the work. The decision was named first, and the metric was
chosen to answer that decision — not pulled off whichever dashboard was
already open. Company revenue was the number the director would have
reached for, and it was the one number on the list that could not have
answered the question.
Every time: name the decision, then pick the metric closest to the thing
that changed, then add the guardrail that would catch the damage.