◀ Course contents Part 3 · Module 3-02

Hypothesis Testing

Turn a guess into something you can prove wrong

Every product decision rests on assumptions. Hypothesis testing makes those assumptions explicit, ranks which ones are riskiest, and designs the cheapest experiment that could prove you wrong, before you spend months building on a guess.

Ready?

1

A Hypothesis Is a Bet You Can Lose

The problem is not the work. It is that the sentence cannot be wrong. There is no threshold, so every possible outcome is arguable — and something that cannot fail has not been tested, whatever anybody called it.

A hypothesis is a prediction specific enough to be proved wrong. That is the whole idea, and it is uncomfortable in exactly the way it needs to be. A forecaster who says “it may or may not rain” is never wrong and never useful. “70% chance of rain tomorrow in this city” can be scored against what actually happens.

The value is not in being right. It is in being able to stop. Adding one sentence — “we are wrong if activation does not reach 34% within three weeks” — changed how one team operated: the feature shipped, activation reached 29%, and they reworked the assumption instead of shipping a second feature on top of the first. The number did not make them correct. It made them able to notice.

So the test to apply to any plan is blunt: what result would tell us this was a bad idea? If nobody can answer, you do not have a hypothesis. You have an intention.

A hypothesis and an intention Two panels. An intention cannot be wrong and so cannot be tested; a hypothesis names a number and a threshold, which is what makes it possible to stop. Intention cannot fail "Improving onboarding will help" No number, no threshold Every outcome is arguable Debate runs for a month Hypothesis can fail "Activation reaches 34% in 3 weeks" A number and a deadline One outcome counts as wrong You know in week three
The two panels describe the same belief about onboarding. Only one of them can lose, and the last bullet in each is why: one ends in a month of argument, the other ends on a Tuesday in week three whether you like the number or not.

Everyday example, the weather forecast

"It might rain" isn't falsifiable, any weather confirms it. "There's an 80% chance of rain by 3pm" is: check the sky at 3pm and you know if the forecast held. A hypothesis needs the same sharpness, a claim specific enough that reality can say yes or no.

The bet that could not be lost

A team wrote "we believe improving onboarding will increase engagement". Six weeks later they shipped, engagement moved by a fraction of a percent, and the debate about whether the bet had paid off ran for a month. No threshold had been written down, so every outcome was arguable, which is another way of saying nothing had been tested.

Quick check

Which of these is a genuine, falsifiable hypothesis?

2

The Anatomy of a Good Hypothesis

A usable hypothesis has four parts, and leaving any of them out produces the vague version that cannot be scored.

Because we believe… — the evidence you already have. This anchors the guess in something real and makes it obvious when a plan rests on nothing but enthusiasm.

If we… — the specific change you will make. One change, described precisely enough that someone else could build it.

Then… will… — the predicted effect, with a number and a direction. Not “engagement will improve” but “week-one activation will rise from 24% to 30%”.

We are wrong if… — the threshold that kills it. This is the line teams leave off, and it is the only one that converts the other three into a hypothesis rather than a well-formatted intention.

Written out: Because 60% of new accounts never import data, if we add a sample dataset on first login, then week-one activation will rise from 24% to 30% within three weeks. We are wrong if it stays below 27%.

Notice what that last clause does. It commits you, in advance and in public, to a result you would find disappointing. That is precisely why it works — the decision gets made while you are still calm, rather than in the meeting where the number is disappointing and everyone is looking for a reason it is fine.

Because we believe…

The evidence you already have. It anchors the guess in something real, and makes it obvious when a plan is resting on nothing but enthusiasm.

If we…

One specific change, described precisely enough that somebody else could build it. One variable, not a bundle.

Then… will…

The predicted effect, with a number and a direction. Not “engagement will improve” but “week-one activation will rise from 24% to 30%”.

We are wrong if…

The threshold that kills it. This is the line teams leave off, and the only one that turns the other three into a hypothesis rather than a well-formatted intention.

The anatomy of a good hypothesis Four stacked parts: the change you will make, who it affects, the measurable effect you predict, and the threshold that would prove you wrong. Because we believe the evidence you already have If we… the specific change Then… will… the predicted, measurable effect We are wrong if the number that kills it
Three of these rows are a plan. The highlighted one is what turns the stack into a hypothesis, and it is the row teams leave off — usually because writing it means agreeing, in advance, to a number you would be disappointed by. A prediction that cannot lose is a plan with better grammar.

Everyday example

A weather forecaster who says "it may or may not rain" is never wrong and never useful. "70% chance of rain tomorrow in this city" can be scored against what happens. A hypothesis has to be able to embarrass you.

The line that made it real

Adding a single sentence — "we are wrong if activation does not reach 34% within three weeks" — changed how a team ran. The feature shipped, activation reached 29%, and the team stopped and reworked the assumption instead of shipping a second feature on top of the first. The number did not make them right; it made them able to stop.

Quick check

A team writes: "We believe redesigning the dashboard will make it better." What's missing?

3

Finding the Leap of Faith Assumption

Every plan rests on a stack of assumptions. Most are safe, one or two are not, and the whole outcome usually depends on which ones you check first.

The dangerous one has a name — the leap of faith assumption — and it is the assumption where being wrong invalidates everything downstream. Not merely inconvenient: fatal.

Finding yours is a two-question sort. For each assumption ask how badly are we hurt if this is false? and how confident are we, honestly, on evidence rather than instinct? Anything that is high-impact and low-confidence belongs at the front of the queue.

Teams reliably test in the wrong order, and the reason is human rather than analytical. We test the assumptions we are most comfortable testing. Confirming something you already believe is pleasant, well-understood work that produces a satisfying result. Checking whether a partner will actually agree to something might end the project this week, so it drifts down the list.

That drift is expensive in a specific way: the cost of a wrong assumption rises with how long you take to find it. An assumption tested in week one costs an afternoon. The same assumption reached in week six has a design, a build and a set of commitments resting on it.

Ending a bad project early is a win, and the order you test in decides whether you are allowed to have it.

Test first

High risk + high uncertainty

You're not sure, and being wrong would sink the idea. This is your leap of faith.

Note, don't test yet

High risk, low uncertainty

Matters a lot but you already have strong evidence. Don't waste a test here.

Skip

Low risk either way

Wouldn't sink the idea even if wrong. Testing it burns time you don't need to spend.

Which assumption to test first A grid crossing how badly you are hurt if an assumption is false against how confident you are. The high-impact, low-confidence quadrant is the leap of faith assumption and belongs first. Fatal, and confident verify cheaply, then move on Fatal, and unsure the leap of faith — test this first Minor, and confident ignore Minor, and unsure test later, if ever confident → unsure minor if wrong → fatal if wrong
The two left-hand cells are the comfortable ones, and they are the two that get tested. The highlighted cell is the only one whose answer can end the project — which is exactly why it drifts down the list, and why an assumption reached in week six costs a design and a build rather than an afternoon.

Everyday example

A business plan for a lemonade stand might assume children like lemonade, that a table can be borrowed, and that anyone walks down that street. Two of those are safe. The whole plan rests on the third, and that is the one to check on a Saturday morning before buying lemons.

Testing the safe assumption first

A team spent five weeks validating that users wanted a faster import — they did, overwhelmingly. The assumption that actually carried the project was that the third-party system would allow bulk reads, which nobody checked until build. It did not. Five weeks had gone into confirming the part that was never in doubt.

Quick check

A marketplace idea assumes: (1) sellers exist in this category, (2) sellers will pay a listing fee, (3) the checkout button should be blue. Which is the leap of faith to test first?

4

Designing a Test You Can Trust

Once you have a hypothesis and a threshold, the remaining question is how to gather evidence you can trust — and “build it and see” is usually the most expensive option available.

Match the test to the risk. If the risk is that nobody wants it, a landing page or a concierge test answers that in days without any product. If the risk is that people cannot use it, a prototype and five sessions will do. If the risk is that it will not scale, you want a technical spike, not a customer conversation. Building the real thing is the last resort, not the default.

And write the result down whichever way it goes, including what you would have done differently. A negative result you can find again in six months is worth more than a positive one nobody recorded, because the second one will simply be re-litigated by whoever joins next.

Match the test to the risk A chooser matching each risk to its cheapest test: a landing page or concierge test for desirability, a prototype for usability, and a technical spike for feasibility. Building the real thing is a last resort. Which risk are you retiring? nobody wants it Landing page or concierge test — days, no product they cannot use it Prototype and five sessions it will not scale Technical spike, not a customer conversation
One team agreed in advance that a landing test needed 200 sign-ups. It got 90, so they shipped nothing — which felt like a wasted fortnight and was the cheapest answer they were ever going to get. None of the three boxes contains the product, and that is the point of the drawing.

Everyday example

Taste-testing a soup by dipping a spoon in the top of an unstirred pot tells you about the top of the pot. A test design is mostly the work of making sure the spoonful represents the soup.

Set the bar before you jump

A team runs a landing-page test for a feature idea and gets a 4% signup rate. Is that good? Nobody decided in advance, so half the room calls it a win, half calls it a flop, and the debate has nothing to do with data. Agree "we need 8% to justify building this" beforehand, and the result answers itself.

Define success criteria before you run the test, not after you see the number.

The bar set before the jump

A landing-page test was agreed in advance to need 200 sign-ups to count. It got 90. The team shipped nothing, which felt like a wasted fortnight and was in fact the cheapest possible answer — the alternative was a quarter of engineering justified by a number nobody had committed to beforehand.

Quick check

Why should success criteria be set before running a test, not after?

Drill what you learned
Scenario 1 easy

A team says: "Our hypothesis is that redesigning the app will improve things."

What's wrong with this hypothesis?

Scenario 2 easy

A test shows a 3% lift and the team debates for a week whether that counts as a win — no threshold was set beforehand.

What should have prevented this debate?

Scenario 3 easy

An idea rests on: users have smartphones (very likely true), users will pay a $5/month fee (unknown), and onboarding uses blue buttons (trivial).

Which should be tested first?

Scenario 4 easy

A landing-page test hits the pre-agreed 8% signup bar. A stakeholder says: "Let's also check three other metrics we didn't originally care about, and if any looks good, call it a bigger win."

What's the risk here?

Scenario 5 easy

A hypothesis reads: "We believe adding SMS reminders for all users will raise 7-day retention by 15%, measured over 3 weeks."

What makes this a strong hypothesis?

Scenario 6 medium

Two teams ran the "same" experiment but got opposite results — one changed button copy and color at once, the other changed only color.

What principle was violated?

Scenario 7 medium

A PM says: "I already know push notifications will boost engagement, so let's skip testing and build the full feature."

What's the risk of skipping the test?

Scenario 8 medium

Success criteria said "10% lift required." The result comes in at 9.4%. A teammate says: "Close enough, let's call it a win."

What's the concern?

Scenario 9 medium

An assumption map shows one item in "high risk, high uncertainty" and five items in "low risk, low uncertainty."

Where should the next test focus?

Scenario 10 medium

A hypothesis says a redesign will help "users." Data later shows it helped power users but hurt new users.

What does this reveal about the original hypothesis?

Scenario 11 hard

A founder wants to skip writing formal hypotheses because "we move too fast for that."

What's the best response?

Scenario 12 hard

A test's sample was 40 users, mostly from one referral source, and the team wants to greenlight a company-wide rollout.

What's the concern with the test design?

Scenario 13 hard

A pricing hypothesis is deemed "too risky and important to test cheaply," so the team plans a six-month rebuild to test it properly.

Is that the right call?

Scenario 14 hard

After a failed test, a PM says: "The hypothesis was wrong, so this whole idea is dead."

Is that the right conclusion?

Scenario 15 hard

A team wants to test a new pricing model but worries a public test could alienate existing customers if it's visible.

What's a reasonable way to test with lower exposure?

Drilled it. Now apply it to a real situation.

Put it to work
From lesson 1

Airbnb rented a camera

Airbnb

Early Airbnb listings in New York were not converting. The data said where the problem was but not why. The founders' belief — which they have described many times since — was that the photographs were so poor that people could not picture staying there.

Rather than build anything, they rented a camera, went to New York, and photographed listings themselves.

What made this a hypothesis rather than an opinion?

Select all that apply — there are 3 to find.

From lesson 2

Four attempts at the same hypothesis

A B2B invoicing tool

The team believes late payments are hurting retention. Four versions of the hypothesis go on the board.

Order them from strongest to weakest.

Drag the rows, or use the arrows, then check.

  1. "We believe that sending a payment reminder three days before the due date will raise on-time payment among freelancers on the Starter plan from 61% to above 70%. We will know we are right when that cohort's on-time rate moves in a four-week test." Strongest. A named population, a specific change, a measurable prediction with a baseline and a threshold, and a stated window. Every part of it can fail.
  2. "We believe payment reminders will improve on-time payment rates." Second. A real causal claim you could test, but with no population, no size and no threshold — so any result at all can be read as support.
  3. "We believe customers want payment reminders." Third. A claim about desire rather than behaviour, and desire is exactly what people misreport. Almost everyone says yes to a free feature in a survey.
  4. "We believe reminders are a best practice we should match." Weakest. Nothing here is about your customers at all, and no outcome could disconfirm it. Competitor behaviour is a prompt to investigate, never a hypothesis.
From lesson 3

Which assumption would sink it

A same-day medicine delivery service

The plan: prescription medicines delivered within two hours. Four assumptions sit underneath it, and the team has budget to test one properly this quarter.

What the plan assumes
#AssumptionIf wrong
1Patients want medicines faster than next-dayThe whole proposition is worthless
2Pharmacies will accept orders through our systemNo supply; the whole proposition is worthless
3Couriers can do two-hour windows profitably at our priceMargin is negative at scale
4The app's tracking screen is clear enoughSupport costs rise modestly

Assume all four are currently untested.

Work out which one to test first.

  1. Step 1 of 3

    Two assumptions would kill the product outright. What separates them?

    The leap of faith is the assumption that is both fatal and genuinely uncertain — not simply the most important one.

  2. Step 2 of 3

    You test pharmacy willingness. What is the cheapest test that would actually settle it?

    Two of the three signed. The third refused on a reason nobody had listed: their dispensing software could not export orders without a manual re-key.

  3. Step 3 of 3

    That objection turns out to apply to about 40% of independent pharmacies. What now?

    A good test usually replaces its own question with a sharper one. That is the sign it was aimed at the right assumption.

From lesson 4

The test that could not have failed

A language-learning app

Hypothesis: a daily streak reminder will improve seven-day retention. The test ran for two weeks and came back at +11%, comfortably significant. Before shipping, someone reads the setup.

How it was run
DecisionWhat they did
Who got the reminderUsers who had opted into notifications
Control groupUsers who had not opted in
MetricSeven-day retention
Duration14 days
GuardrailNone

What is wrong with this test?

Notification