All 29 modules are open from the start — nothing here is
locked, and nothing costs anything. Sign in so your progress, titles and
credentials stay with you, on every device you use.
Free forever, with your Google account. No password, no payment.
Every product decision rests on assumptions. Hypothesis testing makes those assumptions explicit, ranks which ones are riskiest, and designs the cheapest experiment that could prove you wrong, before you spend months building on a guess.
Ready?
1
A Hypothesis Is a Bet You Can Lose
The problem is not the work. It is that the sentence cannot be wrong. There is no threshold, so every possible outcome is arguable — and something that cannot fail has not been tested, whatever anybody called it.
A hypothesis is a prediction specific enough to be proved wrong. That is the whole idea, and it is uncomfortable in exactly the way it needs to be. A forecaster who says “it may or may not rain” is never wrong and never useful. “70% chance of rain tomorrow in this city” can be scored against what actually happens.
The value is not in being right. It is in being able to stop. Adding one sentence — “we are wrong if activation does not reach 34% within three weeks” — changed how one team operated: the feature shipped, activation reached 29%, and they reworked the assumption instead of shipping a second feature on top of the first. The number did not make them correct. It made them able to notice.
So the test to apply to any plan is blunt: what result would tell us this was a bad idea? If nobody can answer, you do not have a hypothesis. You have an intention.
The two panels describe the same belief about onboarding. Only one of them can lose, and the last bullet in each is why: one ends in a month of argument, the other ends on a Tuesday in week three whether you like the number or not.
Everyday example, the weather forecast
"It might rain" isn't falsifiable, any weather confirms it. "There's an 80% chance of
rain by 3pm" is: check the sky at 3pm and you know if the forecast held. A hypothesis
needs the same sharpness, a claim specific enough that reality can say yes or no.
The bet that could not be lost
A team wrote "we believe improving onboarding will increase engagement". Six weeks later they shipped, engagement moved by a fraction of a percent, and the debate about whether the bet had paid off ran for a month. No threshold had been written down, so every outcome was arguable, which is another way of saying nothing had been tested.
Quick check
Which of these is a genuine, falsifiable hypothesis?
2
The Anatomy of a Good Hypothesis
A usable hypothesis has four parts, and leaving any of them out produces the vague version that cannot be scored.
Because we believe… — the evidence you already have. This anchors the guess in something real and makes it obvious when a plan rests on nothing but enthusiasm.
If we… — the specific change you will make. One change, described precisely enough that someone else could build it.
Then… will… — the predicted effect, with a number and a direction. Not “engagement will improve” but “week-one activation will rise from 24% to 30%”.
We are wrong if… — the threshold that kills it. This is the line teams leave off, and it is the only one that converts the other three into a hypothesis rather than a well-formatted intention.
Written out: Because 60% of new accounts never import data, if we add a sample dataset on first login, then week-one activation will rise from 24% to 30% within three weeks. We are wrong if it stays below 27%.
Notice what that last clause does. It commits you, in advance and in public, to a result you would find disappointing. That is precisely why it works — the decision gets made while you are still calm, rather than in the meeting where the number is disappointing and everyone is looking for a reason it is fine.
Because we believe…
The evidence you already have. It anchors the guess in something real, and makes it obvious when a plan is resting on nothing but enthusiasm.
If we…
One specific change, described precisely enough that somebody else could build it. One variable, not a bundle.
Then… will…
The predicted effect, with a number and a direction. Not “engagement will improve” but “week-one activation will rise from 24% to 30%”.
We are wrong if…
The threshold that kills it. This is the line teams leave off, and the only one that turns the other three into a hypothesis rather than a well-formatted intention.
Three of these rows are a plan. The highlighted one is what turns the stack into a hypothesis, and it is the row teams leave off — usually because writing it means agreeing, in advance, to a number you would be disappointed by. A prediction that cannot lose is a plan with better grammar.
Everyday example
A weather forecaster who says "it may or may not rain" is never wrong and never useful. "70% chance of rain tomorrow in this city" can be scored against what happens. A hypothesis has to be able to embarrass you.
The line that made it real
Adding a single sentence — "we are wrong if activation does not reach 34% within three weeks" — changed how a team ran. The feature shipped, activation reached 29%, and the team stopped and reworked the assumption instead of shipping a second feature on top of the first. The number did not make them right; it made them able to stop.
Quick check
A team writes: "We believe redesigning the dashboard will make it better." What's missing?
3
Finding the Leap of Faith Assumption
Every plan rests on a stack of assumptions. Most are safe, one or two are not, and the whole outcome usually depends on which ones you check first.
The dangerous one has a name — the leap of faith assumption — and it is the assumption where being wrong invalidates everything downstream. Not merely inconvenient: fatal.
Finding yours is a two-question sort. For each assumption ask how badly are we hurt if this is false? and how confident are we, honestly, on evidence rather than instinct? Anything that is high-impact and low-confidence belongs at the front of the queue.
Teams reliably test in the wrong order, and the reason is human rather than analytical. We test the assumptions we are most comfortable testing. Confirming something you already believe is pleasant, well-understood work that produces a satisfying result. Checking whether a partner will actually agree to something might end the project this week, so it drifts down the list.
That drift is expensive in a specific way: the cost of a wrong assumption rises with how long you take to find it. An assumption tested in week one costs an afternoon. The same assumption reached in week six has a design, a build and a set of commitments resting on it.
Ending a bad project early is a win, and the order you test in decides whether you are allowed to have it.
Test first
High risk + high uncertainty
You're not sure, and being wrong would sink the idea. This is your leap of faith.
Note, don't test yet
High risk, low uncertainty
Matters a lot but you already have strong evidence. Don't waste a test here.
Skip
Low risk either way
Wouldn't sink the idea even if wrong. Testing it burns time you don't need to spend.
The two left-hand cells are the comfortable ones, and they are the two that get tested. The highlighted cell is the only one whose answer can end the project — which is exactly why it drifts down the list, and why an assumption reached in week six costs a design and a build rather than an afternoon.
Everyday example
A business plan for a lemonade stand might assume children like lemonade, that a table can be borrowed, and that anyone walks down that street. Two of those are safe. The whole plan rests on the third, and that is the one to check on a Saturday morning before buying lemons.
Testing the safe assumption first
A team spent five weeks validating that users wanted a faster import — they did, overwhelmingly. The assumption that actually carried the project was that the third-party system would allow bulk reads, which nobody checked until build. It did not. Five weeks had gone into confirming the part that was never in doubt.
Quick check
A marketplace idea assumes: (1) sellers exist in this category, (2) sellers will pay a listing fee, (3) the checkout button should be blue. Which is the leap of faith to test first?
4
Designing a Test You Can Trust
Once you have a hypothesis and a threshold, the remaining question is how to gather evidence you can trust — and “build it and see” is usually the most expensive option available.
Match the test to the risk. If the risk is that nobody wants it, a landing page or a concierge test answers that in days without any product. If the risk is that people cannot use it, a prototype and five sessions will do. If the risk is that it will not scale, you want a technical spike, not a customer conversation. Building the real thing is the last resort, not the default.
And write the result down whichever way it goes, including what you would have done differently. A negative result you can find again in six months is worth more than a positive one nobody recorded, because the second one will simply be re-litigated by whoever joins next.
One team agreed in advance that a landing test needed 200 sign-ups. It got 90, so they shipped nothing — which felt like a wasted fortnight and was the cheapest answer they were ever going to get. None of the three boxes contains the product, and that is the point of the drawing.
Everyday example
Taste-testing a soup by dipping a spoon in the top of an unstirred pot tells you about the top of the pot. A test design is mostly the work of making sure the spoonful represents the soup.
Set the bar before you jump
A team runs a landing-page test for a feature idea and gets a 4% signup rate. Is that
good? Nobody decided in advance, so half the room calls it a win, half calls it a
flop, and the debate has nothing to do with data. Agree "we need 8% to justify
building this" beforehand, and the result answers itself.
Define success criteria before you run the test, not after you see the number.
The bar set before the jump
A landing-page test was agreed in advance to need 200 sign-ups to count. It got 90. The team shipped nothing, which felt like a wasted fortnight and was in fact the cheapest possible answer — the alternative was a quarter of engineering justified by a number nobody had committed to beforehand.
Quick check
Why should success criteria be set before running a test, not after?
Drill what you learned
Scenario 1
easy
A team says: "Our hypothesis is that redesigning the app will improve things."
What's wrong with this hypothesis?
Scenario 2
easy
A test shows a 3% lift and the team debates for a week whether that counts as a win — no threshold was set beforehand.
What should have prevented this debate?
Scenario 3
easy
An idea rests on: users have smartphones (very likely true), users will pay a $5/month fee (unknown), and onboarding uses blue buttons (trivial).
Which should be tested first?
Scenario 4
easy
A landing-page test hits the pre-agreed 8% signup bar. A stakeholder says: "Let's also check three other metrics we didn't originally care about, and if any looks good, call it a bigger win."
What's the risk here?
Scenario 5
easy
A hypothesis reads: "We believe adding SMS reminders for all users will raise 7-day retention by 15%, measured over 3 weeks."
What makes this a strong hypothesis?
Scenario 6
medium
Two teams ran the "same" experiment but got opposite results — one changed button copy and color at once, the other changed only color.
What principle was violated?
Scenario 7
medium
A PM says: "I already know push notifications will boost engagement, so let's skip testing and build the full feature."
What's the risk of skipping the test?
Scenario 8
medium
Success criteria said "10% lift required." The result comes in at 9.4%. A teammate says: "Close enough, let's call it a win."
What's the concern?
Scenario 9
medium
An assumption map shows one item in "high risk, high uncertainty" and five items in "low risk, low uncertainty."
Where should the next test focus?
Scenario 10
medium
A hypothesis says a redesign will help "users." Data later shows it helped power users but hurt new users.
What does this reveal about the original hypothesis?
Scenario 11
hard
A founder wants to skip writing formal hypotheses because "we move too fast for that."
What's the best response?
Scenario 12
hard
A test's sample was 40 users, mostly from one referral source, and the team wants to greenlight a company-wide rollout.
What's the concern with the test design?
Scenario 13
hard
A pricing hypothesis is deemed "too risky and important to test cheaply," so the team plans a six-month rebuild to test it properly.
Is that the right call?
Scenario 14
hard
After a failed test, a PM says: "The hypothesis was wrong, so this whole idea is dead."
Is that the right conclusion?
Scenario 15
hard
A team wants to test a new pricing model but worries a public test could alienate existing customers if it's visible.
What's a reasonable way to test with lower exposure?
Drilled it. Now apply it to a real situation.
Put it to work
From lesson 1
Airbnb rented a camera
Airbnb
Early Airbnb listings in New York were not converting. The data said
where the problem was but not why. The founders' belief — which they
have described many times since — was that the photographs were so poor
that people could not picture staying there.
Rather than build anything, they rented a camera, went to New York, and
photographed listings themselves.
What made this a hypothesis rather than an opinion?
Select all that apply — there are 3 to find.
What actually happened
Bookings on the professionally photographed listings rose substantially,
and Airbnb built a photography programme off the back of it. The
important part is the order: the belief was made falsifiable and cheaply
tested before any product existed to support it.
A hypothesis is a bet you can lose. If no result would change your mind,
you have a preference, not a hypothesis.
From lesson 2
Four attempts at the same hypothesis
A B2B invoicing tool
The team believes late payments are hurting retention. Four versions of
the hypothesis go on the board.
Order them from strongest to weakest.
Drag the rows, or use the arrows, then check.
"We believe that sending a payment reminder three days before the due date will raise on-time payment among freelancers on the Starter plan from 61% to above 70%. We will know we are right when that cohort's on-time rate moves in a four-week test."Strongest. A named population, a specific change, a measurable prediction with a baseline and a threshold, and a stated window. Every part of it can fail.
"We believe payment reminders will improve on-time payment rates."Second. A real causal claim you could test, but with no population, no size and no threshold — so any result at all can be read as support.
"We believe customers want payment reminders."Third. A claim about desire rather than behaviour, and desire is exactly what people misreport. Almost everyone says yes to a free feature in a survey.
"We believe reminders are a best practice we should match."Weakest. Nothing here is about your customers at all, and no outcome could disconfirm it. Competitor behaviour is a prompt to investigate, never a hypothesis.
What actually happened
Run on the first version, the four-week test came back at 64% — real
movement, well short of the 70% threshold. Because the threshold had
been set in advance, that was a complete "not worth the roadmap slot"
rather than an argument about whether three points counted as a win.
We believe [change] for [who] will produce [effect], and we will know we
are right when [measure] crosses [threshold] by [when].
From lesson 3
Which assumption would sink it
A same-day medicine delivery service
The plan: prescription medicines delivered within two hours. Four
assumptions sit underneath it, and the team has budget to test one
properly this quarter.
What the plan assumes
#
Assumption
If wrong
1
Patients want medicines faster than next-day
The whole proposition is worthless
2
Pharmacies will accept orders through our system
No supply; the whole proposition is worthless
3
Couriers can do two-hour windows profitably at our price
Margin is negative at scale
4
The app's tracking screen is clear enough
Support costs rise modestly
Assume all four are currently untested.
Work out which one to test first.
Step 1 of 3
Two assumptions would kill the product outright. What separates them?
The leap of faith is the assumption that is both fatal and genuinely uncertain — not simply the most important one.
Step 2 of 3
You test pharmacy willingness. What is the cheapest test that would actually settle it?
Two of the three signed. The third refused on a reason nobody had listed: their dispensing software could not export orders without a manual re-key.
Step 3 of 3
That objection turns out to apply to about 40% of independent pharmacies. What now?
A good test usually replaces its own question with a sharper one. That is the sign it was aimed at the right assumption.
What actually happened
They launched with the 60% that could export automatically, and the
re-key question became the following quarter's test. The demand
assumption — the one that felt most important — was never formally
tested, and never needed to be.
Rank assumptions on two axes: how fatal if wrong, and how uncertain
right now. The leap of faith is high on both, and it is rarely the one
that feels most important.
From lesson 4
The test that could not have failed
A language-learning app
Hypothesis: a daily streak reminder will improve seven-day retention.
The test ran for two weeks and came back at +11%, comfortably
significant. Before shipping, someone reads the setup.
How it was run
Decision
What they did
Who got the reminder
Users who had opted into notifications
Control group
Users who had not opted in
Metric
Seven-day retention
Duration
14 days
Guardrail
None
What is wrong with this test?
What actually happened
Re-run as a proper randomised split within the opted-in group only, the
effect was +2.4%. Real, worth shipping, and a quarter of what the first
result claimed — and the guardrail they added, notification opt-outs,
rose enough to cap the reminder at one a day.
Ask who ended up in each group and why. If people sorted themselves, the
test measures the sorting.