All 29 modules are open from the start — nothing here is
locked, and nothing costs anything. Sign in so your progress, titles and
credentials stay with you, on every device you use.
Free forever, with your Google account. No password, no payment.
See the whole system, not just the event in front of you
Products never live in isolation. They sit inside a web of users, teams,
incentives, and metrics that push back on every change. Systems thinking is
the habit of seeing those connections, so you fix root causes instead of
chasing symptoms.
Ready?
1
Events, Patterns & Structure
Picture a support queue on a Monday morning. The same complaint arrives again: “I can’t find the export button.” You answer it, close the ticket, and move on. Next Monday it is there again, from someone else. And the Monday after that.
Most of us are trained to treat each of those tickets as its own small fire. Put it out, move to the next one. That instinct is not wrong exactly — the ticket does need answering — but it means you will be answering that same ticket for as long as the product exists.
Systems thinking is the habit of asking a different question: not “how do I handle this one?” but “what keeps producing these?” A system is simply a set of parts that affect each other — your users, your teams, your features, your metrics, the way people are rewarded. Pull on one part and the others move, whether or not you intended it. Your product is a system in exactly this sense. It is not a to-do list.
The most useful tool for this is the iceberg model. It says that anything you notice sits on top of three layers you cannot see. The event is the ticket you just read. The pattern is that it arrives every week. The structure is the navigation scheme that hides anything outside the top five actions. And underneath all of it sits a mental model — a belief someone holds, like “adding a menu item is a cost we should avoid.”
Here is the part worth remembering: you can fix any of the four layers, and they cost wildly different amounts and last wildly different lengths of time. Answer the ticket and it comes back next week. Change the structure and the tickets stop. Most teams spend their whole week at the top of the iceberg, because that is the only layer that arrives with a notification attached.
Most teams argue at the top of the iceberg, where the news is. The
cheapest lasting fixes are near the bottom — but that is also where
nobody is looking, because structure is invisible until you go
looking for it.
Everyday example
Think of a doctor. A patient keeps getting headaches. A weak doctor
hands out a painkiller each visit (the event). A good doctor
asks why they keep coming back, bad posture? Skipping meals? and
fixes that (the structure). One treats the symptom
forever; the other ends it. PMs face the same choice every week.
The iceberg model, four levels to look
at
1. Events, what just happened. "Checkout crashed
this morning." This is the tip of the iceberg, the only part you see
easily.
2. Patterns, the same event over time. "Checkout
crashes every Monday at peak hours." Spotting the pattern is the first
sign there's something deeper.
3. Structure, the setup that produces the pattern.
"We never sized our servers for the Monday traffic spike." This is
where lasting fixes live.
4. Mental models, the beliefs that built the
structure. "We assumed traffic is roughly flat all week." Change the
belief and you stop building broken structures.
React at the event level and you firefight forever. Change the
structure and the events simply stop happening.
Four levels, one complaint
The event was "users cannot find export". The pattern was that the same ticket arrived weekly for a year. The structure was a navigation scheme that hid anything not in the top five actions. The mental model was that adding a menu item is a cost to be avoided. Each level had a different fix and only the last one was permanent.
Quick check
Users keep filing the same complaint: they can't find the export
button. The team keeps replying to each ticket individually. What's
the systems-thinking move?
2
Reinforcing & Balancing Loops
Two products launch in the same month. One grows steadily for two years and then, over about six weeks, falls apart — faster than it ever grew. The other climbs quickly, flattens out around 40,000 users, and simply refuses to go higher no matter what the team ships.
Those two stories look nothing alike, but they are the same thing running in two different directions. Both are feedback loops — situations where the result of something circles back and becomes the cause of more of it. In plain terms: the output feeds the input.
There are only two kinds, and telling them apart explains most of what products do. A reinforcing loop amplifies. More users create more content, more content means better search results, better search brings more users — and round it goes, each turn bigger than the last. This is a snowball rolling downhill. Crucially, a snowball does not care which way it is pointed: the same loop that compounds your growth will compound your decline just as efficiently, which is what happened to the first product.
A balancing loop does the opposite — it pushes back toward a target. More users mean more load, more load means a slower app, a slower app means people leave, which reduces users. That is a thermostat. It is why the second product stopped at 40,000: nothing had broken, the system had simply found the ceiling where the pushback equalled the push.
When a metric that “should” keep climbing quietly flattens, the useful instinct is not to push harder. It is to go looking for what is pushing back — because something in the system is almost always doing exactly that, and no amount of extra effort at the top will out-run it.
Same shape, opposite behaviour. A reinforcing loop is your growth
engine and your death spiral — direction depends only on
which way it is already turning. A balancing loop is why a metric
that “should” keep climbing quietly flattens:
something in the system is pushing back.
Everyday example
Reinforcing is a snowball rolling downhill, it grabs
more snow, gets bigger, so it grabs even more. It speeds up on its
own. Balancing is the thermostat in your home, when
the room gets too hot it turns off the heat, when it gets too cold it
turns it back on, always pulling back to one target. One accelerates;
the other steadies.
Snowball
Reinforcing loop
Growth feeds more growth (or decline feeds more decline). More users
→ more content → more users. Great for virality, dangerous for death
spirals.
Thermostat
Balancing loop
The system pushes back toward a target. Rising load triggers
slowdowns that reduce load. It resists change and creates stability
, or stubborn plateaus.
The loop that ran both ways
A marketplace's reinforcing loop — more sellers, more selection, more buyers, more sellers — ran beautifully for two years. When a supply shock removed a group of sellers, the same loop ran in reverse at the same speed, and the team was surprised. A reinforcing loop has no preferred direction; it amplifies whatever it is given.
Quick check
On a social app, users who post get likes, which encourages them to
post more, which attracts more users who post. What kind of loop is
this?
3
Stocks, Flows & Delays
You ship a fix for churn. Three weeks pass and the number has not moved. Was the fix wrong? Should you ship something bigger on top of it?
That question is the most expensive one in product work, and answering it well needs two ideas that sound dull and are not.
A stock is a quantity you have right now — active users, cash in the bank, unresolved bugs. A flow is what changes it: sign-ups flow in, churn flows out. The level only rises when the inflow beats the outflow, which sounds obvious until you notice how many growth plans are entirely about the inflow and never mention the drain at all.
Then the part that catches everyone: delay. Cause and effect are usually separated in time, sometimes by months. Your onboarding fix changes what happens to people signing up today — but this month’s churn number is made mostly of people who signed up long before you touched anything. The change has been made; the result has not arrived yet.
This is why so many “this isn’t working” decisions are wrong. The team ships, waits an interval that feels reasonable, sees nothing, and ships something larger on top. When the number finally moves, nobody can attribute it, and they have paid for both changes.
Estimate the delay before you start, write it down, and hold your nerve until it has passed. Most premature conclusions in product are made inside a delay nobody had measured.
A stock is what you have; flows are what change it. The trap is the
gap: you ship the fix, the metric does not move, so you conclude it
failed and ship something else on top.
Most “this isn’t working” calls are made inside the delay.
Everyday example
Picture a bathtub. The water in it is the stock (say,
your active users). The tap pouring in and the drain leaking out are
the flows
(signups and churn). The level only rises if the tap runs faster than
the drain. And a
delay? That's a tap with a long pipe, you turn the
handle now, but the water shows up seconds later. Turn it more because
"nothing's happening" and you'll suddenly flood the room.
Stock, what accumulates
A level that builds up: active users, technical debt, trust, cash.
It changes slowly.
Flow, what changes the stock
The rates in and out: signups vs. churn, debt added vs. paid down.
Delay, the lag between action and result
Effects arrive late. A price change may not hit churn for months, so you can't judge it next week.
Acting inside the delay
A team shipped a fix for churn, saw no movement after three weeks, and shipped a second, larger change on top of it. When the numbers finally moved they could not tell which change had done it, and the second change carried a cost the first did not. The delay was roughly eight weeks; both decisions had been made inside it.
Quick check
You raise prices and churn doesn't move for six weeks, then it jumps.
What trap does this illustrate?
4
Leverage Points & Side Effects
Two teams get the same instruction — “reduce churn” — and the same amount of time. One rewrites onboarding copy and changes a button colour. The other changes what the sales team is paid for. A year later only one of them has moved the number, and it is not the one that worked harder.
A leverage point is a place in a system where a small, well-chosen change produces a large effect. The striking thing about leverage points is that they are rarely where effort naturally goes.
Think of a ladder. At the bottom are numbers: copy, colours, button sizes. Above them, flows — the steps in a process and the order they happen in. Above those, rules — who may do what, what requires approval. Above those, the goal — what the organisation actually rewards. And at the top, the mental model: what everyone believes without ever saying it aloud.
Effort stays roughly constant as you climb. Impact does not.
So why does everyone crowd around the bottom two rungs? Because those are the ones you can change without asking anyone. The higher rungs belong to a manager, a finance lead, a founder — which means leverage is usually a conversation rather than a ticket, and conversations are much harder to put in a sprint.
One warning comes free with this idea: every change ripples. The second-order effect — what happens next, after the obvious result — often matters more than the first. Before pulling any lever, it is worth asking out loud: “and then what happens?”
Teams spend most of their time on the two lowest rungs because those
are the ones they are allowed to change without asking anyone.
The rungs that move the most are owned by someone else
— which is why leverage is usually a conversation, not a ticket.
Everyday example
Weak leverage: nagging people to "use less water." Strong leverage:
putting a water meter on every home so the bill reflects use, behavior changes on its own. Same effort, wildly different result. The
strongest leverage point of all is usually the
goal or the rule you set, not the day-to-day knobs
you fiddle with. As a PM, changing what a team is
rewarded for beats a hundred pep talks.
When a metric fights back
A support team is measured on tickets closed per hour. Closures soar, because agents now close tickets fast without truly solving them.
Customers re-open, re-contact, and overall load rises.
Optimizing one part in isolation (local optimization) can hurt the
whole. Always ask: "and then what happens?"
The metric that fought back
A support team measured on tickets closed per hour saw closures rise and total load rise with them. Agents closed fast, customers re-opened, and each re-opened ticket counted again. The measurement had become a leverage point in the wrong direction.
Quick check
A riddle: "I'm the consequence you didn't plan for, the ripple that
shows up a step later, often undoing the win you just celebrated."
What am I?
Drill what you learned
Scenario 1
easy
Your team is rewarded purely on the number of features shipped per quarter. Shipping is up, but user satisfaction is sliding.
What's the systems-thinking read?
Scenario 2
easy
A referral feature quietly took off: invited friends invite more friends, and signups are compounding week over week.
How should you treat this?
Scenario 3
easy
Two teams each optimize their own metric: the growth team maximizes signups, the infra team minimizes server cost by capping capacity.
What does systems thinking predict?
Scenario 4
easy
Your onboarding funnel drops 20% one week. You add a bigger "Continue" button on the step where people leave. Drop-off barely moves.
What did the button fix miss?
Scenario 5
easy
A marketplace: more buyers attract more sellers, which widens selection, which attracts more buyers. Your CEO asks where one extra engineer would create the most long-term value.
Where does systems thinking point them?
Scenario 6
medium
To hit an aggressive quarterly signup goal, the team leans hard on deep discounts. Signups spike every quarter — but each quarter needs a bigger discount to hit the same number, and margins are thinning.
Which systems pattern is this, and what's the risk?
Scenario 7
medium
You ship infinite scroll. Average session time jumps 30% and everyone celebrates. Three months later, weekly active users are quietly falling.
What likely happened?
Scenario 8
medium
You set a team target: "cut average support-ticket resolution time by 40%." Resolution time drops beautifully. But repeat-contact rate and refunds both climb.
What went wrong with the metric?
Scenario 9
medium
The app slows down under load. Each time, ops adds more servers and it recovers — for a while. But it keeps happening, the bill keeps rising, and nobody has looked at the code in months.
Which archetype is this, and what's the move?
Scenario 10
medium
Your referral loop drove explosive growth for a year. Lately, invites are flat no matter how much you optimize the invite flow. The core audience is niche and mostly on the product already.
What does systems thinking say is happening?
Scenario 11
hard
A shared internal platform team supports every product squad. Each squad, acting sensibly, files "just one urgent request." The platform team is now permanently overloaded and slow for everyone.
Which archetype, and how do you fix it?
Scenario 12
hard
Trust in your brand took three years of consistent quality to build. One botched data-breach response burned a big chunk of it in a week. Signups are now sluggish despite unchanged marketing spend.
What stock-and-flow idea explains this?
Scenario 13
hard
A competitor adds a flashy feature. You rush to match it. They one-up you. You one-up back. Six months later, both products are bloated, neither of you gained share, and both teams are exhausted.
What trap are both companies caught in?
Scenario 14
hard
A leader says: "Our activation problem is simple — the sign-up form is too long. Cut the fields and activation will jump."
What's the systems-thinking caution here?
Scenario 15
hard
You're deciding how to lift retention. Option A: tweak the push-notification copy. Option B: change the core onboarding so new users reach their first real "win" in minute one instead of day three.
Which is the higher-leverage point, and why?
Drilled it. Now apply it to a real situation.
Put it to work
From lesson 1
Adding lanes and getting more traffic
Highway engineering
Widening a congested road reliably fails to fix congestion. The effect
has been studied for decades and has a name — induced demand. Duranton
and Turner's well-known analysis of US interstates found that vehicle
miles travelled rose roughly in proportion to lane miles added.
The event is congestion. The pattern is that congestion returns after
every widening. The structure is what produces both.
Order these from the event at the surface down to the structure underneath.
Drag the rows, or use the arrows, then check.
"The motorway was jammed again this morning"The event. Visible, immediate, and the level at which almost all decisions get made.
"Every widening is followed by a return to the same congestion within a few years"The pattern. Only visible over time, and it is the level where you first notice that the obvious fix is not working.
"Faster journeys attract drivers who previously travelled off-peak, took another route, or took the train"The structure. The mechanism generating the pattern — and note it is a feedback loop, not a chain of causes.
"Roads should be free at the point of use, and travel time is the only thing rationing them"The mental model. The belief underneath the structure, and the level almost nobody argues at — which is exactly why the structure persists.
What actually happened
Interventions aimed at the event — one more lane — are consumed by the
structure. Interventions aimed at the structure, such as congestion
charging, change what the loop does, which is why they work and why they
are politically much harder.
Events, patterns, structure, mental models. Almost every failed fix was
aimed at the event, and almost every durable one was aimed lower down.
From lesson 2
Two loops fighting each other
A marketplace for home services
More tradespeople joining means shorter waits, which attracts more
customers, which attracts more tradespeople. It ran beautifully for two
years and then stopped, with no obvious cause.
Work out what changed.
Step 1 of 3
What kind of loop is "more tradespeople → shorter waits → more customers → more tradespeople"?
Growth that stops without an external cause almost always means a balancing loop has become strong enough to cancel the reinforcing one.
Step 2 of 3
Investigation finds average tradesperson earnings per week falling as more join. What does that create?
Two loops on the same variable, pushing opposite ways. Growth stalls exactly where they balance.
Step 3 of 3
Where should the intervention go?
They capped recruitment by postcode against local demand. Earnings recovered, churn fell, and the reinforcing loop restarted.
What actually happened
Growth resumed within two quarters — from recruiting fewer
tradespeople in the areas that felt most successful, which is the kind of
move that only makes sense once both loops are on the page.
Growth that stalls without an external cause usually means a balancing
loop has caught up. Look for the counterforce before pushing harder.
From lesson 3
The delay that made it worse
A customer support organisation
Support hires when the backlog grows and freezes hiring when it shrinks.
A new hire takes about eleven weeks to be productive. The backlog has
swung between 400 and 4,000 tickets four times in two years.
The cycle, one full swing
Week
Backlog
Action taken
1
3,900
Approve 12 hires
6
4,400
Approve 6 more — it is still climbing
13
2,100
First cohort productive
19
600
Second cohort productive; freeze hiring
31
900
Attrition; still frozen
44
3,700
Approve 14 hires
What is producing the oscillation?
Select all that apply — there are 3 to find.
What actually happened
They changed the rule rather than the effort: hire against a rolling
twelve-week forecast rather than today's backlog, and never adjust more
than once a quarter. The swing narrowed to 900-1,600 within a year, with
no change to headcount budget.
A delay between action and effect turns a balancing loop into an
oscillation. The fix is almost never more effort — it is acting on the
forecast rather than the reading.
From lesson 4
The fix that moved the problem
A food delivery app
Late deliveries are hurting ratings. The team ships a change that pays
couriers a bonus for on-time arrival. On-time rate rises from 71% to 89%
in three weeks.
Over the same period: courier churn up 22%, orders with damaged
packaging up 40%, and complaints about dangerous cycling appear for the
first time.
What does this tell you about the intervention?
Select all that apply — there are 3 to find.
What actually happened
They kept the bonus and added two conditions: it is void on a damaged
delivery, and it is calculated weekly rather than per drop so a single
bad-traffic run costs nothing. On-time settled at 85%, and the damage and
churn numbers returned to baseline.
Before pulling a lever, ask what else the actor controls that they might
trade away. High-leverage points are high-leverage in every direction.