◀ Course contents Part 1 · Module 1-01

Systems Thinking

See the whole system, not just the event in front of you

Products never live in isolation. They sit inside a web of users, teams, incentives, and metrics that push back on every change. Systems thinking is the habit of seeing those connections, so you fix root causes instead of chasing symptoms.

Ready?

1

Events, Patterns & Structure

Picture a support queue on a Monday morning. The same complaint arrives again: “I can’t find the export button.” You answer it, close the ticket, and move on. Next Monday it is there again, from someone else. And the Monday after that.

Most of us are trained to treat each of those tickets as its own small fire. Put it out, move to the next one. That instinct is not wrong exactly — the ticket does need answering — but it means you will be answering that same ticket for as long as the product exists.

Systems thinking is the habit of asking a different question: not “how do I handle this one?” but “what keeps producing these?” A system is simply a set of parts that affect each other — your users, your teams, your features, your metrics, the way people are rewarded. Pull on one part and the others move, whether or not you intended it. Your product is a system in exactly this sense. It is not a to-do list.

The most useful tool for this is the iceberg model. It says that anything you notice sits on top of three layers you cannot see. The event is the ticket you just read. The pattern is that it arrives every week. The structure is the navigation scheme that hides anything outside the top five actions. And underneath all of it sits a mental model — a belief someone holds, like “adding a menu item is a cost we should avoid.”

Here is the part worth remembering: you can fix any of the four layers, and they cost wildly different amounts and last wildly different lengths of time. Answer the ticket and it comes back next week. Change the structure and the tickets stop. Most teams spend their whole week at the top of the iceberg, because that is the only layer that arrives with a notification attached.

The iceberg model Four stacked layers. Above the waterline, Events — what just happened. Below it, in order: Patterns, the trend over time; Structure, the rules and incentives producing the pattern; and Mental models, the beliefs that hold the structure in place. Leverage increases with depth. waterline Events “Churn spiked this week” Patterns “It spikes every quarter after a price change” Structure Sales is paid on new logos, not retention Mental models “Growth means new customers” more leverage
Most teams argue at the top of the iceberg, where the news is. The cheapest lasting fixes are near the bottom — but that is also where nobody is looking, because structure is invisible until you go looking for it.

Everyday example

Think of a doctor. A patient keeps getting headaches. A weak doctor hands out a painkiller each visit (the event). A good doctor asks why they keep coming back, bad posture? Skipping meals? and fixes that (the structure). One treats the symptom forever; the other ends it. PMs face the same choice every week.

The iceberg model, four levels to look at

1. Events, what just happened. "Checkout crashed this morning." This is the tip of the iceberg, the only part you see easily.

2. Patterns, the same event over time. "Checkout crashes every Monday at peak hours." Spotting the pattern is the first sign there's something deeper.

3. Structure, the setup that produces the pattern. "We never sized our servers for the Monday traffic spike." This is where lasting fixes live.

4. Mental models, the beliefs that built the structure. "We assumed traffic is roughly flat all week." Change the belief and you stop building broken structures.

React at the event level and you firefight forever. Change the structure and the events simply stop happening.

Four levels, one complaint

The event was "users cannot find export". The pattern was that the same ticket arrived weekly for a year. The structure was a navigation scheme that hid anything not in the top five actions. The mental model was that adding a menu item is a cost to be avoided. Each level had a different fix and only the last one was permanent.

Quick check

Users keep filing the same complaint: they can't find the export button. The team keeps replying to each ticket individually. What's the systems-thinking move?

2

Reinforcing & Balancing Loops

Two products launch in the same month. One grows steadily for two years and then, over about six weeks, falls apart — faster than it ever grew. The other climbs quickly, flattens out around 40,000 users, and simply refuses to go higher no matter what the team ships.

Those two stories look nothing alike, but they are the same thing running in two different directions. Both are feedback loops — situations where the result of something circles back and becomes the cause of more of it. In plain terms: the output feeds the input.

There are only two kinds, and telling them apart explains most of what products do. A reinforcing loop amplifies. More users create more content, more content means better search results, better search brings more users — and round it goes, each turn bigger than the last. This is a snowball rolling downhill. Crucially, a snowball does not care which way it is pointed: the same loop that compounds your growth will compound your decline just as efficiently, which is what happened to the first product.

A balancing loop does the opposite — it pushes back toward a target. More users mean more load, more load means a slower app, a slower app means people leave, which reduces users. That is a thermostat. It is why the second product stopped at 40,000: nothing had broken, the system had simply found the ceiling where the pushback equalled the push.

When a metric that “should” keep climbing quietly flattens, the useful instinct is not to push harder. It is to go looking for what is pushing back — because something in the system is almost always doing exactly that, and no amount of extra effort at the top will out-run it.

Reinforcing and balancing loops Two circular diagrams side by side. On the left, a reinforcing loop: more users leads to more content, which leads to better search results, which leads back to more users — the loop compounds. On the right, a balancing loop: more users leads to slower load times, which leads to more churn, which reduces users — the loop self-corrects toward a limit. Reinforcing · compounds More users More content Better search Nothing stops it from inside Balancing · self-corrects More users Slower app More churn Finds a ceiling and sits there
Same shape, opposite behaviour. A reinforcing loop is your growth engine and your death spiral — direction depends only on which way it is already turning. A balancing loop is why a metric that “should” keep climbing quietly flattens: something in the system is pushing back.

Everyday example

Reinforcing is a snowball rolling downhill, it grabs more snow, gets bigger, so it grabs even more. It speeds up on its own. Balancing is the thermostat in your home, when the room gets too hot it turns off the heat, when it gets too cold it turns it back on, always pulling back to one target. One accelerates; the other steadies.

Snowball

Reinforcing loop

Growth feeds more growth (or decline feeds more decline). More users → more content → more users. Great for virality, dangerous for death spirals.

Thermostat

Balancing loop

The system pushes back toward a target. Rising load triggers slowdowns that reduce load. It resists change and creates stability , or stubborn plateaus.

The loop that ran both ways

A marketplace's reinforcing loop — more sellers, more selection, more buyers, more sellers — ran beautifully for two years. When a supply shock removed a group of sellers, the same loop ran in reverse at the same speed, and the team was surprised. A reinforcing loop has no preferred direction; it amplifies whatever it is given.

Quick check

On a social app, users who post get likes, which encourages them to post more, which attracts more users who post. What kind of loop is this?

3

Stocks, Flows & Delays

You ship a fix for churn. Three weeks pass and the number has not moved. Was the fix wrong? Should you ship something bigger on top of it?

That question is the most expensive one in product work, and answering it well needs two ideas that sound dull and are not.

A stock is a quantity you have right now — active users, cash in the bank, unresolved bugs. A flow is what changes it: sign-ups flow in, churn flows out. The level only rises when the inflow beats the outflow, which sounds obvious until you notice how many growth plans are entirely about the inflow and never mention the drain at all.

Then the part that catches everyone: delay. Cause and effect are usually separated in time, sometimes by months. Your onboarding fix changes what happens to people signing up today — but this month’s churn number is made mostly of people who signed up long before you touched anything. The change has been made; the result has not arrived yet.

This is why so many “this isn’t working” decisions are wrong. The team ships, waits an interval that feels reasonable, sees nothing, and ships something larger on top. When the number finally moves, nobody can attribute it, and they have paid for both changes.

Estimate the delay before you start, write it down, and hold your nerve until it has passed. Most premature conclusions in product are made inside a delay nobody had measured.

Stocks, flows and delay A tank labelled Stock, for example active users, with an inflow pipe on the left labelled sign-ups and an outflow pipe on the right labelled churn. Between the action taken and its effect on the outflow sits a delay of several weeks, drawn as a gap in the pipe. Stock active users Sign-ups inflow delay Churn outflow your fix lands here …but the number moves weeks later
A stock is what you have; flows are what change it. The trap is the gap: you ship the fix, the metric does not move, so you conclude it failed and ship something else on top. Most “this isn’t working” calls are made inside the delay.

Everyday example

Picture a bathtub. The water in it is the stock (say, your active users). The tap pouring in and the drain leaking out are the flows (signups and churn). The level only rises if the tap runs faster than the drain. And a delay? That's a tap with a long pipe, you turn the handle now, but the water shows up seconds later. Turn it more because "nothing's happening" and you'll suddenly flood the room.

Stock, what accumulates

A level that builds up: active users, technical debt, trust, cash. It changes slowly.

Flow, what changes the stock

The rates in and out: signups vs. churn, debt added vs. paid down.

Delay, the lag between action and result

Effects arrive late. A price change may not hit churn for months, so you can't judge it next week.

Acting inside the delay

A team shipped a fix for churn, saw no movement after three weeks, and shipped a second, larger change on top of it. When the numbers finally moved they could not tell which change had done it, and the second change carried a cost the first did not. The delay was roughly eight weeks; both decisions had been made inside it.

Quick check

You raise prices and churn doesn't move for six weeks, then it jumps. What trap does this illustrate?

4

Leverage Points & Side Effects

Two teams get the same instruction — “reduce churn” — and the same amount of time. One rewrites onboarding copy and changes a button colour. The other changes what the sales team is paid for. A year later only one of them has moved the number, and it is not the one that worked harder.

A leverage point is a place in a system where a small, well-chosen change produces a large effect. The striking thing about leverage points is that they are rarely where effort naturally goes.

Think of a ladder. At the bottom are numbers: copy, colours, button sizes. Above them, flows — the steps in a process and the order they happen in. Above those, rules — who may do what, what requires approval. Above those, the goal — what the organisation actually rewards. And at the top, the mental model: what everyone believes without ever saying it aloud.

Effort stays roughly constant as you climb. Impact does not.

So why does everyone crowd around the bottom two rungs? Because those are the ones you can change without asking anyone. The higher rungs belong to a manager, a finance lead, a founder — which means leverage is usually a conversation rather than a ticket, and conversations are much harder to put in a sprint.

One warning comes free with this idea: every change ripples. The second-order effect — what happens next, after the obvious result — often matters more than the first. Before pulling any lever, it is worth asking out loud: “and then what happens?”

A ladder of leverage Five rungs from weakest to strongest leverage: adjusting numbers such as copy and button colour; changing flows and process; changing the rules; changing the goal; and changing the mental model. Effort stays roughly constant while impact rises sharply up the ladder. impact same effort → Numbers copy, colour Flows process, steps Rules who may do what Goal what we reward Mental model what we believe
Teams spend most of their time on the two lowest rungs because those are the ones they are allowed to change without asking anyone. The rungs that move the most are owned by someone else — which is why leverage is usually a conversation, not a ticket.

Everyday example

Weak leverage: nagging people to "use less water." Strong leverage: putting a water meter on every home so the bill reflects use, behavior changes on its own. Same effort, wildly different result. The strongest leverage point of all is usually the goal or the rule you set, not the day-to-day knobs you fiddle with. As a PM, changing what a team is rewarded for beats a hundred pep talks.

When a metric fights back

A support team is measured on tickets closed per hour. Closures soar, because agents now close tickets fast without truly solving them. Customers re-open, re-contact, and overall load rises.

Optimizing one part in isolation (local optimization) can hurt the whole. Always ask: "and then what happens?"

The metric that fought back

A support team measured on tickets closed per hour saw closures rise and total load rise with them. Agents closed fast, customers re-opened, and each re-opened ticket counted again. The measurement had become a leverage point in the wrong direction.

Quick check

A riddle: "I'm the consequence you didn't plan for, the ripple that shows up a step later, often undoing the win you just celebrated." What am I?

Drill what you learned
Scenario 1 easy

Your team is rewarded purely on the number of features shipped per quarter. Shipping is up, but user satisfaction is sliding.

What's the systems-thinking read?

Scenario 2 easy

A referral feature quietly took off: invited friends invite more friends, and signups are compounding week over week.

How should you treat this?

Scenario 3 easy

Two teams each optimize their own metric: the growth team maximizes signups, the infra team minimizes server cost by capping capacity.

What does systems thinking predict?

Scenario 4 easy

Your onboarding funnel drops 20% one week. You add a bigger "Continue" button on the step where people leave. Drop-off barely moves.

What did the button fix miss?

Scenario 5 easy

A marketplace: more buyers attract more sellers, which widens selection, which attracts more buyers. Your CEO asks where one extra engineer would create the most long-term value.

Where does systems thinking point them?

Scenario 6 medium

To hit an aggressive quarterly signup goal, the team leans hard on deep discounts. Signups spike every quarter — but each quarter needs a bigger discount to hit the same number, and margins are thinning.

Which systems pattern is this, and what's the risk?

Scenario 7 medium

You ship infinite scroll. Average session time jumps 30% and everyone celebrates. Three months later, weekly active users are quietly falling.

What likely happened?

Scenario 8 medium

You set a team target: "cut average support-ticket resolution time by 40%." Resolution time drops beautifully. But repeat-contact rate and refunds both climb.

What went wrong with the metric?

Scenario 9 medium

The app slows down under load. Each time, ops adds more servers and it recovers — for a while. But it keeps happening, the bill keeps rising, and nobody has looked at the code in months.

Which archetype is this, and what's the move?

Scenario 10 medium

Your referral loop drove explosive growth for a year. Lately, invites are flat no matter how much you optimize the invite flow. The core audience is niche and mostly on the product already.

What does systems thinking say is happening?

Scenario 11 hard

A shared internal platform team supports every product squad. Each squad, acting sensibly, files "just one urgent request." The platform team is now permanently overloaded and slow for everyone.

Which archetype, and how do you fix it?

Scenario 12 hard

Trust in your brand took three years of consistent quality to build. One botched data-breach response burned a big chunk of it in a week. Signups are now sluggish despite unchanged marketing spend.

What stock-and-flow idea explains this?

Scenario 13 hard

A competitor adds a flashy feature. You rush to match it. They one-up you. You one-up back. Six months later, both products are bloated, neither of you gained share, and both teams are exhausted.

What trap are both companies caught in?

Scenario 14 hard

A leader says: "Our activation problem is simple — the sign-up form is too long. Cut the fields and activation will jump."

What's the systems-thinking caution here?

Scenario 15 hard

You're deciding how to lift retention. Option A: tweak the push-notification copy. Option B: change the core onboarding so new users reach their first real "win" in minute one instead of day three.

Which is the higher-leverage point, and why?

Drilled it. Now apply it to a real situation.

Put it to work
From lesson 1

Adding lanes and getting more traffic

Highway engineering

Widening a congested road reliably fails to fix congestion. The effect has been studied for decades and has a name — induced demand. Duranton and Turner's well-known analysis of US interstates found that vehicle miles travelled rose roughly in proportion to lane miles added.

The event is congestion. The pattern is that congestion returns after every widening. The structure is what produces both.

Order these from the event at the surface down to the structure underneath.

Drag the rows, or use the arrows, then check.

  1. "The motorway was jammed again this morning" The event. Visible, immediate, and the level at which almost all decisions get made.
  2. "Every widening is followed by a return to the same congestion within a few years" The pattern. Only visible over time, and it is the level where you first notice that the obvious fix is not working.
  3. "Faster journeys attract drivers who previously travelled off-peak, took another route, or took the train" The structure. The mechanism generating the pattern — and note it is a feedback loop, not a chain of causes.
  4. "Roads should be free at the point of use, and travel time is the only thing rationing them" The mental model. The belief underneath the structure, and the level almost nobody argues at — which is exactly why the structure persists.
From lesson 2

Two loops fighting each other

A marketplace for home services

More tradespeople joining means shorter waits, which attracts more customers, which attracts more tradespeople. It ran beautifully for two years and then stopped, with no obvious cause.

Work out what changed.

  1. Step 1 of 3

    What kind of loop is "more tradespeople → shorter waits → more customers → more tradespeople"?

    Growth that stops without an external cause almost always means a balancing loop has become strong enough to cancel the reinforcing one.

  2. Step 2 of 3

    Investigation finds average tradesperson earnings per week falling as more join. What does that create?

    Two loops on the same variable, pushing opposite ways. Growth stalls exactly where they balance.

  3. Step 3 of 3

    Where should the intervention go?

    They capped recruitment by postcode against local demand. Earnings recovered, churn fell, and the reinforcing loop restarted.

From lesson 3

The delay that made it worse

A customer support organisation

Support hires when the backlog grows and freezes hiring when it shrinks. A new hire takes about eleven weeks to be productive. The backlog has swung between 400 and 4,000 tickets four times in two years.

The cycle, one full swing
WeekBacklogAction taken
13,900Approve 12 hires
64,400Approve 6 more — it is still climbing
132,100First cohort productive
19600Second cohort productive; freeze hiring
31900Attrition; still frozen
443,700Approve 14 hires

What is producing the oscillation?

Select all that apply — there are 3 to find.

From lesson 4

The fix that moved the problem

A food delivery app

Late deliveries are hurting ratings. The team ships a change that pays couriers a bonus for on-time arrival. On-time rate rises from 71% to 89% in three weeks.

Over the same period: courier churn up 22%, orders with damaged packaging up 40%, and complaints about dangerous cycling appear for the first time.

What does this tell you about the intervention?

Select all that apply — there are 3 to find.

Notification