Cohort analysis
Cohort analysis is the practice of grouping customers by something fixed about their arrival — the week they signed up, the channel they came from — and following each group separately over time, instead of reading one blended number for everybody. It answers a question a total cannot: not "is our retention good", but "is it better for the people who joined this month than for the people who joined six months ago".
What a cohort actually is
A cohort is a group defined by something that already happened and can never change. The week someone signed up, the campaign that sent them, the plan they started on. That fixedness is the entire point: a customer can move between plans and countries, but nobody ever moves between signup weeks. Because the grouping is stable, you can line up month three for the January arrivals against month three for the June arrivals and know you are comparing the same stage of the same journey.
The reason the technique exists is that a blended average moves for two completely different reasons — the behaviour changed, or the mix changed — and a single number cannot tell you which. A company growing quickly always has a majority of young customers, so its blended retention drifts upward even when nothing about the product improved. A company that stops growing sees the same number fall for the same arithmetic reason. Splitting by cohort is what separates the story about your product from the story about your growth rate.
What cohorts show that totals hide
Once the data is split, four questions become answerable that were not answerable before.
- Whether the product got better. Compare the same elapsed week across successive signup cohorts. This is the only version of "did our change work" that survives a changing growth rate, because both sides of the comparison are at the same age.
- Which channel sends people who stay. Split by acquisition source as well as date. Two channels with an identical customer acquisition cost routinely produce completely different month-three curves, and the cheaper one is often the worse one — a difference a blended cost figure is structurally unable to show you.
- Whether retention has a floor. A curve that flattens means some group found durable value, and the height it flattens at is the real business. A curve that keeps declining toward zero has no floor, which means growth is refilling a bucket rather than building one.
- Whether a fix worked at all. Most onboarding and activation changes only touch people who arrive afterwards, so their effect is diluted into invisibility in a blended number for months. In a cohort view the improvement shows up immediately, in exactly one row.
The two ways a cohort table lies
The first failure is one we walked into ourselves: a cohort table built on a data source that expires will always make older cohorts look worse than they were. We were counting how many customers had actually held a conversation with their agent, and the first version of that number was low by roughly a factor of three — not because customers were quiet, but because the count was reading a store that gets cleared periodically. The further back a cohort sat, the more of its activity had already been deleted. What looked like a behaviour curve was a data-retention curve wearing its clothes, and it only became honest after the count moved to a store that persists. Before reading any cohort chart, ask how long the underlying records live, and whether that is longer than the oldest column on the chart.
The second failure is the opposite: rows that never leave. An audit of our own subscription records found the count of trial accounts overstated by roughly a third, because rows belonging to accounts that had since been deleted were never removed. In a lifetime total that is a mild inflation you might never notice. In a cohort table it is worse, because those rows form a group that never churns and quietly props up the retention of every month they sit in. So the second question to ask is what happens to a row when a customer leaves. If the answer is "nothing", the chart is measuring your database rather than your customers.
How to use it when the numbers are small
Cohort analysis is usually demonstrated as a heat map with thousands of users per row, which is not the situation most founders are in. At small scale that heat map is mostly noise, and reading percentages off it invites confident conclusions from four people changing their minds. The version that works early is deliberately blunt: pick one behaviour that means the product did its job, group by signup week, and look at the first three or four weeks only. You are not reading a percentage, you are asking whether the recent rows look different from the older ones.
Two disciplines make the small-numbers version trustworthy. Write down the definition of the behaviour before you look at the data, because otherwise "active" gets quietly redefined until the chart looks encouraging — the same failure that makes most marketing dashboards useless, discussed in how to measure if your marketing is working. And keep the acquisition source attached to every row from the beginning, because the day you want to compare channels by retention is always a day after the last date you recorded where anyone came from.
FAQ
- What is the difference between cohort analysis and segmentation?
- A segment is defined by an attribute the customer holds right now — plan, country, company size — and it can change, so a customer can drift between segments and blur the comparison. A cohort is defined by something that already happened and is permanent, usually when or how they arrived. In practice you use both together: cohort by signup month, then segment within it by channel or plan.
- How many users do I need before cohort analysis is worth doing?
- Fewer than people assume, as long as you read direction rather than precision. With a few dozen users per cohort you can see whether a curve flattens or falls off a cliff, and whether the newest arrivals behave unlike the oldest. What you cannot do at that size is read a two-point difference as real. Treat small cohorts as a prompt to go and ask someone a question, not as a result.
- What time window should a cohort cover?
- Match the product's natural rhythm. A tool people are meant to open daily wants weekly cohorts, because a month is long enough to hide the entire story. A product used once a month or once a quarter needs monthly cohorts and a lot of patience, since the first meaningful reading is several periods away. Picking a window shorter than the usage cycle produces charts that look like churn but are only calendar noise.
- Is a revenue cohort different from a retention cohort?
- Yes, and they can point in opposite directions at the same time. A revenue cohort can rise while the user cohort falls, because a shrinking group of remaining customers upgraded or grew — which is exactly the effect net revenue retention is designed to capture. Reading only the revenue view can hide the fact that most of the people who arrived are gone.
- What tool do I need to run a cohort analysis?
- A spreadsheet is genuinely enough to start, and starting there is useful because building the table by hand forces you to decide what counts as the arrival event and what counts as the behaviour. Product analytics tools will draw the chart for you later, faster and with more segments. They will not make those two decisions for you, and getting them wrong is the main reason cohort charts mislead.
An AI growth team that runs this for you
AgentCeres is a managed AI marketing team — you approve what ships. 14-day free trial, from $39/month.