Builders Camp

Tools

How to use AI for cohort analysis without reading noise as a trend

Below roughly 100 members per cohort, a weekly retention grid is mostly noise: an observed 30 percent on 40 users carries a 95 percent interval of about 18 to 45 percent, computed with the Wilson score method. AI is worth using here because it makes the definition cheap to change, letting you rebuild the grid on a different anchor event or at account grain in one pass. Read the grid three ways, by row, by column and by diagonal, before you conclude anything.

What decisions does a retention cohort actually settle?

One: whether the thing you changed made later cohorts behave differently from earlier ones at the same age. That is the entire job, and almost every misreading comes from asking the grid a question it cannot answer, usually a question about why.

A cohort grid is a set of groups defined by when they started, measured at the same ages, so that calendar effects and lifecycle effects can be separated. Everything else follows from that. The terminology is in cohort analysis and the shape vocabulary in product retention.

Three decisions to make before you build anything

The anchor event. Cohorts are usually grouped by signup date, which quietly bundles acquisition quality and onboarding quality into one number. If a marketing push changes who is arriving, signup cohorts will show a retention change that has nothing to do with the product. Anchoring on first real use separates the two.

The return event. "Retained" has to mean something specific: logged in, completed the core action, or generated value. The three produce different curves from the same data, and login-based retention is the least informative of the three in almost every product.

The grain. In B2B, user-level retention and account-level retention diverge constantly. An account can lose four of five seats and still renew; a user-level grid records four departures and a healthy account records none. Amplitude and Mixpanel both document their own retention definitions, and the useful move is to read whichever one your team uses and then say out loud which of the three decisions above it made on your behalf.

Where AI genuinely helps: the definition becomes cheap to change

Retention tools are opinionated by design, which is a strength until you need a grid they do not build. A model working from a raw event export will rebuild the whole grid with a different anchor, a different return event, or account-level rollup in one pass, which turns "we should check whether that holds at account level" from a two-week request into a question you answer in the meeting.

That cheapness is the real gain, and it cuts both ways. When any definition is five minutes away, the temptation is to run definitions until one of them shows the story you wanted. Write down which definition you will report before you build the first grid, and record the others as sensitivity checks rather than candidates.

Read the grid three ways before concluding anything

The Product Analytics bootcamp teaches this as three views of the same table, and it is the fastest correction to most cohort misreadings.

  • Row view follows one cohort across time and shows where in the lifecycle people leave.
  • Column view compares cohorts at the same age and is the only honest test of whether an onboarding change worked.
  • Diagonal view cuts across every cohort in the same calendar period and is how outages, releases and seasonality show up.

The diagonal is the one people never look at, and it is the one that explains the most alarming-looking findings. A sharp drop in week three of March across every cohort regardless of signup date is not a lifecycle problem and not an onboarding regression. Something happened in March.

The sample size trap that invalidates the whole grid

A weekly cohort grid on a product with modest volume is mostly noise, and it looks exactly like signal because the cells are coloured. Take a cohort of 40 users where 12 came back: that is 30 percent retention, and the 95 percent confidence interval around it, computed with the Wilson score method, runs from roughly 18 percent to roughly 45 percent. The interval is 27 points wide. The next cohort showing 37.5 percent is not an improvement; the difference between 12 of 40 and 15 of 40 produces a test statistic of about 0.71, nowhere near the conventional threshold.

Run the same comparison at ten times the volume and it inverts. 120 of 400 against 150 of 400 gives a statistic of about 2.24, which does clear the 5 percent threshold, and the interval around 30 percent narrows to roughly 26 to 35 percent. Same percentages, different conclusion, and the only thing that changed was how many people were in the cell.

What that does not mean is that small products cannot use cohorts. It means they should use monthly cohorts, read the row view for lifecycle shape rather than the column view for week-on-week movement, and stop colouring cells that hold fewer than 100 people. A grid that cannot support a comparison should not be drawn as if it can.

Ask the model for the interval, not just the rate

The check is one line added to every request: return the cohort size alongside every retention percentage, and flag any cell computed on fewer than 100 members. A model will not volunteer this, and a dashboard will not either. Once the counts are visible next to the percentages, most of the week-to-week drama in a retention grid disappears on its own.

The second check is the resurrection rule. Decide explicitly whether a user absent in periods 3 to 5 and present in period 6 counts as retained in period 6. Both answers are defensible and they produce visibly different curves, which is why the smiling retention curve is so often reported as a finding when it is a counting rule.

Cohorts do not have to be time-based

The default grid groups by start date, but the more useful cut is often behavioural: everyone who completed a specific action in their first week against everyone who did not, tracked forward at the same ages. That comparison is where the "users who do X retain better" claims come from, and it is also where they go wrong, because the users who did X in week one were different people before they did anything.

Treat a behavioural cohort split as a shortlist of candidate onboarding milestones, not as proof that pushing everyone through X will raise retention. The only way to settle that is to push a randomly chosen half through X and watch what happens. AI will build the split in seconds and cannot supply the randomisation.

What about predicted or modelled retention?

Extrapolating a curve forward is reasonable when you have several cohorts that have already reached that age and the extrapolation matches them. It is guesswork when no cohort has ever lived that long, and a model asked to project 12-month retention from 6 weeks of data will produce a smooth, confident curve with no indication that it is inventing the second half.

If the projection is going into a business case, state the last period you actually observed and treat everything after it as an assumption someone has to sign.

Build the grid your tool refuses to build

The useful first exercise is not the retention chart you already have. It is the one your analytics tool cannot produce: your existing cohorts, re-anchored on first real value instead of signup, at account grain instead of user grain. If the story changes, you have learned something about your onboarding that the default grid was hiding.

See the Product Analytics bootcamp

Product Analytics works through cohort retention directly, including the row, column and diagonal views above and the mistakes teams make reading them, in one week of 2 live sessions and 9 microlessons. Data for Product Managers covers cohorts alongside metrics, funnels and experiments across two weeks, and both sit inside the Data & Analytics Specialist Track. For the significance question underneath the sample size trap, see statistical significance, and for the wider analysis workflow, how to use AI for product analytics.

Bootcamps referred in this Guide

Frequently asked questions

What does AI add to a cohort analysis that the analytics tool does not?

Mostly the ability to change the definition cheaply. Retention tools ship one or two cohorting rules; a model working from a raw export will rebuild the grid with a different anchor event, a different return event, or account-level rather than user-level grain in a single pass, so you can test whether your retention story survives a change of definition.

How small is too small for a weekly cohort?

As a rule of thumb, below about 100 members per cohort a weekly grid is mostly noise. At 40 members, an observed 30 percent retention carries a 95 percent Wilson interval of roughly 18 to 45 percent. At 400 members the same 30 percent narrows to roughly 26 to 35 percent. If you cannot reach 100, move to monthly cohorts.

What are the row, column and diagonal views?

Row follows one cohort across its life and shows where in the lifecycle people leave. Column compares different cohorts at the same age and is the only honest way to tell whether an onboarding change worked. Diagonal cuts across all cohorts in the same calendar period and is how you spot an outage, a release or a seasonal effect.

Should cohorts be grouped by signup date or by first value?

By whichever event you are trying to improve. Signup-date cohorts measure your acquisition and onboarding together. Cohorts anchored on first real use separate the two, which matters when a marketing push changes who is arriving. Run both if the answer differs, because the difference is the onboarding effect.

Why does my retention curve go up?

A curve that declines, flattens and then rises is the smiling retention curve, and it usually means resurrected users are being counted in later periods. Whether that is good news depends entirely on whether your grid counts a user as retained in period 6 if they were absent in periods 3 to 5. Decide that rule before reading the shape.

Can AI tell me why a cohort retained better?

No, and that is the same limit as every other analysis. It can tell you which cohorts differ and by how much, and it can list what else changed in that period. Attributing the difference to a release needs either a holdout group or a change that affected only part of the population.

Which Builders Camp bootcamp teaches cohort retention?

Product Analytics covers cohort retention analysis directly, including the biggest mistakes teams make reading the curves, across one week, 2 live sessions and 9 microlessons. Data for Product Managers covers cohorts alongside metrics, funnels and experiments across two weeks and 4 live sessions.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Mário Araújo

Mário Araújo

Product & Growth Leader | B2B | PLG Expert | Developer-focused products

Product & Growth Leader | B2B | PLG Expert | Developer-focused products

LinkedInMore guides by Mário Araújo

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Product Analytics bootcamp