---
title: "How to Analyse Churn With AI Tools"
description: "Use AI to find the cohort that left, then separate the correlation from the reason: leakage, mix shift, the holdout test and twenty cancellations read by hand."
canonical_url: "https://builderscamp.com/guides/tools/ai-for-churn-analysis"
date_published: "2026-09-18"
date_modified: "2026-09-18"
author: "Andre Albuquerque, Mário Araújo"
publisher: "Builders Camp"
guide_class: "tools"
---

# How to use AI for churn analysis and still know why they left

**TL;DR:** Separate voluntary from involuntary churn first, because a failed card and a disappointed customer need different fixes and a blended number hides both. AI is genuinely fast at segmenting leavers and coding free-text reasons; it is unreliable at cause, because the strongest signals in churn data are usually consequences of the decision rather than causes of it. The two checks that matter are a holdout period and twenty cancellations read by hand.

## What is AI actually good at in a churn analysis?

Segmenting and summarising, not explaining. Point a model at an export of accounts with a churn flag and it will compare leavers against stayers across every dimension in the file in the time it takes you to open a spreadsheet, and it will rank the differences by size. That ranking is a useful place to start looking, and it is not a list of reasons.

The other genuinely strong use is text. Most teams collect cancellation reasons and nobody reads them past the first fifty, so the free-text box quietly becomes decoration. A model will code all of them against a taxonomy you wrote, which is a job that was previously not being done at all rather than one being done slowly.

## Define the churn event before you let anything analyse it

Voluntary and involuntary churn are different products of different failures. Involuntary churn is a payment that did not go through: an expired card, a hard decline, a bank that blocks a cross-border charge. It responds to retry schedules and dunning emails. Voluntary churn is a person deciding the product is not worth the money. Blending them produces a number that neither team can act on, and it is the single most common way a churn analysis wastes a quarter.

Then decide the unit. Account churn and seat churn move independently in B2B, and a product that loses half the seats in an account while keeping the logo looks healthy in one view and terminal in the other. Then fix the observation window, because churn measured over 30 days and churn measured over 90 days are not comparable and someone will compare them anyway. The definitions underneath all of this are in [churn rate](https://builderscamp.com/guides/glossary/churn-rate-product) and [product retention](https://builderscamp.com/guides/glossary/product-retention).

## Why your churn model keeps predicting the cancellation form

The strongest predictors in most churn datasets are consequences of the decision to leave, not causes of it. A support ticket tagged as a cancellation request, a seat removed, a plan downgrade, a login count that fell to zero: all of them separate leavers from stayers beautifully, and all of them arrive after the customer has decided. This is target leakage, and a model reaching for the most predictive feature will find it every time.

The scale of this problem in work that was actually reviewed before publication is worth knowing. Kapoor and Narayanan surveyed literature across research communities that had adopted machine learning and found 17 fields where leakage errors had been identified, collectively affecting 329 papers, and catalogued 8 distinct types of leakage ranging from textbook mistakes to open research problems. Those are peer-reviewed papers, not product teams working under a Thursday deadline, which is exactly why the number is worth citing here: if leakage survives review in published science, it will survive a churn deck nobody reviews at all.

The practical defence is unglamorous. Before you accept any driver, ask when the signal becomes observable relative to the churn date. If the answer is "in the last week," it is a symptom.

## The aggregate is probably lying to you

A blended churn rate can improve while every segment inside it gets worse, because the mix of customers changed. This is Simpson's paradox, and in a subscription business it shows up constantly: a self-serve push brings in a cohort that churns at half the rate of the enterprise base, the blended number falls, and both the self-serve and enterprise segments are individually deteriorating.

Ask for churn by segment and by signup cohort before you ask for the headline. If the two disagree, the segment view is the real one and the blended number is a story about your acquisition mix. Read [cohort analysis](https://builderscamp.com/guides/glossary/cohort-analysis) for the mechanics of building that view.

## The verification step that separates a correlate from a reason

Two checks, run in order, and neither is expensive.

- **Rebuild as of a date.** Cut the data at a point three months back, use only what existed then, and ask whether the analysis predicts who actually left afterwards. A driver that only works when the model can see the future is a symptom.
- **Read twenty cancellations by hand.** Not a summary of them, the raw text. If the factor your analysis ranked first never appears in twenty real conversations, you have found a correlate and not a reason.
- **Check the counterfactual population.** Find accounts with the same supposed driver that did not churn. If there are many, the driver is common rather than causal.

The third one is the check people skip, and it is the one that kills most confident findings. Half the "churn drivers" a model surfaces are properties of your customer base generally, not of the ones who left.

## What about predicting churn before it happens?

Useful, with a narrower claim than usually gets made. A model scoring accounts weekly can flag the ones worth a human conversation, and that is a real operational win for a customer success team with more accounts than hours. What it cannot do is tell you what to say when you get there, and a save play built on a leaked feature will target accounts that were already gone.

Score to prioritise attention. Do not score to explain behaviour. Those are different jobs and the second one needs a person who has talked to customers.

## How do you turn free text into something a team will act on?

Write the taxonomy first, yourself, from twenty or thirty cancellations you have already read. A model asked to invent categories will produce categories that describe the text rather than the decision, and you will end up with a bucket called "pricing" holding both "too expensive for what we got" and "we lost budget," which need opposite responses.

Then hand-code a sample against your own taxonomy, have the model code the same sample, and compare. Where you disagree, the disagreement is usually a definition problem in your categories rather than a model failure. Fix the taxonomy, then run the full set. The same discipline applies to any feedback pile, and [triaging customer feedback with Claude Code](https://builderscamp.com/guides/tools/claude-code-customer-feedback-triage) walks through the mechanics.

## What this changes about the retention conversation

The bottleneck in churn work has never been computation. It has been that nobody had time to read the cancellations, so the retention roadmap got built from the loudest anecdote in the last QBR. Doing the reading at volume is the actual change here, and it only pays off if somebody defined the categories and checked the labels.

## Run it on one quarter you already have an opinion about

Take a quarter where you believe you know why customers left, split voluntary from involuntary, read twenty cancellations, and see whether your belief survives. That exercise costs an afternoon and reorders most retention roadmaps.

[See the Product Analytics bootcamp](https://builderscamp.com/bootcamps/product-analytics?utm_source=guide&utm_medium=organic&utm_campaign=ai-for-churn-analysis)

Product Analytics spends its second half on exactly this: which behavioural signals predict churn, which ones should trigger an expansion play instead, and how to read a retention curve without fooling yourself. Growth for Product Managers puts churn inside the wider retention, acquisition and monetization loop across two weeks, and both sit in the [Data & Analytics Specialist Track](https://builderscamp.com/tracks/data-analytics-specialist) alongside the metrics and experimentation work. If your churn suspicion is really a funnel problem, start instead with [how to analyse a funnel using AI tools](https://builderscamp.com/guides/tools/ai-for-funnel-analysis).

## Frequently asked questions

### What can AI actually do in a churn analysis?

Three things well: segment the accounts that left across dozens of dimensions faster than you would by hand, code free-text cancellation reasons into a taxonomy you defined, and rank which behavioural differences between leavers and stayers are large enough to be worth investigating. None of those three is a statement about cause.

### Why does a churn model keep predicting the cancellation itself?

Target leakage. Features like a support ticket tagged cancellation, a seat downgrade, or a sudden drop in logins exist mostly because the account was already leaving. The model finds them because they are the strongest signal in the data, and they are useless for intervention because by the time they appear the decision is made.

### How do you tell a correlation from a reason?

Two tests. Rebuild the analysis as of a cutoff date, using only data that existed then, and see whether it predicts what happened after. Then read twenty actual cancellations by hand. If the driver your analysis ranked first never appears in those twenty conversations, treat it as a correlate until something else confirms it.

### Our overall churn improved but every segment got worse. Is the number wrong?

Both readings can be true at once, which is Simpson's paradox. If the mix of new customers shifted toward a segment that churns less, the blended rate improves while each segment deteriorates. Always read churn by segment and by cohort before reporting a blended figure.

### Should we separate voluntary and involuntary churn?

Always, and before any other cut. Involuntary churn is failed payments: expired cards, hard declines, currency issues. It responds to dunning logic and retry schedules, not to product changes. Blending it into one number means a billing fix and a product fix compete for the same slide.

### How much free-text feedback does AI coding actually save?

It turns a job nobody does into a job someone reviews. The honest version is that you still hand-code a sample yourself, compare the model's labels against yours, and only trust the full pass once the two broadly agree. The saving is in the remaining hundreds, not in the first thirty.

### Which Builders Camp bootcamp covers churn work?

Product Analytics covers the behavioural signals that predict churn and which ones should instead trigger an expansion play, across one week, 2 live sessions and 9 microlessons. Growth for Product Managers covers the wider retention, acquisition and monetization playbook across two weeks and 4 live sessions.

## Sources

- [Kapoor and Narayanan: Leakage and the Reproducibility Crisis in ML-based Science](https://arxiv.org/abs/2207.07048)
- [Wikipedia: Simpson's paradox](https://en.wikipedia.org/wiki/Simpson%27s_paradox)
- [Builders Camp: Product Analytics bootcamp](https://builderscamp.com/bootcamps/product-analytics)
- [Builders Camp: Growth for Product Managers bootcamp](https://builderscamp.com/bootcamps/growth-product-manager)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
