Builders Camp

Other Guides

Pre/post vs difference in differences vs A/B test: how to measure a product change

Use an A/B test whenever you can randomise users and have enough traffic, because it is the only one of the three methods that isolates the change from everything else. Use difference in differences when a change hits a whole market or group and you have a comparison group that moved in step with it beforehand. Use pre/post only for large, obvious effects in a stable metric, and say plainly that it cannot rule out seasonality or other launches.

Three methods answer "did this change work?", and they are not interchangeable. An A/B test randomises who gets the change. Difference in differences compares the change over time in a group that got it with a group that did not. Pre/post compares a metric before and after, with no comparison group at all. Each one is right in a different situation, and the most common mistake is using pre/post where the effect is small and the calendar is busy.

The method a PM can use is usually decided by the launch, not by preference. Harvard Business Review's account of experimentation at Microsoft, by Ron Kohavi and Stefan Thomke, shows what randomised tests can find: a Bing headline change that increased revenue by 12 percent, worth more than $100 million a year in the United States. That result was only credible because users were randomly split. Many product changes cannot be split that way: a price that must be the same for everyone in a country, a new delivery zone, a policy for a whole marketplace. For those, you need a different method and a clear sense of what it can and cannot prove.

What does each method compare?

Each method compares a different pair of numbers, and what it compares decides what it can prove.

Method What it compares What it can rule out What it cannot rule out
Pre/post Treated group before vs after Nothing systematically Seasonality, trends, other launches, outside events
Difference in differences Change in treated group vs change in comparison group Events that hit both groups equally, shared seasonality Anything that hits only one group, groups that were already diverging
A/B test Randomly assigned groups over the same period Almost all outside factors, since they hit both groups by chance equally Long-term effects beyond the test window, network effects between groups

The table reads left to right as a ladder of evidence. Each step up removes a class of explanation, and each step costs more setup.

When is a pre/post analysis good enough?

A pre/post analysis is good enough when the effect you expect is much larger than the metric's normal movement and nothing else changed at the same time.

A good example, with invented numbers for illustration: a payment provider bug was failing one in five card payments on your checkout. You fix it on a Tuesday. Checkout completion jumps from a steady band around 55 percent to a steady band around 68 percent the next day and stays there. No campaign launched, no season turned. A pre/post comparison is fine here, because the size and timing of the jump leave no competing explanation.

A bad example: a language-learning app ships a new onboarding checklist on 1 September, and week-one retention rises from 40 to 45 percent over the following month. Suppose September is also when your paid campaign for students starts, bringing sign-ups with higher intent. The checklist may have helped. The change in who signed up could produce the same rise. Pre/post cannot tell the two apart, and presenting the 5 points as the checklist's effect overstates what you know.

The test before trusting a pre/post number: look at the metric's week-to-week movement over the previous few months. If the change you measured is not clearly bigger than that normal movement, pre/post cannot support the claim.

How does difference in differences work?

Difference in differences measures the change in the treated group, measures the change in a comparison group over the same period, and treats the gap between those two changes as the effect. Whatever pushed both groups up or down, such as a season, a holiday or a competitor's campaign that hit both markets, cancels out.

The best-known example comes from economics. David Card and Alan Krueger, in NBER Working Paper 4509, studied New Jersey's minimum wage rise from $4.25 to $5.05 per hour in April 1992. They surveyed 410 fast food restaurants in New Jersey and in neighbouring Pennsylvania, where the minimum stayed fixed, before and after the rise. Comparing the change in New Jersey with the change in Pennsylvania, they found New Jersey restaurants increased employment by 13 percent relative to Pennsylvania ones. Pennsylvania did not prove what would have happened in New Jersey. It supplied the best available estimate of it.

A product version, again with invented numbers: a bike-sharing app drops its per-ride start fee, but only in Porto, because pricing is set per city. Rides per active rider per month go from 6.0 to 7.5 in Porto. In Lisbon, which kept the fee, they go from 6.2 to 6.7 over the same weeks, because spring arrived in both cities. The naive pre/post effect in Porto is 1.5 rides. The difference in differences estimate is 1.5 minus 0.5, or 1.0 ride, because Lisbon shows that 0.5 of the rise would have happened anyway.

What is the assumption difference in differences depends on?

Difference in differences depends on parallel trends: without the change, the treated and comparison groups would have moved by the same amount. Columbia University's Mailman School of Public Health states that "The parallel trend assumption is the most critical of the above the four assumptions to ensure internal validity of DID models and is the hardest to fulfill." It also notes there is no statistical test for it, though visual inspection helps when you have many time points.

In practice, plot both groups for as many periods before the change as you can. If Porto and Lisbon moved together for six months and only diverged the week the fee was removed, the estimate is believable. If Porto was already growing faster, the estimate absorbs that head start and calls it an effect.

A bad example: a coworking booking app launches a referral credit in Madrid and compares desk bookings with Barcelona. During the test, Barcelona hosts a large trade fair that fills its desks for a week. Bookings in the comparison city jump for reasons unrelated to the credit, and the difference in differences estimate understates the credit's effect, possibly to zero. An event that hits only one group breaks the method, and no amount of arithmetic fixes it after the fact.

How do you pick a comparison group?

Pick the comparison group before launch, using history, not after launch using the result. Three approaches, from simplest to most formal:

  • Historical match. Choose the single untreated market whose metric tracked the treated market most closely over the previous months.
  • Grouped markets. Combine several similar untreated markets into one comparison group, so a local event in any one of them moves the average less.
  • Synthetic control. Build a weighted blend of untreated markets whose combined history matches the treated market's history as closely as possible, then use the blend's path after launch as the counterfactual.

Whichever you use, write down the choice and the reason in the launch document. A comparison group chosen after seeing which one makes the result look best is not a comparison group.

When should you insist on an A/B test?

Insist on an A/B test when the change can be shown to some users and not others, you have enough traffic to detect the effect you care about, and the decision matters. Random assignment is what makes the Bing result above trustworthy: users in both groups experienced the same week, the same news and the same competitors, so the only systematic difference was the change.

The limits are real. A/B tests need traffic, they capture the effect over the test window rather than over a year, and they break when users in the two groups affect each other, as in a marketplace where buyers in one group compete for the same sellers as buyers in the other. For the step-by-step procedure, see how to run an A/B test, and for the underlying concepts the A/B testing and control group glossary entries.

Which method should a PM choose?

Work down this list and stop at the first yes:

Question If yes
Can you randomise users, and will you reach the needed sample in about four weeks? A/B test
Does the change hit a whole market or group, and do you have an untreated group that moved in step with it beforehand? Difference in differences
Is the expected effect far larger than normal weekly movement, with nothing else changing at the same time? Pre/post, with its limits stated
None of the above Do not claim a measured effect; report the change as directional and plan a cleaner test

The strongest objection to this ladder is that it slows launches down: a team that could ship a city-wide change today now needs a comparison city and a pre-launch plot. The answer is that the plot takes an hour, and the alternative is a quarterly review built on a number that seasonality produced. The decision rule is also where to be honest about statistical versus practical significance: a well-measured effect can still be too small to matter.

Builders Camp runs Data for Product Managers as a 2 week bootcamp with 4 live sessions and 8 self-paced microlessons, directed by Andre Albuquerque. It sits in the Product Management Starter Track, the Discovery Expert Track and the Data & Analytics Specialist Track, and its public topic list covers metrics that matter, funnels and activation, retention and cohorts, segmentation, experimentation and A/B testing basics, and communicating insights.

See the Data for Product Managers bootcamp

Before your next market-level launch, write the name of the comparison group in the launch document next to the metric. If the team cannot agree on one, you have learned before launch that the result will not be measurable, which is the cheapest time to learn it.

Bootcamps referred in this Guide

Frequently asked questions

What is the difference between difference in differences and an A/B test?

An A/B test randomly assigns users to see the change or not, so the two groups differ only by chance and the change. Difference in differences compares the before-and-after change in a group that got the change with the before-and-after change in a group that did not, where you chose the groups rather than a coin flip. The A/B test proves more; difference in differences works where randomising is impossible.

When should I use difference in differences instead of an A/B test?

When the change applies to a whole market, city, store or team and you cannot show it to some users and not others. Pricing that must be the same for everyone in a country, a new warehouse, a regional marketing campaign and a policy change are typical cases. You also need a comparison group that was trending the same way as the treated one before the change.

What is the parallel trends assumption?

The assumption that, without the change, the treated and comparison groups would have kept moving by the same amount. Columbia's Mailman School of Public Health calls it the most critical assumption for difference in differences and the hardest to fulfill. You cannot prove it, but you can check that the two groups moved together for several periods before the change.

Is a pre/post analysis ever good enough?

Yes, when the expected effect is large, the metric is stable week to week, and nothing else changed at the same time. A fix that takes a broken checkout from failing to working needs no control group. A redesign expected to move conversion by two points does, because normal weekly variation can be bigger than that.

How do I choose a control group for difference in differences?

Pick the untreated market or segment whose metric tracked the treated one most closely before the change, over as many periods as you have. Grouping several similar markets into one control smooths out local events, and a synthetic control, a weighted blend of untreated markets built to match the treated one's history, is the more formal version of the same idea.

Can I run an A/B test with low traffic?

You can run one, but it may not be able to detect the effect you care about in a reasonable time. If the sample size calculation says you need months, test a bigger change, pick a metric closer to the change, or fall back to difference in differences with a clear comparison group and state its weaker evidence.

Which method does a Builders Camp bootcamp cover?

Data for Product Managers is a 2 week bootcamp with 4 live sessions and 8 self-paced microlessons, directed by Andre Albuquerque. Its public topic list includes metrics, funnels, cohorts, segmentation, experimentation and A/B testing basics, and communicating insights. A/B Testing for Product Managers goes deeper on experiment design.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque

Last updated 2026-09-27

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Data for Product Managers bootcamp