---
title: "How to Run an A/B Test: A PM's 4 Steps"
description: "How to run an A/B test as a product manager: a testable hypothesis, one primary KPI plus guardrails, a sample size fixed in advance, and a decision rule."
canonical_url: "https://builderscamp.com/guides/other/how-to-run-an-ab-test"
date_published: "2026-09-27"
date_modified: "2026-09-27"
author: "Andre Albuquerque"
publisher: "Builders Camp"
guide_class: "other"
---

# How to run an A/B test as a product manager

**TL;DR:** Run an A/B test in four steps: write a hypothesis about one change, choose one primary KPI plus guardrail metrics, fix the sample size and duration before launch, then analyse against a decision rule you wrote in advance. The two habits that protect every result are refusing to stop early and refusing to ship on the primary KPI alone.

A cheap test can beat an expensive roadmap bet. Ron Kohavi and Stefan Thomke describe in [Harvard Business Review](https://hbr.org/2017/09/the-surprising-power-of-online-experiments) how a 2012 Bing experiment on ad headline display, an idea that had sat in the backlog for more than six months because it looked low priority, increased revenue by 12 percent, worth more than $100 million a year in the United States alone. That is the upside. The downside is that a badly run test produces a confident wrong answer, and Evan Miller's advice in [How Not To Run an A/B Test](https://www.evanmiller.org/how-not-to-run-an-ab-test.html) names the most common way it happens:

"If you run experiments: the best way to avoid repeated significance testing errors is to not test significance repeatedly," Evan Miller writes.

The process below keeps both in view. It uses one original example throughout: a meal kit subscription where 20 percent of visitors to the plan selection page continue to checkout, and the team wants to know whether showing price per serving instead of the weekly total moves that number.

## What does the PM own in an A/B test, and what does the analyst own?

The PM owns the question and the decision: which change is worth testing, which number decides it, which numbers must not break, and what happens for each possible result. The analyst or data scientist owns the machinery: whether users are split correctly, whether the sample size maths is right, and whether the analysis matches the design. A test fails most often at the handoff, when the PM writes "improve checkout" and the analyst has to guess which metric that means.

If you have already read the [definition of A/B testing](https://builderscamp.com/guides/glossary/ab-testing), the rest of this page is the working procedure on top of it.

## Step 1: How do you write a hypothesis you can actually test?

A testable hypothesis names one change, one audience, one measurable outcome and the reason you expect it. Use this shape: "Showing price per serving instead of weekly total on the plan page will raise the share of plan page visitors who start checkout, because first-time visitors compare us with a supermarket shop and a weekly total looks expensive next to a single meal."

Checklist before you move on:

- One change only. If the variant also moves the button and rewrites the headline, a win tells you nothing about which change caused it.
- A number, not a feeling. "Users will like it more" cannot be tested; "checkout start rate rises" can.
- A reason drawn from evidence: interview notes, support tickets, a funnel drop you can point to in your [product funnel](https://builderscamp.com/guides/glossary/product-funnel). The reason is what you learn from when the test loses.

## Step 2: Which KPI should decide the test?

Pick one primary KPI, the closest measurable outcome to the change, and write it down before launch. For the meal kit test that is checkout start rate from the plan page, not monthly revenue, because revenue sits several steps away and moves for reasons your change does not touch. A KPI that is too far downstream needs an enormous sample to register anything; one that is too close (clicks on the price text) can go up without anyone buying more.

Three questions test a candidate KPI. Is it aligned with a goal the business already tracks, ideally traceable up your [metrics tree](https://builderscamp.com/guides/glossary/metrics-tree)? Will you act differently depending on its value? Is it measured reliably, the same way, in both groups?

## What are guardrail metrics, and which should you set?

Guardrail metrics are the numbers the change is not meant to improve but must not damage. They exist because a variant can win the primary KPI by pushing cost somewhere else, and a primary-only readout will call that a win.

For the price per serving test, the realistic ways to win badly are specific:

| Guardrail | Why it could break | Threshold to write down before launch |
|---|---|---|
| First box cancellation rate | Per serving pricing attracts people who then feel misled by the weekly charge | No significant increase |
| Average servings per order | Visitors pick the smallest plan because it looks cheapest per serving | No drop larger than the revenue gain from extra checkouts |
| Pricing related support contacts | Confusion about what is actually charged | No significant increase |
| Plan page load time | New price component slows the page | Within current performance budget |

Keep the list short, three to five. Every extra guardrail is another chance for noise to look like harm, and a list of twenty guarantees something will turn red by chance.

## Step 3: How do you design the experiment?

Design comes down to three decisions: the randomisation unit, the sample size and the duration.

Randomise by user, not by page view, so the same person does not see both prices across two visits. Confirm the split works before trusting any result: if you planned 50/50 and one group ends up clearly larger, something upstream (a redirect, a caching layer, a bot filter) is sorting users, and the comparison is broken. A clean [control group](https://builderscamp.com/guides/glossary/control-group-experimentation) is the only reason the difference between groups means anything.

Size the test from the smallest lift worth shipping, the minimum detectable effect. Using the standard formula behind calculators such as Evan Miller's [sample size calculator](https://www.evanmiller.org/ab-testing/sample-size.html), at 80 percent power and 5 percent two-sided significance:

| Input | Meal kit example |
|---|---|
| Baseline checkout start rate | 20 percent |
| Smallest lift worth shipping | 1 point absolute (20 to 21 percent) |
| Users needed per variant | about 25,600 |
| Plan page visitors per day | 3,000, split 1,500 per variant |
| Days of traffic needed | about 17 |
| Planned run | 21 days (three full weeks) |

The run is rounded up to whole weeks because weekday and weekend visitors behave differently, and a test that covers two Mondays and one Saturday is quietly weighted. If the numbers had come out at ten weeks, that would be the moment to test a bolder change, not to shrink the minimum detectable effect until the calendar looks acceptable.

## Why can't you stop the test when it looks significant?

A significance threshold of 5 percent assumes you look once, at the planned end. Check a running test every morning and stop the first time it crosses the line, and you give chance a new opportunity every day. Evan Miller's worked example makes the size of the error concrete: for a change that does nothing, stopping at the first 5 percent significance reading (or after 150 observations) finds a false winner 26.1 percent of the time, more than five times the error rate the dashboard implies.

Watching is fine, for broken tracking or a guardrail collapsing. Deciding is not. If your team needs the freedom to stop early, use a method designed for it, such as a sequential test or an adaptive approach; the trade-offs are in [multi-armed bandit vs A/B testing](https://builderscamp.com/guides/comparison/multi-armed-bandit-vs-ab-testing).

## Step 4: How do you analyse the result and make the call?

Write the decision rule before launch, then apply it. The rule turns five possible outcomes into five actions, so nobody argues the definition of a win after seeing the numbers:

| Primary KPI | Guardrails | Call |
|---|---|---|
| Significant lift at or above the minimum effect | All hold | Ship |
| Significant lift | One breaks | Do not ship as built; price the damage against the gain, then rework |
| Flat, confidence interval narrow around zero | Hold | Stop; record that this lever does not move this metric |
| Flat, confidence interval wide | Any | Inconclusive; the test was underpowered for this effect, not proof of no effect |
| Significant drop | Any | Stop, and keep the losing variant in the learning log |

Two checks come after the verdict. First, read the size of the effect as well as its [statistical significance](https://builderscamp.com/guides/glossary/statistical-significance): a lift can be real and still too small to pay for the change, which is the subject of [statistical vs practical significance](https://builderscamp.com/guides/other/statistical-vs-practical-significance). Second, compare pre-registered segments only, such as new against returning visitors or mobile against desktop. Slicing twenty ways after the fact will always find one slice that "won", and that slice is almost always noise.

## When should a PM not run an A/B test?

The strongest objection to testing everything is cost. A test needs traffic, engineering time for the variant and the tracking, and weeks of calendar during which the losing half of users gets the worse experience. If your page cannot reach the sample in about four weeks, if the change is a legal or accessibility fix you would ship regardless, or if the decision is irreversible and strategic (a new pricing model, a market entry), a test either cannot answer the question or should not be the thing answering it. When the change has to hit a whole market at once, [pre/post vs difference in differences vs A/B test](https://builderscamp.com/guides/other/pre-post-vs-difference-in-differences-vs-ab-test) covers the fallback methods and what each can prove.

The Bing result cuts the other way, and it is the reason to keep testing cheap ideas. The headline change sat unprioritised for six months because judgment said it was small. Tests are how you find out that judgment was wrong about a specific idea, which is only possible if the ideas you are unsure about get tested rather than argued.

## Where can you practise the full loop?

Reading a result under pressure is the part that takes practice. The [A/B test results challenge](https://builderscamp.com/guides/challenges/ab-testing-for-product-managers-ab-test-results) gives you a mixed readout to call, and the [A/B testing resources roundup](https://builderscamp.com/guides/resources/best-ab-testing-resources-for-product-managers) collects further reading.

Builders Camp runs A/B Testing for Product Managers as a 1 week bootcamp with 2 live sessions and 8 self-paced microlessons, directed by Andre Albuquerque, on testable hypotheses, experiment design with guardrails and duration, metric selection, statistical intuition for PMs, and when to ship, iterate or kill. It sits in the Growth Specialist Track and the Data & Analytics Specialist Track.

[See the A/B Testing for Product Managers bootcamp](https://builderscamp.com/bootcamps/ab-testing-for-product-managers?utm_source=guide&utm_medium=organic&utm_campaign=how-to-run-an-ab-test)

One habit to start with your next test: paste the decision rule table into the ticket before the variant is built. If the team cannot agree on the five rows before launch, the test is not ready, and finding that out costs an hour instead of three weeks.

## Frequently asked questions

### What are the four steps of an A/B test?

Write a testable hypothesis about one change, pick one primary KPI plus a short list of guardrail metrics, design the experiment (randomisation unit, sample size, fixed duration), then analyse against a decision rule you wrote before launch. The PM owns all four; an analyst checks the maths in steps three and four.

### How long should an A/B test run?

As long as it takes to reach the sample size you calculated before launch, rounded up to whole weeks so every weekday is represented equally. In the meal kit example on this page that is 17 days of traffic, run as 21 days.

### What is a guardrail metric?

A metric you do not expect the change to improve but that must not get worse, such as cancellations, refunds, revenue per order, support contacts or page load time. Guardrails catch a variant that wins the primary KPI by pushing cost somewhere else.

### Can I stop an A/B test early if the result is already significant?

Not with a standard fixed sample test. Evan Miller's worked example shows that stopping a test of a change that does nothing the moment it first looks significant produces a false winner 26.1 percent of the time, against the 5 percent you think you are accepting.

### How much traffic do I need to run an A/B test?

Enough to detect the smallest lift worth shipping. At a 20 percent baseline, detecting a 1 point absolute lift with 80 percent power at 5 percent significance needs roughly 25,600 users per variant. If you cannot reach that in about four weeks, test a bigger change or use a different method.

### What should the decision rule say?

What you will do for each outcome: ship if the primary KPI wins and guardrails hold, rework if the primary wins but a guardrail breaks, stop if the result is flat with a narrow interval, and treat a flat result with a wide interval as inconclusive rather than as proof of no effect.

### Do I need a data scientist to run an A/B test?

Not to design one. The hypothesis, the KPI choice, the guardrails and the decision rule are product decisions. You do want someone who can check the randomisation and the analysis, because a broken split quietly invalidates every number that follows.

## Sources

- [Harvard Business Review: The Surprising Power of Online Experiments (Kohavi and Thomke, 2017)](https://hbr.org/2017/09/the-surprising-power-of-online-experiments)
- [Evan Miller: How Not To Run an A/B Test](https://www.evanmiller.org/how-not-to-run-an-ab-test.html)
- [Evan Miller: Sample Size Calculator](https://www.evanmiller.org/ab-testing/sample-size.html)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
