---
title: "Statistical vs Practical Significance"
description: "Statistical significance vs practical significance for PMs: why a real lift can be too small to ship, when a p of 0.06 can still ship, and when that is spin."
canonical_url: "https://builderscamp.com/guides/other/statistical-vs-practical-significance"
date_published: "2026-09-27"
date_modified: "2026-09-27"
author: "Andre Albuquerque"
publisher: "Builders Camp"
guide_class: "other"
---

# Statistical significance vs practical significance: why a winning test is not a ship decision

**TL;DR:** Statistical significance tells you a difference is probably real; practical significance tells you whether it is big enough to be worth shipping. Decide on the confidence interval compared with a minimum worthwhile effect you set before launch: ship when the whole interval clears it, stop when the whole interval sits below it, and treat everything in between as a judgment you write down.

The American Statistical Association settled the core point in 2016, when it published six principles on the use of p-values. Principle 5 is the one product teams forget most often:

"A p-value, or statistical significance, does not measure the size of an effect or the importance of a result," the [American Statistical Association's statement](https://www.amstat.org/asa/files/pdfs/p-valuestatement.pdf) says.

Principle 3 goes further and names business directly: "Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold." For a PM, that turns every test readout into two separate questions. Is the difference real? And if it is, is it worth what it costs to ship and keep?

## What is the difference between statistical and practical significance?

Statistical significance is a claim about noise: how surprising the observed difference would be if the change did nothing. The [statistical significance glossary entry](https://builderscamp.com/guides/glossary/statistical-significance) covers the p-value and the 0.05 convention. Practical significance is a claim about value: whether the size of the effect, at the low end of what the data allows, beats the cost of building, running and maintaining the change.

The two questions come apart because sample size moves one and not the other. More users shrink the noise, so a smaller and smaller effect becomes statistically significant. Nothing about more users makes a small effect more valuable.

## Why can a tiny effect be statistically significant?

Take an original example. A signup page converts 8.0 percent of visitors. The team tests an embedded product tour, a third-party component with a monthly licence. After 400,000 visitors per variant, the variant converts 8.2 percent.

| Readout | Value |
|---|---|
| Control | 32,000 signups from 400,000 visitors (8.0 percent) |
| Variant | 32,800 signups from 400,000 visitors (8.2 percent) |
| Difference | +0.2 points (2.5 percent relative) |
| p-value | about 0.001 |
| 95 percent confidence interval | +0.08 to +0.32 points |

Statistically, that is a clean win. Now price it. Suppose the page gets 800,000 visitors a month and a signup is worth €3 in expected revenue. The point estimate is 1,600 extra signups, about €4,800 a month. The low end of the interval is 640 extra signups, about €1,920 a month. If the tour licence and its upkeep cost €2,500 a month, the honest reading is that the change probably pays, but the data cannot rule out that it loses money. That is a different conversation from "p = 0.001, ship it".

## How do you decide whether an effect is big enough to ship?

Set a minimum worthwhile effect before the test starts, the smallest lift that pays for the change. Add up what shipping costs: build time, licences, the performance budget it spends, the extra code path every future change has to respect, and the support load if it confuses anyone. Divide by the value of one extra conversion. The answer is a lift in points, and it does double duty: it sizes the sample (see [how to run an A/B test](https://builderscamp.com/guides/other/how-to-run-an-ab-test)) and it becomes the bar the result has to clear.

Then read the result against that bar using the confidence interval, not the p-value. The interval is the range of effects the data is compatible with, which is the thing you actually compare with cost.

## What if a large effect misses p < 0.05?

The opposite case appears on low-traffic pages. A B2B pricing page sees 1,200 visitors per variant over four weeks. Control converts 10.0 percent to a demo request, the variant 12.5 percent. That is a 25 percent relative lift, and the p-value is about 0.053, just over the line. The 95 percent confidence interval runs from roughly minus 0.03 points to plus 5.0 points.

Reading that as "no effect" throws away real information. The interval says the change is somewhere between harmless and very good, with almost none of it below zero. If the change is a copy edit that costs nothing to keep and one click to revert, and no guardrail moved, shipping it as a monitored bet is defensible. The HBR account of Bing's ad headline test is a reminder of what small, cheap changes can carry: Ron Kohavi and Stefan Thomke report in [Harvard Business Review](https://hbr.org/2017/09/the-surprising-power-of-online-experiments) that one such change raised revenue by 12 percent after sitting unprioritised for more than six months.

## When is shipping a non-significant result self-deception?

Shipping on a near miss turns into spin as soon as any of these is true:

- The analysis changed after the numbers came in: a new metric, a new time window, a new outlier rule, anything that turned 0.08 into 0.049.
- The "winning" segment was found after launch by slicing until something looked good.
- The low end of the interval is well below zero, so the data is as compatible with harm as with gain.
- The change is expensive or hard to reverse, so being wrong costs real money or trust.
- You would have shipped it whatever the result said, which means the test was theatre and the decision was made elsewhere.

When one of those applies, the right move is to extend the test if the plan allowed for it, rerun with a larger sample, or stop, and to write down which.

## How do segment checks change the call?

Pre-registered segments are a legitimate second lens. If you said before launch that you would compare new and returning visitors, and the lift comes entirely from new visitors while returning visitors are flat, that changes what you ship (the variant for first visits only) and what you learn (the change fixes a first impression, not a habit).

Unplanned segments are a trap. With twenty independent slices and no real effect, the chance that at least one clears 5 percent significance is about 64 percent (1 minus 0.95 to the power of 20), and that slice will come with a plausible story attached. Treat a post-hoc segment as a hypothesis for the next test, never as the verdict of this one. [Confirmation bias](https://builderscamp.com/guides/glossary/confirmation-bias-data) does most of its damage at exactly this step.

## Ship, iterate or kill: what does the decision rule look like?

| Confidence interval compared with your minimum worthwhile effect | Guardrails | Call |
|---|---|---|
| Entirely above it | Hold | Ship |
| Above zero, straddles it | Hold | Judgment: ship if cheap and reversible, else extend or iterate for a bigger effect |
| Straddles zero, lower end close to zero | Hold | Monitored bet only if cheap and reversible; otherwise inconclusive |
| Entirely below it, including zero | Hold | Kill; the lever is too weak here |
| Any | A guardrail breaks | Do not ship as built; iterate |

Write the chosen row into the decision log with the numbers. A team that records its judgment calls can check, a quarter later, whether its monitored bets actually held up, which is the only way to learn whether its judgment deserves the latitude.

## Doesn't relaxing the 0.05 line invite p-hacking?

The p-hacking objection is the strongest one against this page, and it is right about the risk. A team that ships anything "close enough" will drift toward shipping whatever it wanted. The answer is not a looser threshold; it is moving the judgment to before the test, where the minimum worthwhile effect, the guardrails and the decision rule are set while nobody knows the result yet.

The ASA statement puts the principle plainly: "The p-value was never intended to be a substitute for scientific reasoning," said Ron Wasserstein, the ASA's executive director, in the same release. The p-value is one input. Cost, reversibility and the full interval are the others, and the decision is yours.

## Where can you practise making these calls?

Builders Camp runs A/B Testing for Product Managers as a 1 week bootcamp with 2 live sessions and 8 self-paced microlessons, directed by Andre Albuquerque. Its public syllabus includes statistical intuition for PMs (p-values, confidence, power and sample size) and interpreting results and making calls: when to ship, iterate or kill, and how to handle mixed outcomes.

[See the A/B Testing for Product Managers bootcamp](https://builderscamp.com/bootcamps/ab-testing-for-product-managers?utm_source=guide&utm_medium=organic&utm_campaign=statistical-vs-practical-significance)

For the wider habit of writing decisions down before the evidence arrives, the [product decision-making framework](https://builderscamp.com/guides/other/product-decision-making-framework) applies the same discipline to calls that no test can settle. And if your traffic is too thin for either kind of significance, [multi-armed bandit vs A/B testing](https://builderscamp.com/guides/comparison/multi-armed-bandit-vs-ab-testing) covers when a different method fits better.

## Frequently asked questions

### What is the difference between statistical and practical significance?

Statistical significance asks whether a difference is likely to be real rather than noise. Practical significance asks whether a real difference is large enough to be worth the cost, risk and complexity of shipping it. A test result needs both answers before it becomes a decision.

### Can a result be statistically significant but not practically significant?

Yes, and large samples make it common. In the example on this page, 400,000 users per variant turn a 0.2 point lift on an 8 percent signup rate into a p-value near 0.001, while the low end of its confidence interval is worth less than the component costs to run.

### Can a result that is not statistically significant still be worth shipping?

Sometimes: when the change is cheap, easy to reverse, harms no guardrail, and the confidence interval's lower end sits close to zero rather than well below it. Ship it as a monitored bet, say so in writing, and keep watching the metric after launch.

### What is a minimum worthwhile effect?

The smallest lift that pays for building, running and maintaining the change. Set it before the test, from cost and value, and use it both to size the sample and to judge the result. A lift below it can be real and still not worth shipping.

### Should I look at the p-value or the confidence interval?

Both, but decide on the interval. The p-value tells you whether zero is a plausible effect; the interval tells you the range of plausible effects, which is what you compare with cost. The ASA's 2016 statement says a p-value does not measure the size of an effect.

### Is checking segments after a test a good idea?

Only for segments you named before launch, such as new against returning users or mobile against desktop. Searching every slice after the fact will surface one that looks like a winner by chance, and shipping to that slice is how teams convince themselves of effects that are not there.

## Sources

- [American Statistical Association: Statement on Statistical Significance and P-Values (press release, March 7, 2016)](https://www.amstat.org/asa/files/pdfs/p-valuestatement.pdf)
- [Harvard Business Review: The Surprising Power of Online Experiments (Kohavi and Thomke, 2017)](https://hbr.org/2017/09/the-surprising-power-of-online-experiments)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
