---
title: "A/B Testing Results Practice Exercise"
description: "Practice reading a real A/B test with mixed results, calculating revenue impact, and deciding to ship, kill, or extend. A 60 minute exercise from Builders Camp."
canonical_url: "https://builderscamp.com/guides/challenges/ab-testing-for-product-managers-ab-test-results"
date_published: "2026-09-16"
date_modified: "2026-09-16"
author: "Andre Albuquerque"
publisher: "Builders Camp"
guide_class: "challenges"
---

# A/B test results practice exercise for product managers

**TL;DR:** This exercise hands you a real-looking A/B test at the midpoint of its run: one metric significant and positive, two others moving the wrong way. You reconstruct the hypothesis, classify each metric, calculate the actual revenue impact, and decide whether to ship, kill, or extend, in writing, to a stakeholder who wants to ship now.

## The scenario

A booking platform is running a test on its hotel listing page. The change adds a "Price Match Guarantee" badge next to the price, on the theory that removing pricing anxiety increases the chance someone completes a booking. The test is 14 days into a planned 21 day run, split 50/50, with roughly 40,000 users landing in each arm every day.

This morning's read shows booking conversion up, with a p-value under 0.05, the usual bar for statistical significance. That is the number the engineering lead has already seen. Their message: "We're at 95 percent confidence on the primary metric. Can we just ship it?"

Underneath that headline number are two others moving the wrong way. Average booking value is down, though not yet at standard significance. Add-on purchases, the extras customers add at checkout, are down sharply and are significant. A 30-day repeat booking metric is inconclusive either way.

Read as a single number, the test looks like a clean win. Read in full, it looks like a trade: more people book, but each one is worth less, and a chunk of ancillary revenue disappears. Whether that trade is worth taking is not a statistics question. It is a business decision that needs the statistics to be read correctly first.

## What you are asked to do

You work through five pieces of the same decision, each building on the last:

- Reconstruct the original hypothesis in a standard "we believe that X will cause Y, resulting in Z" format, and name the assumption inside it that the current data is already calling into question.
- Classify every metric in the results table as green, yellow, red, or grey, based on direction and significance, not on which number is largest.
- Calculate the net revenue impact per 10,000 visitors, combining the conversion gain, the booking value loss, and the add-on revenue loss, showing your working rather than a single final figure.
- Make a ship, kill, or extend call, and defend it using the specific numbers from your own calculation.
- Write a direct, non-defensive reply to the engineering lead's Slack message, explaining in four to six sentences why a single significant p-value is not, by itself, a decision.

The deliverable is a short written document: the hypothesis, the classification table, the revenue math, the decision, and the reply. There is no code and no spreadsheet macro required, just arithmetic you can show.

## What a strong answer covers

The bootcamp's own difficulty rating on this exercise is advanced, and the objectives it is built around double as the questions a strong answer has to hold up under:

- Does your hypothesis reconstruction name the specific behavioral assumption the badge was supposed to change, not just restate the feature that shipped?
- Does your metric classification treat "significant and moving the wrong way" as seriously as "significant and moving the right way," instead of only checking the primary metric?
- Does your revenue calculation account for both the gain from higher conversion and the loss from lower average value and lower add-on attach rate, netted against each other, not reported as three separate positive-sounding numbers?
- Does your decision name the specific number that would have to change for you to switch it, rather than reading as a fixed conclusion?
- Does your reply to the engineering lead correct the record without talking down to a peer who is not wrong to ask, just working from one number instead of four?

## Skills this exercise practises

Reading a partial experiment result under pressure to declare a winner early. Separating a metric that is merely moving from one that is moving and statistically significant. Turning a percentage change into an actual revenue number instead of stopping at the percentage. Writing a decision memo that survives someone else checking your math. These are the same skills covered across the bootcamp's own curriculum: hypothesis design, experiment guardrails, statistical intuition for decision makers, and making the ship, iterate, or kill call on a mixed result, which is exactly what most A/B testing content skips in favor of the clean-win case. For a broader look at prioritizing what to build once your experiments are in, see the [RICE prioritization framework](https://builderscamp.com/guides/other/rice-prioritization-framework) and how to pick the [north star metric](https://builderscamp.com/guides/other/north-star-metric-examples) that experiments like this one should ultimately move. If this is the kind of judgment call you want more practice with before your next promotion cycle, [how to get promoted to senior product manager](https://builderscamp.com/guides/path/how-to-get-promoted-to-senior-pm) covers the broader case for it.

## Which bootcamp this comes from

This exercise is the practical challenge from [A/B Testing for Product Managers](https://builderscamp.com/bootcamps/ab-testing-for-product-managers), a one week, two live session bootcamp on Builders Camp covering hypothesis design, experiment design, metrics selection, statistical intuition, and interpreting results under real constraints. Completing the practical challenge counts toward the bootcamp's completion requirement and its certificate, alongside the certification quiz. If you want the retention-side version of reading ambiguous data under a deadline, the [retention cohort diagnosis exercise](https://builderscamp.com/guides/challenges/growth-for-product-managers-retention-cohort-diagnosis) from Growth for Product Managers runs the same muscle on a different data set.

Builders Camp runs this bootcamp both live and self-paced, and it is included with the Builders Camp Membership alongside every other bootcamp, track, and masterclass.

## Frequently asked questions

### How long does this A/B testing exercise take?

About 60 minutes. It is rated intermediate difficulty, and most of the time goes into the revenue calculation and the written reply to the engineering lead, not the reading.

### Do I need a statistics background to attempt it?

No. You need to know what a p-value and a confidence level are at a basic level. The exercise is built around a decision, not a formula derivation, and it tells you the significance threshold you need.

### What makes this test result 'ambiguous' instead of just a clear win?

The primary metric, conversion, moved in the right direction and hit significance. Two other metrics, average booking value and add-on revenue, moved the wrong way. A single green number never tells the whole story, and this exercise is built around that gap.

### Is there a single correct answer to ship, kill, or extend?

No. The exercise is scored on whether your decision follows from your own revenue math and results reading, not on which of the three options you pick. A defensible extend and a defensible kill both pass; an undefended ship does not.

### Does completing this challenge count toward a certificate?

Yes, if you complete it as part of the A/B Testing for Product Managers bootcamp on Builders Camp. Submission counts toward the bootcamp's completion requirement and its certificate.

### What is the difference between this and the bootcamp's certification quiz?

The quiz is 10 multiple-choice questions with a 70 percent pass mark, testing recall of A/B testing concepts. This exercise is a single applied scenario with a written submission, testing whether you can apply those concepts under a specific, messy set of numbers.

## Sources

- [Builders Camp: A/B Testing for Product Managers bootcamp page](https://builderscamp.com/bootcamps/ab-testing-for-product-managers)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
