---
title: "AI Triage Bias Audit Practice Exercise"
description: "Work through a live AI bias incident: a triage feature deprioritizing lower-income patients, an angry nurse, a board meeting in 10 days. A 90 minute exercise."
canonical_url: "https://builderscamp.com/guides/challenges/ai-product-management-ai-bias-audit"
date_published: "2026-09-16"
date_modified: "2026-09-16"
author: "Andre Albuquerque"
publisher: "Builders Camp"
guide_class: "challenges"
---

# AI bias audit practice exercise for product managers

**TL;DR:** This exercise puts you in charge of an AI triage feature that has just failed a bias audit, deprioritizing lower-income patients at more than double the rate of others. You diagnose where the bias entered the model, decide whether to ship, pause, or mitigate, write the board briefing, and answer the nurse who flagged it and wants a straight answer.

## The scenario

A mid-size healthcare platform shipped an AI feature six months ago that reads patient intake forms and flags who should get a same-day appointment, meant to catch urgent cases faster and reduce nurse workload. A nurse noticed something: patients who described symptoms in plain language, like "my chest feels funny," scored lower urgency than patients using clinical vocabulary for the same complaint.

An audit confirms it. Patients from the lowest income zip codes are deprioritized at more than twice the rate of patients from the highest, even after controlling for actual symptom severity. The model is 81 percent accurate overall, above the threshold set at launch. The likely cause is not the algorithm itself but the four years of historical human triage decisions it was trained on, which appear to carry the same bias, plus the fact that patients with less exposure to medical terminology tend to describe symptoms differently, not less seriously.

Nothing about this is simple to act on. The feature is live, processing about 1,200 intakes a week. Nurse workload has dropped 30 percent since launch, and the clinical team does not want it rolled back. No confirmed harm has occurred, though the audit only covers patients who returned for a follow-up visit, meaning it cannot see what happened to anyone who did not. Legal has said nothing publicly can go out without their approval. A board meeting is in 10 days. And the nurse who flagged the problem has asked directly whether this is going to be fixed or quietly patched.

## What you are asked to do

You produce a package covering five connected decisions:

- Diagnose where bias most plausibly entered the system, across training data, labeling, feature selection, and pre-launch evaluation, and name the single point where catching it would have been cheapest.
- Choose and defend one of three options: pause the feature immediately, keep it running with a mandatory human review layer, or keep it running unchanged while a fix is prepared, anticipating and answering the strongest objection to your choice.
- Write the board briefing: what happened, how it was found, what is and is not yet known, and what you need the board to decide, without speculating beyond what is confirmed or writing anything that reads as an admission of fault.
- Define the retraining requirements as specific, testable acceptance criteria, not a general instruction to "be less biased."
- Write what you would actually say to the nurse who flagged this, in a real conversation, and draft the patient communication you would want to send, with a stated reason if you believe a more limited version is genuinely justified rather than just safer.

## What a strong answer covers

The exercise's own objectives are the questions a strong submission has to answer directly:

- Does the diagnosis go past "the training data was biased" to name where in the pipeline, training data, labeling, feature choice, or evaluation, catching this would have been cheapest?
- Does the ship, pause, or mitigate decision engage honestly with what it gives up, not just what it protects, and name the specific objection a reasonable person would raise against it?
- Does the board briefing separate what is confirmed from what is not yet known, instead of either downplaying the finding or speculating about harm the data cannot show?
- Are the retraining acceptance criteria specific and measurable, the kind a data scientist could act on directly, rather than a restated good intention?
- Does the answer to the nurse read as an honest, human response rather than a corporate one, and does the patient communication actually tell patients something, or explain plainly why it does not?

## Skills this exercise practises

Tracing a model failure to its actual point of origin instead of stopping at "the AI got it wrong." Making a defensible call when the evidence is incomplete and every option has a real cost. Writing for two audiences at once, a board that needs facts and a person who flagged a real problem and deserves honesty. These map onto the AI Product Management bootcamp's own curriculum on evals, agent and model design, and capturing product value responsibly. If you want to work the same kind of diagnosis on a much smaller, single-prompt classifier, [fixing an AI classifier](https://builderscamp.com/guides/challenges/foundations-of-ai-fixing-an-ai-classifier) from Foundations of AI is a good next exercise, and the [AI agent pipeline postmortem](https://builderscamp.com/guides/challenges/ai-agents-agent-pipeline-postmortem) covers the same kind of incident writing for a multi-agent failure. For the broader groundwork, see [how to build an AI product](https://builderscamp.com/guides/tools/how-to-build-an-ai-product) and [how to write evals for AI products](https://builderscamp.com/guides/tools/how-to-write-evals-for-ai-products).

## Which bootcamp this comes from

This exercise is the practical challenge from [AI Product Management](https://builderscamp.com/bootcamps/ai-product-management), a two week, four live session bootcamp on Builders Camp covering AI-native experience design, evals, agent and environment design, and capturing value responsibly. Completing the practical challenge counts toward the bootcamp's completion requirement and its certificate, alongside the certification quiz.

Builders Camp runs AI Product Management both live and self-paced, included with the Builders Camp Membership alongside every other bootcamp, track, and masterclass.

## Frequently asked questions

### What is the AI bias audit exercise actually about?

A healthcare AI feature that prioritizes patients for same-day appointments is found, six months after launch, to deprioritize patients from lower-income zip codes at more than twice the rate of others, after controlling for symptom severity. You diagnose why, and decide what to do about a feature that is live and in daily use.

### Do I need a data science background to complete it?

No. The audit findings are given to you in plain language. The exercise tests product judgment: where bias could have entered the system, whether to ship, pause, or mitigate, and how to write about it honestly under legal constraints, not statistical modeling.

### Why is there no clearly right answer between pausing and keeping the feature running?

Because there isn't one in the real situation this is modeled on. Pausing costs the clinical team a real workload reduction and risks missed cases during the transition. Continuing means knowingly running a biased feature while a fix is built. The exercise scores the reasoning and the acknowledgment of the strongest objection, not which option you pick.

### What makes the board briefing and the patient communication hard to write?

You have to be honest about what happened without writing anything that could read as an admission of liability before legal reviews it. That tension, between transparency and legal exposure, is a real constraint any product leader shipping AI into a regulated space has to work inside, not a hypothetical one.

### Does this exercise count toward a certificate?

Yes, when completed as part of the AI Product Management bootcamp on Builders Camp. It counts toward the bootcamp's completion requirement alongside the certification quiz, and the bootcamp issues a LinkedIn-integrated certificate.

### How long does the exercise take?

About 90 minutes. It is rated advanced difficulty and has five parts: root cause diagnosis, the ship or pause decision, the board briefing, the retraining requirements, and the personal response to the nurse who flagged the problem.

## Sources

- [Builders Camp: AI Product Management bootcamp page](https://builderscamp.com/bootcamps/ai-product-management)

## How this guide was made

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
