Practice challenges
AI classifier prompt practice exercise for product managers
This exercise gives you a working AI review classifier that has been quietly misclassifying app store reviews for three weeks. You diagnose the one or two root causes behind six specific wrong answers, rewrite the prompt to fix them, predict how the new prompt reclassifies those same reviews, and define what to monitor so the next drift gets caught early.
The scenario
A field service app used by technicians gets 60 to 80 app store reviews a week across iOS and Android. Three weeks ago, a PM automated the weekly triage process: every review gets sent to GPT-4 with a single instruction to classify it as Bug Report, Feature Request, Praise, or Churn Signal, and the category gets written back into a shared sheet.
The prompt itself reads reasonably. It names the four categories and asks for accuracy and concision. It has been running unreviewed for three weeks. This morning, ahead of sprint planning, someone opens the sheet and notices half the classifications look wrong, and the PM who built it is on holiday.
Several specific errors stand out: a crash report on photo upload classified as praise, a request for parts inventory classified as a churn signal, a review about switching to a competitor classified as a bug report, a cancelled subscription classified as a feature request, and a complaint about missed job alerts classified as praise. These are not random noise. They share one or two underlying causes in how the prompt was written, and finding those causes, not just relabeling the six examples by hand, is the actual job.
What you are asked to do
The exercise has five connected steps:
- For each of the five worst misclassifications, state what the correct category should have been and why the prompt's wording produced the wrong one instead.
- Name the one or two root causes shared across the errors, specifically what was missing or ambiguous in the prompt, not a general statement that it was "unclear," and name the one thing the original author likely assumed the model would figure out on its own.
- Rewrite the system prompt to address both root causes, including context about the product and its users, a precise definition of each category with the ambiguous boundaries resolved, and instructions for reviews that could plausibly fit more than one category.
- Predict how your new prompt would reclassify each of the five reviews, explaining what specifically in your new category definitions makes the new answer correct.
- Define a monitoring plan naming a specific, measurable signal to track, how often to check it, and the threshold that should trigger a manual review of the prompt.
What a strong answer covers
The exercise's own objectives are what a strong submission has to satisfy:
- Does the diagnosis for each misclassified review name the specific gap in the prompt's category definitions, rather than a general note that the model got it wrong?
- Are the named root causes shared across multiple errors, showing that the diagnosis found the pattern rather than treating six mistakes as six unrelated problems?
- Does the rewritten prompt resolve the specific ambiguous boundary that caused the original errors, such as the line between a bug report and a churn signal?
- Do the predicted reclassifications follow logically from the new prompt's own definitions, rather than simply asserting the correct answer?
- Is the monitoring signal specific and measurable, with a stated check frequency and a concrete threshold, rather than a general instruction to "keep an eye on it"?
Skills this exercise practises
Tracing a run of bad AI outputs back to the specific prompt language causing them, instead of relabeling each mistake by hand. Writing category definitions precise enough to resolve a genuinely ambiguous case. Predicting how a prompt change will behave before deploying it, which is what makes a fix testable rather than hopeful. Defining a monitoring threshold specific enough to actually catch the next drift. These map onto the bootcamp's own curriculum on prompting and ideation workflows and scaling LLMs and workflows reliably. If you want to work a comparable diagnosis on a much higher-stakes classifier, the AI bias audit exercise from AI Product Management runs the same kind of root cause analysis in a healthcare setting. For the broader groundwork, see how to build an AI product and how to become an AI product manager.
Which bootcamp this comes from
This exercise is the practical challenge from Foundations of AI, a two week bootcamp on Builders Camp with four live sessions covering generative AI and agents explained for builders, AI across the product lifecycle, and prompting and ideation workflows. Completing the practical challenge counts toward the bootcamp's completion requirement and its certificate, alongside the certification quiz.
Builders Camp runs this bootcamp both live and self-paced, included with the Builders Camp Membership alongside every other bootcamp, track, and masterclass.
Bootcamps referred in this Guide
Frequently asked questions
What is broken in this Foundations of AI exercise?
A weekly automation that classifies app store reviews into Bug Report, Feature Request, Praise, or Churn Signal using a single-instruction GPT-4 prompt has been misclassifying reviews for three weeks, and the person who built it is on holiday when it is discovered.
Do I need machine learning experience to complete this exercise?
No. This is a prompt design exercise, not a model training one. You are given specific misclassified reviews and asked to trace the failure back to what the prompt did not specify, then fix the prompt itself.
Why do a few specific wrong classifications matter more than the overall pattern?
Because the exercise asks you to find the one or two root causes producing most of the errors, not to treat each wrong answer as a separate, unrelated mistake. Naming the actual root cause is what makes the rewritten prompt durable instead of a patch for one example.
What does the final deliverable look like?
A rewritten system prompt addressing the identified root causes, a prediction of how it reclassifies the specific reviews that were wrong before, and a monitoring plan naming a specific metric, a check frequency, and a threshold that would trigger a manual prompt review.
How long does the exercise take?
About 60 minutes, rated advanced difficulty, inside a two week bootcamp with four live sessions.
Does completing it count toward a certificate?
Yes. Finishing the practical challenge counts toward completing the Foundations of AI bootcamp on Builders Camp, alongside the certification quiz, and the bootcamp issues a certificate on completion.
Sources

Andre Albuquerque
CEO of Builders Camp, SuperOperator, and other companies. Building products.
CEO of Builders Camp, SuperOperator, and other companies. Building products.
LinkedInMore guides by Andre AlbuquerqueLast updated 2026-09-16
Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
Related guides
How to Build an AI Product from Scratch
Building an AI product from scratch starts with naming what a wrong answer costs, not what a right one looks like...
Andre AlbuquerqueHow to Become an AI Product Manager
Becoming an AI product manager means adding an evaluation and guardrail layer on top of the core PM skills you already...
Andre AlbuquerqueBest AI Tools for Product Managers in 2026
The best AI tools for product managers in 2026 are not one tool but four categories: a reasoning and writing assistant...
Andre AlbuquerqueAI bias audit practice exercise for product managers
This exercise puts you in charge of an AI triage feature that has just failed a bias audit, deprioritizing lower-income...
Andre Albuquerque
