Builders Camp

Practice challenges

AI workflow automation case study exercise

This exercise hands you a support ticket automation that has misclassified real tickets for two weeks, including a double-charge complaint buried under a 47 word summary. You diagnose why each output failed, rewrite the prompt in 200 tokens or fewer, predict its new outputs, and design edge case tests to catch the next failure before it ships.

The scenario

A product manager at a SaaS company built an automation that reads every new support ticket and asks an AI model to output four things: a category, the product area, a priority score from one to three, and a one sentence internal summary. The prompt reads reasonably: it names the four categories, asks for a priority score, and asks for conciseness. It ran for two weeks before the support team complained.

Four real outputs show what went wrong. A customer asking to download invoices from their account page, with an accountant waiting on them by end of week, got classified as a low priority feature request. A customer with a broken export function and a board presentation the next morning also got marked low priority, filed as a bug rather than treated as urgent. A question about webhook retry documentation got labeled a feature request instead of a docs question. And a customer who was charged twice, already disputing it with their bank and threatening to cancel, got correctly flagged urgent, but with a 47 word summary where one sentence was asked for.

Each failure has a different cause. Some come from the prompt's category definitions being too loose to draw a clean line between a bug report and a churn signal. Others come from the model reading the tone of a message instead of its actual business impact. The team has already spent 40 percent over its monthly AI budget, so the fix has a hard constraint: the rewritten prompt has to fit in 200 tokens, which forces a real choice about what to prioritize and what to let go.

What you are asked to do

The exercise moves through four connected steps:

  • Diagnose each of the four failed outputs individually, naming whether the root cause sits in the prompt's instructions, its category definitions, its priority criteria, its output format, or a limitation of the model itself.
  • Identify which single failure is hardest to fix with better prompt wording alone, and explain what would actually be needed to solve it properly.
  • Rewrite the prompt in 200 tokens or fewer, naming explicitly which failures you are prioritizing and which you are knowingly leaving unresolved.
  • Predict, without running it, how your rewritten prompt would classify all four original tickets, and design three new edge case tickets meant to break your prompt or expose its remaining limits.

What a strong answer covers

The exercise's own objectives are what a strong submission has to hold up against:

  • Does the diagnosis treat the four failures as distinct problems with distinct causes, rather than one generic complaint that "the prompt is unclear"?
  • Is the hardest-to-fix failure correctly identified as the one that needs more than better wording, such as a category definition that genuinely overlaps two real cases?
  • Does the rewritten prompt name its trade-off explicitly, stating which failure it is not fully solving, rather than claiming to fix everything in 200 tokens?
  • Are the predicted outputs specific enough to actually be checked against the real prompt, not just a restatement of what you hope will happen?
  • Do the three edge case tickets target genuinely different failure modes, such as a ticket spanning two categories or one where tone contradicts urgency, rather than three small variations on the same idea?

Skills this exercise practises

Diagnosing an AI output failure down to its specific mechanism instead of stopping at "it got it wrong." Making an explicit trade-off under a hard constraint rather than pretending you can fix everything at once. Designing adversarial test cases that would have caught a production failure before it shipped. These map onto the bootcamp's own curriculum on workflow mapping, reliable AI steps, and monitoring for degradation over time. For the tools this kind of automation typically runs on, see Make.com for product managers and n8n for product managers. If you want to build a reusable, repeatable version of a workflow like this instead of a one-off fix, the reusable AI skill design exercise from Building your AI Operating System is the natural next step.

Which bootcamp this comes from

This exercise is the practical challenge from Automate Workflows with AI, a one week bootcamp on Builders Camp with two live sessions covering workflow mapping, automation building blocks, reliable AI steps, human-in-the-loop approvals, and monitoring. Completing the practical challenge counts toward the bootcamp's completion requirement and its certificate, alongside the certification quiz.

Builders Camp runs this bootcamp both live and self-paced, included with the Builders Camp Membership alongside every other bootcamp, track, and masterclass. For a broader look at no-code automation options, see no-code tools for product managers.

Bootcamps referred in this Guide

Frequently asked questions

What is broken in this automation exercise?

A Make.com automation that classifies inbound support tickets with GPT-4 has been running for two weeks and produced four bad outputs: a billing complaint tagged as a feature request, a critical bug with a board deadline marked low priority, a docs question mislabeled, and a 47 word summary that was supposed to be one sentence.

Why does the fix have to fit in 200 tokens?

Because the exercise mirrors a real constraint: the team building this automation is already 40 percent over its monthly AI budget. Every extra token is a real cost, so you have to prioritize which of the four failures to fix and say plainly which one you are deprioritizing.

Do I need to know Make.com or a specific automation tool to attempt this?

No. The exercise is entirely about prompt diagnosis and rewriting under a token budget. No automation platform account or setup is required.

What is the hardest part of this exercise?

Step two, naming which of the four failures cannot be fixed by a better prompt alone. Three are prompt problems. One needs something more than clearer instructions, and identifying which one, and why, is where most people get it wrong.

How long does the exercise take?

About 90 minutes, rated intermediate difficulty, as part of a one week bootcamp with two live sessions.

Does finishing it count toward a certificate?

Yes. Completing the practical challenge counts toward finishing the Automate Workflows with AI bootcamp on Builders Camp, alongside the certification quiz, and the bootcamp issues a certificate on completion.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque

Last updated 2026-09-16

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Automate Workflows with AI bootcamp