Builders Camp

Practice challenges

AI agent pipeline failure case study exercise

This exercise gives you a working three agent pipeline that just sent a false churn alert to an account manager, who then made an awkward call to a healthy customer. You diagnose which agent did what wrong, design a human review gate with concrete trigger conditions, and write the postmortem the non-technical customer success team is waiting to read.

The scenario

A customer success team automated its weekly account health report with three agents working in sequence. The first pulls 30 days of CRM activity per account and classifies it green, yellow, or red for churn risk, with a confidence score. The second writes a plain language summary of why an account got its classification. The third formats every summary into a document and tags the account owner on anything flagged red. It runs every Friday morning with no human review, because the entire point was to save four hours of manual work each week.

Last week it flagged the company's fourth largest account, worth over $300,000 in yearly revenue, as high risk with 81 percent confidence, citing three recent billing support tickets and a drop in login activity. The account manager, following the process, called the customer on Monday to check in. The customer was not struggling. They were mid-rollout to 40 new users, and the billing tickets were about upgrading their plan. The login dip was two power users on vacation. The customer's own procurement contact asked directly if the call meant the vendor thought they were about to cancel.

The first agent had learned, from historical data, that billing tickets usually meant trouble, because in the training period they mostly did. Nobody updated it to tell the difference between a billing problem and a billing upgrade. The second agent wrote an accurate summary of a wrong classification. The third agent sent it, because sending is what it does. Nobody reviewed any of it first, because review was the four hours the automation was built to remove.

What you are asked to do

You write a postmortem package covering four connected pieces:

  • Diagnose each agent's specific role in the failure, in terms of the information it had and the decision it made with that information, not just "it got the wrong answer," then name the single change that would have most likely prevented this specific outcome.
  • Design a human review gate as a concrete rule, specific enough that an engineer could implement it as a conditional check, stating exactly which combinations of account value, confidence score, and signal type should hold an output for human review before it sends.
  • Redesign the pipeline architecture, including what changes about what the first agent is trained or prompted on, and define what happens to a flagged account if the assigned human reviewer does not respond by the end of the day.
  • Write the automation policy, in plain internal-handbook language, and the postmortem itself, addressed to the non-technical customer success team, that explains what happened, why, what is changing, and what to do in the meantime, without either blaming the technology as broken or minimizing what happened to a real customer relationship.

What a strong answer covers

The exercise's own objectives map directly onto what a strong submission has to demonstrate:

  • Does the diagnosis separate what each of the three agents actually did, at the level of inputs and decisions, rather than treating the pipeline as one undifferentiated failure?
  • Does the review gate rule name specific, checkable conditions, such as account value thresholds and confidence score bands, instead of a general instruction like "review anything risky"?
  • Does the redesigned architecture change what the first agent is trained or prompted on, not just add a review step on top of the same underlying classifier?
  • Does the fallback behavior for an unreviewed flagged account get a defined answer, rather than being left open?
  • Does the postmortem read as honest to a non-technical, upset audience without either scrapping the case for automation or brushing past what went wrong?

Skills this exercise practises

Diagnosing a multi-agent failure at the level of training data, prompt design, and architecture, not just the visible output. Designing where a human checkpoint belongs in an otherwise automated workflow, and writing that checkpoint as an implementable rule rather than a principle. Writing an incident communication for a non-technical audience that is honest without destroying trust in the system it is defending. These map onto the AI Agents bootcamp's own curriculum on tool use, planning and orchestration, and safety with human-in-the-loop design. If you want the version of this same judgment call applied to a single model instead of an agent pipeline, the AI bias audit exercise from AI Product Management runs a comparable diagnosis-then-decide structure. For the underlying design patterns behind any of the three agents here, see how to design an AI agent and how to write evals for AI products.

Which bootcamp this comes from

This exercise is the practical challenge from AI Agents, a two week, three live session bootcamp on Builders Camp covering agent fundamentals, tool use and integrations, planning and orchestration, memory and context management, safety and human-in-the-loop design, and evaluation. Completing the practical challenge counts toward the bootcamp's completion requirement and its certificate, alongside the certification quiz.

Builders Camp runs AI Agents both live and self-paced, included with the Builders Camp Membership alongside every other bootcamp, track, and masterclass.

Bootcamps referred in this Guide

Frequently asked questions

What actually goes wrong in this AI agents exercise?

A three agent pipeline, one to classify account risk, one to write a summary, one to compile and send the report, flags a large account as high churn risk based on billing tickets. The tickets were about a plan upgrade, not a problem, and the account manager makes an awkward call to a customer who was never at risk.

Do I need to know how to build multi-agent systems to attempt this?

No. The exercise gives you the full architecture and the failure. Your job is diagnosis and redesign, in writing, not implementation. It fits the AI Agents bootcamp's own coverage of tool use, orchestration, and human-in-the-loop design.

Why does the exercise insist the fix can't be 'always review everything'?

Because that defeats the reason the pipeline exists. The automation's whole value was replacing a weekly manual report. A policy that routes every output to a human is not a fix, it is turning the automation off, and the exercise scores you on finding the middle ground.

What is the hardest part of this exercise?

Writing the automation policy in step four. It has to name specific trigger conditions for mandatory human review, not general principles, while still preserving the time savings for most cases. Vague policies are easy to write and easy to see through.

Does this exercise count toward the AI Agents certificate?

Yes. Submitting the practical challenge counts toward completing the AI Agents bootcamp on Builders Camp, alongside the certification quiz, and the bootcamp issues a LinkedIn-integrated certificate on completion.

How is this different from a general prompt engineering exercise?

It is not about writing a better single prompt. The failure spans three agents and a missing human checkpoint, so the fix has to be architectural: where does a human gate sit, what triggers it, and what happens if nobody responds in time.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque

Last updated 2026-09-16

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the AI Agents bootcamp