Builders Camp

Practice challenges

Claude Code operating model practice exercise

This exercise starts with a broken deploy that shipped because nobody defined what 'done' meant for an AI coding agent. You diagnose the missing pieces, write a 30 line CLAUDE.md as a project constitution, design exactly three agents with explicit boundaries, define objective acceptance criteria, and build a governance layer of hooks and policies that enforce good behavior automatically.

The scenario

A senior engineer at a 40 person fintech company watched her team adopt an AI coding agent and get faster, until something else started happening alongside the speed. The agent repeated mistakes nobody had told it to stop making. Context degraded mid-session until it was acting on decisions that were no longer current. One Friday afternoon, a junior engineer accepted a diff that passed every automated test but broke the staging environment, because the agent sounded confident and there was no agreed definition of what "done" actually meant beyond that.

The CTO's response was direct: he wanted a proper operating model before the next sprint, not another one-off prompt fix. The team already has the pieces of a framework available to it, six layers, each with a distinct job, but nobody has actually built them out. The task is to diagnose exactly what failed in each of the three visible symptoms, repeated mistakes, degraded context, and no shared definition of done, and then hand the team something that enforces good behavior automatically rather than relying on people remembering to check.

What you are asked to do

The exercise has five connected parts:

  • Diagnose the root cause and the missing framework layer behind each of the three symptoms, then name the single highest-priority fix and explain why it comes before the others.
  • Write a CLAUDE.md file under 30 lines containing only rules that are always true project-wide, including the project's purpose, its source of truth, at least one explicit stop line, the git workflow rule, and the testing rule.
  • Design an agent architecture using exactly three agents, each with a defined role, its responsibilities, and what it is explicitly not allowed to do.
  • Define what "done" means for a real task, as specific acceptance criteria, an automated quality gate that runs before a human sees the diff, and a stop condition that tells the agent when to pause for approval.
  • Design a governance layer with one automated hook, one policy governing what external tools or data the agents may access, and one improvement loop that captures what went wrong in a session and feeds it back into the system.

What a strong answer covers

The exercise's own objectives are what a strong submission has to satisfy:

  • Does the diagnosis name a specific missing layer for each symptom, rather than one general complaint that "the agent made mistakes"?
  • Does the CLAUDE.md actually fit under 30 lines and contain only rules that hold true in every session, with anything only sometimes true correctly cut?
  • Do the three agents have genuinely distinct, non-overlapping roles, with an explicit statement of what each one may not do, rather than three descriptions of the same general-purpose assistant?
  • Are the acceptance criteria specific enough to verify, going beyond "tests pass" to name the exact conditions that must hold, with a quality gate that runs as a command or hook rather than a reminder?
  • Does the governance layer name a real trigger for its hook, a real access boundary for its tool policy, and a concrete mechanism for turning one session's lesson into a permanent system change?

Skills this exercise practises

Diagnosing why an AI-assisted workflow breaks down at the level of missing structure, not just the visible symptom. Writing a persistent project constitution instead of a one-off prompt. Designing agent roles with real boundaries instead of one agent doing everything. Turning "looks good" into an actual, checkable standard before a human ever reviews the output. These map onto the bootcamp's own curriculum on context architecture, skills and agents, and hooks and quality gates. If you want to practice the reviewing side of this same problem, the AI code diff review exercise from Claude Code for Product Managers has you catch a regression an agent introduced while making an otherwise clean change. For the broader integration layer this governance design touches, see how to build an AI assistant with MCP.

Which bootcamp this comes from

This exercise is the practical challenge from Building with Claude Code, a one week bootcamp on Builders Camp with two live sessions covering CLAUDE.md and context architecture, skills, agents, and memory, MCP and tool integrations, and hooks, CI/CD, and quality gates, taught by Guilherme Salgueiro. Completing the practical challenge counts toward the bootcamp's completion requirement and its certificate, alongside the certification quiz.

Builders Camp runs this bootcamp both live and self-paced, included with the Builders Camp Membership alongside every other bootcamp, track, and masterclass. If you are newer to agentic coding tools generally, how to build a prototype with Cursor covers a comparable workflow in a different editor.

Bootcamps referred in this Guide

Frequently asked questions

What triggers this Building with Claude Code exercise?

A junior engineer at a fintech accepts an AI-generated diff that passes every test but breaks staging, because nobody had defined what 'done' actually meant. The CTO asks for a proper operating model before next sprint, not another prompt.

What does a CLAUDE.md file need to contain to pass this exercise?

Only rules that are always true and persistent across sessions, such as the project's source of truth, an explicit stop line the agent must never cross without approval, the git workflow, and the testing rule, all inside a 30 line limit that forces you to cut anything only sometimes true.

Why does the exercise cap the agent architecture at exactly three agents?

Because an unlimited number of agents lets you avoid the hard part: deciding what each one is explicitly not allowed to do. Forcing exactly three means real boundaries, not infinite decomposition.

Is this exercise only relevant if I already use Claude Code?

The specific tool is Claude Code, but the skill, defining acceptance criteria, quality gates, and stop conditions before letting an AI agent touch production code, applies to any AI-assisted development workflow a product manager oversees.

How long does the exercise take?

About 60 minutes, rated intermediate difficulty, inside a one week bootcamp with two live sessions.

Does completing it count toward a certificate?

Yes. Finishing the practical challenge counts toward completing the Building with Claude Code bootcamp on Builders Camp, alongside the certification quiz, and the bootcamp issues a certificate on completion.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Guilherme Salgueiro

Guilherme Salgueiro

Builder and AI systems practitioner. Guilherme helps developers, PMs, and founders move beyond prompting into structured AI system design -- building with Claude Code, agents, and automation pipelines to ship products faster and more reliably.

Builder and AI systems practitioner. Guilherme helps developers, PMs, and founders move beyond prompting into structured AI system design — building with Claude Code, agents, and automation pipelines to ship products faster and more reliably.

LinkedInMore guides by Guilherme Salgueiro

Last updated 2026-09-16

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Building with Claude Code bootcamp