Tools
When not to use an AI coding agent
An AI coding agent earns its cost on work that is large, mechanical, well specified and reversible, which is a narrower band than the tooling suggests. The seven cases below are where the review cost exceeds the writing cost, or where being wrong is not recoverable. The deciding question is not difficulty, it is who gets paged when the output is wrong.
The deciding question is blast radius, not difficulty
Most advice about when to reach for an AI coding agent sorts tasks by how hard they are. That is the wrong axis. Agents are frequently excellent at hard things and dangerous at easy ones, because the risk in agent-produced code is not that it fails to work, it is that it works and is wrong in a way nobody checks. The sort that actually predicts trouble is consequence: what happens if this is subtly incorrect and ships.
Builders Camp's AI Product Management curriculum states the same idea from the design side. How much rope to give a model is treated as a product choice: what it can see, what it may use, when a human has to approve, and what happens when it fails. The cases below are the ones where the honest answer to that last question is bad enough to change the decision.
The seven cases where an agent costs you more than it saves
You cannot state the acceptance test. If you cannot write down what correct output looks like before you start, the agent will produce something plausible and you will spend the afternoon negotiating with it instead of thinking. The fix is not a better prompt. It is spending twenty minutes deciding what you actually want, which was the real work all along.
The change is one line in code you do not understand. The economics invert on small changes. Writing the line takes three minutes. Reading the agent's forty-line diff to confirm it only did the thing you asked takes longer, and skipping that review is how an unrelated regression arrives with your name on the commit.
The binding constraint lives outside the repository. A rate limit in a vendor contract, a retention rule your legal team applied, an undocumented quirk of a partner API that someone learned the hard way in March. The agent reads code, not institutional memory, and it will generate something that compiles and breaks the constraint with total confidence.
The failure is unbounded. Payments, authentication, permissions, data deletion, and anything that fans out to customers, such as an email send or a webhook. These are not hard problems. They are problems where being wrong once is not recoverable by editing the file again.
You are learning the system. Delegating exploration removes the only reason the task was worth your time. Asking an agent to explain a codebase is a genuinely good use of it; asking it to change the codebase while you are still forming a mental model leaves you owning something you never read.
The real disagreement is a decision, not code. When two teams want different scopes, an agent that builds both versions has not resolved anything, it has doubled the thing you now have to throw away. Wikipedia's own account of the practice notes that AI-assisted building is generally considered risky when output is accepted without review, and a decision nobody has made is the review step you are skipping.
The task is genuinely mechanical and small. Renaming one variable, bumping one dependency, fixing one typo in a string. A free Cursor Hobby tier or a paid seat at 20 dollars a month is not the constraint here; your attention is. Routing trivial work through a review loop costs more attention than doing it.
What the strongest objection to this page is
That most of these are solved by setup rather than avoidance, and the objection is largely correct. Claude Code's documentation describes hooks that run shell commands before or after an agent action, plan review inside the editor, and explicit control over tool access and permissions, which is the machinery for turning "do not do this" into "do this, but the linter blocks you and a human approves the merge". Builders Camp's Building with Claude Code teaches that machinery directly, covering context files, skills, agents, and hooks with CI/CD-style quality gates, across 1 week, 6 hours total, 4 of them taught in 2 live sessions of 120 minutes.
So the narrower and more defensible claim is this: none of the seven cases is permanently off limits, but each one requires a control you have to build before you delegate. The mistake is not using an agent in a high-consequence area. The mistake is using one there with the same setup you use for a throwaway script.
What to do instead in each case
For the unstated acceptance test, write the test first and let the agent implement against it. For the small change in unfamiliar code, ask the person who owns the file, which costs one message and buys you the constraint you did not know about. For the constraint outside the repository, write it into the project's context file so the agent can see it at all, which is precisely what a CLAUDE.md-style instruction file exists to do.
For unbounded failure, add the human approval step before the action rather than after it. Builders Camp's AI Agents bootcamp teaches this as its safety module: approvals, constraints and monitoring on high-stakes actions, with autonomy treated as a dial rather than a switch. It runs 2 weeks, 6 hours total, 4 of them taught across 3 live sessions of 90 minutes. The design question it asks is the useful one to borrow, which is not whether the agent can do the task but what it is allowed to do without asking.
How do you spot blast radius before you delegate?
Ask who gets paged. Not who wrote the code, not how complicated it is: who is woken up, who answers the support ticket, whose quarter it ruins. That question resolves almost every ambiguous case in about five seconds, and it produces the right answer even for changes that look trivial, which are exactly the ones that slip through.
The second question is whether the failure is visible. A broken build announces itself. A permissions check that now returns true for the wrong user does not, and will not, until someone reports it from outside. Anything in the second category deserves a human pass regardless of how confident the agent sounded, and confidence is not a signal here: agent output reads the same whether it is right or wrong.
Who this is for, and who it is not for
This fits a product manager, tech lead or founder who has already adopted an agent and is now deciding where to stop, rather than someone still deciding whether to start. It assumes you have seen at least one agent-produced change that was wrong in a way nobody caught, because that experience is what makes the rest of this useful rather than theoretical.
It is a weaker fit if your current problem is the opposite one: a team that will not use these tools at all. Blanket refusal costs more than careless adoption in most product organisations, and the argument for that is a different piece than this one.
Learn to design the environment, not just the prompt
The pattern under all seven cases is the same, and it is a product design problem rather than a tooling problem: deciding what the model sees, what it may touch, when a person has to approve, and how you find out when it was wrong. AI Product Management covers exactly that ground, including evaluations treated as the launch gate rather than as a testing afterthought, across 2 weeks, 12 hours total, 8 of them taught in 4 live sessions of 120 minutes. Its practical challenge is rated advanced and opens on an unresolved bias incident in a live triage feature, which is the realistic version of this problem rather than the tidy one.
See the AI Product Management bootcamp
For the design side of the same question, see how to design an AI agent and how to write evals for AI products. For the two controls named most often above, see AI guardrails and human in the loop. For the automated version of the approval step, see quality gates in an AI pipeline.
Bootcamps referred in this Guide
Frequently asked questions
Is there a simple test for whether to use an agent?
Two questions, in order. Can you state, before you start, what correct output looks like? And if the agent is wrong in a way you miss, what is the worst thing that happens? A clear answer to the first and a boring answer to the second means go ahead. Anything else means slow down.
Are agents bad at small changes?
They are fine at them and often uneconomic. A one-line change in a file you already understand takes three minutes to write and one to check. The same change routed through an agent produces a diff you now have to read in full, because agents routinely touch adjacent code while implementing the thing you asked for.
Does this mean AI coding agents are overrated?
No. It means their value is concentrated in a narrower band than the marketing suggests: work that is large, mechanical, well specified and reversible. The cases below are the edges of that band, not an argument against the middle of it.
What about compliance or regulated code?
The problem there is rarely the code and usually the constraint that lives outside it. An agent reads the repository, not your data processing agreement or the rule your legal team applied last quarter. It will generate something that compiles and violates the constraint, confidently, with no signal that it did.
How does a product manager know the blast radius of a change?
Ask who gets paged if it is wrong. If the answer is nobody, the change is safe to delegate. If the answer is an on-call engineer, a support queue or a customer, the change belongs to whoever owns that consequence, regardless of how small the diff is.
Do guardrails fix most of these cases?
Several of them, and that is the honest counterargument to this page. Builders Camp's AI Product Management curriculum puts it directly: most failures trace to missing guardrails rather than to weak models. Permissions, approval steps and automated checks move several cases from do not to proceed carefully.
Is it wrong to use an agent to learn an unfamiliar codebase?
Using it to explain the codebase is one of the better uses available. Using it to make the change for you while you are still learning removes the reason you were doing the task, and you end up owning a system you never actually read.
Sources

Andre Albuquerque
CEO of Builders Camp, SuperOperator, and other companies. Building products.
CEO of Builders Camp, SuperOperator, and other companies. Building products.
LinkedInMore guides by Andre Albuquerque
Inês Lourenço
CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.
CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.
LinkedInMore guides by Inês LourençoLast updated 2026-09-18
Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
Related guides
How to Design an AI Agent
Designing an AI agent means deciding what it can see, what it may use, when a human has to approve its action, and what...
Andre AlbuquerqueHow to Write Evals for AI Products
An eval set is a fixed list of real inputs and the pass or fail rubric you score them against before any prompt or...
Andre AlbuquerqueHow to Build an AI Assistant with MCP
Building an AI assistant with MCP means adding one or more MCP servers to an AI tool like Claude Code, scoping each...

Andre Albuquerque & Guilherme SalgueiroBest AI Tools for Product Managers in 2026
The best AI tools for product managers in 2026 are not one tool but four categories: a reasoning and writing assistant...
Andre Albuquerque

