Builders Camp

Tools

How to use AI for estimation when the model has never met your team

Ask a model to decompose a feature, not to size it: most estimation error comes from forgotten work rather than mis-sized work, and decomposition is the part that automates well. No model knows your team's velocity, because velocity is a property of your codebase and this quarter, not of the feature description.

What should you ask a model to estimate?

Nothing. Ask it to decompose instead.

The instinct is to paste a feature description and ask how long it will take, and the answer that comes back will be fluent, plausible and detached from your codebase. A better use of the same two minutes is to ask what work this feature implies, exhaustively, given your stack and your Definition of Done. That output is a checklist, and the checklist is where the money is, because most estimation error is not mis-sizing. It is omission.

Teams do not usually estimate the coding badly by a factor of three. They forget the data migration, the feature flag they will have to clean up, the second review round because the first one surfaced an architectural objection, the analytics events product asked for in a Slack thread, and the fact that the screen being changed has no test coverage so every change requires manual verification. Each omission is small. Together they are the gap between a two week estimate and a five week delivery. For a whole launch rather than one feature, a product breakdown structure catches the omissions outside engineering before any of this sizing starts.

What inputs turn a generic decomposition into a useful one?

Four, and the fourth is the one nobody supplies.

Input Effect on the output
Feature description Baseline. Alone, produces a happy path list
Acceptance criteria Adds the edge cases somebody already thought about
Stack and architecture, in a paragraph Adds the migration, the cache, the queue, the client update
Your Definition of Done Adds the tests, the accessibility check, the analytics, the docs

The Definition of Done is what converts a list of coding tasks into a list of work. The Scrum Guide treats it as one of the three things that make a team confident in a forecast, alongside past performance and upcoming capacity, and it is the only one of the three you can hand a model directly. Write it out, paste it in every time, and the decomposition stops stopping at "build the endpoint".

A worked example: "let users export their data"

Take a one line request and run the decomposition with all four inputs supplied. A reasonable output lists work in four groups rather than one.

Build: the export endpoint, the job queue entry because a synchronous export will time out on large accounts, the file generation, the storage location and its retention rule, the download link and its expiry, the UI entry point, and the empty and error states.

Constraint work: permissions, since not every role should export everything; rate limiting, because this is an obvious way to hammer the database; and a decision about whether exports include deleted records, which is a product question disguised as a technical one.

Definition of Done work: tests at whatever level you require, the analytics event, the accessibility pass on the new UI, the support documentation, and the internal runbook for when an export fails.

Unknowns to resolve before sizing: what the largest account's export looks like, whether the current schema can produce the format customers actually want, and whether legal has a view on what a data export contains.

That fourth group is the output most worth having. Three unknowns identified before the estimate is given are three sources of a three week overrun removed, and identifying them is exactly the kind of reading work that scales badly for humans and well for a model. It is the same evidence-first move that makes acceptance criteria useful rather than decorative.

Why can no model know your team's velocity?

Because velocity is not a property of the work. It is a property of this team, this codebase and this quarter, and all three change.

The same export feature costs one team four days and another team five weeks, and the difference is not skill. It is whether the codebase already has a job queue, whether the test suite runs in four minutes or forty, whether code review takes an hour or three days, how many people are on call, and whether the one person who understands the storage layer is available. The Scrum Guide's framing is exact here: confidence in a forecast comes from what the Developers know about their past performance, their upcoming capacity and their Definition of Done. A model has no access to the first two and only sees the third if you paste it.

There is a sharper version of this. Even your own historical velocity is a weak predictor across different kinds of work, which is why Atlassian's estimation guidance treats relative sizing as a team-local practice rather than a transferable unit. Story points from your team do not mean anything on another team, and they certainly do not mean anything to a model trained on other people's repositories.

The anchoring problem, and why you should withhold the numbers

If the model does attach numbers to its decomposition, do not show them to the team.

Anchoring is well documented in forecasting generally, and Lovallo and Kahneman's 2003 Harvard Business Review analysis of executive optimism describes the related pattern where teams reason from inside a project rather than from the record of similar projects. A number on a shared screen moves everyone toward it, and the effect does not go away because people know the number was machine generated. If anything it gets worse: a number with no visible author is harder to argue with than a number Sofia said out loud.

Share the checklist. Let the team size it with whatever method they already use. Then compare the total against what similar work actually took, which is the outside view that corrects for optimism, and re-forecast rather than defending the first figure.

Is a drafted decomposition just more process?

It can be. A team that already ships predictably does not need a machine to tell it that an export needs rate limiting, and adding a pre-estimation step to a working process is pure overhead. The honest read is that this helps most where estimates are currently wrong in the same direction every time, which usually means the team is repeatedly forgetting a category of work rather than misjudging effort.

Beyond Agile makes the wider argument: keep the practices that produce a feedback loop, kill the ones that only produce artefacts. A decomposition pass earns its place if it changes what the team commits to. If the checklist gets generated, nodded at, and ignored, it is a document, and the right move is to stop producing it.

What to do with a better estimate

Estimates do not get good; they get honest. The change worth aiming for is a team that states a range with its assumptions attached, names the three unknowns that would move the range, and re-forecasts when those unknowns resolve rather than at the end. That is a communication skill as much as an analytical one.

Project Management for Product runs one week on defining milestones, slicing work and setting expectations without false certainty, plus the reporting side: writing updates that name blockers and decisions rather than activity. It sits in the Product Delivery Specialist Track with the rest of the delivery curriculum.

See the Project Management for Product bootcamp

For the steps either side, AI for sprint planning covers what happens once the sizes exist, AI for capacity planning covers the quarter above that, and how to write acceptance criteria with AI covers the input that makes a decomposition sharp.

Bootcamps referred in this Guide

Frequently asked questions

Can AI estimate how long a feature will take?

Not credibly, and that is the wrong thing to ask it for. It can decompose a feature into the pieces of work it implies, which is where most estimation error actually comes from. The sizing of those pieces belongs to the people who will build them.

Why can a model not know our velocity?

Velocity is a property of a specific team, a specific codebase and a specific quarter. The Scrum Guide ties forecast confidence to what the Developers know about their past performance, their upcoming capacity and their Definition of Done. None of those three exists in a model's training data, and two of them change month to month.

Where does estimation error actually come from?

Forgotten work rather than mis-sized work. Teams rarely estimate the coding badly by a factor of three. They forget the migration, the feature flag cleanup, the second review cycle and the fact that the change touches a screen that has no tests. Decomposition is the fix, and decomposition is mechanical enough to automate.

What should I paste in to get a useful decomposition?

The feature description, the acceptance criteria, a short description of your stack and architecture, and your Definition of Done. That last one is what makes it list the accessibility check and the analytics event rather than stopping at the happy path.

How do I use the output without anchoring the team?

Share the checklist of work items, not any numbers the model attached to them. Anchoring is the documented risk with estimation, and a number on a page moves people toward it even when everyone knows where it came from.

Does this work for a team with no historical data?

It helps most there, because the decomposition gives a new team a shared list to argue about instead of a shared silence. It still cannot supply the missing history, so expect wide ranges and re-forecast often rather than pretending to precision.

Does Builders Camp teach a specific estimation tool?

No. Project Management for Product covers milestone definition, slicing work and setting expectations without false certainty, taught as a method rather than through a named product.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Tiago Pedro da Costa

Tiago Pedro da Costa

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

LinkedInMore guides by Tiago Pedro da Costa

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Project Management for Product bootcamp