Builders Camp

Tools

How to use AI for bug reports without inventing the reproduction steps

A bug report becomes actionable at six fields: reproduction steps, expected, actual, environment, frequency and first seen. AI is good at reformatting a support thread into that shape and at flagging which fields the thread does not actually support, and it is dangerous at filling the gaps, because an invented reproduction step costs an engineer a day. Severity stays a human call scored on blast radius, separate from priority.

What separates a bug report from a complaint?

Six fields. Steps to reproduce, expected result, actual result, environment and build, frequency, and when it was first seen. Chromium's public bug reporting guidelines ask for close to that set, and Bugzilla's bug writing guidelines add the two rules that get skipped most: be precise about what you saw, and file one bug per report. Everything else in a ticket is context. These six are the difference between work starting today and a thread that runs for a week.

AI helps with the shape of that report and almost nothing else. Give a model a support conversation and it will produce a clean, well formatted ticket in seconds. The question worth asking before you paste it into the tracker is which of those six fields came from the conversation and which came from the model.

Where does AI genuinely help, and where does it fabricate?

It helps at four jobs, all of them retrieval or reformatting. Turning a rambling support thread into the six field structure. Telling you which fields the thread does not support. Rewriting a title like "checkout broken" into "checkout returns to cart after applying a promo code on mobile Safari". Searching the existing backlog for tickets with similar symptoms so you can check for a duplicate before filing.

It fabricates at exactly one job, and that job is the one people most want handed over: producing the steps a user actually took. A language model completes patterns, so given a partial account of a failure it will supply the missing steps in the most plausible order. Plausible is not observed. An engineer who spends a morning following a reproduction path the user never walked has lost a morning, and the ticket comes back marked as not reproducible while the real bug stays open.

The fix is a prompt instruction, not a better model. Ask for the six fields, and ask it to mark every element it inferred rather than read directly from the input. Then go and confirm the marked ones with the person who reported it. That round trip takes ten minutes.

How do you get a minimal reproduction instead of a long one?

A minimal reproduction is the shortest sequence that still produces the failure. Every step you remove is a hypothesis eliminated for free, which is why a four step reproduction is worth far more than the twelve step version it came from, even though both technically work.

Here the model is a useful sparring partner rather than an author. Give it the long version and ask which steps are plausibly incidental and why, then test its suggestions yourself by actually removing them. You will often find that the promo code matters and the browser does not, or the reverse, and either answer narrows the search before an engineer opens the file. This is the one part of bug reporting where a wrong AI suggestion is cheap: you disprove it in thirty seconds by trying it.

Two conditions turn out to matter more often than people expect, so check them explicitly rather than waiting for a model to raise them. The first is state: a fresh account behaves differently from one with three years of data. The second is timing, which is where a race condition hides and where "sometimes" in a report usually comes from.

Severity is not priority, and a model will merge the two

Ask any general purpose model to assign severity to a bug and it will hand back a number that quietly encodes how annoying the bug sounds. That is not severity. Atlassian's severity guidance frames severity around the impact on customers and the business, while priority is a separate scheduling decision that also weighs the release calendar, the cost of the fix and who is available to do it. A cosmetic defect on the signup page the week before a launch can be low severity and high priority at the same time, and a ticket that fuses them into one number loses that.

So put them in two fields and fill them separately. Severity is answered by four blast radius questions:

  • How many users hit it, measured from logs or support volume rather than from how loud the report was.
  • Is data being lost or corrupted, which raises severity regardless of how few people are affected.
  • Does the failure look like success to the user, which is the question most reports never ask.

That third one carries more weight than its position suggests. A loud error gets reported within minutes by the first person who sees it. A silent failure that shows a success message produces no reports at all, and the damage runs until somebody checks the data. Default silent failures above loud ones and argue the exception.

Priority you set in the room, with the release plan in front of you. No model has the inputs for that call, and handing it over produces a queue nobody defends.

What does a workable prompt look like?

Give the model the raw material, name the six fields, forbid invention, and ask for the uncertainty back. Something close to this: here is a support thread. Produce a bug report with steps to reproduce, expected result, actual result, environment and build, frequency and first seen. Use only what is in the thread. Mark any element you inferred rather than read. List separately every field the thread does not support, and suggest the exact question I should ask the reporter to get it.

The last clause is the one that changes the output. Without it you get a confident, complete looking ticket. With it you get a shorter ticket and a list of three questions, which is a more honest description of what you actually know. Writing acceptance criteria with AI runs on the same discipline from the other end of the cycle: state the observable condition, never the intention behind it.

What about logs, screenshots and session recordings?

Check them before they leave your machine. Production logs carry email addresses, session tokens, internal hostnames and sometimes whole customer payloads, and a screenshot of an admin panel carries more than the bug. Strip identifiers, or use a tool your company has already cleared for that data class. Nothing in the analysis depends on knowing which customer it was.

Once stripped, logs are where a model earns its place. Pattern matching across a few hundred lines to find the first anomalous entry, or to tell you whether this error appears elsewhere in the same window, is fast and checkable. Treat file paths and error classes as reliable, and treat any causal story the model tells about them as a hypothesis for the engineer who owns the code. If you want the agentic version of that workflow, with a repository open alongside the trace, triaging bugs with Claude Code as a PM covers the routing decision in detail, and this page stays on the report itself. The failure mode both pages are guarding against has a name: AI hallucination.

Where this goes wrong in practice

The first failure is volume. A team that can generate a tidy ticket in twenty seconds files more tickets, and a backlog of 400 well formatted reports is harder to work than a backlog of 90 rough ones. Formatting was never the constraint. Decide what you are not going to file before you make filing cheap.

The second is the confident duplicate. Similarity search across an issue tracker is genuinely useful as a shortlist and unreliable as a verdict, because two bugs with near identical symptoms often have unrelated causes. Read the three candidates it surfaces; do not close on its say so.

The third is tone. A generated report reads smoothly, and smooth prose implies a level of certainty the evidence rarely supports. "The promo code validation fails on mobile Safari" is a claim. "The user reported that the promo code did not apply on an iPhone, browser not confirmed" is what you actually have. Keep the second one until somebody has reproduced it.

Where the delivery system around bug reports lives

Bug reports are one artefact inside a delivery cadence, and the cadence is what decides whether they get fixed or accumulate. Project Management for Product covers that layer in 1 week, with 1 live session and 8 self-paced microlessons on planning and scoping, dependency management, risk and the kind of status communication that names blockers and decisions rather than listing activity.

If the deeper problem is that your team files careful reports into a process that never gets to them, Beyond Agile is the closer match: 1 week, 1 live session and 11 microlessons on outcome driven planning, ownership, roadmap governance and which rituals to keep or kill. The Product Delivery Specialist Track bundles both alongside the rest of the delivery path, and Builders Camp tracks are curated sequences of 6 to 10 bootcamps.

Try this on the three oldest tickets in your backlog

Take the three that have sat untouched longest and run only the field check on each: which of the six fields does the ticket actually support, and which were assumed by whoever wrote it. If two of the three turn out to be missing reproduction steps entirely, the reason they never got picked up was never priority. It was that nobody could start.

See the Project Management for Product bootcamp

For the wider delivery path, see the Product Delivery Specialist Track, and for the boundary question that decides how much of this to hand over at all, when not to use an AI coding agent.

Bootcamps referred in this Guide

Frequently asked questions

What does a bug report need before anyone can work on it?

Steps to reproduce, expected result, actual result, environment and build, frequency, and when it was first seen. Chromium's public bug reporting guidelines ask for essentially that set, and Bugzilla's writing guidelines add two rules people skip: be precise, and file one bug per report. A ticket missing reproduction steps is a conversation, not a report.

Can AI write the reproduction steps for me?

It can draft them from something you give it, such as a support transcript, a session recording summary or your own rough notes. It cannot observe. Anything it produces that was not in your input is a guess, and a guessed step sends an engineer down a path the user never took. Ask it to mark every step it inferred, then go and confirm those with the reporter.

What is a minimal reproduction, and why does it matter more than a long one?

The shortest sequence that still produces the failure, with everything incidental removed. It matters because each removed step is a hypothesis eliminated for free. A twelve step reproduction that only needs four is eight chances for an engineer to wonder whether a detail is load bearing.

How is severity different from priority?

Severity describes the damage the defect does. Priority describes when the team will act. Atlassian's severity level guidance frames severity around impact on customers and the business, and priority is a separate scheduling decision that also weighs release timing, cost and who is free. A model asked for one will usually return a blend of both, so ask for them in two separate fields.

How do I set severity without guessing?

Score blast radius, not effort. How many users hit it, is data being lost or corrupted, is there a workaround support can give out, and does the failure look like success to the user. That last question usually outranks the other three, because a silent failure generates no reports at all while the damage accumulates.

Should I paste production logs into a general AI tool?

Only after checking what is in them. Logs routinely contain email addresses, tokens, internal hostnames and customer payloads. Strip identifiers first, or use a tool your company has already approved for that data class. The analysis does not get worse for losing a customer name.

What does AI genuinely speed up in a bug report?

Turning a messy support thread into the fixed field set, spotting which of those fields the thread does not support, rewriting a vague title into a specific one, and checking a new report against existing tickets for possible duplicates. All four are formatting and retrieval jobs, which is where the technology is strongest.

Which Builders Camp bootcamp covers the delivery side of this?

Project Management for Product runs 1 week with 1 live session and 8 self-paced microlessons on planning, dependency management, risk and status communication. Beyond Agile runs 1 week with 1 live session and 11 microlessons on the operating model around those rituals, including which ceremonies to keep and which to kill.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Tiago Pedro da Costa

Tiago Pedro da Costa

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

LinkedInMore guides by Tiago Pedro da Costa

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Project Management for Product bootcamp