Builders Camp

Tools

How to use AI for incident reviews without losing the blameless part

Use a model for the one part of an incident review that is mechanical: merging deploy logs, alerts and chat into a single ordered timeline with every entry labelled by source. Do not let it write the analysis unsupervised, because a model handed a timeline will assign blame to named people without being asked, and blameless review exists precisely because punished teams stop reporting. Ban names in the prompt, require each action to be paired with what the person knew at the time, and put an owner on every follow up.

What is an incident review supposed to produce?

Not a document. A set of changes to the system, each with an owner and a date, plus a shared understanding of why a reasonable person made the decision that contributed to the outage. Google's SRE guidance frames the write up as a learning artefact and is direct about the condition that makes it work: the review stays blameless, focused on contributing causes rather than indicting an individual or a team, because engineers who expect to be punished stop reporting the near misses that would have prevented the next one.

That condition is fragile, and it is exactly what an AI drafted review puts at risk. The model is useful, but the specific thing it does badly is the specific thing the whole practice depends on.

What is a model actually good at during a review?

Reconstructing the timeline. An incident leaves evidence in five or six places at once: deploy history, alert firing times, a status page, an incident channel, a support queue, sometimes a customer email. They are in different formats, often in different time zones, and putting them into a single ordered sequence takes an hour of dull work that nobody wants at the end of a bad week.

Hand a model the raw exports and ask for one chronological list where every entry carries its timestamp, its source and a quoted fragment. Ask it explicitly to mark gaps, meaning stretches where nothing was recorded, rather than smoothing over them. Those gaps are usually the most interesting part of the timeline, because the window between the failure starting and the first alert is where your detection problem lives, and a narrative that flows nicely from event to event hides it.

The output is checkable line by line, which is what separates this from the rest of the process. Every entry points at a source you can open. A wrong entry costs ten seconds to catch.

The model will assign blame, and it will sound like analysis

Give a model a clean timeline and ask what happened, and you will get back sentences shaped like this: the on-call engineer deployed the configuration change without running the full test suite, which caused the outage. That is a grammatical sentence, a fluent one, and an analytically empty one. It stops at the person.

This is not a prompt failure you can fix by being polite. Narrative prose about events defaults to naming an agent for every action, because that is how almost all writing about events works, and a model completing that pattern will produce blame whether or not anyone asked for it. It will also produce it with more confidence than a human writer would, because it has no stake in the room and no memory of what the deploy pipeline looked like at 2am.

The blameless version of the same fact is longer and more useful: a configuration change reached production without passing the full test suite, because the pipeline's required checks did not cover configuration files, and the on-call runbook described the manual step as optional. Same event. One of those sentences produces a defensive meeting, and the other produces two pull requests.

How do you keep the draft blameless?

Two constraints in the prompt, and one pass with your own eyes.

  • Ban personal names and role titles in the analysis, and require system actors instead: the deploy pipeline, the alerting rule, the runbook, the feature flag. Names belong in the timeline, where they record who was present, and nowhere else.
  • Pair every action with what the person knew at that moment, not with what the review now knows. The difference between those two is the entire content of hindsight bias.
  • Require contributing factors in the plural, because a review that finds one cause has usually stopped at the first plausible one.

Then read the draft yourself with a single test: would this sentence still make sense in a performance review. If it would, rewrite it as a statement about the system. That test catches more than any instruction you can write into the prompt, and it takes about four minutes on a two page document.

PagerDuty's published postmortem process and Atlassian's incident handbook both treat the write up as a step inside a habit rather than as the deliverable, and that framing is the other half of keeping a generated review honest. The document is not the point. The follow up actions with owners are.

What about contributing factors and action items?

Use the model to widen the list, never to close it. Ask it to propose every factor that plausibly contributed, including detection, escalation, tooling, documentation and the decision environment, and to say what evidence in the timeline supports each one. Expect some of them to be wrong. A list of nine candidates where four survive scrutiny is a better input to a meeting than a confident list of two, because it makes the team argue about the system instead of ratifying a story.

Then stop. Cause selection belongs to the people who own the systems, and the model has an obvious failure mode here: a fluent causal explanation is precisely what it generates well, and a plausible wrong one is expensive because it ends the search. The same asymmetry runs through triaging bugs with Claude Code as a PM, where the instruction that earns its place in every prompt is do not propose a fix. Here the equivalent is do not select a cause.

Action items are where generated reviews most often go quietly wrong. A model will produce items like improve monitoring and add more tests, which are not actions because nobody can be assigned them and nobody can tell when they are done. Require each item to name a person, a specific change and a date, and delete anything that fails all three. A review with two real actions beats one with eight aspirations.

What should never be pasted in?

Incident channels are an unusually bad data source to hand over unchecked. Under pressure people paste API keys, customer email addresses, database rows, internal hostnames and the occasional exasperated comment about a colleague, and all of that sits in the same export you are about to upload. Strip identifiers and secrets before anything leaves your environment, or use a tool your company has already cleared for that data class.

The exasperated comments deserve their own mention. A model summarising a transcript will faithfully carry the tone across into the write up, and a single irritated line from hour three of an outage reads very differently in a permanent document than it did in the moment. Cut them. They are not evidence about the system.

When the incident is an AI feature misbehaving

The review changes shape when the thing that failed was a model. There is no stack trace, the behaviour is probabilistic, and "it worked in testing" is often literally true, which makes the timeline harder to build and the contributing factors harder to name. Builders Camp's AI Product Management bootcamp covers that version directly across 2 weeks and 4 live sessions, treating evaluations as the launch gate that replaces acceptance criteria, and treating guardrails as the product decision that most failures trace back to rather than model quality. Its practical challenge is rated advanced and takes about 90 minutes, and it opens on a shipped AI feature that has just failed a bias audit with nobody having done anything wrong on purpose, which is the honest shape of most AI incidents.

For the everyday delivery layer around reviews, Project Management for Product runs 1 week with 1 live session and 8 self-paced microlessons on planning, dependencies, risk and status communication. The Product Delivery Specialist Track collects that with the rest of the delivery path, as a curated sequence of 6 to 10 bootcamps.

Run the timeline step alone before you automate anything else

Take the last incident you reviewed and rebuild only its timeline from the raw sources, with the gaps marked. Compare it to the write up you published. If the reconstructed version shows a twenty minute window between the failure starting and the first alert, and your original document opened at the alert, then the review you ran was never about the incident. It was about the response, and the detection problem underneath it is still there, unowned, waiting for the next one. Working with engineers on that kind of finding is its own skill, covered in how PMs work with engineers using AI agents.

See the Project Management for Product bootcamp

For synthesising a fixed set of documents with citations back to the source, see NotebookLM for product managers.

Bootcamps referred in this Guide

Frequently asked questions

What does blameless actually mean in an incident review?

It means the write up explains the system that made a decision look reasonable at the time, instead of naming who made it. Google's SRE guidance is explicit that a blameless postmortem focuses on the contributing causes without indicting an individual or a team, on the grounds that people who expect punishment stop reporting the things you most need reported.

What is AI genuinely good at here?

Merging sources into one ordered timeline. Deploy logs, alert history, a Slack thread and a status page update all describe the same hour in different formats and time zones, and stitching them into a single sequence with every entry labelled by source is tedious, mechanical and checkable. That is the best use of a model in this process.

Why does a model add blame when nobody asked for it?

Because narrative prose about events defaults to naming an agent for each action, and the training data is full of writing that does exactly that. Given a timeline it will produce lines like a named engineer deployed without running the tests. That is a grammatically natural sentence and an analytically useless one, because it stops at the person instead of asking why the pipeline allowed it.

How do you stop the draft from turning into an accusation?

Two prompt rules and one review pass. Ban personal names and role titles in favour of system actors such as the deploy pipeline or the alerting rule, and require every action to be paired with what the person knew at that moment. Then read the draft looking for any sentence that would still make sense as a performance review comment, and rewrite it.

Can a model identify the cause of an incident?

It can propose candidates, and you should treat them as candidates. Fluent causal stories are exactly what a language model produces well, and a plausible wrong one is expensive because the team stops looking. Use it to widen the list of contributing factors, then let the people who own the systems narrow it.

What should never go into a general purpose AI tool during a review?

Raw logs and chat transcripts before anyone has checked them for customer data, tokens, internal hostnames and personal messages. An incident channel is an unusually rich source of all four, because people paste things under pressure. Strip first, or use a tool your company has already approved for that data class.

Does automating the write up make reviews better?

It makes them likelier to happen, which matters, and it does nothing for the part that creates value. Both the Google and Atlassian guidance treat the review as an organisational habit rather than a document: the write up is worth the time only if the follow up actions get owners and get done. A faster document with no owner on any action is a faster way to produce nothing.

Which Builders Camp programme covers the surrounding practice?

Project Management for Product covers the delivery layer in 1 week, with 1 live session and 8 self-paced microlessons on planning, dependency management, risk and status communication. AI Product Management covers the harder version for AI features, in 2 weeks across 4 live sessions, including evaluation, guardrails and what to do when a shipped model behaves badly.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Tiago Pedro da Costa

Tiago Pedro da Costa

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

As Co-founder & CTO of Zumer, Tiago builds platforms that leverage AI to automate knowledge, improve collaboration, and accelerate sustainability in the construction industry. His work ranges from Abaqus, a platform for project and site management, to an AI-powered assistant supporting BREEAM certification.

LinkedInMore guides by Tiago Pedro da Costa

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Project Management for Product bootcamp