Builders Camp

Tools

How to use AI for prioritization when an AI score is an argument, not an answer

A model will score 60 backlog items on RICE in about two minutes, and at least two of the four inputs it uses will be invented. The output is worth having as an argument that exposes disagreement fast, and it is never the decision, because deciding what to cut means telling someone no.

What does AI do well when you are ranking work?

It applies the same rubric to item 60 that it applied to item 1. Human scoring drifts, and it drifts in a predictable direction: the items scored in the first hour get real thought, the items scored in the last twenty minutes get whatever number ends the meeting. A model scoring 60 items in one pass is consistent by construction, and consistency is what makes an outlier visible. When an item lands three times higher than everything near it, that is now a signal about the item rather than a signal about when it got scored.

The second thing it does well is speed on the parts that are genuinely mechanical. Intercom's RICE framework asks for reach as a number of people or events per time period, effort as person months, impact on a small scale, and confidence as a percentage. Two of those four, reach and effort, are estimable from a ticket plus your product data. Get a first pass on all 60 in the time it takes to make coffee, then spend the meeting arguing about the ones that matter.

Why is an AI generated RICE score an argument and not an answer?

Because the arithmetic is trivial and the inputs are the whole game. RICE multiplies reach, impact and confidence, then divides by effort, so the output inherits every weakness of what went in. Intercom's own description sets confidence at 100 percent for high, 80 percent for medium and 50 percent for low, explicitly as a brake on enthusiasm for exciting but ill defined ideas. A model has no data about your product, so when it assigns 80 percent confidence it is expressing the shape of a plausible answer, not a measurement.

Impact is worse. It is the factor that most moves a ranking and the one with no external anchor at all. A model estimating impact is pattern matching against how similar features have been described on the internet, which is a reasonable prior and a terrible substitute for knowing that your activation rate is already near its practical ceiling and this particular improvement has almost nothing left to win.

Read the output accordingly. The score is a position in an argument, with its reasoning attached. That is more useful than most meetings produce, and it is not a decision.

Which numbers should you supply rather than let it guess?

Three, and supplying them changes the character of the exercise entirely.

Factor Who should produce it Why the model gets it wrong
Reach Your analytics, as a real count per period It has no access to how many users touch that screen
Effort The engineers who would build it, in person months It cannot see your codebase, your test suite or your review queue
Impact You, against a stated goal for the period It is pattern matching on descriptions, not on your funnel
Confidence You, per Intercom's 100, 80 and 50 scale A percentage it invents reads as evidence and is not

With reach and effort supplied by people who know, the model does clustering and arithmetic, which it does reliably. Without them, you have automated the part of prioritization that was never the bottleneck and dressed up guesses as a ranked list. ProductPlan's description of the same model makes the point that RICE exists to force explicit trade offs rather than to produce a number, and an invented number defeats the purpose it was built for.

Score problems, not requested features

The move that pays more is upstream of scoring. If you rank 60 feature requests against each other, you have accepted whatever solutions people happened to type into the tracker as the universe of options, which is how teams end up building the fourth filter on a table that should have been a different view.

Group by underlying problem first. Fifteen tickets asking for different exports are one problem about getting data out. Score that problem against the other problems, and the ranking becomes a statement about where to invest attention rather than a list of features with numbers next to them. A model is genuinely good at that clustering step, and it is the step most teams skip because it feels like extra work before the real work. This is the same discipline described in AI for backlog refinement, applied one level up.

What about the bias the score is supposed to remove?

Scoring frameworks exist partly to interrupt optimism, and the optimism is documented. Dan Lovallo and Daniel Kahneman's 2003 Harvard Business Review analysis of executive forecasting describes how teams systematically overrate their own initiatives by reasoning from the inside of the project rather than from the record of similar projects. A RICE score is a small structural defence against that, and a model producing the score removes one specific bias: the person doing the scoring is no longer the person who wants the feature.

It introduces a different one. A model produces the median plausible answer, which flattens anything unusual. A genuinely novel bet, the kind with thin evidence and high upside, scores badly against a rubric built from averages, and it scores badly consistently. If your process is score, rank, build the top five, a drafted score will quietly filter out exactly the work that makes a difference to your strategy. That is a real cost, and it is why the ranking arrives at the meeting rather than replacing it.

Which call never gets delegated?

What gets cut. Ranking is arithmetic; cutting is telling a named person that the thing they asked for is not happening this quarter, and then holding that position when they escalate. No score carries that. The decision depends on which commitments are already made, whose trust you need in six months, and how much disappointment the team can absorb alongside everything else going on.

Two more stay human. Whether the goal the scores are measured against is the right goal, which is a strategy question rather than a prioritization one, and whether an item that scores well is something the team can actually take on given what is already in flight. How to Design OKRs treats the first of those directly, including the failure mode where teams score initiatives enthusiastically against objectives nobody believes.

Use the score to find the disagreement

Here is the move worth stealing. Score the list twice, once with the model told the goal is growth and once with the goal set to retention, and put the two rankings side by side. The items that move most between the two are the ones where your team does not actually agree about what the quarter is for, and that disagreement is usually invisible until something has to be cut in week eight.

That trick costs five minutes and surfaces a strategy problem that would otherwise show up as a delivery problem. Product Strategy spends 2 weeks on the underlying skill, choosing where to compete and defending the trade off, and Product Manager Foundations covers the decision toolkit underneath it for anyone earlier in the path.

See the Product Strategy bootcamp

For the frameworks themselves, the RICE prioritization framework and the ICE scoring framework cover the mechanics, and AI for roadmap planning covers what happens to the ranked list once it becomes a sequence.

Bootcamps referred in this Guide

Frequently asked questions

Can AI score a backlog with RICE?

It can produce a score for every item in minutes, which is genuinely useful. Reach and effort are the two factors it handles least badly, because both are estimable from the ticket and from product data. Impact and confidence are where it fills gaps with plausible numbers, and those two drive most of the ranking.

Why is an AI generated score an argument rather than an answer?

Because the score is arithmetic on four inputs, and a model invents at least two of them. Intercom's own RICE description sets confidence at 100 percent for high, 80 for medium and 50 for low, based on whether you have data. A model has no data about your product, so its confidence value is a guess wearing a percentage sign.

What should I give it so the numbers mean something?

Real reach figures from your analytics rather than estimates, effort figures from the engineers who would build each item, and your actual goal for the period. With those three supplied, the model is doing arithmetic and clustering rather than invention.

Where does AI beat a human at prioritization?

Consistency across volume. A person scoring 60 items drifts: the items scored at 9 AM get treated differently from the ones scored at 5 PM. A model applies the same rubric to item 60 as to item 1, which makes the outliers easier to spot.

What is the decision that never gets delegated?

What gets cut. Scoring ranks things; cutting means telling a stakeholder no and living with it. That call depends on commitments, relationships and what the team can absorb, none of which is in the ticket.

Does this work for prioritising between problems rather than features?

It works better there. Scoring feature requests against each other entrenches whatever solutions people happened to propose. Group the requests by underlying problem first, then score the problems, and the ranking becomes a statement about where to look rather than what to build.

Does Builders Camp teach a specific prioritization tool?

No. Product Strategy covers strategic choices and trade offs, and Product Manager Foundations covers the core PM decision toolkit, both taught as methods rather than through a named product.

Sources

Written by

Andre Albuquerque

Andre Albuquerque

CEO of Builders Camp, SuperOperator, and other companies. Building products.

CEO of Builders Camp, SuperOperator, and other companies. Building products.

LinkedInMore guides by Andre Albuquerque
Inês Lourenço

Inês Lourenço

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.

LinkedInMore guides by Inês Lourenço

Last updated 2026-09-18

Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.

See the Product Strategy bootcamp