Tools
How to use AI for OKRs that survive contact with the quarter
The highest value use of AI on OKRs is the stress test, not the draft: ask for the data source and query behind every key result and the unmeasurable ones fail in about a minute. The ambition level stays human, because how hard a target should be depends on what the team has just been through.
What is AI actually good at when you are writing goals?
Demolition. Give a model a set of OKRs and ask it to name, for each key result, the exact data source and the query that would produce that number today, and the fake ones collapse immediately. "Improve customer satisfaction" has no query. "Increase developer happiness" has no query. "Reduce time to first value" has a query only if somebody has defined first value, which they usually have not. This test costs a minute and it catches the single defect that kills most OKR cycles: a key result nobody can measure becomes a key result nobody scores, which becomes a quarter with no feedback loop.
Drafting is the weaker use, and it is the one people reach for. A model asked to write OKRs for a product team produces competent, generic goals about activation, retention and time to value, because that is what the internet contains. They are not wrong. They are also not yours, and an objective that is not yours will not survive the first week where somebody has to choose between it and a customer escalation.
What shape should a drafted set come back in?
Constrain it before you ask. The standard structure, as What Matters describes it, is one objective with 3 to 5 supporting key results, expressible as a single sentence: I will achieve this objective as measured by these key results. Google's re:Work guidance on goal setting adds the discipline that key results should be measurable and gradable rather than binary.
Give the model both rules plus one more that it will otherwise break: a key result states a change in a number, not the completion of a piece of work. Then ask it to label each drafted key result as outcome or activity before you read them. That labelling step is the difference between a useful draft and a task list with percentage signs.
| Draft key result | Label | Fixed version |
|---|---|---|
| Ship the new onboarding flow by week 6 | Activity | Onboarding completion rises from its current rate to a stated target |
| Improve activation | Outcome, unmeasurable | Activation, defined as a first successful import, rises to a stated rate |
| Run 12 customer interviews | Activity | Interviews are an initiative under this objective, not a key result |
| Reduce support tickets by 20 percent | Outcome, gameable | Tickets per active account falls 20 percent with resolution time flat or better |
That last row is the one worth studying. A raw ticket count falls if your product loses users, which is why the fix names a denominator and a guardrail.
The stress test that finds the gaming risk
Three questions, asked of every key result, in this order.
- What data source produces this number today, and what is the query?
- What could a team do to hit this number while making the product worse?
- What would have to be true in week three for this to still be achievable in week eleven?
Question two is the one most reviews never ask, and a model answers it well precisely because it is not invested in the answer. Point it at "reduce average response time" and it will tell you, without embarrassment, that closing tickets faster with worse answers achieves that. Point it at "increase signups" and it will name the acquisition channels that inflate the number with users who never return. Ask a team the same question about their own goals and you get defensiveness, because nobody enjoys describing how they would cheat.
Question three converts a quarterly goal into something checkable in week three, which is the only way a missed OKR gets caught early enough to do anything about.
Why does the ambition level stay a human call?
Because it is not a property of the goal. It is a property of the team and the company around it. A target that is right for a team coming off a strong quarter is demoralising for a team that just absorbed two departures and an incident, and the number itself is identical in both cases. A model has no read on that. Asked for an ambitious target it produces a figure that looks ambitious in a document, which is a different thing from a figure that produces effort rather than resignation.
The ambition call also depends on how your company treats a missed goal. In a place where 70 percent on a stretch OKR is a success, stretch targets work. In a place where any miss becomes a performance conversation, the same targets produce sandbagging, and everyone quietly sets goals they have already achieved. That is organisational knowledge, it is rarely written down anywhere a model can read, and getting it wrong wastes a full quarter.
Lovallo and Kahneman's 2003 analysis of executive forecasting in Harvard Business Review describes the structural pull toward overoptimism in exactly this kind of planning. Useful to remember in both directions: a model will not correct your optimism, because it was trained on documents written by optimistic people.
Where a drafted OKR set fails in practice
The common failure is alignment theatre. A team asks a model to cascade a company objective into team objectives, and it does, fluently. Every team gets an objective that visibly references the company one and every key result sounds plausible. Nothing in the output is wrong, and nothing in it required anyone to argue about who owns what.
Then week four arrives, two teams discover they both believed the other one owned the shared metric, and the cascade turns out to have distributed words rather than accountability. How to Design OKRs names unowned key results as one of the failure modes it teaches teams to spot, alongside vanity objectives and metric overload, and a drafted cascade produces all three faster than a human could.
The correction is small and unpopular: after the draft, every key result needs a named person who agrees out loud, in the room, that they own it. If nobody will say the sentence, the key result does not exist regardless of what the document says.
What to do with the draft on Monday
Take your current OKRs, not hypothetical ones, and run the three stress questions against them before you write anything new. Most teams find at least one key result with no query behind it and at least one that can be gamed. Fixing those two is worth more than a fresh drafted set, and it takes under an hour.
How to Design OKRs runs one week on the full cycle: crafting objectives tied to customer and business outcomes, choosing key results that are leading enough to manage, connecting company goals to team goals without copy paste, and running check-ins that surface blockers rather than status. Setting Powerful Objectives covers the step before it for teams whose problem is that nobody agrees what the objective should be. Both sit alongside the wider leadership curriculum in the Product Leadership Track.
See the How to Design OKRs bootcamp
The glossary entry on OKRs defines the terms, leading versus lagging indicators covers the choice inside a key result, and AI for roadmap planning covers the work that ladders into these goals.
Bootcamps referred in this Guide
Frequently asked questions
What is AI genuinely good at with OKRs?
Catching a key result that cannot be measured. Ask it to state, for each key result, the exact data source and the query that would produce the number today, and the unmeasurable ones fail loudly. That test takes about a minute and catches the defect that kills most OKR cycles.
Should AI write the objective?
It can draft one and the draft will be competent and generic. The objective encodes what your company is choosing to care about this quarter, which is the part that cannot be pattern matched. Use the draft as a starting position to react against, not as the answer.
What is the standard shape of an OKR?
An objective with 3 to 5 supporting key results, which What Matters also expresses as the sentence: I will achieve this objective as measured by these key results. If a drafted set has nine key results, the objective is doing too many jobs.
How do I stop AI producing key results that are really tasks?
Give it the rule explicitly: a key result states a change in a number, not the completion of a piece of work. Then ask it to mark every draft key result as outcome or activity and rewrite the activities. Shipping the redesign is activity. Onboarding completion moving from one figure to another is an outcome.
Can AI set the ambition level?
No, and this is the call to keep. How hard a target should be depends on what your team has survived recently, what is already committed, and how your company treats a missed goal. A model has no read on any of that and will default to numbers that look ambitious in a document.
What does a good stress test prompt look like?
Ask three questions of each key result: what data source produces this number today, what would a team do to hit this number while making the product worse, and what would have to be true in week three for this to still be achievable. The second question finds the gaming risk that most reviews miss.
Does Builders Camp teach a specific OKR tool?
No. How to Design OKRs runs one week on objective crafting, choosing key results that measure progress, alignment across teams and the check-in cadence, without naming a product.
Sources

Andre Albuquerque
CEO of Builders Camp, SuperOperator, and other companies. Building products.
CEO of Builders Camp, SuperOperator, and other companies. Building products.
LinkedInMore guides by Andre Albuquerque
Inês Lourenço
CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.
CPTO and founder at Compound Works, Inês helps product leaders build AI-powered operating systems for their teams. She designs context layers, agent workflows, and decision frameworks that let PMs move faster, think clearer, and execute at a higher level.
LinkedInMore guides by Inês LourençoLast updated 2026-09-18
Researched from Builders Camp's bootcamp, track and masterclass material and the sources listed on this page, drafted with AI, and fact-checked against every source cited.
Related guides
How to use AI for roadmap planning without outsourcing the sequence
The useful AI move in roadmap planning is synthesis, not sequencing: paste six categories of scattered input and get...

Andre Albuquerque & Inês LourençoHow to use AI for prioritization when an AI score is an argument, not an answer
A model will score 60 backlog items on RICE in about two minutes, and at least two of the four inputs it uses will be...

Andre Albuquerque & Inês LourençoHow to write a product one pager with AI
Amazon caps the press release half of a PR/FAQ at less than one page and its teams write ten drafts or more before...

Andre Albuquerque & Inês LourençoNorth Star Metric Examples
A North Star Metric is the single number a company treats as the clearest signal that its product is delivering real...
Andre Albuquerque

