Turbo AI PM
Kits / product scenario · full framework

What should we show this user next?

Recommendations & Personalization. Feeds, "you might also like," homepages, next episodes, product carousels. The model is one component of a recommendation product. Most of the decisions that determine whether it works — what can be shown, what "good" means, what must never be shown, and how you'll know it helped — are product decisions.

Every decision below, with checkboxes.

The system

Four components, and only one of them is "the model." Each raises a product question before it raises a technical one.

Candidate generationWhat could this user possibly see?
→
RankingWhat appears first — optimized for what?
→
Business & product constraintsWhat must the model never override?
→
User experienceHow is it shown, and what does the user do?
↺ What users do next becomes the training data for the next recommendation — see step 9.
Same discipline, different model. This is a ranking system, not a chatbot — and the PM job doesn't change. Whether the model is a ranker, a classifier, a retriever, or an LLM, you still define the objective, the system's behavior, how you'll evaluate it, how it fails, how you'll experiment, what it may never do, what you'll monitor, and what it's worth.
01

Define the product objective

What user and business outcome are we trying to improve?

Decide
  • Which user outcome: finding something faster, discovering something new, or coming back more often?
  • Which business outcome — and how directly does a better recommendation move it?
  • Which proxy will you optimize day to day, and how far is it from the outcome you care about?
02

Define the recommendation problem

What are we recommending, to whom, and in what context?

Decide
  • What is being recommended: items, content, people, actions?
  • Where does it appear: homepage, product page, email, post-purchase?
  • How many slots, and what is the user doing in that moment?

How personal do recommendations need to be?

Generic (popular for everyone), segment-level (popular with people like you), and individual personalization each need more data and carry more cold-start and privacy risk. On some surfaces, well-chosen popular items perform close to a personalized model.

Decide
  • Is there enough data to personalize individually on this surface?
  • Does individual personalization beat segment-level by enough to justify it?
  • Which personal data are we using — and would users be comfortable knowing?
03

Design candidate generation

Where do recommendations come from before anything is ranked?

What can this user possibly see — and what gets thrown out before ranking?

Ranking only reorders what candidate generation lets through. An item that never enters the pool can never be recommended, however relevant it is.

Decide
  • What is the pool: the full catalog, in-stock items, items available in the user's region?
  • Which sources feed it: similar items, co-purchases, trending, editorial picks, new arrivals?
  • How large should it be — big enough to keep good options, small enough to rank quickly?
  • What stops new or niche items from being cut before they get a chance?

What do we show someone we know nothing about — and how does a brand-new item ever get seen?

New users and new items have no history, so a purely personalized system has nothing to work with at exactly the moments that matter most: a user's first session and an item's launch.

Decide
  • New user: popular items, a few onboarding questions, or context such as location, device, and referral source?
  • New item: reserved exposure, similarity to existing items, or editorial placement?
  • How much history does a new user need before switching to personalized results?
04

Design ranking

How do we decide what appears first?

What are we actually optimizing when we decide what appears first?

Relevance, clicks, conversion, engagement, retention, and revenue each produce a different ranking. Choosing the objective — or the blend — is a product and business decision. The model optimizes whatever it is given.

Decide
  • One objective, or a weighted blend — and who signs off on the weights?
  • What must ranking never trade away to improve that objective?
  • How often are the weights revisited?

Are we recommending what's good, or what's already popular?

Popular items collect more interactions, which makes them look more relevant, which gets them shown more. Exposure concentrates on a small head of the catalog, and new or niche items never gather the interactions they need to compete — a real problem for marketplaces and content ecosystems that depend on a healthy supply side.

Decide
  • What share of impressions goes to the top 1% of items — and is that acceptable?
  • Do new items and smaller sellers or creators get a guaranteed path to exposure?
  • Is catalog coverage tracked alongside relevance?

How much variety should one screen of recommendations have?

Ranking purely on relevance often returns ten versions of the same thing. Users tire of repetition, and near-duplicates waste slots that could teach you something new about what the user wants.

Decide
  • Should similar items be grouped, capped, or spread across the page?
  • How much relevance are we willing to give up for discovery?
  • Is novelty a goal on this surface, or a distraction?
05

Define constraints

What should the model never override?

Which recommendations should never be shown, whatever the model scores them?

The highest-scoring item is not always one you can or should show. Recommendation systems are decision systems embedded in a product, and the product has rules the model does not know about.

Decide
  • Eligibility: inventory, region, age, account type
  • Commercial: contracts, sponsored placements, margin floors
  • Safety and fairness: excluded content, sensitive categories
  • Experience: frequency caps, not recommending what they just bought
  • Where do the rules run — before ranking, after it, or both?
06

Define feedback signals

What evidence will the system learn from — and can we trust it?

Which user behavior actually tells us a recommendation was useful?

Explicit signals (ratings, likes, saves) are clear but rare. Implicit signals (clicks, views, purchases, skips) are plentiful but ambiguous: a click can mean interest, curiosity, or a misleading thumbnail. They often disagree — heavily clicked items can be poorly rated — so the signal you train on shapes what the system learns to produce.

Decide
  • Which signal is the training target, and which ones validate it?
  • How do a click, a save, and a purchase compare in weight?
  • What does a skip or an immediate bounce count as?

If users click the first result because it's first, are we teaching the model that position equals relevance?

Top slots get more clicks regardless of quality. Train on raw clicks and the system learns to keep whatever it already ranked first — and an experiment can credit a new model for what is really a position effect.

Decide
  • Do we log the position each item was shown in, so the effect can be corrected for?
  • Do we shuffle a small share of results to measure it?
  • When comparing two models, are their items shown in comparable positions?

What if the outcome that matters shows up weeks after the click?

Purchases, subscriptions, retention, and repeat usage arrive long after the recommendation. Optimizing what is measurable today — clicks — can pull against what matters later: clickbait wins the week and loses the quarter.

Decide
  • Which short-term signal best predicts the long-term outcome?
  • How long must an experiment run to observe the outcome that matters?
  • Which long-term metric guards against short-term optimization?
07

Define evaluation

How do we know the system works before real users see it?

Why can a model look better on historical data and still not improve the product?

Offline evaluation replays past behavior — but that behavior was shaped by the previous system's recommendations, so it cannot show how users react to things they were never shown. It is cheap and fast for ruling models out. Only online evaluation — A/B tests and holdouts — measures real outcomes.

Decide
  • Which offline metrics are gates a model must pass before an online test?
  • What is the baseline: popularity ranking, the current system, or random?
  • Which slices — new users, new items, small categories — must be checked separately?
08

Experiment

Does it improve the actual product?

What must be defined before an experiment result can be trusted?

An A/B test only answers the question it was designed to answer. Decide these before launch — not after you have seen the numbers.

Decide
  • Control and treatment: exactly what differs between them
  • Primary metric: the one number that decides it
  • Guardrails: what must not get worse — retention, complaints, catalog coverage
  • Segments: new vs. returning users and key categories, chosen up front rather than mined afterward
  • Duration and stopping: long enough for weekly cycles and delayed outcomes, and no stopping the moment it looks good

Do we always show what we know the user likes, or reserve some traffic for discovering something new?

Pure exploitation maximizes today's metrics while slowly starving the system of new information: it never learns about items or interests it does not already show. A small exploration budget costs some short-term performance and keeps personalization, new-item discovery, and future training data healthy.

Decide
  • What share of traffic or slots is reserved for exploration?
  • Which surfaces can tolerate it — the homepage, but probably not checkout?
  • How do we measure what exploration taught us?
09

Monitor

What can deteriorate after launch?

Is the system reinforcing its own past decisions?

Recommendation → user behavior → training data → next recommendation. Exposed items get clicked, clicks make them look popular, and popular items get exposed again. Over time the system can narrow what users see and mistake its own influence for user preference. Treat it as a production risk to monitor, not an ML detail.

Decide
  • Are exposure concentration and catalog coverage monitored over time?
  • Is some training data collected from exploration or randomized traffic?
  • Who reviews whether recommendations are narrowing for users?

Where the kits come in

A scenario is the product problem and the decisions it raises. Kits are reusable system-building workflows — one scenario draws on several, and each kit serves several scenarios.

Existing kits
RAG / knowledge assistantCandidate generation is retrieval: the same Recall@k, Precision@k, MRR, and NDCG apply, and the same "if it isn't retrieved, it can't be ranked" rule.
Classifier / triage & moderationClick and purchase prediction are classification problems underneath — imbalance, thresholds, and slice gaps behave the same way.
Existing frameworks & tools
Baselines & slicesPopularity ranking is the baseline a recommender must beat; new users and new items are the slices to check.
Evaluation methodsEval set design and holdouts for the offline gate.
Monitoring & flywheelDrift, feedback loops, and exploration traffic after launch.
Go / No-GoWhere a homepage feed vs. a checkout recommendation sits on blast radius and reversibility.
ROI calculatorServing cost per recommendation against revenue per session.
Not built yet
Recommendation / Ranking kitA reusable system-building workflow, if this scenario proves the pattern.
Experimentation frameworkA/B design, guardrails, and stopping rules — reusable beyond recommendations.
Personalization frameworkData, privacy, and segment-vs-individual decisions across products.

Other product scenarios