What RICE is really for
The output is a ranked list, but the value is in the argument it forces. Without a framework, roadmap discussions run on conviction and seniority: the person who cares most, or outranks everyone else, tends to win. RICE replaces that with four estimates, each of which can be investigated.
When two people disagree about a score, the disagreement is now specific. They differ on reach, which analytics can settle, or on effort, which engineering can estimate, or on confidence, which usually means one of them knows something the other does not. All three are more productive arguments than whether a feature feels important.
That is why the framework survives contact with reality better than its arithmetic deserves. The maths is trivial; the discipline of writing down what you believe and why is the part that changes decisions.
Estimating each input honestly
Reach is people affected per time period, usually a quarter. Use a number you can defend — monthly actives who touch the relevant screen, support tickets on the topic, the size of the affected segment. If your only available figure is total users, you are almost certainly overstating it, because most features touch a fraction of the base.
Impact uses a fixed scale rather than a free number precisely to stop inflation: 3 for massive, 2 high, 1 medium, 0.5 low, 0.25 minimal. Resist the urge to invent a 5.
Effort is person-months, and it is where optimism does the most damage. Include design, testing, review, documentation and rollout — not just the coding. A feature estimated at two weeks of development frequently costs six weeks of team time once everything around it is counted.
Confidence is the part teams skip
Confidence exists to stop enthusiasm dominating the ranking. A feature with a huge projected impact and no evidence behind it should not outrank a modest one you are certain about, and without a confidence term it always will.
The scale is deliberately coarse: 100% when you have data, 80% when you have strong reasoning, 50% when it is an informed guess, 20% for a moonshot. Assigning 50% is not an admission of weakness — it is information, and it often prompts the cheapest possible action, which is to go and find out.
A useful discipline is to ask what would have to be true for the reach and impact estimates to hold, and whether anyone has checked. If nobody has, the confidence score should reflect that rather than the strength of the advocacy.
Where the score should not be trusted
RICE cannot see dependencies. A low-scoring piece of infrastructure that unlocks three high-scoring features is worth more than its score, and the framework has no way to represent that.
It cannot see strategy either. Commitments already made to customers, regulatory deadlines, competitive table stakes and deliberate bets on a new market all sit outside the arithmetic, and a team that follows the ranking mechanically will systematically underinvest in all of them.
The right use is diagnostic. When the ranking matches your intuition, you have a defensible way to explain the roadmap. When it contradicts your intuition, one of the two is wrong — and finding out which is the most valuable thing the exercise produces.