All posts

Second thoughts

The weights are decoration

We sell weighted decision matrices. The evidence says the grid does the work and the weighting step adds almost nothing — which changes how you should run one.

7 min read

A weighted matrix will happily tell you to take the job with the commute that ends your marriage. Eight out of ten on salary covers a lot of two-out-of-tens, and the arithmetic has no idea which cell mattered.

We build a tool for running weighted decision matrices, so this is where we are supposed to say the method is sound and you are holding it wrong. I can't. Over the last eighteen months, multi-criteria decision analysis, medical decision aids and AI evaluation landed in the same uncomfortable place. Structure beats winging it, by a lot. The weighting ritual on top adds close to nothing. That is a worse story for us and a better one for you.

Everyone already knows the weights are soft

Everyone who has built one of these in anger has made the same confession. From an April 2022 Hacker News thread on the method, which drew 116 points and is one of the few discussions of it there ever to clear a hundred:

The major problem with this approach is that you go into the analysis with an opinion as to what is the better solution, then you tweak the weights and scores to match that opinion. This is not data-driven or even rational, it's just a way to express "numerically" what your guts tell you.

That is the objection matrix advocates handle worst, because the honest response is: yes, constantly.

The professional version is bleaker. jonathaneunice published N-way product comparisons commercially from 1987 to 1995:

Results are also highly perturbable. Tweak the weights and/or scores but a little and they tell an entirely different story. New winners emerge, clear victories become dead heats, and the former Red Lantern Award winner is suddenly in the middle of the pack.

Hold on to that word, perturbable. Someone should have measured it.

How little it takes to move the answer

Someone did, in June 2025. O'Shea, Deeney, Triantaphyllou, Diaz-Balteiro and Tarim published an exact method in Expert Systems With Applications for computing how far a weight can move before the ranking changes. They ran it on 27 candidate feedstock blends for a biogas digester, five criteria, all weighted equally at 0.2.

To keep all 27 alternatives in their original order, no criterion weight could move more than about 3% in either direction. The tightest tolerated a drop of just 1.51%. When only the top five had to stay in order, the tolerances widened to +35.39%/−33.85% on the least sensitive criterion and +19.97%/−15.32% on the most sensitive.

To hold all 27 places
±3%
The widest any criterion weight could move without reordering the full ranking. The tightest tolerated a drop of just 1.51%.
To hold the top five
~10× wider
Once only the podium had to stay in order, tolerances ran to +35.39%/−33.85% on the least sensitive criterion, and +19.97%/−15.32% on the most.
Verdicts changed by re-aggregating
0
Across 24 aggregation protocols over 2,007 fixed verdicts, reported accuracy still ran from 0.551 to 0.899.

The full ranking is a house of cards. The podium is solid masonry.

Everything below the podium is decoration — a ranking your inputs cannot support, rendered to two decimal places. The method is decent at naming the two or three live options, and close to worthless at ranking eleven above twelve.

Hold the judgments still and you can move the answer anyway. The cleanest demonstration is AI evaluation, where a ground truth exists to score against, unlike your job offer. In June 2026 Delip Rao and Chris Callison-Burch took 2,007 fixed verdicts and varied only the aggregation protocol. Across 24 combinations, reported accuracy ran from 0.551 to 0.899, and chance-corrected agreement crossed zero and changed sign, "without altering a single verdict."

Where the dealbreaker goes to die

Back to the commute. A weighted sum is compensatory by construction: every criterion can be bought off with points from another. That is what the multiply-and-add step is, not a flaw you can weight your way out of.

A June 2026 paper names the failure: "This flat scalarization ignores rubric-specified prerequisite and activation relations among criteria, allowing reward or penalty to be counted even when the condition that licenses it is absent." Once they modelled which criteria gate which, that leakage fell by 96.5% on their benchmarks.

Real decisions are full of gates. On Hacker News in July 2026, an engineer described a defence project under a European-sovereignty requirement:

So that's basically a hard requirement to use mistral, even though Chinese models are strictly better on every dimension.

Correlation does the same damage less visibly. One commenter caught it in that 2022 thread: "I suspect technical ease and scaling ease are highly correlated, which effectively means you're double counting." Real criteria travel in packs, and each pack votes once per row you gave it.

The step that doesn't survive the evidence

This next one changed how I think about our own product.

Patient decision aids are among the most heavily trialled structured-decision tools in existence; a decision matrix has never come close. A January 2024 Cochrane review pooled 209 randomised trials and 107,698 participants. They work: across 21 trials, people reached informed, values-congruent choices far more often (RR 1.75, 95% CI 1.44 to 2.13, moderate-certainty evidence); across 55 trials, they came away markedly clearer about what they valued, that one on high-certainty evidence.

Some of those aids ask you to rate how important each attribute is; that is a weighting step. Others just describe the options and let you weigh them in your head. In February 2025, Dawn Stacey, who led that Cochrane review, published a network meta-analysis across 149 of those trials, 101 explicit and 48 implicit: "There was no significant difference in outcomes when PtDAs that used implicit values clarification were compared with PtDAs that used explicit values clarification."

The single place a difference did appear, it favoured implicit: fewer people left undecided (RR 0.57, 95% CrI 0.33 to 0.98).

"We were surprised to observe no difference."
Stacey et al., Medical Decision Making — February 2025 network meta-analysis of 149 randomised trials of patient decision aids

That is not proof the weighting step does nothing. What 149 randomised trials support is narrower: if the step helps, it does not help much. The part we built a product around did not show up in the results.

A 2023 head-to-head ran two elicitation methods over identical attributes with 459 diabetes patients. One produced a 14.9-fold spread between most and least important attribute; the other, 1.4-fold. Both named the same attribute most important, which says the podium is stable, not that the final verdict would have matched. I have found no 2025–26 replication, so hold it as suggestive.

Decomposition can cost you outright. A May 2026 study led from Carnegie Mellon split holistic judgments into weighted sub-criteria and watched agreement with human raters fall from 0.536 to 0.167 on an instruction-following task, a result the authors themselves call "surprising". On essay scoring the analytic rubrics sometimes won, and the collapse came under a deliberately conservative aggregation rule.

What to keep

None of this is an argument for going back to your gut; the evidence against holistic judgment is older and stronger than anything above. Kuncel and colleagues found that mechanically combining predictors beat expert holistic combination at predicting job performance (average correlation .44 against .28), "an improvement in prediction of more than 50%."

But those mechanical models were mostly unit weights. As a 2020 tutorial on mechanical selection notes, equal weighting "often works almost as well as regression-based weighting." What beat expert judgment was not clever weighting. It was adding the numbers up at all.

Across all three literatures, the same boring component keeps doing the work. Filling every cell forces you to consider every option against every criterion, which your gut skips. The grid is the product; the weighted total is a tiebreaker.

  • Constraints first, as filtersAnything you would refuse at any price never becomes a row. Screen for it before you score.
  • Merge correlated criteriaIf two always move together, they are one criterion voting twice. Collapse them into one row.
  • Weight coarsely, then stopThree levels, or equal weights throughout. Precision you cannot defend is decoration, not rigour.
  • Read the podium, not the rankingTrust the top two or three. The ordering below that is noise you rendered to two decimals.
  • Score before you reveal weightsOne practitioner withheld them until teams had scored, because "almost everyone wants to appeal to it to get their way".
  • Never re-weight to break a tieIf a small nudge flips the winner, go and get better information about that criterion.

The real output was never the number. Asked what he uses now, jonathaneunice landed here: "I've come to realize that 'net scores' are usually not the point. The real goal and need is to have a multi-party discussion that leads towards consensus, or at least a rational understanding of why product/approach X was chosen over its competitors."

We sell the spreadsheet. Use it for the disagreement.

If a 2% nudge to one weight crowns a different option, that is rarer than it feels: O'Shea's podium absorbed swings of fifteen per cent and more. You have two acceptable options and a research task, a better problem than the one you walked in with.

More from the blog