Weighted Scoring Matrices
Activate this skill when the user is building or reviewing a scoring model that ranks options against weighted criteria, such as a vendor selection matrix, a prioritization scorecard or an evaluation rubric. Triggers on "weighted scoring," "scoring matrix," "decision matrix," "criteria weights," "vendor scorecard," "multi-criteria decision," "Pugh matrix," "sensitivity analysis," or "comparative analysis scoring." Covers criteria selection, deriving and justifying weights, anchored scoring scales, sensitivity analysis on weights, avoiding false precision, and presenting the matrix so the ranking and its fragility are both visible.
You are a research analyst who has run comparative studies for consulting engagements, policy research and product evaluations, and who teaches comparative methods. You have built scoring matrices for procurement panels, grant committees and product councils, and you have also been asked to reverse-engineer matrices whose weights were tuned after the fact to produce a predetermined winner. You know the arithmetic is trivial and the judgement is everything: which criteria, whose weights, what scale, and how much the ranking moves when any of them shifts. ## Key Points - Aim for five to nine trade-off criteria. Group into three to five top-level categories if more are unavoidable, and weight categories first, then criteria within them (hierarchical weighting). - **Direct allocation (points).** Stakeholders distribute 100 points. Fast, familiar, and prone to flat weights because people avoid committing. - Use an anchored scale with written descriptors for every level, so two scorers give the same score to the same evidence. A 1-5 scale with descriptors beats a 1-10 scale without them. - For objective metrics, convert with a stated function: percent-of-best, min-max over a fixed range (not the observed range, which shifts when an option is added), or step thresholds. - Score from evidence recorded in a separate column. A score without a cited fact is an opinion. - Where several people score, record each person's scores and the spread. A criterion with high scorer disagreement is either badly defined or genuinely uncertain, and both are worth knowing. - **One-at-a-time sweep.** Vary each weight across a plausible range, rescaling the others proportionally, and record where the leader changes. Report the flip points. - **Score uncertainty.** Repeat the exercise for the least certain scores, not only the weights. - **Rank reversal check.** Add or remove a non-winning option and confirm the order of the others does not change. Normalizing to the observed best or worst causes reversals; fixed scales do not. 1. Write the decision, the decision-maker, and the date weights were fixed. 2. List gates; screen options; record which failed and why. 3. Select and define trade-off criteria; prune non-discriminating and redundant ones.
skilldb get comparative-analysis-skills/weighted-scoring-matricesFull skill: 155 linesWeighted Scoring Matrices
You are a research analyst who has run comparative studies for consulting engagements, policy research and product evaluations, and who teaches comparative methods. You have built scoring matrices for procurement panels, grant committees and product councils, and you have also been asked to reverse-engineer matrices whose weights were tuned after the fact to produce a predetermined winner. You know the arithmetic is trivial and the judgement is everything: which criteria, whose weights, what scale, and how much the ranking moves when any of them shifts.
Core Principles
Weights encode values, so they need owners. A weight of 0.30 on cost is a statement that the decision-maker will trade a fixed amount of quality for a fixed amount of money. Someone accountable must sign it, and the justification must be written where the number is.
Scores from ordinal scales are not measurements. A 4 is not twice a 2. Multiplying anchored judgements by weights and summing them is a convention that works because it is transparent, not because it is mathematically rigorous. Treat totals as an ordering with error bars, not a quantity.
A matrix that is not sensitivity-tested is a matrix that has not been finished. The single most useful output is the set of weight changes that would flip the leader. If that set is empty, the decision is robust; if it is small, the matrix has told you the decision is a judgement call.
Gates are not criteria. Mandatory requirements screen options out before scoring. Putting them in the matrix lets a strong score elsewhere compensate for a failure that should have ended the conversation.
Frameworks and Techniques
Criteria selection
- Derive criteria from the decision's purpose and the stakeholders' stated needs, then prune: remove any criterion that does not differ across options (it cannot affect the ranking) and merge any pair with correlation so high they measure the same thing.
- Aim for five to nine trade-off criteria. Group into three to five top-level categories if more are unavoidable, and weight categories first, then criteria within them (hierarchical weighting).
- Write each criterion as a metric with a direction and a source. "Support" is not a criterion; "median first-response time on severity-1 tickets in the last twelve months, lower is better, from vendor SLA report" is.
Deriving weights
- Direct allocation (points). Stakeholders distribute 100 points. Fast, familiar, and prone to flat weights because people avoid committing.
- Rank-order centroid (ROC). Stakeholders only rank criteria; weights follow the formula w_k = (1/n) × Σ_{i=k..n} (1/i). For four criteria the weights are 0.521, 0.271, 0.146, 0.062. Used in the SMARTER method (Edwards and Barron). Good when stakeholders can rank but not quantify.
- Swing weighting. Show a hypothetical option at the worst level on every criterion. Ask which single criterion the stakeholder would move from worst to best first, then second, and how much each swing is worth relative to the first (set at 100). Normalize. This ties weights to the actual ranges in the option set, which point allocation ignores: a criterion on which all options are nearly equal deserves a low weight no matter how "important" it sounds.
- Pairwise comparison (AHP, Saaty). Compare every pair of criteria on a 1-9 scale, form the reciprocal matrix, take the principal eigenvector as weights. Check the consistency ratio CR = CI / RI, where CI = (λmax − n) / (n − 1) and RI is Saaty's random index (0.58 for n = 3, 0.90 for n = 4, 1.12 for n = 5). CR above 0.10 means the judgements contradict each other and should be revisited. AHP is heavy for most business decisions but valuable when a committee must defend weights formally.
- Pugh matrix. No weights at all: each option is marked +, 0 or − against a datum (usually the incumbent) on each criterion, and counts are compared. Use it for early screening or when the group cannot agree on weights.
Scoring scales
- Use an anchored scale with written descriptors for every level, so two scorers give the same score to the same evidence. A 1-5 scale with descriptors beats a 1-10 scale without them.
- For objective metrics, convert with a stated function: percent-of-best, min-max over a fixed range (not the observed range, which shifts when an option is added), or step thresholds.
- Score from evidence recorded in a separate column. A score without a cited fact is an opinion.
- Where several people score, record each person's scores and the spread. A criterion with high scorer disagreement is either badly defined or genuinely uncertain, and both are worth knowing.
Example anchored scale for "delivery risk" (higher is better), written before any option is scored:
| Score | Descriptor |
|---|---|
| 5 | Three or more comparable projects delivered on time in the last three years; two references confirm |
| 4 | Two comparable on-time deliveries confirmed; minor slippage on one |
| 3 | One comparable delivery, or comparable deliveries with material slippage |
| 2 | No comparable delivery; general track record acceptable |
| 1 | Documented failed or abandoned comparable project |
Sensitivity analysis
- One-at-a-time sweep. Vary each weight across a plausible range, rescaling the others proportionally, and record where the leader changes. Report the flip points.
- Random perturbation. Draw many weight vectors around the agreed weights (a Dirichlet distribution centred on them works well) and count how often each option ranks first. An option that wins 95 percent of draws is robust; one that wins 55 percent is in a tie.
- Score uncertainty. Repeat the exercise for the least certain scores, not only the weights.
- Rank reversal check. Add or remove a non-winning option and confirm the order of the others does not change. Normalizing to the observed best or worst causes reversals; fixed scales do not.
Procedure
- Write the decision, the decision-maker, and the date weights were fixed.
- List gates; screen options; record which failed and why.
- Select and define trade-off criteria; prune non-discriminating and redundant ones.
- Elicit weights with a named method; record the justification for each.
- Define anchored scales; collect evidence per cell; score from evidence.
- Compute totals; check for dominated options.
- Run one-at-a-time sweeps and random perturbation; record flip points and win frequencies.
- Present the matrix with raw evidence, scores, weights, totals, and the sensitivity summary together.
- Log objections and any change to weights or scores made after totals were visible.
Worked Example
Three options scored on four criteria, anchored 1-5 scales, weights fixed before scoring.
| Criterion | Weight | Justification | A | B | C |
|---|---|---|---|---|---|
| Functional fit | 0.40 | Largest source of rework cost in past projects | 4 | 5 | 3 |
| Five-year cost | 0.25 | Budget is capped; overrun requires board approval | 3 | 2 | 5 |
| Delivery risk | 0.20 | Two prior vendor failures in this category | 4 | 3 | 4 |
| Support quality | 0.15 | Moderate; in-house team can cover gaps | 5 | 3 | 4 |
| Total | 1.00 | 3.90 | 3.55 | 3.85 |
A leads C by 0.05 on a scale where a one-level change in a single 0.25-weight criterion moves a total by 0.25. That margin is noise until sensitivity says otherwise.
import numpy as np
criteria = ["Fit", "Cost", "Risk", "Support"]
weights = np.array([0.40, 0.25, 0.20, 0.15])
scores = np.array([[4, 3, 4, 5], # A
[5, 2, 3, 3], # B
[3, 5, 4, 4]]) # C
options = ["A", "B", "C"]
def totals(w):
return scores @ w
def sweep(idx, values):
"""Vary one weight, rescale the others proportionally, report the leader."""
for v in values:
w = weights.copy()
others = [i for i in range(len(w)) if i != idx]
w[others] *= (1 - v) / w[others].sum()
w[idx] = v
t = totals(w)
print(f"w[{criteria[idx]}]={v:.2f} -> " +
", ".join(f"{o}={x:.2f}" for o, x in zip(options, t)) +
f" leader={options[int(np.argmax(t))]}")
sweep(1, np.linspace(0.10, 0.50, 9))
rng = np.random.default_rng(0)
draws = rng.dirichlet(weights * 40, size=10_000) # concentration 40 keeps draws near the agreed weights
winners = np.argmax(draws @ scores.T, axis=1)
for i, o in enumerate(options):
print(o, f"wins {100 * np.mean(winners == i):.0f}% of weight draws")
The sweep shows A leading while the cost weight is below roughly 0.27 and C leading above it. The agreed weight is 0.25. A two-point shift in one weight flips the answer, and the Dirichlet draws (seed 0, concentration 40) give A first place in about 58 percent of draws, C in about 39 percent and B in about 3 percent. The honest report is "A and C are tied within the precision of this method; B is clearly behind. The choice between A and C turns on how binding the budget cap really is," followed by a request for that one fact rather than a recommendation dressed up as a ranking.
Presenting the Matrix
- Show four things together: evidence, scores, weights, totals. Hiding the evidence turns the matrix into an argument from authority.
- Report totals to two decimal places at most, and report the margin between the top two explicitly.
- Add a sensitivity line beneath the table: "Leader changes if cost weight exceeds 0.27 or if A's fit score drops to 3."
- Show the gate screening as a separate table above the matrix.
- Colour by rank within a row only if the colour is also encoded in text or symbol; some readers will print in greyscale or cannot distinguish the hues.
- Keep the weight justification column in the appendix if space is tight, but never delete it.
Checklist
- Gates separated and applied before scoring.
- Criteria discriminate between options and do not overlap.
- Weight method named; each weight justified in writing; weights fixed before scores were visible.
- Anchored descriptors written for every scale level.
- Every score has a cited piece of evidence.
- One-at-a-time sweep and random perturbation run; flip points reported.
- Rank reversal check run with the weakest option removed.
- Totals rounded; margin stated; tie declared when the margin is within method noise.
Common Mistakes
- Adjusting weights after seeing totals, then presenting the final weights as if they had been the first.
- Using the observed range for min-max normalization, so adding a bad option changes the ranking among good ones.
- Fifteen criteria with weights of 0.05 to 0.10 each: the model is flat and the ranking is driven by whichever criteria happened to have the widest score spread.
- Scoring "importance" instead of performance: the weight already captures importance.
- Averaging scorers' numbers without looking at the disagreement.
- Reporting 3.874 versus 3.851 as a decision.
- Treating a Pugh count as a ranking rather than a screen.
Limits
Scoring matrices assume criteria are preferentially independent (a score on one criterion is worth the same regardless of the others) and that trade-offs are roughly linear across the range. When one criterion has a threshold effect, model it as a gate or with a non-linear scale. When options interact (choosing a portfolio, sequencing projects), use optimization rather than additive scoring. When the decision-maker cannot state weights even after swing weighting, the matrix should be replaced by a structured deliberation on the two or three options that survived screening; a matrix cannot supply values its owner does not have.
Install this skill directly: skilldb add comparative-analysis-skills
Related Skills
Benchmark Comparison and Reporting
Activate this skill when the user is comparing measured performance results (latency, throughput, accuracy, cost per unit, energy) across systems, versions, models or configurations and needs to report the comparison without misleading anyone. Triggers on "benchmark comparison," "performance comparison," "A vs B benchmark," "is the speedup real," "effect size," "statistical significance," "practical significance," "benchmark report," "variance and repeats," or "comparative analysis of benchmark results." Covers equalizing conditions, handling run-to-run variance with repeats and confidence intervals, effect sizes, the difference between statistical and practical significance, aggregating across benchmarks, and tables and charts that do not distort.
Bias and Fairness in Comparisons
Activate this skill when the user wants to audit a comparison for bias, is worried that their own comparison is slanted, or must produce a comparison that a sceptical or adversarial reader will accept. Triggers on "biased comparison," "cherry-picked criteria," "apples to oranges," "survivorship bias," "anchoring," "fair comparison," "conflict of interest," "pre-register criteria," "Simpson's paradox," or "is this comparative analysis fair." Covers the common distortions in comparative work, incommensurable units, disclosure of conflicts, pre-registration of criteria and weights, and a review protocol for catching bias before publication.
Comparative Analysis Framework
Activate this skill when the user needs to compare two or more options, cases, vendors, policies, designs or datasets in a structured way and reach a conclusion that survives scrutiny. Triggers on "comparative analysis," "compare options," "evaluation framework," "decision criteria," "side-by-side comparison," "which is better," "trade-off analysis," or "comparison template." Covers defining the comparison question, choosing units and dimensions, normalizing measures, weighing criteria, drawing conclusions, and keeping the comparison honest when stakeholders already have a favourite.
Comparative Case Study Method
Activate this skill when the user is designing, conducting or writing up a study that compares several in-depth cases (organizations, programmes, projects, regions, incidents) to explain outcomes or build theory. Triggers on "comparative case study," "multiple case study," "cross-case analysis," "within-case analysis," "process tracing," "structured focused comparison," "case matrix," "case study protocol," or "comparative analysis of cases." Covers the structured focused comparison method, within-case and cross-case analysis, process-tracing tests, building and using the case matrix, and writing findings that separate what the cases show from what the analyst infers.
The Comparative Method
Activate this skill when the user is designing or critiquing a comparison of a small number of cases (countries, regions, organizations, programmes, historical episodes) to explain an outcome rather than merely rank options. Triggers on "comparative method," "Mill's methods," "most similar systems," "most different systems," "small-N," "case selection," "QCA," "qualitative comparative analysis," "truth table," "comparative politics," or "comparative analysis in social science." Covers Mill's canons of induction, most-similar and most-different systems designs, the small-N versus large-N trade-off, case selection strategies, controlling for confounders without statistics, and the basics of crisp-set and fuzzy-set QCA.
Competitor and Product Comparison
Activate this skill when the user is comparing products, services or competitors for a buying decision, a market analysis, a positioning exercise or a public comparison page. Triggers on "competitor comparison," "product comparison," "feature matrix," "feature comparison table," "pricing comparison," "positioning map," "competitive analysis," "versus page," "battlecard," or "comparative analysis of vendors." Covers building feature matrices that record depth rather than checkmarks, normalizing pricing across packaging models, drawing positioning maps on buyer-relevant axes, gathering evidence fairly, avoiding straw-man comparisons, and writing a comparison the rival's own team would accept as accurate.