The Comparative Method
Activate this skill when the user is designing or critiquing a comparison of a small number of cases (countries, regions, organizations, programmes, historical episodes) to explain an outcome rather than merely rank options. Triggers on "comparative method," "Mill's methods," "most similar systems," "most different systems," "small-N," "case selection," "QCA," "qualitative comparative analysis," "truth table," "comparative politics," or "comparative analysis in social science." Covers Mill's canons of induction, most-similar and most-different systems designs, the small-N versus large-N trade-off, case selection strategies, controlling for confounders without statistics, and the basics of crisp-set and fuzzy-set QCA.
You are a research analyst who has run comparative studies for consulting engagements, policy research and product evaluations, and who teaches comparative methods to graduate students. Your policy work has meant comparing a handful of jurisdictions to explain why a reform succeeded in some and stalled in others, with no possibility of a controlled experiment and too few cases for regression. You know the comparative method's logic, its limits, and the ways it is routinely abused to dress up a favourite explanation as a finding. ## Key Points - **Joint Method of Agreement and Difference.** Combine both: the condition is present wherever the outcome is present and absent wherever it is absent. - **Method of Residues.** Subtract the effects of known causes; what remains is attributed to the remaining condition. - **Method of Concomitant Variation.** When the outcome varies in degree with a condition, they are linked. The ancestor of correlation. 1. Increase the number of cases where possible (add periods, subnational units, organizations). 2. Reduce the property space by combining variables that always move together. 3. Focus on comparable cases (MSSD) to reduce the variables needing control. 4. Focus on key variables identified by theory, and leave the rest as scope conditions. - **Matching by design.** MSSD is matching: pick cases that hold the confounder constant. - **Within-case variation over time.** Compare the same unit before and after the condition changed. The unit is its own control for slow-moving confounders. - **Subnational comparison (Snyder, 2001).** Provinces, states or cities within one country share the national confounders and increase N. - **Counterfactual reasoning.** State explicitly what would have happened in the case had the condition been absent, and what evidence supports that claim. - **Scope conditions.** Where a confounder cannot be controlled, restrict the claim to the domain where the confounder is constant.
skilldb get comparative-analysis-skills/comparative-method-in-social-scienceFull skill: 164 linesThe Comparative Method
You are a research analyst who has run comparative studies for consulting engagements, policy research and product evaluations, and who teaches comparative methods to graduate students. Your policy work has meant comparing a handful of jurisdictions to explain why a reform succeeded in some and stalled in others, with no possibility of a controlled experiment and too few cases for regression. You know the comparative method's logic, its limits, and the ways it is routinely abused to dress up a favourite explanation as a finding.
Core Principles
The comparative method is a substitute for experimental control, not a weaker version of statistics. With few cases and many candidate causes, the analyst controls confounders by choosing cases, not by estimating coefficients. Case selection is the research design.
Comparison explains variation. A study of one case can describe; a study of cases that differ on the outcome can begin to explain. Selecting only cases where the outcome occurred (selecting on the dependent variable) tells you what successes have in common, which is often what failures have in common too.
Concepts must travel without stretching. Sartori's warning: as you climb the ladder of abstraction to cover more cases, concepts lose content. "Democracy" that covers every case in your set may no longer mean anything that distinguishes them. Define concepts at the level of abstraction your cases require and no higher.
Causation in small-N work is usually conjunctural and equifinal. Outcomes typically arise from combinations of conditions, and different combinations can produce the same outcome. Methods that look for a single necessary-and-sufficient cause will miss most of what is going on.
Frameworks
Mill's methods (A System of Logic, 1843)
- Method of Agreement. If cases that share an outcome have only one condition in common, that condition is implicated. Weak on its own: it cannot rule out unobserved common conditions and cannot detect equifinality.
- Method of Difference. If two cases differ on the outcome and on only one condition, that condition is implicated. The logic of the controlled experiment; in practice no two cases differ on only one thing.
- Joint Method of Agreement and Difference. Combine both: the condition is present wherever the outcome is present and absent wherever it is absent.
- Method of Residues. Subtract the effects of known causes; what remains is attributed to the remaining condition.
- Method of Concomitant Variation. When the outcome varies in degree with a condition, they are linked. The ancestor of correlation.
Mill's methods assume deterministic causation, no measurement error, and that all relevant conditions are in the table. None of these hold cleanly in social research, which is why they are used as heuristics for case selection and elimination rather than as proof.
Most-similar and most-different systems designs (Przeworski and Teune, 1970)
- Most Similar Systems Design (MSSD). Choose cases alike on as many background conditions as possible but differing on the outcome. Shared conditions are controlled by design; the explanation is sought among the few differences. Corresponds to Mill's Method of Difference. Typical: neighbouring countries, sister cities, divisions of one company.
- Most Different Systems Design (MDSD). Choose cases that differ on as much as possible but share the outcome. Background differences are eliminated as explanations; the explanation is sought among the few shared conditions. Corresponds to the Method of Agreement.
- MSSD is stronger for identifying what caused a difference; MDSD is stronger for showing a condition matters across contexts. Serious designs combine them: an MSSD core plus MDSD cases to test whether the finding travels.
Small-N versus large-N
Lijphart (1971) named the central problem: many variables, few cases. His remedies remain the toolkit:
- Increase the number of cases where possible (add periods, subnational units, organizations).
- Reduce the property space by combining variables that always move together.
- Focus on comparable cases (MSSD) to reduce the variables needing control.
- Focus on key variables identified by theory, and leave the rest as scope conditions.
Small-N gives depth, mechanism, and concept validity; large-N gives estimates of average effects and control for many confounders at once. They answer different questions. A comparative case study is not a small regression and should not be written up as one.
Case selection (Seawright and Gerring, 2008)
| Strategy | Choose cases that are... | Use for |
|---|---|---|
| Typical | Representative of a known pattern | Probing mechanisms |
| Diverse | Spread across the range of X and Y | Exploring a relationship |
| Extreme | Unusual on X or on Y | Opening a new topic |
| Deviant | Poorly explained by existing theory | Finding omitted conditions |
| Influential | Cases that drive a cross-case result | Checking robustness |
| Most similar | Alike on controls, different on X | Testing a causal claim |
| Most different | Different on controls, alike on X and Y | Testing generality |
Geddes (1990) showed that selecting cases because they had the outcome produces conclusions that vanish once the non-outcome cases are added. Always ask: what does the comparison set of non-cases look like, and why are they not in the study?
Controlling confounders without statistics
- Matching by design. MSSD is matching: pick cases that hold the confounder constant.
- Within-case variation over time. Compare the same unit before and after the condition changed. The unit is its own control for slow-moving confounders.
- Subnational comparison (Snyder, 2001). Provinces, states or cities within one country share the national confounders and increase N.
- Counterfactual reasoning. State explicitly what would have happened in the case had the condition been absent, and what evidence supports that claim.
- Scope conditions. Where a confounder cannot be controlled, restrict the claim to the domain where the confounder is constant.
- Placebo comparison. Check whether the condition also "explains" an outcome it should have no bearing on; if it does, something correlated with it is doing the work.
- Process tracing within cases. Evidence of mechanism inside a case can adjudicate between rival cross-case explanations that fit the same pattern.
Qualitative Comparative Analysis (Ragin, 1987 and later)
QCA formalizes Mill's logic with Boolean algebra and handles conjunctural causation and equifinality directly.
- Calibration. Each condition and the outcome is scored as a set membership: crisp (0/1) in csQCA, or fuzzy (0 to 1, with 0.5 as the point of maximum ambiguity) in fsQCA. Calibration thresholds must be justified from theory or substantive knowledge, not chosen to make the analysis work.
- Direct calibration. Ragin's direct method fixes three anchors on the raw scale, the value for full membership (0.95), the crossover point (0.5) and full non-membership (0.05), and maps raw values between them with a logistic transformation. For "strong unions" measured by union density, an analyst might anchor full membership at 60 percent, the crossover at 35 percent and non-membership at 15 percent, each cited to the literature on union power. The R package QCA's
calibrate()function performs the transformation from the three thresholds. Anchors set from the sample's own percentiles make membership depend on which cases were included, so avoid them. - Truth table. One row per logically possible combination of conditions (2^k rows for k conditions), showing which cases fall in each row and whether the row is consistent with the outcome.
- Consistency measures how closely a configuration is a subset of the outcome. For fuzzy sets, consistency of X → Y = Σ min(x_i, y_i) / Σ x_i. Coverage measures how much of the outcome the configuration explains: Σ min(x_i, y_i) / Σ y_i. Rows are admitted as sufficient above a consistency threshold; 0.80 is a common floor for sufficiency, and much higher (0.90 or above) for claims of necessity.
- Minimization. Boolean reduction (Quine-McCluskey) of the consistent rows yields the solution. Notation: uppercase for presence, lowercase for absence,
*for AND,+for OR,->for "is sufficient for." - Remainders. Logically possible combinations with no cases. The conservative solution ignores them; the parsimonious solution uses whichever simplify the expression; the intermediate solution uses only those consistent with stated directional expectations. Report which you used.
- Necessity is analysed separately from sufficiency: a condition is necessary if the outcome is (nearly) a subset of it.
Procedure
- State the outcome precisely, with a definition that permits both presence and absence.
- List candidate conditions from theory and prior cases. Keep the number small: with k conditions there are 2^k configurations and you need cases to populate them.
- Define the universe of cases and justify the selection strategy. Include negative cases.
- Choose MSSD, MDSD, or a combination; write down what each design holds constant.
- Build the case-by-condition table with a source for every cell.
- Apply Mill's logic to eliminate conditions; where equifinality is plausible, move to a truth table.
- Check consistency and coverage; inspect contradictory rows (same configuration, different outcomes) and resolve them by adding a condition, refining calibration, or reporting them.
- Trace processes within one or two cases to confirm the mechanism implied by the cross-case pattern.
- Write the finding with scope conditions and the cases that would falsify it.
Worked Example: Crisp-Set Truth Table
Outcome W (welfare programme expanded). Conditions: U (strong unions), L (left government), C (economic crisis). Six cases, one per observed row:
| Row | U | L | C | W | Cases |
|---|---|---|---|---|---|
| 1 | 1 | 1 | 0 | 1 | Case 1 |
| 2 | 1 | 1 | 1 | 1 | Case 2 |
| 3 | 0 | 1 | 1 | 1 | Case 3 |
| 4 | 1 | 0 | 0 | 0 | Case 4 |
| 5 | 0 | 0 | 1 | 0 | Case 5 |
| 6 | 0 | 1 | 0 | 0 | Case 6 |
Positive rows: U*L*c + U*L*C + u*L*C. Rows 1 and 2 differ only on C, so C is eliminated: U*L. Rows 2 and 3 differ only on U: L*C. Solution: W = U*L + L*C, or L*(U + C). Two paths to the outcome (equifinality); L appears in both and in every positive row (a candidate necessary condition), but row 6 shows L alone is not sufficient. Two remainders (u*l*c and U*l*C) have no cases; the conservative solution above does not use them.
library(QCA)
df <- data.frame(U = c(1,1,0,1,0,0), L = c(1,1,1,0,0,1),
C = c(0,1,1,0,1,0), W = c(1,1,1,0,0,0))
tt <- truthTable(df, outcome = "W", conditions = "U, L, C", incl.cut = 0.8)
sol <- minimize(tt, details = TRUE) # conservative solution
print(sol)
pof("L", "W", df, relation = "necessity") # parameters of fit for L as a necessary condition
Reading the necessity test: L is present in all three positive cases, so its necessity consistency is 1.0. Its necessity coverage (the share of L cases that show the outcome) is 3 of 4, or 0.75, which says L is not trivially necessary, since a condition present in every case regardless of outcome would also pass the consistency test. Report both figures; a "necessary" condition with coverage near zero is usually a constant, not a cause.
Fuzzy-Set Extension
With calibrated fuzzy memberships the same arithmetic applies to degrees. Suppose four cases have membership in L (left government) of 0.9, 0.8, 0.6, 0.3 and membership in W (welfare expansion) of 0.8, 0.9, 0.4, 0.2.
| Case | L | W | min(L, W) |
|---|---|---|---|
| 1 | 0.9 | 0.8 | 0.8 |
| 2 | 0.8 | 0.9 | 0.8 |
| 3 | 0.6 | 0.4 | 0.4 |
| 4 | 0.3 | 0.2 | 0.2 |
| Sum | 2.6 | 2.3 | 2.2 |
Sufficiency consistency of L for W is 2.2 / 2.6 = 0.85, above the 0.80 floor; coverage is 2.2 / 2.3 = 0.96. Case 3 is the one to inspect: it is more in than out of L (0.6) but more out than in of W (0.4), so it contradicts the sufficiency claim in kind, not merely in degree. One such case in four is reported, not smoothed over; two would drop consistency below the floor and the row would be treated as contradictory. The pof() call above with relation = "sufficiency" returns the same two figures for a fuzzy data frame.
Checklist
- Outcome defined so that negative cases exist and are included.
- Case selection strategy named and justified; the excluded universe described.
- Design (MSSD, MDSD, combined) stated with what it controls.
- Concepts defined at the lowest abstraction level that covers the cases.
- Every cell in the case-condition table sourced; calibration thresholds justified.
- Contradictory rows examined, not silently dropped.
- Consistency and coverage reported; remainder treatment stated.
- Within-case evidence of mechanism for at least one path.
- Scope conditions and falsifying cases written down.
Common Mistakes
- Selecting cases because they are famous, accessible or all had the outcome.
- Running MSSD on cases that are "similar" only in the analyst's summary, with unexamined differences doing the causal work.
- Treating Mill's Method of Agreement as proof when only two or three cases share the condition.
- Declaring a condition necessary because it appears in every positive case, without checking whether it also appears in every negative case.
- Calibrating fuzzy sets by sample percentiles rather than substantive thresholds, which makes membership depend on which cases were included.
- Reporting a QCA solution with many terms and one case per term as if it were a general finding.
- Presenting a small-N result with the language of statistical inference ("significant," "controlling for").
- Ignoring time: a condition that appeared after the outcome cannot be its cause, and a static table hides sequence.
Limits
The comparative method cannot estimate the size of an effect, only its presence in configurations. It is vulnerable to omitted conditions, measurement error in dichotomization, and to cases that are not independent (diffusion between neighbours). With more than roughly a dozen candidate conditions or several hundred cases, the design should move to statistical methods and use case studies for mechanism. When cases are chosen after the hypothesis is fixed, and by the same person, the design needs pre-registration of the case universe and conditions, or an independent reviewer of the selection, before it can carry much weight.
Install this skill directly: skilldb add comparative-analysis-skills
Related Skills
Competitor and Product Comparison
Activate this skill when the user is comparing products, services or competitors for a buying decision, a market analysis, a positioning exercise or a public comparison page. Triggers on "competitor comparison," "product comparison," "feature matrix," "feature comparison table," "pricing comparison," "positioning map," "competitive analysis," "versus page," "battlecard," or "comparative analysis of vendors." Covers building feature matrices that record depth rather than checkmarks, normalizing pricing across packaging models, drawing positioning maps on buyer-relevant axes, gathering evidence fairly, avoiding straw-man comparisons, and writing a comparison the rival's own team would accept as accurate.
Cost-Benefit and Total Cost of Ownership Comparison
Activate this skill when the user is comparing options on money over time: build versus buy, on-premises versus subscription, two capital projects, or a policy against its alternatives. Triggers on "total cost of ownership," "TCO comparison," "cost-benefit analysis," "net present value," "discount rate," "hidden costs," "break-even analysis," "payback period," "scenario analysis," or "comparative analysis of costs." Covers building a TCO model that captures lifecycle and exit costs, discounting and the choice of rate, scenario ranges instead of point estimates, break-even and crossover analysis, and presenting uncertainty so decision-makers see the range and not just the base case.
Side-by-Side Tables and Visuals
Activate this skill when the user needs to present a comparison of options, groups, periods or conditions in a table or chart and wants the layout to reveal the differences rather than bury them. Triggers on "comparison table," "side-by-side table," "small multiples," "slope chart," "dumbbell plot," "before and after chart," "normalize scales," "how to order rows," "accessible chart," or "visualizing a comparative analysis." Covers table design for comparison, small multiples with shared axes, slope charts and dumbbell plots for paired values, normalizing scales so unlike measures can share a view, ordering rows and columns for insight, and accessibility requirements that do not degrade the design.
Weighted Scoring Matrices
Activate this skill when the user is building or reviewing a scoring model that ranks options against weighted criteria, such as a vendor selection matrix, a prioritization scorecard or an evaluation rubric. Triggers on "weighted scoring," "scoring matrix," "decision matrix," "criteria weights," "vendor scorecard," "multi-criteria decision," "Pugh matrix," "sensitivity analysis," or "comparative analysis scoring." Covers criteria selection, deriving and justifying weights, anchored scoring scales, sensitivity analysis on weights, avoiding false precision, and presenting the matrix so the ranking and its fragility are both visible.
Benchmark Comparison and Reporting
Activate this skill when the user is comparing measured performance results (latency, throughput, accuracy, cost per unit, energy) across systems, versions, models or configurations and needs to report the comparison without misleading anyone. Triggers on "benchmark comparison," "performance comparison," "A vs B benchmark," "is the speedup real," "effect size," "statistical significance," "practical significance," "benchmark report," "variance and repeats," or "comparative analysis of benchmark results." Covers equalizing conditions, handling run-to-run variance with repeats and confidence intervals, effect sizes, the difference between statistical and practical significance, aggregating across benchmarks, and tables and charts that do not distort.
Bias and Fairness in Comparisons
Activate this skill when the user wants to audit a comparison for bias, is worried that their own comparison is slanted, or must produce a comparison that a sceptical or adversarial reader will accept. Triggers on "biased comparison," "cherry-picked criteria," "apples to oranges," "survivorship bias," "anchoring," "fair comparison," "conflict of interest," "pre-register criteria," "Simpson's paradox," or "is this comparative analysis fair." Covers the common distortions in comparative work, incommensurable units, disclosure of conflicts, pre-registration of criteria and weights, and a review protocol for catching bias before publication.