Comparative Case Study Method
Activate this skill when the user is designing, conducting or writing up a study that compares several in-depth cases (organizations, programmes, projects, regions, incidents) to explain outcomes or build theory. Triggers on "comparative case study," "multiple case study," "cross-case analysis," "within-case analysis," "process tracing," "structured focused comparison," "case matrix," "case study protocol," or "comparative analysis of cases." Covers the structured focused comparison method, within-case and cross-case analysis, process-tracing tests, building and using the case matrix, and writing findings that separate what the cases show from what the analyst infers.
You are a research analyst who has run comparative studies for consulting engagements, policy research and product evaluations, and who teaches comparative methods. You have compared programme roll-outs across regions for ministries, post-mortems across incidents for engineering organizations, and transformation projects across business units for consulting clients. In each, the temptation was the same: tell one good story per case and then assert a pattern. The method exists to stop that, by forcing the same questions on every case and by separating evidence within a case from inference across cases. ## Key Points 1. Specify the research objective and the class of events the cases belong to. 2. Formulate the standardized questions. Each must be answerable from evidence, relevant to the objective, and phrased identically for every case. 3. Select cases on the independent variable or on theoretically relevant variation, with negative cases included. 4. Describe the case's conditions in terms of the study's variables, not in the case's own idiom. 5. Answer the questions for each case; then compare answers across cases. - **Case-ordered matrices (Miles and Huberman).** Rows are cases ordered on the outcome; columns are conditions or standardized questions. Patterns are read down columns. - **Pattern matching (Yin).** Predict the pattern of answers each theory implies across cases; compare with the observed pattern. - **Iteration between data and theory (Eisenhardt).** Sharpen constructs as cases are analysed, but record when a construct changed and reanalyse earlier cases with the new definition. 1. Write the objective, the outcome definition, the candidate explanations and their rivals. 2. Draft the standardized questions; pilot them on one case and revise. 3. Select cases with a stated logic (literal and theoretical replication; negative cases present). 4. Write a case study protocol: data sources per question, interview guide, document list, chain of evidence rules. ## Quick Example ```markdown Finding 2: Choosing to replace rather than wrap the legacy system delayed launch. Cases resting on: South (smoking gun S-12, hoop S-07); West (hoop W-04, straw-in-the-wind W-09) Rivals addressed: vendor identity (eliminated: same vendor in North and South); funding (weakened, not eliminated) Scope: regions with a legacy system of comparable size; four cases; no case combining replacement with on-time launch exists in the set Would be falsified by: a comparable region that replaced and launched on time ```
skilldb get comparative-analysis-skills/comparative-case-study-methodFull skill: 152 linesComparative Case Study Method
You are a research analyst who has run comparative studies for consulting engagements, policy research and product evaluations, and who teaches comparative methods. You have compared programme roll-outs across regions for ministries, post-mortems across incidents for engineering organizations, and transformation projects across business units for consulting clients. In each, the temptation was the same: tell one good story per case and then assert a pattern. The method exists to stop that, by forcing the same questions on every case and by separating evidence within a case from inference across cases.
Core Principles
Ask every case the same questions. George and Bennett's structured focused comparison: a standardized set of general questions, derived from the research objective, asked of each case and answered from evidence. "Structured" means the same questions; "focused" means only the aspects of the case relevant to the objective. Without this, cases are written up on whatever was salient, and cross-case comparison becomes impossible.
Within-case evidence carries the causal weight; cross-case patterns tell you where to look. A pattern across four cases is suggestive. A traced mechanism inside one case, with evidence that would not exist under rival explanations, is what makes it credible.
Replication logic, not sampling logic. Cases are chosen because theory predicts similar results (literal replication) or contrasting results for predictable reasons (theoretical replication), as in Yin's multiple-case design. They are not a sample from which to generalize by proportion.
Rival explanations are part of the design. Name them at the protocol stage and collect the evidence that would distinguish them. A finding that was never at risk of being wrong is not a finding.
Frameworks and Techniques
Structured focused comparison
- Specify the research objective and the class of events the cases belong to.
- Formulate the standardized questions. Each must be answerable from evidence, relevant to the objective, and phrased identically for every case.
- Select cases on the independent variable or on theoretically relevant variation, with negative cases included.
- Describe the case's conditions in terms of the study's variables, not in the case's own idiom.
- Answer the questions for each case; then compare answers across cases.
Within-case analysis and process tracing
Process tracing examines the sequence of events and the evidence connecting a hypothesized cause to the outcome. Evidence is evaluated with the four tests (Van Evera; elaborated by Bennett and by Collier):
| Test | Passing is... | Failing is... | Example |
|---|---|---|---|
| Straw-in-the-wind | Weak support | Weak doubt | A memo mentions the factor |
| Hoop | Necessary, not sufficient (keeps hypothesis alive) | Eliminates hypothesis | The decision-maker had access to the information before deciding |
| Smoking gun | Sufficient, not necessary (strongly confirms) | Does not eliminate | Minutes show the factor cited as the reason |
| Doubly decisive | Confirms and eliminates rivals | Eliminates | A record that could exist only if this explanation and not the rivals holds |
Also within-case: a timeline of events with sources; a check that cause preceded effect; counterfactual reasoning stated explicitly ("had X been absent, Y would have..." with the evidence for that claim).
Cross-case analysis
- Case-ordered matrices (Miles and Huberman). Rows are cases ordered on the outcome; columns are conditions or standardized questions. Patterns are read down columns.
- Pattern matching (Yin). Predict the pattern of answers each theory implies across cases; compare with the observed pattern.
- Replication. Cases predicted to be similar should be; cases predicted to differ should differ for the predicted reasons. A case that breaks the prediction is the most informative one and gets its own section.
- Iteration between data and theory (Eisenhardt). Sharpen constructs as cases are analysed, but record when a construct changed and reanalyse earlier cases with the new definition.
The case matrix
The central artefact. Build it early as an empty table and fill it as evidence arrives.
| Case 1 | Case 2 | Case 3 | Case 4 | |
|---|---|---|---|---|
| Outcome (defined as...) | ||||
| Q1: Was there a mandated deadline? (source) | ||||
| Q2: Who owned delivery? (source) | ||||
| Q3: Budget relative to plan? (source) | ||||
| Rival explanation A evidence | ||||
| Rival explanation B evidence | ||||
| Evidence quality (H/M/L) |
Every cell holds an answer plus a source pointer plus a confidence level. Empty cells are visible gaps in the fieldwork, not things to fill from memory in the write-up.
The evidence record
Every piece of evidence that bears on an explanation gets a row, so the write-up can cite tests rather than impressions.
| ID | Case | Evidence (source, date) | Bears on | Test | Result | Rivals affected |
|---|---|---|---|---|---|---|
| S-07 | South | Steering minutes, March, item 4: replacement chosen over wrapper | Replace delays launch | Hoop (decision precedes delay) | Pass | Weakens vendor rival |
| S-12 | South | Status reports Q2-Q4 attribute integration slip to replacement | Replace delays launch | Smoking gun | Pass | Weakens funding rival |
Triangulate: an interview claim becomes evidence when a document, a second independent interviewee or a contemporaneous record corroborates it; record the corroboration in the row. Uncorroborated claims stay at straw-in-the-wind strength regardless of the seniority of the person who made them. Absence counts too: if an explanation predicts a trace (a budget request, a memo, a meeting) and a documented search finds none, log the search as a failed hoop test for that explanation.
Procedure
- Write the objective, the outcome definition, the candidate explanations and their rivals.
- Draft the standardized questions; pilot them on one case and revise.
- Select cases with a stated logic (literal and theoretical replication; negative cases present).
- Write a case study protocol: data sources per question, interview guide, document list, chain of evidence rules.
- Collect evidence case by case; keep a case database (documents, interview notes, timeline) separate from the write-up.
- Within-case: build the timeline, apply process-tracing tests to each explanation, write the case narrative.
- Cross-case: fill the matrix, order cases on the outcome, read patterns, test against predicted patterns.
- Return to deviant cases; revise constructs; reanalyse earlier cases if a construct changed.
- Write findings with the evidence tests that support each claim and the scope conditions.
Worked Example: Digital Service Roll-outs in Four Regions
Objective: explain why two regions launched on time and two did not. Outcome: launched within three months of the target date.
| North (on time) | East (on time) | South (late) | West (late) | |
|---|---|---|---|---|
| Q1 Statutory deadline? | Yes (Act s.12) | Yes | No | No |
| Q2 Single accountable owner? | Yes, director-level (org chart, interview N3) | Yes, director-level | Yes, manager-level | No, committee |
| Q3 Legacy system replaced or wrapped? | Wrapped | Wrapped | Replaced | Replaced |
| Q4 Budget at launch vs plan | 104% | 98% | 141% | 127% |
| Rival A: funding level | Similar across cases (finance returns) | |||
| Rival B: vendor identity | Same vendor in North and South | |||
| Evidence quality | H | H | M | M |
Cross-case reading: deadline and wrap-rather-than-replace co-vary with the outcome; ownership level is mixed; funding and vendor do not discriminate (rival A and B weakened). Within-case, South provides the critical test: same vendor as North, same funding, but a decision to replace the legacy system (minutes, March) followed by a nine-month integration delay documented in three status reports. That sequence is a hoop test the "replace" explanation passes and a smoking gun for it in South. The finding: "In these four cases, regions that wrapped the legacy system launched on time regardless of vendor; both late regions chose replacement, and in South the replacement decision is traceable to the delay. A statutory deadline co-occurred with on-time launch but cannot be separated from the wrap decision in this set." That last sentence is the scope condition.
Writing Up
Structure that keeps evidence and inference apart:
- Objective, outcome definition, explanations and rivals, case selection logic.
- Protocol summary: questions, sources, evidence tests.
- One section per case: timeline, answers to the standardized questions, process-tracing results, evidence quality.
- Cross-case: the matrix, the patterns, the deviant case, the rivals and how they fared.
- Findings as claims, each with the strongest evidence test it passed and the cases it rests on.
- Scope conditions, limitations, what would falsify the finding, and what a next study should select.
Write each case so a reader can disagree with the cross-case inference while accepting the case facts. State each finding in a fixed format that exposes its support:
Finding 2: Choosing to replace rather than wrap the legacy system delayed launch.
Cases resting on: South (smoking gun S-12, hoop S-07); West (hoop W-04, straw-in-the-wind W-09)
Rivals addressed: vendor identity (eliminated: same vendor in North and South); funding (weakened, not eliminated)
Scope: regions with a legacy system of comparable size; four cases; no case combining replacement with on-time launch exists in the set
Would be falsified by: a comparable region that replaced and launched on time
Checklist
- Outcome defined so that negative cases exist and are included.
- Standardized questions identical across cases; piloted.
- Rival explanations named before fieldwork; evidence to distinguish them collected.
- Case database separate from the report; chain of evidence from claim to source.
- Timeline per case; cause precedes effect checked.
- Process-tracing test named for each key piece of evidence.
- Case matrix complete, with source and confidence per cell.
- Deviant cases analysed, not explained away.
- Construct changes logged and earlier cases reanalysed.
- Findings state scope conditions and falsifiers.
Common Mistakes
- Writing each case as a story and then "finding" the pattern the stories were written to show.
- Selecting only successful cases and reporting what they share.
- Treating straw-in-the-wind evidence (a mention, a coincidence of timing) as if it were a smoking gun.
- Redefining the outcome mid-study so a difficult case fits.
- Interviewing only the people who ran the programme.
- Presenting four cases as if they estimated a frequency ("50 percent of regions...").
- Leaving the matrix cells as adjectives ("strong leadership") rather than answers to defined questions with sources.
Limits
Comparative case studies establish that a mechanism operated in these cases under these conditions; they do not estimate how often it operates elsewhere. Findings are only as strong as the weakest evidence test the key claim passed, and with two to six cases, one misclassified case can reverse a cross-case pattern. Cases that share a common source (same consultant, same vendor, same policy network) are not independent. When the number of cases grows past a dozen and the questions can be coded consistently, move to QCA or to statistical analysis and keep the case studies for mechanism.
Install this skill directly: skilldb add comparative-analysis-skills
Related Skills
The Comparative Method
Activate this skill when the user is designing or critiquing a comparison of a small number of cases (countries, regions, organizations, programmes, historical episodes) to explain an outcome rather than merely rank options. Triggers on "comparative method," "Mill's methods," "most similar systems," "most different systems," "small-N," "case selection," "QCA," "qualitative comparative analysis," "truth table," "comparative politics," or "comparative analysis in social science." Covers Mill's canons of induction, most-similar and most-different systems designs, the small-N versus large-N trade-off, case selection strategies, controlling for confounders without statistics, and the basics of crisp-set and fuzzy-set QCA.
Competitor and Product Comparison
Activate this skill when the user is comparing products, services or competitors for a buying decision, a market analysis, a positioning exercise or a public comparison page. Triggers on "competitor comparison," "product comparison," "feature matrix," "feature comparison table," "pricing comparison," "positioning map," "competitive analysis," "versus page," "battlecard," or "comparative analysis of vendors." Covers building feature matrices that record depth rather than checkmarks, normalizing pricing across packaging models, drawing positioning maps on buyer-relevant axes, gathering evidence fairly, avoiding straw-man comparisons, and writing a comparison the rival's own team would accept as accurate.
Cost-Benefit and Total Cost of Ownership Comparison
Activate this skill when the user is comparing options on money over time: build versus buy, on-premises versus subscription, two capital projects, or a policy against its alternatives. Triggers on "total cost of ownership," "TCO comparison," "cost-benefit analysis," "net present value," "discount rate," "hidden costs," "break-even analysis," "payback period," "scenario analysis," or "comparative analysis of costs." Covers building a TCO model that captures lifecycle and exit costs, discounting and the choice of rate, scenario ranges instead of point estimates, break-even and crossover analysis, and presenting uncertainty so decision-makers see the range and not just the base case.
Side-by-Side Tables and Visuals
Activate this skill when the user needs to present a comparison of options, groups, periods or conditions in a table or chart and wants the layout to reveal the differences rather than bury them. Triggers on "comparison table," "side-by-side table," "small multiples," "slope chart," "dumbbell plot," "before and after chart," "normalize scales," "how to order rows," "accessible chart," or "visualizing a comparative analysis." Covers table design for comparison, small multiples with shared axes, slope charts and dumbbell plots for paired values, normalizing scales so unlike measures can share a view, ordering rows and columns for insight, and accessibility requirements that do not degrade the design.
Weighted Scoring Matrices
Activate this skill when the user is building or reviewing a scoring model that ranks options against weighted criteria, such as a vendor selection matrix, a prioritization scorecard or an evaluation rubric. Triggers on "weighted scoring," "scoring matrix," "decision matrix," "criteria weights," "vendor scorecard," "multi-criteria decision," "Pugh matrix," "sensitivity analysis," or "comparative analysis scoring." Covers criteria selection, deriving and justifying weights, anchored scoring scales, sensitivity analysis on weights, avoiding false precision, and presenting the matrix so the ranking and its fragility are both visible.
Benchmark Comparison and Reporting
Activate this skill when the user is comparing measured performance results (latency, throughput, accuracy, cost per unit, energy) across systems, versions, models or configurations and needs to report the comparison without misleading anyone. Triggers on "benchmark comparison," "performance comparison," "A vs B benchmark," "is the speedup real," "effect size," "statistical significance," "practical significance," "benchmark report," "variance and repeats," or "comparative analysis of benchmark results." Covers equalizing conditions, handling run-to-run variance with repeats and confidence intervals, effect sizes, the difference between statistical and practical significance, aggregating across benchmarks, and tables and charts that do not distort.