Research Design
A reading guide for lecture p-07 — ten competing hypotheses that could explain your result instead of the program, and the standard of proof each one is held to.
Section one
The ten items, grouped, with the fix for each.
Key concepts
A Campbell Score is a structured search for competing hypotheses.
Two statements about the same result:
This is the identification problem stated in plain language. Getting a significant result establishes that the outcome changed. It does not establish that the program caused the change.
The ten items fall into four families.
Selection / omitted variables
Trends in the data
Study calibration
Contamination
Items 1 and 2 are guilty until proven innocent. The other eight are not.
This is the deck’s most important procedural rule. Selection and attrition are so common and so damaging that a rigorous evaluation must affirmatively demonstrate it has handled them before it earns any baseline of internal validity.
The remaining eight are potential concerns. You cannot simply assume they are present — you have to make a reasonable argument from the data and evidence in the study, or from sound reasoning that goes beyond speculation.
Selection into a program (#1)
If people choose whether to enroll, enrollees differ from non-enrollees. This is a source of omitted variable bias. The fix: randomization into treatment and control, or a rigorous matching process — and the randomization must be happy, which you demonstrate with a balance table and a Bonferroni-corrected decision rule (α / number of contrasts).
Non-random attrition (#2)
If leavers differ from stayers, the effect calculation is biased. The fix: compare the characteristics of those who stay against those who leave, using only measures taken before the treatment occurred.
A key qualification: if attrition is non-random but occurs equally across groups, it typically will not bias the results. That reprieve is unavailable in reflexive designs, where there is no second group to absorb it.
Maturation (#3) and secular trends (#4)
The same problem at two scales. Maturation is growth expected naturally within individuals — children’s cognitive ability improves whether or not you intervene. Secular trends are global processes outside the individual — economic or cultural change. The fix for both: a comparison group, so the trend can be differenced out.
Seasonality (#5)
Data with cycles have natural highs and lows, so a pre–post comparison across different points in the cycle produces an invalid impact claim. The fix: compare observations from the same period, or average over a full cycle.
Testing (#6)
Repeated exposure to the same questions or tasks improves performance independently of any training. The fix: change the test, use a post-test-only design, or use a control group that also takes the test. Applies to a narrow set of programs.
Regression to the mean (#7)
Any extreme observation is more likely to be followed by one closer to the mean than by one equally extreme. Quality-improvement programs aimed at low performers therefore have a built-in improvement bias regardless of whether they work. The fix: do not select the study group from the top or bottom of the distribution based on a single time period.
Measurement error (#8)
Significant measurement error in the dependent variable biases effects toward zero, making programs look less effective than they are. The fix: better measures.
Study time-frame (#9)
Too short and a real effect looks like nothing. Too long and attrition takes over. The fix: use prior research in the domain to choose the window.
Intervening events (#10)
Something happens during the study that affects one group and not the other — the treatment school burns down, prices change for a substitute good in the control area. The fix: there often isn’t one. Intervening events can be very hard to remove.
Section two
The mental map. Why two items are guilty until proven innocent and eight are not.
The deck opens with “Can Ants Count?” as the inspiration for the assignment. The experimental logic there is the logic here: you do not prove ants count by showing they walk the right distance. You prove it by systematically eliminating every other explanation for why they walked the right distance. A Campbell Score is that elimination process turned into a checklist.
Selection and attrition get the harsher standard because they are near-universal in observational work — most observational studies will be significantly affected by them. Assuming innocence there would let almost every weak study pass.
The other eight are held to the ordinary standard because they are situational. Seasonality is irrelevant to most studies; testing effects apply to a small set of programs. Asserting them without evidence would let a critic dismiss any study by reciting the list.
T2 − C2 is not always enoughThree cases, and they are worth studying closely:
C1 = T1 — groups equivalent at baseline. T2 − C2 removes the trend correctly.C1 ≠ T1, comparison below treatment — T2 − C2 does not fully remove the trend.C1 ≠ T1, comparison above treatment — T2 − C2 removes too much trend.The note attached to both failure cases is the point: diff-in-diff separates trends even when the groups are not equivalent. This is the Campbell Score view of what lecture p-03 established formally.
Two examples make it stick. The batting slump: players are sent to the batting coach only when performing badly, so performance improves afterward whether or not the coach helps. And the FiveThirtyEight piece on sham surgery: chronic pain peaks and wanes, patients seek treatment at the worst moment, so improvement afterward may be the condition’s natural course rather than the operation.
The surgery example is doing double duty — invasive procedures also produce stronger placebo effects than pills or injections — which is a good reminder that these ten items are not mutually exclusive.
And the deck’s aside is as important as the main lesson: “consumption” here is measured as wholesale volume, not consumer consumption. So the study has a serious measurement problem (#8) sitting underneath its time-frame problem (#9).
Your job is to make a strong case, using the item definitions and the evidence in the case study. For items 1 and 2, look for what the study did to rule them out — a balance table, an attrition analysis — and score down if it is absent. For items 3–10, look for positive evidence in the study that the threat is live. Both directions require argument from evidence, not assertion.
Section three
The checklist to run against a study before crediting the program.
T2 − C2 only removes the trend cleanly when the groups started equal; diff-in-diff works
even when they did not.