Research Design
A reading guide for lecture p-06 — what a number actually measures, how to tell whether a scale is reliable, and what happens when a metric becomes a target.
Section one
Types of measures, instruments, validity, reliability, and the statistic that checks reliability.
Key concepts
The same poverty rate can describe two very different communities.
Harlem and Chinatown both sit in the 40–80% poverty band on the Manhattan map. In Harlem the residents are mostly US-born, with high rates of inter-generational poverty. Chinatown has many new immigrants with few financial assets but strong social capital, and their children are highly mobile.
The deck makes the point with two communities that each have a 25% poverty rate:
The poverty rate measures a stock. It says nothing about the flows in and out, and the flows are what separate a trap from a way station.
A measure is only as good as the construct behind it.
Before you can build an index you have to decide what it is an index of. The deck’s list for poverty: lack of money, lack of opportunity, limited access to healthcare, lack of education, lack of mobility, or position in a caste. It also lists “lack of character” as one popular theory. And is a college student on a fixed budget poor?
A richer poverty measure might combine several dimensions: financial capital, human capital, social capital, physical health, and public goods. If you have good parks, free libraries, and public art, do you need as much money?
There are three types of measures.
An instrument is the tool that turns a construct into a number.
For direct measures the instruments are microscopes, spectrometers, and scales. For latent constructs they are survey questions, observational protocols for coding data, and standardized exams. The Oxford Happiness Questionnaire is an example: agree/disagree items on a six-point scale, some of them reverse-coded (R) so that agreement means less happiness.
Validity asks whether you measure the right thing. Reliability asks whether you measure it consistently.
The deck’s four-item “good dancer” scale scores each item 0–4, for a total of 0 to 16. One item is “I am athletic.” It is plausibly related to dancing, but it measures a different construct, and the correlation structure shows it.
Cronbach’s alpha measures internal consistency.
Alpha measures how closely related a set of items are as a group. It is the standard measure of scale reliability, and it runs from 0 to 1:
where N is the number of items, is the average inter-item covariance, and is the average variance per item.
| α | Reliability |
|---|---|
| 0.9 – 1.0 | Excellent |
| 0.8 – 0.9 | Good |
| 0.7 – 0.8 | Acceptable |
| 0.6 – 0.7 | Questionable |
| below 0.6 | Poor / inadequate |
Section two
The mental map. How measurement connects to the regression and design problems from earlier units.
The two worked examples in the deck follow the same logic. Read the correlation matrix, find the items that don’t correlate with the rest, and remove them.
The formula also shows the other lever. Holding the average covariance fixed, adding items raises alpha. A long scale can look reliable even when its items are only weakly related, so a high alpha is not a substitute for looking at the correlations.
The three-item bro-culture scale is highly reliable. The items hang together, because people who like beer pong also like Family Guy. Whether that cluster actually measures the traits on the slide (entitlement, disregard for others, self-destructive behavior) is a validity question, and alpha can’t answer it. A scale can consistently measure the wrong thing.
The reverse doesn’t hold. An unreliable instrument can’t be valid. The Myers-Briggs reading is the cautionary case. Retake the test after five weeks and there is roughly a 50% chance you land in a different type, which is no better than a coin toss.
A low-alpha scale is a noisy measure, and you already know what noise does to a regression:
Improving an instrument’s reliability improves every estimate built on top of it.
The examples section shows three practical designs:
Tie a metric to rewards and penalties and it starts to change the behavior it was meant to record. The deck lists five challenges:
The Humans of New York teacher makes the fourth point concrete. Forty percent of their job rating depends on test scores, which also reflect abuse at home, missed breakfasts, and late arrivals. A rating built on raw outcomes, without a counterfactual, holds a person accountable for things they don’t control. The closing slides, drawn from Soss, Fording, and Schram’s The Organization of Discipline, argue that performance management doesn’t just measure agents. It reshapes what they do.
The measurement lab app is the hands-on companion. Work through it after slides 14–35.
Section three
The questions to ask before you trust an outcome variable.