Research Design

Measurement Theory

A reading guide for lecture p-06 — what a number actually measures, how to tell whether a scale is reliable, and what happens when a metric becomes a target.

Section one

Key concepts

Types of measures, instruments, validity, reliability, and the statistic that checks reliability.

Key concepts

The same poverty rate can describe two very different communities.

Harlem and Chinatown both sit in the 40–80% poverty band on the Manhattan map. In Harlem the residents are mostly US-born, with high rates of inter-generational poverty. Chinatown has many new immigrants with few financial assets but strong social capital, and their children are highly mobile.

The deck makes the point with two communities that each have a 25% poverty rate:

  • Community A: 1% of the population enters poverty each year and 1% leaves. That leaves 24% in chronic poverty.
  • Community B: 20% enter each year and 20% leave. Only 5% are chronically poor.

The poverty rate measures a stock. It says nothing about the flows in and out, and the flows are what separate a trap from a way station.

A measure is only as good as the construct behind it.

Before you can build an index you have to decide what it is an index of. The deck’s list for poverty: lack of money, lack of opportunity, limited access to healthcare, lack of education, lack of mobility, or position in a caste. It also lists “lack of character” as one popular theory. And is a college student on a fixed budget poor?

A richer poverty measure might combine several dimensions: financial capital, human capital, social capital, physical health, and public goods. If you have good parks, free libraries, and public art, do you need as much money?

There are three types of measures.

  • Direct measures count the thing itself, like the number of windshields a factory worker installs.
  • Markers or predictors are direct measures that stand in as proxies for something harder to measure.
  • Latent constructs can’t be observed at all, only inferred. Intelligence (IQ tests), depression (surveys), and health (surveys) are examples.

An instrument is the tool that turns a construct into a number.

For direct measures the instruments are microscopes, spectrometers, and scales. For latent constructs they are survey questions, observational protocols for coding data, and standardized exams. The Oxford Happiness Questionnaire is an example: agree/disagree items on a six-point scale, some of them reverse-coded (R) so that agreement means less happiness.

Validity asks whether you measure the right thing. Reliability asks whether you measure it consistently.

  • Validity: do the items measure the latent construct they claim to?
  • Reliability: how consistently do the items measure the same construct?

The deck’s four-item “good dancer” scale scores each item 0–4, for a total of 0 to 16. One item is “I am athletic.” It is plausibly related to dancing, but it measures a different construct, and the correlation structure shows it.

Cronbach’s alpha measures internal consistency.

Alpha measures how closely related a set of items are as a group. It is the standard measure of scale reliability, and it runs from 0 to 1:

α= N · c v + (N − 1) · c

where N is the number of items, c is the average inter-item covariance, and v is the average variance per item.

Reading an alpha score
αReliability
0.9 – 1.0Excellent
0.8 – 0.9Good
0.7 – 0.8Acceptable
0.6 – 0.7Questionable
below 0.6Poor / inadequate

Section two

Key relationships

The mental map. How measurement connects to the regression and design problems from earlier units.

Dropping the item that doesn’t belong can raise alpha

The two worked examples in the deck follow the same logic. Read the correlation matrix, find the items that don’t correlate with the rest, and remove them.

The formula also shows the other lever. Holding the average covariance fixed, adding items raises alpha. A long scale can look reliable even when its items are only weakly related, so a high alpha is not a substitute for looking at the correlations.

Reliability is necessary but not sufficient for validity

The three-item bro-culture scale is highly reliable. The items hang together, because people who like beer pong also like Family Guy. Whether that cluster actually measures the traits on the slide (entitlement, disregard for others, self-destructive behavior) is a validity question, and alpha can’t answer it. A scale can consistently measure the wrong thing.

The reverse doesn’t hold. An unreliable instrument can’t be valid. The Myers-Briggs reading is the cautionary case. Retake the test after five weeks and there is roughly a 50% chance you land in a different type, which is no better than a coin toss.

Unreliable measures are the measurement error from Unit 06

A low-alpha scale is a noisy measure, and you already know what noise does to a regression:

Improving an instrument’s reliability improves every estimate built on top of it.

Good instruments are built for the people who use them

The examples section shows three practical designs:

Measurement is inherently political

Tie a metric to rewards and penalties and it starts to change the behavior it was meant to record. The deck lists five challenges:

  1. What if we can’t measure what we care about?
  2. Data collection costs time and money.
  3. Perverse incentives. Under multi-tasking, what gets measured gets done at the expense of everything else. Under creaming, providers avoid serving the poorest and hardest cases.
  4. Poor counterfactuals. Most agencies lack the capacity to evaluate properly.
  5. Fatigue from reporting to multiple stakeholders.

The Humans of New York teacher makes the fourth point concrete. Forty percent of their job rating depends on test scores, which also reflect abuse at home, missed breakfasts, and late arrivals. A rating built on raw outcomes, without a counterfactual, holds a person accountable for things they don’t control. The closing slides, drawn from Soss, Fording, and Schram’s The Organization of Discipline, argue that performance management doesn’t just measure agents. It reshapes what they do.

How to read the slides

The measurement lab app is the hands-on companion. Work through it after slides 14–35.

Section three

What should be clear in my mind?

The questions to ask before you trust an outcome variable.

Questions to ask of any outcome measure

Key takeaways