Research Design
A short lecture note on measurement theory — why every latent construct is measured with error, how multi-item indices reduce it, and how to build a reliable index of your own.
Section one
Observed scores, true scores, and the signal-to-noise ratio called reliability.
Key concepts
Every measure of a latent construct is a true score plus error.
Measurement theory defines a quantitative measure of a latent construct as:
where M is the observed measurement (the recorded scores on an IQ test), T is the true score (each individual’s actual intelligence quotient relative to the population), and e is random error. M, T, and e are vectors of scores, one per person.
Reliability is a signal-to-noise ratio.
Reliability is the ratio of the variance of the true scores to the variance of the true scores plus measurement error:
Alpha tells us how reliably our observed measure captures the true level of the latent construct. Cronbach’s alpha is the usual way to estimate this ratio from item-level data.
Instruments combine multiple items to triangulate the construct.
Measures like weight or height only require one measurement. Instruments designed to measure latent constructs usually combine several. All of the verbal reasoning questions on an exam might be combined into a single reading comprehension score, and responses to several survey items might be combined into a single index.
Each item can be decomposed into a component X that accurately captures the latent construct and a component e that is random noise:
An index built from three items combines them:
Its reliability is measured in exactly the same way, as signal over signal plus noise: α = var(T) / var(T + e).
Summing and averaging give the same reliability.
If three appraisers value a painting using slightly different methods, adding their three appraisals would not make sense. You would average them to preserve the scale:
Dividing by a constant rescales the signal and the noise by the same factor, so the alpha of an averaged index is the same as the alpha of a summed index.
The takeaways, if the math is not crystal clear
Section two
How experts turn a fuzzy idea like neighborhood quality into a validated index.
The American Human Development Index. Measure of America’s index is modeled on the UN’s Human Development Index for nations. It measures well-being for America’s 435 congressional districts plus Washington, D.C., which shows how uneven socio-economic well-being is not just between but within many of the largest metros. It combines three dimensions:
The overall index is graded on a one-to-ten scale, with ten the highest.
Toxic and punishing neighborhoods. Manduca and Sampson (2019) ask how social and physical environments beyond concentrated poverty predict children’s long-term well-being. They focus on neighborhoods that are harsh on child development: high violence, incarceration, and lead exposure. Their measures come from four sources:
These are examples of how scholars operationalize neighborhood quality in order to understand how neighborhoods affect the people who live in them. The measurement is not always straightforward. When data are sparse or research is poorly implemented, neighborhood quality is often measured with a single proxy such as the poverty rate. Proxies like that can be overly simplistic, and as a result not very informative.
A better approach, from a measurement theory perspective, is to ask whether we can measure the personality of a neighborhood the same way psychologists measure personality types: identify the most salient dimensions, then develop reliable instruments to measure each one precisely.
Section three
Build a reliable index of community well-being in the measurement lab.
A community is healthy if we expect its residents to achieve a good quality of life and high economic stability while living there. Your task is to create a reliable index that measures it.
Cronbach’s alpha runs from 0 (all noise) to 1 (an index that measures the construct very precisely). By convention an index needs alpha above 0.6 to be considered reliable, and above 0.8 it is considered highly reliable. Items on a well-designed index produce highly correlated responses because they measure the same construct. The higher the correlations, the higher the alpha.
You can build an index where a high score means well-being, or one built from items that measure the absence of well-being. The high and low ends of a scale are easy to reverse. Either way, the goal is a reliable index.