In the debate prompted by my recent post, The Genetic Alibi, the loudest disagreement was not really about whether people have genes, whether human populations have histories, or whether environments matter.
Those propositions are not especially controversial.
The central disagreement concerned a more fundamental issue: the conditions required to establish the cause of an observed difference.
That is the question causal inference is meant to answer. It is not a fashionable phrase for putting equations next to common sense. It is the discipline of refusing to confuse a pattern with the mechanism that produced it.
“Correlation is not causation” is the bumper-sticker version. This is true, fundamentally, but almost comically incomplete. The real question is this: what comparison would make causation plausible?
This matters when we study school outcomes, crime, drug policy, incarceration, employment, health, policing, poverty, or group differences of any kind. It matters whenever someone looks at a difference in outcomes and says, “There it is. That right there! That is the explanation.”
In most cases, it is not.
The Missing World
Consider a few familiar claims:
People who attend college tend to earn more than people who do not.
Neighborhoods with more police may have lower crime rates.
Counties that expand addiction treatment may see fewer overdose deaths.
One group may score differently from another on a standardized test.
Each of these statements identifies a pattern, but none provides evidence of causation.
To say that college causes higher earnings, for example, we would need to compare a person’s actual earnings after attending college with that same person’s earnings had they not attended college. But we cannot observe both lives. The person gets one history, not two.
That unobserved alternative history is the counterfactual.
Causal inference is the systematic attempt to construct a credible substitute for the unobserved counterfactual. While it is not possible to observe alternative histories for individuals, cities, or institutions, researchers can seek reasonable comparisons in other persons, places, or time periods.
The challenge is that people and places do not sort themselves randomly.
Students who attend college may differ from non-attendees before enrollment. They may have different academic preparation, family resources, health, career expectations, access to credit, neighborhood conditions, or social networks. Those factors can affect both college attendance and later income.
A simple comparison of average earnings does not isolate the effect of college attendance. Instead, it conflates the effect of college with all the factors that influence the likelihood of attending college.
Economists call this selection. Other disciplines call closely related problems confounding or selection bias. The label matters less than the logic: if the groups differed before one received the alleged cause, a later difference in outcomes cannot automatically be attributed to that cause. Observational research has to grapple directly with selection and confounding rather than assume them away.
The Cleanest Comparison
The ideal solution is random assignment.
Suppose a city wants to evaluate a new reentry-support program but has funds for only half of the eligible applicants in its first year. If applicants are assigned by lottery, those selected and those not selected should, on average, have similar characteristics before the program starts—both the things researchers can measure and, in expectation, the things they cannot.
If employment, housing stability, rearrest, or substance-use outcomes differ later, the lottery gives us a strong reason to attribute at least some of that difference to access to the program.
This is why randomized controlled trials are powerful. Randomization does not make people identical. It makes the treatment and control groups comparable on average before treatment. It attacks the selection problem at its source.
But even a randomized trial is not magic.
People can drop out. Some assigned to treatment may never receive it. Staff may implement the intervention unevenly. Participants may behave differently because they know they are being studied. And a program that works in one jurisdiction may work differently elsewhere.
A randomized study can provide robust evidence regarding a specific intervention, within a defined population and context. It does not justify generalizations beyond the scope of the study.
That limitation is not a defect. It is intellectual honesty.
Most Policy Is Not a Lottery
For many social questions, randomization is infeasible, unethical, or impossible.
States do not randomly decide whether to expand Medicaid. Cities do not randomly experience housing crises. People do not randomly become poor, incarcerated, uninsured, sick, exposed to lead, or discriminated against.
That does not mean causal analysis stops. It means the work gets harder.
Researchers look for institutional rules, policy timing, administrative thresholds, or outside events that create comparisons resembling random assignment. These are commonly called quasi-experiments or natural experiments. They are not shortcuts around causal inference. They are attempts to build a credible comparison from the world as it is.
Several designs recur across economics, public policy, political science, sociology, criminology, and public health. Regression discontinuity, difference-in-differences, instrumental variables, and related designs are among the major quasi-experimental tools used when a conventional experiment is unavailable.
The names can sound forbidding. The basic intuition is not.
Finding Comparisons in the Real World
Difference-in-differences: compare changes, not snapshots
Imagine that one state expands Medicaid while a neighboring state does not. The expansion state may already have a different uninsured rate, income distribution, health-care market, or political culture. A simple post-policy comparison would tell us little.
Difference-in-differences instead asks whether the uninsured rate changed more in the expansion state than it changed in the non-expansion state.
The method’s central assumption is called parallel trends: absent the policy, the two places would have moved along similar trajectories. That assumption cannot be proven directly, because it concerns the unobserved world without the policy. But researchers can examine pre-policy trends, use multiple comparison places, test for implausible early effects, and conduct placebo analyses.
The point is not that difference-in-differences makes skepticism unnecessary. It tells us exactly where skepticism should be aimed.
Regression discontinuity: rules create useful edges
Now imagine a diversion program that accepts defendants with a risk score below 50 and excludes those at 50 or above.
A defendant scoring 49 is likely more comparable to one scoring 50 than to one scoring 12. If outcomes jump sharply at the threshold, and if nobody can precisely manipulate their score, the eligibility rule can approximate random assignment for people near the cutoff.
That is regression discontinuity.
The principal strength of this method is its ability to generate a credible local comparison. Its principal limitation is that it estimates effects only for cases near the threshold, and does not necessarily generalize to individuals far from the cutoff or to other jurisdictions.
A good causal claim always has a jurisdiction. It tells you for whom, where, and under what conditions the estimate applies.
Matching: useful, but not a time machine
Matching tries to improve an observational comparison by pairing treated and untreated people who looked similar beforehand: comparable ages, prior records, education, employment history, family circumstances, health indicators, neighborhoods, or assessed risks.
That is usually better than comparing everyone who received a program with everyone who did not.
But matching only balances what researchers observe. It cannot account for an unmeasured factor that affects both treatment participation and the outcome: motivation, an informal job lead, a supportive relative, an undocumented health condition, a relationship with a case manager, or simply the ability to navigate bureaucracy.
For this reason, controlling for numerous variables is not equivalent to establishing a causal effect. Statistical adjustment can enhance the credibility of a comparison, but it cannot compensate for unmeasured information.
Instrumental variables: an outside push
Instrumental variables use an outside event or institutional feature that changes someone’s exposure to a treatment for reasons plausibly unrelated to the outcome.
Suppose that eligible defendants are more likely to enter treatment court when randomly assigned to a judge who tends to refer people there. If judge assignment is genuinely as-good-as random, it may provide an outside nudge into treatment-court participation.
The difficult question is whether the judge affects later outcomes through anything else. Does the judge also set different bail conditions? Should he impose different sentences? Maybe she should use different supervision practices? If so, the instrument is not isolating treatment-court access.
The method is powerful when its assumptions are credible. It is fragile when you assert them rather than defend them.
Synthetic control: build the comparison you do not have
Sometimes a policy changes in one city or one state, and no single comparison jurisdiction is adequate.
Synthetic control constructs a weighted composite of untreated places to create the closest available counterfactual. If Oregon changes a drug-policy regime, for example, a researcher might combine several other states into a “synthetic Oregon” that closely matches Oregon’s prior overdose trends, demographics, and related indicators.
If actual Oregon diverges from synthetic Oregon after the policy change, that divergence can be informative.
The credibility of this method depends on the quality of the pre-policy match and the absence of contemporaneous events uniquely affecting the treated unit. Synthetic control does not generate causality automatically; rather, it is a transparent effort to approximate the counterfactual.
The Sentence That Begins With “Because”
Every causal design rests on a sentence that begins with because.
We can attribute a difference to the program because access was randomly assigned.
We can attribute a change to a policy because the treated and comparison jurisdictions had been moving in parallel beforehand.
We can attribute a discontinuity to eligibility because cases immediately on both sides of the cutoff were otherwise comparable.
We can attribute a result to treatment because the instrument changed treatment exposure but did not independently affect the outcome.
The entire argument lives or dies in that “because.”
If it is vague, if it assumes what needs to be shown, if the data contradict it, or if obvious alternatives remain unaddressed, the causal claim weakens. It does not necessarily disappear. It simply deserves less confidence.
This is why causal inference does not require researchers to eliminate every imaginable confounder. No serious social scientist can do that. The standard is not omniscience.
The standard is a credible design, clear assumptions, transparent evidence, reasonable robustness checks, and conclusions calibrated to what the evidence can actually bear.
“We cannot know with perfect certainty” does not mean “a plausible story is now a causal finding.”
Back to Group Differences
This brings us back to the dispute that prompted this note.
Suppose someone offers three propositions:
Human populations exhibit ancestry-related genetic variation.
Some traits are substantially heritable within populations.
Broad, socially defined groups differ in an observed outcome.
Each proposition may be true. None, separately or together, establishes a genetic explanation for the observed group difference.
The missing step is causal identification.
Within-group heritability describes variation among individuals in a particular population and environment. It does not tell us what proportion of a difference between group averages is genetic. A 2024 Proceedings of the National Academy of Sciences article states the point in unusually direct terms: aggregate within-group genetic and phenotypic information cannot, by itself, separate group differences into genetic and environmental components.
This does not establish that all observed group differences are environmental in origin. Rather, it demonstrates that within-group heritability is insufficient to resolve the question.
Similarly, enumerating potential environmental mechanisms does not, by itself, constitute a causal explanation. Factors such as school quality, nutrition, health, stress, discrimination, labor markets, family circumstances, neighborhood conditions, and institutional exposure may all be causally relevant, but merely listing them does not provide a causal decomposition. A rigorous argument must distinguish between what is plausible and what is established, regardless of the direction of the claim.
Genetic association research also has its own identification problems. Population stratification can create associations when ancestry is correlated with social and environmental structure; subtle stratification can accumulate in polygenic scores and reproduce environmental patterns. Indirect genetic effects further complicate interpretation, because parental genotypes may influence the environment that children experience.
This is not a demand to declare genetic influences impossible. It is a demand to distinguish a possibility from a demonstrated explanation.
A Reader’s Replication Checklist
Before accepting a headline, a paper, or a confident social-media claim as causal, ask:
What is the alleged cause? “Policy,” “race,” “culture,” and “the environment” are often too vague to test.
What is the outcome? Is it measured consistently, and does it capture what the argument says it captures?
Compared with what? What stands in for the missing counterfactual?
Why are the groups comparable? Random assignment, policy timing, a cutoff, an external shock, matching, or something else?
What assumption does the design need? Parallel trends? No manipulation? No unobserved confounding? An exclusion restriction?
Could cause and effect run both ways? Do more police lower crime, or do police deployments rise where crime is expected to increase?
What did the researchers do to stress-test the result? Look for placebo tests, pre-trend checks, alternative specifications, sensitivity analyses, and replication.
How large is the effect? Statistical significance is not the same as substantive importance.
For whom does it apply? A result at an eligibility cutoff or in one city may not generalize everywhere.
What remains unknown? Good research names its limits instead of hiding them.
Uncertainty Is Sometimes the Result
Causal inference is not an instruction to remain forever agnostic. Strong designs can reveal important facts about interventions, institutions, harms, incentives, and opportunities. They can tell policymakers what worked, for whom, and at what cost.
But observed differences are not self-interpreting.
A correlation provides a basis for further investigation, but it does not justify asserting a specific causal mechanism. The higher the stakes of the proposed explanation, particularly in matters such as criminality, intelligence, poverty, race, or human worth, the greater the responsibility to provide rigorous evidence.
When the evidence cannot supply a credible comparison, uncertainty is not evasion.
It is the result.


