Causal inference is the branch of epidemiology concerned with determining whether a specific exposure, intervention, or treatment genuinely produces a change in a health outcome, and if so, how large that change is. Its central question is deceptively simple: What would have happened to the same people if they had been exposed differently? Because that counterfactual scenario can never be directly observed, the field's intellectual work consists of developing frameworks, study designs, and analytical methods that make reliable answers to this question possible despite the absence of direct evidence.
The stakes are high. Epidemiology informs public health policy, clinical guidelines, and individual medical decisions. A claim that a drug reduces mortality, that air pollution increases asthma attacks, or that a screening program saves lives is a causal claim. If the underlying inference is flawed, the resulting policies can be ineffective or harmful. Causal inference in epidemiology is therefore not merely a technical specialty but the methodological core of evidence-based medicine and public health.
The deep difficulty of causal inference was recognized long before the field acquired its modern name. In the nineteenth century, physicians and statisticians grappled with whether specific exposures—contaminated water, poor housing, occupational dust—caused disease. The British physician John Snow's investigation of the 1854 London cholera outbreak is often cited as an early example of causal reasoning: by comparing cholera rates among customers of different water companies, he provided evidence that contaminated water, not miasma, transmitted the disease. Yet Snow's work was observational and relied on a natural experiment—the fact that two water companies drew from different sources and served overlapping populations. This pattern, exploiting a naturally occurring comparison to approximate a controlled experiment, remains a core strategy in the field.
The formal statistical foundations emerged in the early twentieth century. The English statistician Ronald Fisher developed randomized experiments and analysis of variance, establishing that random assignment could balance unknown confounders between treatment groups. The Austrian-born biologist and statistician Jerzy Neyman, working in the 1920s, introduced a formal notation for potential outcomes—the idea that each unit has a set of possible outcomes corresponding to different treatments, only one of which is observed. This potential outcomes framework, later developed extensively by Donald Rubin and others, became one of the two dominant formal languages of modern causal inference.
A second major intellectual stream came from epidemiology itself. In 1965, the British epidemiologist Austin Bradford Hill published a set of considerations for judging whether an observed association is causal. Hill's criteria—including strength of association, consistency across studies, temporality, biological gradient, plausibility, and coherence—were never intended as a rigid checklist, but they became widely taught as a practical guide. They represent a tradition of causal reasoning based on accumulating evidence across multiple studies and biological knowledge, rather than on a single formal model.
The potential outcomes framework, also called the Rubin causal model, provides the most widely used formal language for causal questions in epidemiology. For each individual and each possible exposure level, there exists a potential outcome: the health status that person would experience if exposed, and the health status they would experience if not exposed. The causal effect for that individual is the difference between these two potential outcomes. The fundamental problem is that only one of these outcomes is ever observed; the other is counterfactual.
Because individual-level effects are unobservable, the framework shifts attention to average effects. The average treatment effect is the mean difference between potential outcomes across a population. The central challenge is that the observed difference in outcomes between exposed and unexposed groups is not generally equal to the average treatment effect, because the groups may differ systematically in ways that affect the outcome. This imbalance is called confounding.
The framework makes explicit the conditions under which an observed association equals a causal effect. The key condition is exchangeability: conditional on measured covariates, the potential outcomes must be independent of actual exposure. In a randomized trial, random assignment guarantees this condition on average. In observational studies, the condition must be assumed, and the analyst's task is to make it plausible by measuring and adjusting for confounders—variables that influence both exposure and outcome.
This framework has several practical consequences. It clarifies that adjustment for confounders is not merely a statistical procedure but an attempt to reconstruct the conditions of a randomized experiment. It motivates methods such as propensity score matching, in which individuals with different exposures but similar probabilities of exposure are compared, and inverse probability weighting, in which the sample is reweighted to create a pseudo-population in which exposure is independent of measured covariates. It also makes explicit the problem of unmeasured confounding: if an important confounder is not measured, no adjustment method can fully remove bias.
A second major formal language, developed primarily in the 1980s and 1990s by computer scientist Judea Pearl and epidemiologists such as James Robins and Sander Greenland, uses directed acyclic graphs (DAGs) to represent causal assumptions. A DAG is a diagram in which nodes represent variables and arrows represent direct causal effects, with the constraint that no path loops back to its starting point. The graph encodes the analyst's beliefs about the causal structure of the system under study.
DAGs provide a systematic way to identify which variables must be controlled to estimate a causal effect and which must not be. The key concept is the backdoor path: a path from exposure to outcome that begins with an arrow pointing into the exposure. Such paths create non-causal associations. Confounders are variables that lie on backdoor paths; adjusting for them blocks those paths. But the graph also reveals subtleties that intuition alone often misses. A variable that is affected by exposure and also influences the outcome—a mediator—should generally not be adjusted for, because doing so blocks part of the causal effect. A variable that is a common effect of two other variables—a collider—creates a spurious association between those variables when conditioned upon. Adjusting for a collider can induce bias rather than remove it.
The DAG framework also gave rise to the concept of the front-door path and to methods for estimating effects when some confounders are unmeasured, provided the graph has a particular structure. More broadly, DAGs made explicit the distinction between association and causation and provided a calculus for moving between them. They are now standard tools in epidemiologic teaching and practice, used primarily for clarifying assumptions and guiding variable selection rather than for estimation itself.
A major advance in the 1980s and 1990s addressed a class of problems that neither simple adjustment nor standard regression could handle: exposures that change over time and are themselves influenced by earlier health states. Consider a study of the effect of a drug on mortality in a chronic disease. Patients who become sicker may be more likely to start the drug, and the drug may affect subsequent disease progression. Standard adjustment for time-varying confounders—variables like disease severity that affect both later treatment and later outcome—can introduce bias, because those confounders are themselves affected by earlier treatment.
James Robins developed a class of methods, including marginal structural models and g-estimation, to address this problem. Marginal structural models estimate the causal effect of a treatment regime by weighting observations according to the inverse probability of receiving the treatment actually received, given past treatment and confounders. This creates a pseudo-population in which treatment is independent of time-varying confounders, allowing a direct estimate of the causal effect. G-estimation, a related approach, models the effect of treatment on the outcome while accounting for the fact that treatment decisions depend on time-varying covariates.
These methods are now standard for analyzing longitudinal data in epidemiology, particularly in studies of chronic diseases, HIV, and other conditions where treatment is dynamic. They represent a significant departure from earlier approaches because they require the analyst to specify a model for the treatment process, not just for the outcome, and they make the assumptions about sequential exchangeability explicit.
A distinct tradition within causal inference seeks to exploit natural experiments—situations in which exposure is assigned by forces outside the researcher's control in a way that mimics randomization. The classic example is the instrumental variable: a variable that affects the exposure but has no direct effect on the outcome except through the exposure. If such a variable exists, it can be used to estimate the causal effect of the exposure even in the presence of unmeasured confounding.
In epidemiology, instrumental variables have been used in various forms. Genetic variants that influence a risk factor but are not otherwise associated with the outcome have been used in Mendelian randomization studies, an approach that has become prominent in recent decades. Policy changes, such as the introduction of a new screening program in some regions but not others, can also serve as natural experiments. The key advantage of instrumental variable methods is that they can address unmeasured confounding, which is the primary limitation of standard adjustment methods. The key limitation is that valid instruments are difficult to find: the assumption that the instrument affects the outcome only through the exposure is often questionable and cannot be fully tested.
A recent development that has reshaped the field is the target trial framework, articulated by Miguel Hernán and James Robins in the 2010s. The idea is straightforward: to answer a causal question using observational data, the analyst should first specify the randomized experiment that would ideally be conducted to answer the question—the target trial—including the eligibility criteria, treatment assignment mechanism, and outcome definition. The observational analysis is then designed to emulate that target trial as closely as possible.
This framework has several consequences. It forces researchers to be explicit about the causal question, including the treatment regime and the time zero at which follow-up begins. It highlights common errors in observational analyses, such as adjusting for variables measured after treatment initiation or failing to align the start of follow-up with the start of treatment. It also provides a structured way to handle issues like treatment switching and non-adherence. The target trial framework has been widely adopted, particularly in pharmacoepidemiology, where it has become a standard approach for designing observational studies of drug effects.
Contemporary causal inference in epidemiology is characterized by a convergence of these traditions. The potential outcomes framework provides the formal language for defining effects and stating assumptions. DAGs provide a graphical tool for representing and communicating those assumptions. Marginal structural models and related methods address time-varying exposures. The target trial framework provides an overarching design principle. Instrumental variable methods, including Mendelian randomization, offer a complementary approach for addressing unmeasured confounding.
The field is also marked by ongoing debates and open questions. One concerns the role of machine learning in causal inference. Modern methods can estimate propensity scores and outcome models using flexible algorithms that do not require the analyst to specify functional forms, but they do not solve the fundamental problem of unmeasured confounding. Another debate concerns the interpretation of effects in the presence of interference—when one person's exposure affects another person's outcome, as in infectious disease transmission or vaccine herd effects. Standard methods assume no interference, and extending causal inference to settings where this assumption fails is an active area of research.
A further tension concerns the relationship between causal inference and traditional epidemiologic practice. Some epidemiologists argue that the formal frameworks have become too dominant, and that careful descriptive epidemiology, biological plausibility, and triangulation of evidence across multiple study designs remain essential. Others counter that formal frameworks are necessary to avoid the logical errors that informal reasoning permits. This is not a dispute about whether causal inference matters—both sides agree it does—but about how much weight to place on formal methods versus substantive knowledge and diverse evidence.
The field's practical influence is substantial. Regulatory agencies, clinical guideline developers, and public health bodies increasingly expect causal claims to be supported by analyses that explicitly address confounding, selection bias, and measurement error. The target trial framework has become a standard for observational drug safety studies. Methods for time-varying exposures are routinely applied in chronic disease epidemiology. At the same time, the field's central message—that causal claims require either randomization or strong, often untestable assumptions—remains a sobering reminder of the limits of observational data. Causal inference in epidemiology is thus best understood not as a set of techniques that guarantee truth, but as a disciplined way of making assumptions explicit, testing their consequences, and communicating the degree of certainty that evidence can support.