Causal inference is the subfield of econometrics concerned with using data to estimate the effect of one variable on another—that is, to answer questions about what would happen if a particular cause were changed. While much of statistics focuses on describing associations or making predictions, causal inference aims to uncover the underlying mechanisms that generate those associations. The central challenge is that we almost never observe the same unit of observation both with and without a treatment or intervention; we must infer the counterfactual outcome from data on other units, under assumptions that are often strong and difficult to verify.
The core of modern causal inference in econometrics is the potential outcomes framework, sometimes called the Rubin Causal Model after its development by Donald Rubin, though it has roots in earlier work by Jerzy Neyman and others. For each unit (a person, firm, region, etc.), we imagine two potential outcomes: the outcome if the unit receives a treatment, and the outcome if it does not. The causal effect for that unit is the difference between these two potential outcomes. The fundamental problem of causal inference is that we observe only one of these potential outcomes for each unit—the one corresponding to the treatment actually received. The other is missing.
The goal of causal inference is to estimate some summary of the individual causal effects, most often the average treatment effect (ATE) across a population. Because we cannot observe the counterfactual for any given unit, we must rely on comparisons between treated and untreated units, or between the same units at different times. The validity of these comparisons depends on assumptions about how units came to receive the treatment. The most important assumption is that of unconfoundedness (or ignorability): that, conditional on observed covariates, treatment assignment is independent of the potential outcomes. This assumption cannot be tested from data alone; it must be justified by knowledge of the assignment mechanism.
Before the potential outcomes framework became dominant, econometric causal inference was largely conducted within the structural equation modeling (SEM) tradition. This approach, associated with the Cowles Commission in the mid-20th century, treats causal relationships as systems of linear equations. A causal effect is a parameter in a structural equation—a coefficient that describes how a change in one variable directly changes another, holding all other variables in the system fixed. The key idea is that the causal structure is specified by the researcher through a set of equations and assumptions about which variables are exogenous (determined outside the system) and which are endogenous (determined within it).
The structural equation approach differs from the potential outcomes framework in several ways. It is more explicit about the functional form of relationships (typically linear) and about the role of theory in specifying the causal model. It also naturally handles systems with multiple outcomes and feedback loops. However, it has been criticized for requiring strong, often untestable assumptions about the form of the equations and the exclusion of variables. The potential outcomes framework, by contrast, focuses on the assignment mechanism and can accommodate nonparametric or semiparametric estimation, making fewer functional form assumptions. Despite these differences, the two traditions are not incompatible; many modern methods combine insights from both.
A major shift in applied econometric causal inference began in the 1990s, driven by the recognition that credible causal estimates often require a source of exogenous variation in the treatment—variation that is "as good as randomly assigned" with respect to the potential outcomes. This movement, sometimes called the "credibility revolution," emphasized research designs that exploit natural experiments: events or policies that create variation in treatment that is plausibly unrelated to other determinants of the outcome.
The most prominent of these designs are:
These designs are not mutually exclusive; they can be combined or used to check the robustness of results. Their shared emphasis is on transparent, defensible assumptions rather than on complex statistical modeling. The credibility revolution has been enormously influential, shifting the focus of applied work from estimation within a fully specified structural model to the careful construction of a credible counterfactual.
A separate but related tradition, developed primarily in computer science and statistics by Judea Pearl and others, uses directed acyclic graphs (DAGs) to represent causal relationships. In a DAG, nodes represent variables, and arrows represent direct causal effects. The graph encodes a set of conditional independence relationships that can be read off using graphical criteria (such as d-separation). This framework provides a systematic way to identify which causal effects can be estimated from observational data, and what variables must be controlled for (or must not be controlled for) to avoid bias.
The DAG approach has been adopted by many econometricians, particularly for clarifying the assumptions underlying identification strategies. For example, it can show why controlling for a common effect of treatment and outcome (a "collider") can induce bias, or why an instrumental variable must satisfy the exclusion restriction. The graphical framework is complementary to the potential outcomes framework; many results can be translated between the two languages. However, the DAG approach is more explicit about the structure of confounding and selection bias, and it provides algorithms for identifying causal effects in complex settings.
Contemporary causal inference in econometrics is characterized by a productive tension between the design-based approach of the credibility revolution and the model-based approach of the structural tradition. Many researchers now see these as complementary rather than competing. The design-based approach excels at estimating the average effect of a well-defined intervention in a specific setting, but it may not generalize to other populations or policies. The structural approach can provide estimates of deeper parameters (such as elasticities or discount factors) that are more portable across contexts, but it requires stronger assumptions.
Several active areas of research reflect this synthesis. One is the development of methods for combining multiple quasi-experimental designs, or for using machine learning to estimate treatment effects that vary across individuals (heterogeneous treatment effects). Another is the extension of causal inference to settings with multiple treatments, dynamic treatments, or interference between units (where one unit's treatment affects another's outcome). A third is the integration of causal inference with economic theory, using models of individual behavior to guide the choice of instruments or the interpretation of reduced-form estimates.
There is also ongoing debate about the role of formal testing of assumptions. Some researchers advocate for placebo tests, sensitivity analyses, and other checks that can reveal violations of identifying assumptions. Others argue that the most important assumptions (such as unconfoundedness or the exclusion restriction) are fundamentally untestable and must be defended on substantive grounds. The field has not settled this question, but there is broad agreement that transparency about assumptions is essential.
A further area of contention concerns the use of "big data" and machine learning. While these tools can improve the precision of estimates and allow for more flexible modeling of covariates, they also raise new challenges. For example, using machine learning to select instruments or to choose a functional form can lead to overfitting and invalid inference unless proper sample-splitting or regularization techniques are used. The development of methods that combine the flexibility of machine learning with the rigor of causal inference is an active frontier.
Finally, the field continues to grapple with the limits of causal inference from observational data. Even the most credible quasi-experimental design rests on assumptions that may be violated in practice. Replication across different settings, designs, and data sources is increasingly recognized as essential for building reliable knowledge. Causal inference in econometrics is not a set of recipes that guarantee truth, but a disciplined way of reasoning about what can be learned from data and what cannot.