Impact evaluation is the systematic attempt to determine the causal effect of a deliberate intervention—a program, policy, or project—on outcomes of interest. It is a subfield of development economics, but its methods are shared with public health, education, political science, and other social sciences. The defining question is not simply "Did things change?" but "What would have happened in the absence of the intervention?" This counterfactual question is what separates impact evaluation from monitoring, which tracks implementation and outputs, and from descriptive studies, which document correlations between activities and conditions.
The central intellectual challenge is that the counterfactual—the state of the world had the intervention not occurred—is never directly observable. A researcher can observe a village after a new cash-transfer program, but cannot observe that same village had the program never existed. Any comparison between those who received the intervention and those who did not is vulnerable to selection bias: the treated and untreated groups likely differ in ways that affect the outcome, independent of the intervention itself.
For example, farmers who adopt a new seed variety may be more entrepreneurial or better connected to extension services than non-adopters. Comparing their yields to non-adopters conflates the seed's effect with these pre-existing differences. Similarly, comparing outcomes before and after an intervention confounds the intervention with any other changes occurring simultaneously—a price shock, a drought, a new road. Impact evaluation is fundamentally a battle against these confounds, and its history is largely the development of increasingly credible strategies for constructing a valid comparison.
Early development evaluation, from the post-war period through the 1970s, relied heavily on qualitative case studies, expert judgment, and simple before-and-after comparisons. These approaches could describe what happened but could not convincingly attribute changes to the intervention. The rise of rigorous quantitative evaluation in economics began in earnest with the "credibility revolution" of the 1990s, which pushed researchers to move beyond observational methods that required strong, often implausible, assumptions.
The watershed moment was the introduction of randomized controlled trials (RCTs) into development economics. While randomization had long been used in agriculture and medicine, its systematic application to social programs in low-income countries—pioneered by researchers working on education, health, and poverty in the 1990s and 2000s—transformed the field. RCTs offered a transparent, mechanical solution to the selection problem: if assignment to treatment is truly random, the treatment and control groups are statistically identical on average, and any difference in outcomes can be attributed to the intervention.
This shift was not merely technical. It changed the kinds of questions asked, the relationship between researchers and implementing agencies, and the standards of evidence for what counted as "what works." The rise of RCTs also generated intense debate about external validity (whether results from one context generalize), the ethics of withholding interventions from control groups, and whether the focus on average treatment effects crowded out questions about mechanisms, heterogeneity, and long-term institutional change.
The field today is organized less around competing schools of thought than around a shared toolkit of identification strategies, each with distinct assumptions and trade-offs. These are best understood as a spectrum from the most credible to the most assumption-dependent.
Randomized Controlled Trials are the benchmark. In an RCT, the researcher (or implementing agency) uses a lottery to assign units—individuals, villages, schools—to treatment or control. The key assumption is that randomization balances all observed and unobserved characteristics across groups. The estimate of impact is simply the difference in mean outcomes between groups. RCTs are powerful because their internal validity is high and their logic is transparent. Their limits are practical: they are expensive, slow, and often impossible for large-scale or national policies. They also answer a narrow question—the effect of a specific program as implemented in a specific context—and their results may not travel well. Ethical constraints also bind: one cannot randomize a policy believed to be beneficial without justification, nor withhold a universally available service.
Quasi-experimental methods attempt to approximate randomization using observational data. These are used when an RCT is infeasible or unethical, and they exploit natural or administrative features of the world that create plausibly exogenous variation in treatment.
These methods are not rivals in the sense of competing paradigms; they are tools chosen based on the structure of the problem. A researcher studying a national policy change might use DiD; one studying a program with an eligibility cutoff would use RD; one with a natural experiment might use IV. The field's methodological pluralism reflects the reality that no single method is universally superior.
Impact evaluation is not method for its own sake; it is driven by substantive questions about what reduces poverty, improves health, increases learning, and strengthens institutions. The field has produced a large body of evidence on specific interventions: cash transfers (conditional and unconditional), microfinance, school inputs, teacher incentives, deworming, agricultural extension, and many others.
A notable feature of the modern field is the emphasis on mechanisms—understanding why an intervention works, not just whether it does. This has led to the integration of behavioral economics, which examines how psychological factors like present bias, social norms, and limited attention shape responses to programs. For example, an impact evaluation might find that a savings program works, but a deeper question is whether it works because it provides a commitment device, because it changes social expectations, or because it simply provides information. This focus on mechanisms has blurred the line between impact evaluation and theory testing.
Another important development is the shift from average effects to heterogeneity. An intervention may help some people and hurt others; the average effect can mask this. Modern evaluations increasingly examine effects by gender, initial wealth, location, and other characteristics. This is not merely descriptive: it informs targeting and helps explain why a program that works in one setting fails in another.
The field today is mature but contested. The dominance of RCTs has produced a large, credible evidence base, but it has also generated a backlash. Critics argue that the focus on small-scale, researcher-managed trials has diverted attention from macro-level questions—structural adjustment, trade policy, institutional reform—that are not amenable to randomization. They also point to the "file drawer problem" (the tendency for null results to go unpublished), the difficulty of scaling up pilot programs, and the risk that evidence from one context is applied uncritically elsewhere.
In response, the field has evolved in several ways. There is growing interest in generalizability and external validity, including efforts to replicate studies across contexts and to use data from multiple sites to understand when and why effects vary. There is also a push toward open science practices—pre-registration of hypotheses, sharing of data and code—to increase transparency and reduce researcher degrees of freedom.
A second tension concerns the relationship between impact evaluation and policy making. The original promise was that rigorous evidence would lead to better policy. In practice, the translation is imperfect. Evidence is often ignored for political reasons, and programs that work in trials may fail when implemented at scale by governments with weaker capacity. The field has responded by studying implementation, "scaling up" processes, and the political economy of evidence use, but these remain underdeveloped areas.
A third tension is between the quantitative, experimental tradition and other research traditions. Qualitative research, participatory evaluation, and theory-based approaches (such as process tracing or contribution analysis) ask different questions—about meaning, context, and causal mechanisms in complex systems—and are sometimes dismissed by experimentalists as unscientific. In practice, many evaluations now use mixed methods, combining quantitative impact estimates with qualitative fieldwork to understand implementation and interpretation. The relationship is not one of replacement but of complementary strengths and unresolved disagreements about what counts as evidence.
Despite these debates, the core of impact evaluation is stable. It is the discipline of asking the counterfactual question rigorously, of being explicit about assumptions, and of acknowledging uncertainty. The field's contribution is not a set of findings but a way of thinking: that claims about effectiveness must be tested against the best available comparison, that correlation is not causation, and that the credibility of an estimate depends on the design that produced it. This way of thinking has spread beyond economics into global health, education, and social policy, and it is now a standard part of how governments and international organizations assess their programs.
The future of the field will likely involve better integration of methods, more attention to external validity and long-term effects, and a continued struggle to make evidence relevant to the messy realities of policy. But the fundamental question—what would have happened without this intervention?—will remain, because it is the only honest way to know whether efforts to improve the world actually do so.