Health policy evaluation is the systematic assessment of the design, implementation, and effects of policies intended to influence health, health care, and the systems that deliver them. It sits at the intersection of health economics and public policy analysis, drawing on economic theory, statistical methods, and political science to answer a deceptively simple question: did this policy do what it was supposed to do, and was it worth it?
The field is defined less by a single method than by a commitment to evidence-based judgment about collective decisions that affect health. Its practitioners evaluate policies ranging from insurance expansions and drug pricing regulations to public health mandates and payment reforms. The stakes are high because health policies allocate scarce resources, redistribute risk across populations, and can mean the difference between life and death for identifiable groups.
At its core, health policy evaluation asks three families of questions. The first concerns effectiveness: did the policy change the outcomes it targeted? This requires distinguishing the policy's effect from other forces operating simultaneously—economic trends, demographic shifts, medical advances, or unrelated programs. The second concerns efficiency: did the benefits justify the costs? This is where health economics enters most directly, applying frameworks like cost-effectiveness analysis and cost-benefit analysis to compare policies against each other or against doing nothing. The third concerns distribution: who gained, who lost, and was the trade-off acceptable? Policies that improve average health may worsen inequality, and evaluation must surface these patterns rather than bury them in aggregate statistics.
These questions are complicated by the peculiar nature of health as an economic good. Health care markets feature severe information asymmetries between patients and providers, substantial uncertainty about treatment outcomes, third-party payment that insulates consumers from true costs, and ethical constraints on treating health as an ordinary commodity. Health policy evaluation must therefore grapple with the fact that the object of study is not a well-behaved market but a hybrid institution shaped by professional norms, government regulation, and social values.
A further complication is that policies are not natural experiments. They are chosen by political processes, implemented by bureaucracies, and responded to by strategic actors—hospitals, insurers, physicians, patients—who adjust their behavior in ways that may amplify or undermine the policy's intended effects. Evaluation must account for these behavioral responses, which is why the field has become methodologically sophisticated about causal inference and why its findings often challenge simple narratives about what works.
The intellectual roots of health policy evaluation lie in several mid-twentieth-century developments. Operations research, developed during World War II, brought systematic quantitative analysis to complex organizational decisions. Cost-benefit analysis, first applied to water resources and public works, provided a framework for valuing outcomes in monetary terms. And the randomized controlled trial, established in medicine and agriculture, offered a gold standard for establishing causal effects.
The modern field took shape in the 1960s and 1970s, when governments in wealthy countries expanded health programs—Medicare and Medicaid in the United States, national health services in Britain and elsewhere—and began demanding evidence about whether these programs worked. The RAND Health Insurance Experiment, conducted in the United States from the 1970s to the 1980s, became a landmark: it randomly assigned families to health insurance plans with different levels of cost-sharing and measured effects on health care use and health outcomes. Its finding that higher cost-sharing reduces health care spending without, on average, harming health remains influential, though its limits—it studied one region, one era, and a non-elderly population—are now better understood.
The 1980s and 1990s brought a methodological revolution. Economists and statisticians developed techniques for estimating causal effects from observational data: difference-in-differences, instrumental variables, regression discontinuity, and panel data methods. These tools allowed evaluators to study policies that could not be randomized—national reforms, regulatory changes, payment system overhauls—by exploiting natural variation in who was exposed and when. The Oregon Health Insurance Experiment, which used a lottery to allocate Medicaid coverage to low-income adults in 2008, revived the randomized approach for a major policy question and found that Medicaid improved financial security and self-reported health but did not significantly improve measured physical health outcomes over the study period.
The field has also expanded geographically. While much early work focused on the United States and Western Europe, health policy evaluation now spans the globe, addressing questions specific to low- and middle-income countries: how to finance primary care, whether user fees deter utilization, how to design health insurance for informal workers, and how to evaluate large donor-funded programs. This global expansion has brought attention to implementation realities—supply constraints, governance failures, and measurement challenges—that complicate the transfer of findings across settings.
Health policy evaluation is organized less by rival schools than by complementary approaches that answer different questions. However, three broad traditions can be distinguished, each with its own assumptions, methods, and characteristic strengths.
The first tradition focuses on establishing whether a policy caused the outcomes attributed to it. Its organizing problem is selection bias: people who receive a policy intervention typically differ from those who do not in ways that correlate with outcomes. The approach treats policy evaluation as an exercise in constructing a credible counterfactual—what would have happened to the treated population in the absence of the policy.
The methods of this tradition are now highly developed. Randomized trials, when feasible, provide the strongest evidence. Natural experiments exploit policy variation that is plausibly as-if random: a cutoff in eligibility, a sudden implementation date, a lottery. Difference-in-differences compares changes over time between a treated group and a comparison group, assuming parallel trends in the absence of treatment. Instrumental variables use an exogenous source of variation in treatment exposure to identify causal effects. Regression discontinuity compares outcomes for individuals just above and below an eligibility threshold.
The strength of this tradition is its rigor about causal claims. Its limitation is that it often answers narrow questions—did this specific policy, in this specific context, at this specific time, affect this specific outcome—without providing generalizable knowledge about what works. A policy that works in one setting may fail elsewhere because of differences in implementation, population, or context. The tradition has also been criticized for prioritizing internal validity over external validity, and for treating policy as a well-defined treatment when real policies are complex bundles of components that interact.
The second tradition asks whether a policy's benefits justify its costs. Its organizing problem is scarcity: resources devoted to one policy cannot be devoted to another, so choices must be made. The approach translates policy effects into a common metric to enable comparison.
Cost-effectiveness analysis measures outcomes in natural units—life years gained, cases averted, quality-adjusted life years (QALYs) gained—and reports the cost per unit of outcome. Cost-benefit analysis goes further, valuing outcomes in monetary terms using willingness-to-pay methods or revealed preferences, allowing comparison across entirely different domains. Cost-utility analysis, a variant of cost-effectiveness analysis, uses preference-weighted health outcomes like QALYs to capture both quantity and quality of life.
The strength of this tradition is its explicit framework for trade-offs. Its limitations are well documented. Valuing health outcomes requires controversial assumptions about whose preferences count and how to compare them. QALYs, for example, embed particular views about disability and age that may discriminate against certain groups. Monetary valuation of health is even more contentious, raising questions about whether willingness to pay reflects ability to pay rather than the intrinsic value of health. The tradition also tends to focus on efficiency at the expense of distribution, though newer methods like distributional cost-effectiveness analysis attempt to incorporate equity concerns.
The third tradition examines how policies actually operate in practice. Its organizing problem is that policies are not self-executing: they are interpreted by administrators, adapted by frontline workers, resisted or embraced by target populations, and shaped by existing institutional arrangements. The approach draws on qualitative methods, organizational theory, and political science to understand the gap between policy as written and policy as delivered.
This tradition includes implementation science, which studies the processes by which evidence-based interventions are adopted and sustained; policy process research, which examines how agendas are set, policies formulated, and decisions made; and comparative health systems research, which analyzes how different institutional arrangements produce different outcomes. Its methods include case studies, process tracing, interviews, document analysis, and mixed-methods designs.
The strength of this tradition is its attention to context and mechanism. It explains not just whether a policy worked but why and under what conditions. Its limitation is that its findings are often less generalizable and less amenable to quantitative summary. It has also historically been less prominent in economics departments, though health policy evaluation as practiced in schools of public health and public policy increasingly integrates qualitative and quantitative approaches.
These traditions are not rivals in the sense of competing paradigms; they are complementary tools for different questions. A complete evaluation typically requires all three. Causal inference establishes whether a policy had an effect. Economic evaluation determines whether that effect was worth its cost. Implementation analysis explains how the effect came about and whether it can be reproduced elsewhere.
In practice, however, the traditions have different statuses. Causal inference methods dominate the most prestigious journals and receive the most methodological attention. Economic evaluation has become institutionalized in health technology assessment agencies—bodies like the National Institute for Health and Care Excellence (NICE) in England—that formally use cost-effectiveness analysis to decide which treatments and technologies to fund. Implementation research is often seen as a later-stage activity, relevant after a policy has been shown to work, though this sequential view is increasingly challenged.
The relationship has also been marked by productive borrowing. Economists have incorporated insights from implementation research about behavioral responses to incentives. Implementation researchers have adopted quasi-experimental methods to study implementation strategies. And the growing field of behavioral health economics has blurred the boundaries by studying how psychological factors—present bias, loss aversion, social norms—shape responses to health policies, a topic that fits uneasily in all three traditions.
Several durable features characterize health policy evaluation today. The first is methodological pluralism combined with a hierarchy of evidence that privileges randomized and quasi-experimental designs. This hierarchy is contested but persistent, and it shapes what gets studied, what gets published, and what gets funded.
The second is the growing importance of data infrastructure. Administrative data—insurance claims, hospital records, vital statistics—have become the primary raw material for evaluation, enabling studies of entire populations rather than samples. Linked data across domains (health, employment, education, criminal justice) allow evaluators to trace policy effects into non-health outcomes. The expansion of electronic health records and the development of privacy-preserving record linkage have accelerated this trend, though data access remains uneven across countries and settings.
The third is the rise of predictive modeling and machine learning as complements to traditional evaluation. These methods are used to identify heterogeneous treatment effects—who benefits from a policy and who does not—and to improve targeting. They also raise new challenges about transparency, fairness, and the validity of predictions when policies change the environment that generated the data.
The fourth is the increasing attention to equity. Evaluations now routinely report effects by income, race, geography, and other dimensions of disadvantage. Methods for distributional analysis have advanced, and funders increasingly require equity-focused evaluation. This shift reflects both political pressure and a recognition that average effects can conceal consequential differences.
The fifth is the globalization of the field. Health policy evaluation is no longer dominated by high-income countries. Researchers in low- and middle-income countries conduct evaluations of their own policies, and international organizations like the World Health Organization and the World Bank commission and synthesize evaluations across settings. This has enriched the field with diverse questions and contexts, while also exposing the limits of generalizing findings across very different health systems.
The field also faces persistent challenges. Publication bias—the tendency to publish positive or surprising results—distorts the evidence base. The pressure to produce policy-relevant findings can conflict with the slow, careful work that credible evaluation requires. And the translation of evaluation findings into policy change remains imperfect, as political considerations, stakeholder interests, and institutional inertia often outweigh evidence.
Health policy evaluation is thus a mature but unsettled field. Its methods are powerful, its questions are consequential, and its findings regularly shape public policy. But it remains aware of its limits: the difficulty of establishing causation in complex systems, the contestability of value judgments embedded in economic analysis, and the gap between what evaluation can establish and what decision-makers need to know. Its practitioners are trained to be skeptical of simple answers, including their own.