Global health metrics and evaluation is the subfield of global health dedicated to measuring the health of populations, understanding the causes and patterns of health loss, and assessing the impact of health policies, programs, and interventions. It is the quantitative backbone of global health governance, providing the evidence base for priority-setting, resource allocation, and accountability. The field is defined less by a single method than by a shared commitment to making health states comparable across time, place, and population, and to using those comparisons to guide action.
The work of this subfield is organized around a cluster of enduring questions. First, how many people are sick, disabled, or dying, and from what causes? This requires estimating the incidence and prevalence of diseases and injuries, often in settings where vital registration systems are incomplete or absent. Second, how does health loss compare across populations and over time? Comparing a death from malaria in a child with a decade of disability from depression in an adult requires a common unit of measurement. Third, what are the risk factors—behavioral, environmental, metabolic, and social—that drive health loss, and how much of that loss is attributable to each? Fourth, are health interventions and programs actually working? This involves evaluating whether a vaccination campaign, a sanitation program, or a new treatment protocol produced its intended effect, and at what cost.
The stakes are high because measurement shapes action. A disease that is poorly measured may be invisible to policymakers; a risk factor that is underestimated may receive insufficient funding; an intervention that is not rigorously evaluated may continue despite being ineffective or harmful. Conversely, the metrics themselves can become contested, because they determine which populations and conditions receive attention and resources. The field therefore operates at the intersection of science, policy, and ethics, and its practitioners are often called upon to defend the assumptions embedded in their numbers.
The modern subfield emerged from several converging traditions. Nineteenth-century vital statistics, pioneered in European states, established the practice of registering births, deaths, and causes of death and using those records to compute mortality rates. Early twentieth-century epidemiology added the systematic study of disease distribution and determinants, while demography contributed techniques for estimating population size and structure when censuses were unreliable. The mid-twentieth century saw the rise of health economics and the use of cost-effectiveness analysis to compare interventions, particularly in low-income countries where resources were scarcest.
A decisive development came in the 1970s and 1980s with the search for a summary measure of population health that could combine mortality and morbidity. The World Health Organization (WHO) and the World Bank, among others, needed a way to compare the overall burden of different diseases. This led to the development of the disability-adjusted life year (DALY), a metric that quantifies health loss as the sum of years of life lost due to premature death and years lived with disability, weighted by the severity of the disability. The first Global Burden of Disease (GBD) study, published in the early 1990s, applied this framework to produce comprehensive, comparable estimates of disease burden for all major causes across the world. This study, and the ongoing GBD enterprise that grew from it, transformed the field by making global health measurement a continuous, standardized, and highly visible activity.
The evaluation side of the field developed in parallel, drawing on clinical trial methodology, biostatistics, and the program evaluation traditions of public health and development economics. Randomized controlled trials became the gold standard for assessing intervention efficacy, while quasi-experimental designs—such as interrupted time series, difference-in-differences, and instrumental variable analyses—were developed for situations where randomization was impractical or unethical. The rise of implementation science in the late twentieth and early twenty-first centuries added a focus on how interventions perform in real-world settings, including questions of coverage, fidelity, and sustainability.
The most influential framework within the field is the burden of disease approach, operationalized through the GBD study. Its organizing assumption is that all health loss can be expressed in a common unit—the DALY—allowing direct comparison across diseases, injuries, and risk factors. The approach is explicitly global and comparative: it produces estimates for every country, age group, sex, and year, using standardized case definitions and statistical models to fill gaps in primary data.
The GBD study is a massive, ongoing collaborative enterprise that synthesizes thousands of data sources, including vital registration, verbal autopsies, surveys, disease registries, and published studies. Its methods are complex and continually evolving. For causes without direct measurement, it uses statistical models that borrow strength from covariates and neighboring locations. For risk factors, it estimates the burden attributable to each by combining exposure distributions with relative risks derived from the literature. The results are published regularly and are used by governments, international agencies, and researchers to set priorities and track progress toward health targets.
The burden of disease approach has been enormously influential, but it is not without critics. Some argue that the DALY's disability weights—the severity values assigned to different health states—are culturally biased or insufficiently sensitive to the lived experience of illness. Others contend that the emphasis on aggregate burden can obscure distributional concerns, such as inequality within countries or the concentration of disease among marginalized groups. The GBD's reliance on statistical modeling has also been questioned, particularly when estimates for countries with sparse data are presented with a precision that the underlying evidence does not support. Proponents respond that the models are transparent, uncertainty intervals are published, and the alternative—no estimates at all—is worse for decision-making. These debates are ongoing and are themselves a central part of the field's intellectual life.
A second major area of work concerns the development and validation of health indicators. This includes the design of survey instruments to measure health status, functional limitation, and well-being; the construction of composite indices such as the Human Development Index or the Universal Health Coverage service coverage index; and the harmonization of definitions across data sources so that comparisons are meaningful. A key challenge is ensuring that indicators are comparable across cultures and languages. A question about mobility or pain may be interpreted differently in different settings, and instruments must be translated and culturally adapted without losing their measurement properties.
This area also includes the measurement of health system inputs and outputs: health expenditure, workforce density, facility availability, service utilization, and quality of care. These indicators are often used for benchmarking and for tracking progress toward universal health coverage. The field has increasingly recognized that measuring service coverage alone is insufficient; the quality of care delivered matters enormously for health outcomes, and quality metrics are now a major focus of research and policy.
The evaluation side of the subfield asks whether specific policies, programs, and interventions work. This work is grounded in the counterfactual logic of causal inference: to estimate an intervention's effect, one must construct a credible comparison between what happened with the intervention and what would have happened without it. Randomized trials provide the strongest basis for this comparison, and many global health interventions—vaccines, insecticide-treated bed nets, micronutrient supplements, cash transfers—have been evaluated in this way.
However, randomization is not always possible or appropriate. Evaluating a national health policy, a mass media campaign, or a health system reform often requires observational or quasi-experimental methods. These methods exploit natural variation—for example, the timing of a policy change, geographic differences in program rollout, or threshold rules for eligibility—to approximate a randomized comparison. Each method has assumptions that must be defended, and the field has developed a sophisticated literature on when each design is credible and how to test its validity.
A further distinction is between efficacy and effectiveness. Efficacy trials measure whether an intervention works under ideal, controlled conditions; effectiveness studies measure whether it works in routine practice. The gap between the two is often substantial, and understanding why interventions lose their impact when scaled up is a central concern of implementation science. This includes studying barriers to uptake, fidelity of delivery, and the adaptations needed for different contexts. Economic evaluation adds a further layer, asking whether an intervention's benefits justify its costs relative to alternatives. Cost-effectiveness analysis, expressed in dollars per DALY averted or per quality-adjusted life year gained, is a standard tool for priority-setting, though its use raises ethical questions about how to value a year of life and whose preferences should count.
Underpinning all of this work is the challenge of producing reliable estimates from imperfect data. In many low- and middle-income countries, vital registration is incomplete, causes of death are not medically certified, and disease surveillance is patchy. The field has developed a range of methods to address these gaps. Verbal autopsy—interviewing family members about the symptoms and circumstances preceding a death—is used to assign probable causes of death when medical certification is absent. Demographic surveillance systems track health and vital events in defined populations over time, providing high-quality longitudinal data in selected sites. Household surveys, such as the Demographic and Health Surveys and Multiple Indicator Cluster Surveys, provide nationally representative data on health behaviors, service utilization, and child health outcomes.
Statistical estimation methods are used to combine these disparate sources into coherent national and global estimates. These methods include small-area estimation, which borrows strength across space to produce subnational estimates; Bayesian hierarchical models, which formally combine prior information with observed data; and nowcasting and forecasting techniques for tracking trends in real time. A persistent tension in the field is between the desire for comprehensive, comparable estimates and the need to respect the uncertainty inherent in sparse or biased data. The field's response has been to make uncertainty intervals a standard part of reporting, though these intervals are often wide and are not always well understood by users.
The contemporary field is characterized by several durable features. The GBD study has become a permanent, continuously updated global enterprise, and its estimates are widely used for monitoring progress toward the Sustainable Development Goals and other international health targets. At the same time, there is growing recognition that global averages can mask profound inequalities, and the field has shifted toward producing subnational estimates to reveal within-country disparities. The COVID-19 pandemic underscored both the importance and the fragility of health metrics: the need for timely, accurate data on cases, deaths, and health system capacity was urgent, yet many countries lacked the systems to produce it, and even well-resourced settings struggled with data quality and comparability.
The field is also becoming more pluralistic in its methods and concerns. There is increasing attention to health equity, with metrics disaggregated by sex, age, wealth, education, and geography. The measurement of non-communicable diseases, mental health, and injuries has expanded, reflecting the epidemiological transition in low- and middle-income countries. The social determinants of health—the conditions in which people are born, grow, live, work, and age—are now a recognized focus, though measuring their contribution to health loss remains methodologically challenging. And the field is engaging with new data sources, including mobile phone data, satellite imagery, and electronic health records, which offer the promise of more timely and granular measurement but also raise questions about privacy, representativeness, and algorithmic bias.
The relationship between measurement and action remains the field's defining tension. Metrics are produced to inform decisions, but they are also social constructs that embody assumptions about what counts as health, whose health matters, and how health loss should be valued. The most sophisticated practitioners in the field are those who understand both the technical machinery of estimation and the interpretive and political dimensions of their work. Global health metrics and evaluation is thus not merely a set of tools but a way of seeing the world—one that has become indispensable to how the global community understands and responds to health.