Survey sampling is the branch of statistics concerned with selecting a subset of individuals from a finite population in order to estimate characteristics of that entire population. Its central problem is deceptively simple: how can a relatively small number of observations yield reliable knowledge about a much larger group? The answer, developed over roughly a century, rests on the deliberate use of chance. By giving every member of the population a known, nonzero probability of being selected, statisticians can make precise statements about the uncertainty of their estimates—statements that are impossible when samples are chosen haphazardly or by convenience.
The fundamental challenge of survey sampling is that a sample is never a perfect miniature of the population. Even with flawless execution, the particular individuals selected will differ from those not selected, producing sampling error. The discipline's central achievement is not eliminating this error but quantifying it. When a sample is drawn using random methods, the laws of probability allow researchers to calculate how much estimates are likely to vary from the true population values, and to construct confidence intervals that capture that uncertainty.
This matters because surveys are used to make consequential decisions. Governments use sample surveys to measure unemployment, inflation, and population demographics. Businesses use them to gauge customer satisfaction and market demand. Public health researchers use them to estimate disease prevalence and risk factors. In each case, the cost of surveying everyone would be prohibitive, and in many cases impossible. Sampling makes large-scale measurement feasible, but only if the sample is designed so that its results can be trusted.
The stakes extend beyond accuracy to fairness. A poorly designed sample can systematically exclude certain groups, producing estimates that are not merely imprecise but biased. The history of survey sampling is in large part a history of learning to recognize and control the many ways a sample can go wrong—not just through bad luck, but through flawed design, nonresponse, and measurement error.
The intellectual foundations of survey sampling were laid in the early twentieth century, though the practice of taking censuses and partial enumerations is far older. The key insight—that random selection allows probability-based inference—emerged from the broader development of mathematical statistics, particularly the work of Ronald Fisher on experimental design in the 1920s. Fisher showed that randomization in agricultural experiments allowed valid inference about treatment effects. The application of this idea to human populations followed.
The Norwegian statistician Anders Kiaer advocated for "representative" sampling in the 1890s, arguing that carefully selected samples could substitute for full censuses. However, early representative sampling often relied on purposive selection—choosing individuals judged to be typical—which provided no way to quantify error. The crucial shift came with the recognition that probability sampling, where selection is governed by chance, was both more defensible and more useful.
The 1930s and 1940s saw the consolidation of the field. The Indian statistician P. C. Mahalanobis developed large-scale sample surveys for agricultural statistics, introducing concepts like interpenetrating subsamples to measure survey error. In the United States, the statistician and sociologist Paul Lazarsfeld and others developed survey methods for social research, while the U.S. Bureau of the Census, under the leadership of Morris Hansen and William Hurwitz, pioneered the practical application of probability sampling to national statistics. The 1940 U.S. census included a sample-based supplement, and by the 1950s, probability sampling had become the standard for official statistics in many countries.
The theoretical foundations were formalized in a series of influential texts, most notably by William Cochran and by Leslie Kish. These works established the core vocabulary of the field—sampling frames, strata, clusters, weights—and provided the mathematical machinery for analyzing complex survey designs. The field has since expanded to address the practical challenges of real-world surveys: nonresponse, missing data, and the increasing difficulty of reaching respondents in an era of declining participation rates.
The modern field is organized around a family of probability sampling designs, each addressing a different practical constraint. The simplest is simple random sampling, where every subset of the population of a given size has an equal chance of being selected. This design is the conceptual baseline, but it is rarely used in practice because it requires a complete list of the population and can be inefficient.
Stratified sampling divides the population into homogeneous subgroups, or strata, and draws independent samples from each. This approach guarantees representation of each stratum and can reduce sampling error when the strata differ in the characteristic being measured. For example, a national survey might stratify by region or by urban-rural status to ensure that small but important subgroups are adequately covered. Stratification is most effective when the strata are internally similar with respect to the outcome of interest.
Cluster sampling takes the opposite approach. When no complete list of individuals exists, but a list of groups does—such as households, schools, or villages—researchers can sample clusters first and then sample individuals within selected clusters. This design is often far more practical and cost-effective than simple random sampling, because it reduces travel and listing costs. However, it is typically less precise, because individuals within a cluster tend to be similar to one another. The loss of precision can be quantified by the design effect, which measures how much the variance of an estimate under a complex design exceeds that of a simple random sample of the same size.
Systematic sampling selects every kth element from an ordered list, with a random starting point. It is simple to implement and can be more convenient than simple random sampling, but it can produce biased results if the list has a periodic structure that aligns with the sampling interval. In practice, it is often used as a convenient approximation to simple random sampling when the list is effectively random.
Multistage sampling combines these approaches, typically using cluster sampling at early stages and stratification or simple random sampling at later stages. Most large national surveys use multistage designs: first selecting geographic areas, then households within those areas, then individuals within households. This flexibility allows survey designers to balance cost, precision, and operational feasibility.
Probability proportional to size sampling is a refinement used in multistage designs. When clusters vary greatly in size, giving larger clusters a higher probability of selection can improve efficiency, provided that the analysis accounts for the unequal selection probabilities. This technique is standard in establishment surveys, where businesses vary enormously in employment or revenue.
Once a sample is drawn, the task shifts to producing estimates and quantifying their uncertainty. The key concept is the sampling weight, which reflects the probability of selection. Each sampled unit's weight is the inverse of its selection probability, and weighted estimates are unbiased for population totals and means under the design. Weights also incorporate adjustments for nonresponse and for coverage errors, where the sampling frame does not perfectly match the population.
The variance of an estimate from a complex design cannot be computed using the simple formulas of elementary statistics, because observations are not independent. Instead, survey statisticians use methods that respect the design structure. Linearization (also called the Taylor series method) approximates the variance of a nonlinear statistic, such as a ratio or a regression coefficient. Replication methods, including jackknife and bootstrap, repeatedly recompute the estimate from subsets of the data to empirically assess variability. These methods are now standard in statistical software, allowing analysts to correctly account for stratification, clustering, and weighting.
A central distinction in the field is between design-based and model-based inference. Design-based inference treats the population values as fixed and the randomness as arising solely from the sampling design. This approach is the traditional foundation of survey sampling, and it has the advantage of requiring few assumptions: the validity of estimates rests on the known selection probabilities, not on a model of how the data were generated. Model-based inference, by contrast, treats the population values as realizations of a stochastic process and uses models to predict the unobserved values. This approach can be more efficient when the model is correct, and it is essential for handling missing data and for small-area estimation, where sample sizes are too small for direct design-based estimates.
The two approaches are not mutually exclusive. Modern practice often blends them, using models to improve precision while relying on design-based principles to guard against model misspecification. The tension between them has been a productive source of methodological development, particularly in the treatment of nonresponse and in the construction of calibration weights.
No survey achieves complete response. Some sampled individuals cannot be contacted, and others refuse to participate. If nonresponse is related to the outcomes being measured, the resulting estimates can be biased. The field has developed a range of strategies to address this problem.
The first line of defense is prevention: careful questionnaire design, repeated contact attempts, and incentives can reduce nonresponse. The second line is adjustment. Post-stratification and raking adjust weights so that the sample matches known population totals for demographic variables, such as age, sex, and region. These methods reduce bias if the variables used for adjustment are related to both response propensity and the outcomes of interest.
More formal approaches treat nonresponse as a missing-data problem. Imputation fills in missing values using models based on observed data, allowing analysts to use standard complete-data methods. Weighting adjustments model the probability of response and increase the weights of respondents to compensate for nonrespondents. The choice between these approaches depends on the pattern of missingness and the goals of the analysis. Modern practice increasingly uses multiple imputation, which creates several plausible completed datasets and combines results across them to reflect the uncertainty due to missing values.
A related challenge is coverage error, which occurs when the sampling frame does not include everyone in the target population. People without telephones, without stable addresses, or who live in institutions may be systematically excluded. Frame construction is often the most difficult and costly part of a survey, and its quality directly affects the validity of results.
Survey sampling today is shaped by several converging trends. Response rates have declined in many countries, making nonresponse adjustment more important and more difficult. At the same time, new data sources—administrative records, social media, and other digital traces—offer potential alternatives or supplements to traditional surveys. These sources are not probability samples, and their use raises fundamental questions about inference. The field has responded by developing methods for nonprobability sampling, including techniques that use propensity scores or calibration to adjust nonprobability samples to resemble probability samples. These methods are promising but require strong assumptions, and their validity remains an active area of research.
Another major development is the integration of surveys with auxiliary data. Census records, administrative databases, and satellite imagery can be used to improve sampling frames, to calibrate weights, and to support small-area estimation. The increasing availability of such data has made model-based methods more attractive, as they can leverage rich covariates to predict outcomes for unsampled areas or groups.
The field has also become more attentive to measurement error—the gap between what respondents report and the true values. Cognitive testing, mode effects (telephone versus web versus in-person), and questionnaire design are now recognized as central to survey quality, not merely as practical concerns. The total survey error framework, which organizes all sources of error—sampling, coverage, nonresponse, measurement, and processing—has become the standard way of thinking about survey quality.
Finally, the discipline has expanded geographically and institutionally. National statistical offices in many countries conduct sophisticated probability surveys, and international organizations coordinate cross-national studies. Academic training in survey methodology is well established, and the field's professional societies and journals are global. While the core principles of probability sampling remain the foundation, the contemporary practice of survey sampling is a pragmatic blend of design-based theory, model-based methods, and a deep engagement with the practical realities of data collection.