Employment discrimination is the study of how labor market outcomes—hiring, pay, promotion, job assignment, and dismissal—are shaped by group characteristics such as race, sex, ethnicity, religion, age, or disability, rather than by individual productivity alone. As a subfield of labor economics, it asks not only whether such disparities exist, but why they persist, how they can be measured, and what policies might reduce them. The field sits at the intersection of economic theory, empirical methods, and public policy, and it has produced one of the most extensive bodies of evidence in applied economics.
The core puzzle of employment discrimination is deceptively simple: if employers are profit-maximizers, why would they systematically favor or disfavor workers on the basis of characteristics that are unrelated to productivity? The stakes are high. Discrimination affects lifetime earnings, occupational segregation, unemployment rates, and intergenerational mobility. It also raises fundamental questions about whether markets, left to themselves, will erode prejudice or entrench it.
The field distinguishes between several forms of discrimination. Market discrimination refers to differential treatment that occurs in the labor market itself—an employer paying a woman less than a man for the same work. Pre-market discrimination refers to differences in skills, education, or health that arise before workers enter the labor market, often due to earlier discrimination in schooling, family investment, or neighborhood resources. Post-market discrimination includes differences in treatment after hiring, such as promotion rates or access to training. A central methodological challenge is separating these channels: observed wage gaps between groups may reflect employer bias, differences in human capital, or choices made in response to anticipated discrimination.
Systematic economic study of employment discrimination began in earnest in the mid-twentieth century, though earlier observers had noted racial and gender disparities in pay and hiring. The field's founding moment is usually associated with the publication of The Economics of Discrimination by Gary Becker in 1957. Becker treated discrimination not as a moral failing but as a measurable economic behavior with costs. His framework—the "taste for discrimination" model—assumed that some employers, workers, or customers simply preferred not to associate with certain groups and were willing to pay a price to indulge that preference.
Becker's work was followed by a second foundational contribution: the statistical discrimination model, developed independently by Kenneth Arrow and Edmund Phelps in the early 1970s. This approach did not require animus. Instead, it showed that rational employers with imperfect information might use group averages as a proxy for individual productivity. If, for example, employers believe that women are more likely to quit for family reasons, they may offer all women lower wages or fewer training opportunities, even if a particular woman has no intention of leaving. Statistical discrimination is thus a product of rational inference under uncertainty, not prejudice—though it can produce equally harmful outcomes.
These two models—taste-based and statistical—remain the twin pillars of the theoretical literature. They are not mutually exclusive, and much subsequent work has attempted to test which one better explains observed patterns. The field's empirical turn came later, driven by the availability of large datasets and the development of quasi-experimental methods. By the 1990s and 2000s, economists were using audit studies, natural experiments, and decomposition techniques to measure discrimination in ways that Becker and Arrow could only theorize about.
Becker's model treats discrimination as a preference. An employer with a "taste for discrimination" acts as if there is a psychic cost to hiring a member of a disfavored group. This cost can be modeled as a wedge between the wage paid and the worker's marginal product. If an employer dislikes hiring Black workers, for example, they will only hire them if their wage is low enough to compensate for the employer's discomfort.
The model's key insight is that discrimination is costly to the discriminator. In a competitive market, employers who indulge their tastes will face higher labor costs and therefore lower profits. Over time, they should be driven out of business by non-discriminating competitors. This led Becker to predict that competitive markets would erode discrimination over the long run—a prediction that has been only partially borne out. The model's limits are now well understood: it assumes perfect competition, ignores the possibility that discrimination may be profitable in the short run, and cannot easily explain why discrimination persists in highly competitive industries. Nevertheless, the taste-based framework remains influential because it provides a clear, testable prediction: discriminating firms should have higher costs and lower profits.
Statistical discrimination begins from a different premise. Employers do not know a worker's true productivity with certainty; they observe noisy signals such as education, experience, and interview performance. When signals are imperfect, it can be rational to use group averages as additional information. If, on average, one group has lower productivity—whether due to pre-market discrimination, differences in education quality, or any other cause—then an employer who cannot perfectly observe individual productivity may rationally offer lower wages to all members of that group.
The model has two important implications. First, it can produce self-fulfilling prophecies. If employers believe that women are less committed to careers, they will invest less in training women, which makes women's careers less rewarding, which in turn makes women more likely to leave—confirming the original belief. Second, statistical discrimination can persist even in fully competitive markets, because it is not costly to the discriminator. The employer is simply using available information efficiently.
The distinction between taste-based and statistical discrimination matters for policy. Taste-based discrimination might be addressed by enforcing anti-discrimination laws or by increasing market competition. Statistical discrimination is more intractable, because it can persist even without animus, and it may require interventions that change the underlying information structure—such as providing better signals of individual productivity, or reducing the real group differences that make group averages informative.
A third wave of theoretical work, emerging in the 2000s, has moved beyond the taste-versus-statistics dichotomy. This work draws on behavioral economics and social identity theory. It asks how discrimination can arise from in-group favoritism, implicit bias, or the desire to maintain social hierarchies, even when no one holds explicit prejudiced beliefs. Some models incorporate the idea that workers themselves may prefer to work with members of their own group, creating segregation that is not driven by employer animus. Others explore how discrimination can be a "norm" that persists because individuals who deviate from it face social sanctions.
These newer approaches do not replace the earlier models so much as enrich them. They recognize that real-world discrimination is likely a mixture of animus, rational inference, and social dynamics. The challenge is that these mechanisms are difficult to distinguish empirically, and they may interact in complex ways.
The oldest empirical approach is to compare average wages between groups and then attempt to "explain" the gap using observable characteristics such as education, experience, and occupation. The Blinder-Oaxaca decomposition, developed in the 1970s, divides the wage gap into an "explained" component (due to differences in measured characteristics) and an "unexplained" component (often interpreted as discrimination, though it also captures unobserved productivity differences). This method is intuitive but has a fundamental limitation: it can only control for what is measured. If important productivity-relevant characteristics are omitted, the unexplained component will overstate discrimination; if employers discriminate in ways that affect the measured characteristics themselves—for example, by steering women into less demanding jobs—the explained component will understate it.
A more direct approach is the audit study, in which researchers send matched pairs of applicants—identical in all relevant respects except for the characteristic of interest—to apply for real jobs. The earliest such studies, conducted in the 1970s and 1980s, used trained actors who were matched on appearance, speech, and credentials. More recent "correspondence studies" send fictional résumés by mail or email, which allows for larger samples and tighter control. The landmark study in this genre, by Marianne Bertrand and Sendhil Mullainathan in 2004, found that résumés with typically White names received roughly 50 percent more callbacks than identical résumés with typically Black names.
Correspondence studies have a major advantage: they isolate discrimination at the hiring stage with a clean experimental design. Their limitations are equally clear. They measure only the initial callback, not offers, wages, or on-the-job treatment. They cannot easily distinguish between taste-based and statistical discrimination. And they may underestimate discrimination in settings where bias operates through networks, referrals, or informal channels that résumés do not capture.
A third approach exploits natural experiments—situations where discrimination-relevant conditions change for reasons unrelated to the outcomes being studied. The classic example is the "blind" orchestra audition, where musicians perform behind a screen. Studies of American orchestras found that the adoption of blind auditions increased the probability that women would be hired, suggesting that gender bias had been operating in the hiring process. Other natural experiments include changes in anti-discrimination law, the random assignment of judges or caseworkers, and the timing of policy interventions.
These designs offer stronger causal inference than decompositions, but they are limited by the availability of suitable natural experiments. They also tend to measure the effect of a specific intervention or context, making it difficult to generalize to the broader labor market.
Economists have also used controlled experiments, both in the laboratory and in the field, to study discrimination. Lab experiments can manipulate information and incentives precisely, allowing researchers to distinguish between taste-based and statistical discrimination. Field experiments go further by embedding the experiment in a real market. For example, researchers have sent testers to inquire about housing, applied for jobs with different names, or interacted with firms as customers. These methods have produced a robust body of evidence showing that discrimination persists across many contexts, though the magnitude varies considerably.
One of the field's most important findings is that discrimination has not disappeared, despite decades of anti-discrimination law and the competitive pressures that Becker's model suggested should erode it. The evidence from correspondence studies, in particular, shows that racial discrimination in hiring remains substantial. This persistence has led economists to ask why markets have not eliminated bias.
Several explanations have been offered. One is that discrimination is not as costly as the taste-based model assumes—if the supply of prejudiced employers is large, or if customers prefer to deal with certain groups, discriminating firms may not face a competitive disadvantage. Another is that statistical discrimination is self-reinforcing, as described above. A third is that discrimination operates through networks and referrals, which are difficult for outsiders to penetrate. Finally, some economists argue that the persistence of discrimination reflects the fact that it is embedded in institutional practices—such as word-of-mouth recruiting or subjective performance evaluations—that are not easily changed by market forces alone.
The subfield is closely connected to anti-discrimination policy. In the United States, Title VII of the Civil Rights Act of 1964 prohibits employment discrimination on the basis of race, color, religion, sex, or national origin, and subsequent legislation extended protections to age and disability. Economists have studied the effects of these laws, asking whether they have reduced disparities and whether they have unintended consequences.
The evidence is mixed. Some studies find that Title VII and affirmative action policies increased employment of minority workers, particularly in the years immediately following their enactment. Others find that anti-discrimination laws can have perverse effects, such as reducing hiring of protected groups if employers fear costly lawsuits. The economic analysis of these policies is complicated by the fact that the laws themselves are endogenous—they were enacted in response to discrimination, and their effects are difficult to isolate from broader social changes.
A key policy debate concerns the relative merits of "equal treatment" versus "affirmative action" approaches. Equal treatment policies prohibit discrimination but do not require preferential treatment. Affirmative action policies go further, requiring or encouraging employers to take positive steps to increase representation of underrepresented groups. Economists have analyzed both approaches, with some arguing that affirmative action can correct for statistical discrimination by providing better information, and others arguing that it can create stigma or backlash.
The contemporary field is characterized by methodological pluralism and a growing emphasis on understanding the mechanisms behind discrimination. Several trends are notable.
First, the availability of large administrative datasets and machine learning techniques has opened new avenues for measuring discrimination. Researchers can now study pay gaps within firms, promotion patterns, and the effects of algorithmic decision-making. The rise of algorithms in hiring has created a new set of questions: can algorithms reduce discrimination by removing human bias, or do they reproduce and even amplify existing biases encoded in training data?
Second, there is increasing attention to intersectionality—the idea that discrimination operates differently for individuals who belong to multiple groups. A Black woman may face discrimination that is not simply the sum of racial and gender discrimination, but is qualitatively distinct. This complicates both measurement and policy.
Third, the field has expanded beyond the United States. Cross-country comparisons reveal that the patterns and magnitudes of discrimination vary considerably, and that the legal and institutional context matters. Discrimination against ethnic minorities, immigrants, and religious groups is a global phenomenon, and the economic tools developed in the American context have been adapted to study it elsewhere.
Finally, there is a growing recognition that discrimination is not only a matter of individual employers' preferences or beliefs, but is embedded in the structure of labor markets. Occupational segregation, differences in access to networks, and the design of jobs themselves can produce discriminatory outcomes even in the absence of discriminatory intent. This has led some economists to call for a broader focus on "structural" discrimination, though the concept remains contested.
The open questions are substantial. How much of observed group differences in labor market outcomes is due to discrimination, as opposed to pre-market factors? What is the relative importance of taste-based versus statistical discrimination? Why does discrimination persist in competitive markets? And what policies are most effective at reducing it without creating new inefficiencies or injustices? These questions are unlikely to be settled definitively, but the field's combination of rigorous theory, creative empirical methods, and direct policy relevance ensures that they will continue to be studied.