Teacher effectiveness is the branch of education economics concerned with measuring, explaining, and improving the contribution of individual teachers to student outcomes. At its core lies a deceptively simple question: how much do teachers matter, and what makes one teacher better than another? The field exists because teaching is both the largest single input in most education systems and one of the most difficult to observe directly. Unlike textbooks, class sizes, or school facilities, a teacher's quality cannot be read off a balance sheet or counted in a supply order. It must be inferred from what students subsequently know and can do.
The stakes are considerable. Because teachers are numerous and salaries dominate education budgets, even modest improvements in average teacher effectiveness translate into large aggregate gains in student learning and later earnings. Conversely, the field's findings have fueled contentious policy debates about teacher tenure, performance pay, dismissal procedures, and the assignment of teachers to schools. The economic lens adds a distinctive perspective: it treats teaching not primarily as a craft or a calling, but as a production process with inputs, outputs, and measurable productivity differences across workers.
The central technical challenge is that teacher effectiveness cannot be observed directly. It must be estimated from data, and every estimation strategy carries assumptions. The earliest systematic attempts, dating to the 1960s and 1970s, used simple correlations between teacher characteristics—degrees earned, years of experience, certification status—and student test scores. These studies generally found weak or inconsistent relationships. A teacher with a master's degree was not reliably better than one without; an experienced teacher was not always superior to a novice. This "no systematic relationship" finding was itself important, because it suggested that the observable credentials used to set teacher salaries bore little connection to actual classroom performance.
The measurement problem became more tractable with the advent of large longitudinal datasets linking individual students to their teachers over multiple years. These data made possible value-added models, which estimate a teacher's contribution by comparing the test-score growth of students in that teacher's classroom to the growth of similar students taught by other teachers. The key innovation is the use of growth rather than levels: a teacher whose students enter the year far behind and leave near grade level is credited with high effectiveness, even if absolute scores remain low. Value-added models statistically adjust for prior achievement and, in more elaborate specifications, for student demographics, class size, and school fixed effects.
Value-added models represented a genuine breakthrough, but their limitations are equally important. They require tests that measure growth reliably across the full achievement distribution, which most standardized tests only imperfectly do. They are noisy: a teacher's estimated value-added in one year is only moderately correlated with the same teacher's estimate in another year, meaning that classification of individual teachers as above or below average is subject to substantial error. They are also sensitive to the specific model specification—whether prior-year scores are included, how missing data are handled, and whether classroom-level variables are controlled. Perhaps most fundamentally, value-added models assume that students are assigned to teachers in ways that are conditionally random once prior achievement and observable characteristics are accounted for. If principals systematically assign the most challenging students to the most skilled teachers, or if parents sort into schools based on unobserved factors, the estimates will be biased. Researchers have developed various corrections, including student fixed effects and experimental or quasi-experimental designs, but each solution introduces its own assumptions.
The field is organized around three broad approaches that answer the effectiveness question in different ways and, importantly, produce different kinds of evidence.
The production-function tradition treats the classroom as a small factory. Its practitioners estimate statistical relationships between teacher inputs and student outputs, typically using large administrative datasets. The earliest work in this tradition focused on measurable teacher attributes; the modern version centers on value-added estimation. This approach prizes external validity—the ability to generalize from a sample to a population—and is the natural tool for evaluating system-wide policies. Its weakness is internal validity: because teachers are not randomly assigned to students, the estimated effects may reflect selection rather than causation. The tradition has responded by developing increasingly sophisticated econometric techniques, but the fundamental identification problem remains.
The experimental tradition sidesteps the selection problem by creating it. In a typical experiment, teachers or classrooms are randomly assigned to receive a particular intervention—a new curriculum, a coaching program, a financial incentive—and student outcomes are compared across conditions. Random assignment ensures that, in expectation, the treatment and control groups differ only in the intervention, allowing causal claims. The experimental tradition has been especially influential in evaluating specific programs designed to improve teacher effectiveness, such as performance-pay schemes, teacher training models, and classroom observation systems. Its strength is internal validity; its weakness is external validity. An experiment showing that a particular coaching program works in one district tells us little about whether it will work elsewhere, and experiments are expensive and slow to mount at scale.
The observational and qualitative tradition approaches effectiveness from inside the classroom. Drawing on educational psychology, classroom research, and sometimes anthropology, this work asks what effective teachers actually do: how they manage discussion, give feedback, structure tasks, and build relationships. The most influential products of this tradition are observation protocols—structured rubrics that trained observers use to rate classroom practice on dimensions such as instructional clarity, classroom management, and cognitive demand. These protocols have become widely used in teacher evaluation systems, often in combination with value-added scores. The observational tradition captures dimensions of teaching that tests miss, including non-cognitive outcomes and long-term effects. Its limitations are the cost and subjectivity of observation, the difficulty of training reliable observers, and the unresolved question of whether the practices the rubrics reward actually cause student learning or merely correlate with it.
These three approaches are not mutually exclusive, and the field's most influential work often combines them. A common design uses value-added models to identify effective and ineffective teachers, then observes their classrooms to identify distinguishing practices, then tests those practices experimentally. This triangulation has produced the field's most robust findings: that teachers differ substantially in their effects on student achievement; that these differences persist across years and subjects; and that a teacher's effect on test scores is only weakly related to observable credentials.
The accumulated research supports several conclusions that are widely accepted, though each carries qualifications.
First, teachers vary enormously in their effectiveness. Studies using value-added models consistently find that the difference between a highly effective and a highly ineffective teacher is equivalent to several months of learning per year. The magnitude varies by subject, grade level, and model specification, but the general finding is robust: teacher quality is the largest school-based factor in student achievement, dwarfing class size, school spending, and most other policy levers.
Second, the characteristics that predict effectiveness are largely unobservable at the point of hire. Experience matters, but only in the first few years: teachers improve substantially during their first three to five years, after which additional experience yields little measurable gain. Advanced degrees, certification status, and college selectivity show weak or inconsistent relationships with student outcomes. This finding has profound policy implications, because salary schedules that reward credentials and seniority are not aligned with the characteristics that actually predict performance.
Third, teacher effects extend beyond test scores. Research linking value-added estimates to long-term outcomes has found that students assigned to more effective teachers are more likely to attend college, earn higher wages, and are less likely to become teenage parents. These effects are larger for disadvantaged students. The mechanisms are not fully understood—they may operate through non-cognitive skills, aspirations, or behaviors that tests do not capture—but the finding suggests that test-score-based measures capture only part of what effective teachers produce.
Fourth, the distribution of teacher effectiveness is unequal. Low-income and minority students are systematically more likely to be taught by novice teachers, teachers with lower value-added scores, and teachers who are less effective by observational measures. This "teacher quality gap" is a persistent feature of the landscape and a central concern of equity-oriented research.
The field's findings have generated a set of policy proposals that remain contested. Performance-based pay, which ties compensation to value-added or observational measures, has been implemented in various forms and evaluated experimentally. The results are mixed: some studies find modest positive effects on student achievement, others find none, and the effects appear to depend heavily on program design. Teacher evaluation systems that combine value-added scores with classroom observation have been adopted widely, but their implementation has been uneven, and research on their effects is still accumulating.
The most contentious policy question concerns dismissal. If effectiveness can be measured, the argument runs, then persistently ineffective teachers should be removed. The counterargument emphasizes measurement error: because value-added estimates are noisy, any dismissal system will inevitably remove some teachers who are actually average or better, and the threat of dismissal may deter talented candidates from entering the profession. The field has not resolved this tension, and the debate is as much about values as about evidence.
A separate strand of policy research asks whether effectiveness can be improved rather than merely measured. Teacher preparation programs, induction and mentoring for novices, and ongoing professional development have all been studied, with generally disappointing results: most interventions produce small or null effects on student achievement. The most promising findings come from intensive, content-focused coaching models, but these are expensive and difficult to scale. The field's honest conclusion is that we are better at identifying effective teachers than at creating them.
Contemporary research on teacher effectiveness is characterized by methodological pluralism and a growing attention to context. The early confidence in value-added models as a standalone measure has given way to a more cautious stance that treats them as one input among several. Many researchers now advocate for "multiple measures" systems that combine value-added scores, observation ratings, and student surveys, on the grounds that each measure captures a different dimension of effectiveness and that their combination is more reliable than any single one.
The field has also expanded beyond test scores. Researchers increasingly study teacher effects on attendance, suspensions, grades, and social-emotional outcomes, recognizing that schools pursue multiple goals. The measurement of these non-cognitive outcomes is less developed than test-score measurement, and the field is still working out how to incorporate them into effectiveness estimates.
International comparisons have broadened the empirical base. The same value-added methods developed in the United States have been applied to administrative data in dozens of countries, with broadly similar findings: teachers vary, credentials predict little, and the distribution of effectiveness is unequal. Cross-country work has also highlighted the importance of institutional context—hiring practices, union rules, and accountability systems shape both the distribution of teacher quality and the feasibility of policy interventions.
The field's unresolved questions are as important as its established findings. How much of teacher effectiveness is stable across contexts, and how much depends on the match between teacher and school? Can effectiveness be reliably identified at the point of hire, or only after years of classroom data? Do the practices that raise test scores also improve long-term outcomes, or are there trade-offs? These questions are unlikely to receive definitive answers, because they depend on values as much as on evidence. What the field has achieved is a rigorous, quantitative language for discussing teacher quality—a language that has permanently changed how education systems think about their most important human resource.