Educational assessment is the systematic process of gathering, interpreting, and using evidence about student learning to make decisions about education. It is the field of practice and inquiry concerned with questions like: What have students learned? How well have they learned it? How do we know? And what should happen next, for the student, the teacher, the school, or the system?
At its core, assessment is about inference. Educators rarely have direct access to what a person knows or can do; they observe performances—answers on a test, a written essay, a lab demonstration, a classroom discussion—and draw conclusions from those observations. The central intellectual challenge of the field is to make those inferences as accurate, fair, and useful as possible. This involves designing tasks that elicit relevant evidence, developing methods for scoring or judging that evidence, and interpreting the results in ways that support sound decisions. The stakes are high: assessment results can determine a student's grade, a diploma, or admission to a university; they can influence a teacher's evaluation, a school's funding, or a nation's educational policy.
Assessment serves distinct purposes that shape how it is designed and used. The most common distinction is between formative and summative assessment. Formative assessment is intended to support learning while it is happening. It involves gathering evidence of student understanding during instruction and using it to adjust teaching and learning—for example, through classroom questioning, feedback on drafts, or short quizzes that reveal misconceptions. Summative assessment, by contrast, is intended to judge learning after instruction has occurred. End-of-course exams, standardized state tests, and final projects are summative because they summarize what a student has achieved.
A related distinction concerns the locus of decision. Classroom assessment is designed and used by teachers for their own students; it is typically embedded in instruction and can be highly responsive to local contexts. Large-scale assessment is designed by external agencies—state or national education departments, testing companies, or international organizations—and is used to compare students, schools, or systems. Large-scale assessments often serve accountability purposes: they may determine whether a school is meeting standards, whether a student is promoted, or whether a country's education system is competitive internationally.
Another important distinction is between selected-response and constructed-response formats. Selected-response items (multiple-choice, true/false, matching) require students to choose among provided options. They are efficient to score and can cover broad content, but they constrain how students can demonstrate knowledge. Constructed-response items (essays, short answers, problem-solving tasks, portfolios) require students to produce their own answers. They can capture complex thinking and authentic performance but are more time-consuming to score and raise questions about the consistency of judgment.
Assessment practices are as old as formal education, but the modern field took shape in the late nineteenth and early twentieth centuries. The rise of mass schooling created a need to sort and certify large numbers of students efficiently. Early intelligence testing, developed in France by Alfred Binet and later adapted in the United States, provided a model for measuring individual differences with standardized procedures. This influenced the development of standardized achievement tests—tests administered and scored under uniform conditions so that results can be compared across students and schools.
The mid-twentieth century saw the emergence of educational measurement as a technical discipline. Drawing on statistics and psychology, measurement specialists developed methods for quantifying test reliability (the consistency of scores) and validity (the degree to which a test measures what it claims to measure). The work of researchers such as Lee Cronbach and Samuel Messick broadened the concept of validity from a property of the test itself to an argument about the interpretations and uses of test scores. In this view, a test is not simply "valid" or "invalid"; rather, the validity of a proposed interpretation must be evaluated with evidence and theory.
A significant shift occurred in the late twentieth century with the rise of standards-based reform. Governments began to define explicit content standards—statements of what students should know and be able to do at each grade level—and to align assessments with those standards. This movement tied assessment directly to curriculum and instruction, making tests a lever for improving schools. It also intensified debates about the consequences of high-stakes testing, including concerns about teaching to the test, narrowing the curriculum, and the fairness of using the same test for all students regardless of background.
Several distinct approaches to assessment have developed, each addressing different problems and resting on different assumptions.
The oldest and most technically developed approach treats assessment as a branch of psychological measurement. Its central problem is precision: how to construct tests that yield scores with minimal error and maximal comparability. This tradition developed classical test theory, which models an observed score as the sum of a true score and random error, and later item response theory (IRT), a more sophisticated family of models that describes the probability of a correct answer as a function of the test-taker's ability and the item's characteristics. IRT underlies most modern large-scale testing, including computer-adaptive tests that adjust difficulty to the test-taker's performance level.
The measurement tradition has contributed rigorous methods for evaluating tests: reliability coefficients, standard errors of measurement, and procedures for detecting biased items. Its limits are equally clear. It tends to focus on what is easily measurable—typically factual knowledge and discrete skills—and it assumes that the construct being measured is stable and unidimensional. Critics argue that this tradition can reduce learning to a set of decontextualized performances and that its technical sophistication does not guarantee educational value.
In response to the limitations of summative, measurement-oriented assessment, a movement known as assessment for learning (or formative assessment) gained prominence in the 1990s. Its central problem is not measurement precision but instructional usefulness. Drawing on research about feedback and self-regulated learning, proponents argue that assessment should be designed primarily to help students learn, not merely to judge them.
Key practices include sharing learning goals with students, using questioning and classroom tasks to elicit evidence of understanding, providing feedback that helps students close the gap between current and desired performance, and involving students in self-assessment and peer assessment. The influential review by Paul Black and Dylan Wiliam, published in 1998, synthesized research suggesting that such practices can produce substantial learning gains, particularly for low-achieving students. This movement has had a major influence on teacher education and classroom practice, though it has also faced criticism for being difficult to implement faithfully and for relying on research evidence that is less conclusive than sometimes claimed.
A third tradition challenges the assumption that tests must consist of questions with right or wrong answers. Performance assessment requires students to demonstrate their knowledge and skills through complex tasks—writing an essay, conducting a science experiment, solving a real-world problem, or assembling a portfolio of work. The term authentic assessment is used when these tasks resemble the kinds of performances valued in adult life, such as writing for a real audience or designing a product.
This approach addresses a problem that multiple-choice tests cannot solve: the need to assess higher-order thinking, creativity, collaboration, and the ability to apply knowledge in novel situations. Performance assessments are typically scored with rubrics—descriptive scoring guides that define levels of quality for each dimension of the performance. The approach has been influential in fields like writing instruction, the arts, and project-based learning. Its challenges include the time and cost of scoring, the difficulty of achieving consistent judgments across raters, and the risk that students' scores depend as much on the rater as on the performance itself.
A more recent body of work draws on socio-cultural learning theory and critical theory to question the assumptions underlying mainstream assessment. Socio-cultural perspectives argue that learning is situated in social contexts and that assessment should therefore attend to how students participate in authentic communities of practice, not just what they know individually. This has led to interest in dynamic assessment, where an examiner interacts with the learner, providing prompts and support to discover not just what the learner can do independently but what they can do with assistance—a measure of learning potential rather than current achievement.
Critical perspectives examine assessment as a social and political practice. They ask who benefits from particular assessment arrangements, whose knowledge is valued, and how tests reproduce or challenge social inequalities. Researchers in this tradition have documented how standardized tests can disadvantage students from non-dominant linguistic and cultural backgrounds, how test preparation can crowd out meaningful learning, and how assessment labels can shape students' identities and opportunities. These perspectives do not offer a single alternative method but rather a lens for interrogating the purposes, consequences, and fairness of assessment practices.
Two concepts organize much of the field's technical and ethical work. Validity is the degree to which evidence and theory support the interpretations and uses of assessment results. Modern validity theory, following Messick, treats validity as a unified concept: one validates not the test itself but the proposed interpretation and use. A test score might be valid for one purpose (e.g., diagnosing a student's reading difficulties) and invalid for another (e.g., ranking schools). Validation is therefore an ongoing argument that requires evidence about content coverage, response processes, internal structure, relations to other variables, and the consequences of test use.
Fairness concerns whether assessment results are comparable and meaningful across different groups of students. A test is unfair if it systematically underestimates what certain students know and can do because of factors unrelated to the construct being measured—for example, unfamiliar vocabulary, cultural references, or test-taking experience. The field has developed statistical methods for detecting differential item functioning (whether items behave differently for different groups after controlling for ability) and procedural standards for ensuring that all students have an opportunity to learn the tested content. Fairness also raises deeper questions about whether the same assessment can be appropriate for students with different backgrounds, languages, and disabilities, leading to practices such as accommodations (e.g., extended time, translated directions) and alternate assessments.
Current assessment practice is shaped by several ongoing developments. Technology has transformed both the delivery and the nature of assessment. Computer-based testing allows for adaptive item selection, immediate scoring, and the use of interactive tasks that simulate real-world problems. It also enables the collection of process data—keystrokes, response times, and navigation patterns—that can provide new insights into how students solve problems, though methods for interpreting such data are still developing.
Learning progressions—empirically grounded descriptions of how understanding develops over time—have become an important framework for designing assessments that can locate a student's current level and suggest next steps. This approach connects assessment more closely to cognitive research and supports the goal of using assessment to guide instruction.
International assessments such as the Programme for International Student Assessment (PISA) and the Trends in International Mathematics and Science Study (TIMSS) have made cross-national comparison a prominent feature of educational policy. These assessments have generated both valuable comparative data and controversy about what they measure, whether their results justify policy changes, and whether the pressure to perform well distorts national curricula.
The field also faces persistent tensions. There is a continuing gap between the technical sophistication of large-scale testing and the relatively simple assessment practices common in many classrooms. There is debate about the optimal balance between summative accountability and formative support. And there is unresolved disagreement about whether assessment should primarily serve the goal of sorting and certifying students or the goal of helping every student learn. These tensions are not likely to be resolved; they reflect competing values about what education is for. The enduring contribution of educational assessment as a field is to make those values visible, to develop methods that serve chosen purposes as well as possible, and to keep asking whether the evidence we collect is worth the costs of collecting it.