Quality and safety is a subfield of health services research concerned with understanding, measuring, and improving the performance of healthcare delivery. It asks a deceptively simple pair of questions: Do patients receive the care they should, and are they harmed by the care they receive? Beneath these questions lies a complex body of knowledge about how healthcare organizations function, how human error occurs, and how systems can be designed to protect patients.
The field is built on a fundamental observation: healthcare frequently fails to deliver what evidence says it should, and it sometimes injures the very people it aims to help. This gap between ideal care and actual care takes two related but distinct forms. The first is underuse—patients not receiving interventions known to be effective, such as a beta-blocker after a heart attack. The second is harm—patients experiencing adverse events, from medication errors to surgical complications to hospital-acquired infections.
These problems are not rare anomalies but systematic features of how care is organized. Landmark studies in the late 1990s and early 2000s, most notably the U.S. Institute of Medicine's report To Err Is Human (1999), estimated that tens of thousands of Americans died annually from preventable medical errors. Similar studies in other countries produced comparable findings. The shock of these numbers was not that individual clinicians were incompetent—most were dedicated and skilled—but that the systems in which they worked made errors predictable. This reframing, from blaming individuals to examining systems, is the intellectual foundation of the modern field.
The concern with healthcare quality is older than the subfield itself. In the early twentieth century, the American surgeon Ernest Codman advocated for "end result" tracking—following patients to see whether treatments actually worked—and was ostracized for his efforts. The formal quality assurance movement in hospitals, which began mid-century, focused on inspection and credentialing: setting minimum standards, reviewing records, and identifying practitioners who fell below acceptable performance. This approach was largely regulatory and retrospective; it asked whether care met predetermined criteria, but it did not systematically ask how to make care better.
A crucial shift came from outside medicine. In the mid-twentieth century, statisticians and engineers, building on the work of Walter Shewhart and W. Edwards Deming, developed methods for improving industrial processes. Their insight was that quality is not achieved by inspecting finished products but by understanding and improving the production process itself. Variation in output, they argued, is caused by flaws in the system, not by lazy or careless workers. In the 1980s and 1990s, these ideas—continuous quality improvement, total quality management, statistical process control—were imported into healthcare. Hospitals began using run charts to track infection rates, forming multidisciplinary teams to redesign care processes, and applying the Plan-Do-Study-Act cycle to test changes on a small scale before implementing them broadly.
This importation created a permanent tension in the field. On one side are those who see quality as a technical problem of measurement and improvement, solvable through better data and better processes. On the other are those who emphasize the professional and cultural dimensions—the need to change how clinicians think about error, hierarchy, and accountability. The field has never fully resolved this tension, and its history is partly a story of oscillation between these emphases.
Patient safety emerged as a distinct concern within the broader quality movement in the 1990s, driven by two converging streams. The first was the growing recognition, from the Harvard Medical Practice Study and similar investigations, that adverse events were common and often preventable. The second was the influence of human factors engineering—the discipline that studies how people interact with complex systems, originally developed in aviation, nuclear power, and the military.
The central insight of the safety movement is that human error is not a moral failing but a predictable consequence of poorly designed systems. When a nurse administers the wrong medication, the question is not "Who is to blame?" but "What features of the environment made this error likely?" The answer often lies in look-alike drug packaging, confusing labels, interruptions during medication administration, or inadequate staffing. The remedy, therefore, is not to retrain or punish the individual but to redesign the system: barcode scanning, standardized protocols, independent double-checks, and forcing functions that make it impossible to proceed with a dangerous action.
This systems approach has several practical consequences. It shifts the focus from retrospective blame to prospective risk reduction. It encourages reporting of errors and "near misses" so that latent hazards can be identified before they cause harm. It treats safety as a property of the whole organization, not of individual practitioners. And it recognizes that safety must be designed into the work environment, not achieved through exhortation or vigilance alone.
A key concept in this tradition is the distinction between active failures—the immediate errors made by frontline workers—and latent conditions—the underlying organizational factors, such as poor training, inadequate equipment, or production pressures, that make those errors more likely. The British psychologist James Reason, whose work on this distinction was highly influential, argued that most accidents are not caused by a single reckless act but by the accumulation of latent conditions that line up, like holes in a Swiss cheese, to allow harm to occur. The implication is that improving safety requires finding and fixing the latent conditions, not just punishing the active failures.
A central preoccupation of the field is measurement. To improve care, one must first be able to assess it. But measuring healthcare quality is surprisingly difficult. The most common framework, developed by Avedis Donabedian in the 1960s, distinguishes three types of measures: structure (the attributes of the setting in which care occurs, such as staffing levels, equipment, and training), process (what is actually done to the patient, such as whether a diabetic receives an eye exam), and outcome (the results of care, such as mortality, complication rates, or patient-reported well-being).
Each type has strengths and weaknesses. Outcome measures are the most meaningful to patients, but they are influenced by factors outside the healthcare system's control, such as the patient's underlying health and socioeconomic circumstances. A hospital with sicker patients may appear to have worse outcomes even if its care is excellent. Process measures are more directly under the control of providers and easier to act on, but they assume that the process in question is actually linked to better outcomes—an assumption that is not always well supported. Structural measures are the easiest to collect but the furthest from what patients actually experience.
The field has developed sophisticated statistical techniques to address these challenges. Risk adjustment attempts to account for differences in patient populations so that outcomes can be compared fairly across institutions. Composite measures combine multiple indicators into a single score. Patient-reported outcome measures (PROMs) and patient-reported experience measures (PREMs) attempt to capture the patient's perspective, which is often poorly correlated with clinical measures. Yet measurement remains contested. Critics argue that the emphasis on measurement has led to "gaming"—providers focusing on what is measured rather than what matters—and to a proliferation of indicators that burden clinicians without improving care.
Knowing what should be done is not the same as getting it done. A large portion of the field is devoted to the study of how to change clinical practice. This is the domain of implementation science and improvement science, which ask why evidence-based practices are adopted slowly and unevenly, and how they can be spread more effectively.
The dominant approach to improvement in healthcare has been the Model for Improvement, developed by the Associates in Process Improvement and popularized by the Institute for Healthcare Improvement. It consists of three questions—What are we trying to accomplish? How will we know that a change is an improvement? What changes can we make that will result in improvement?—followed by iterative cycles of testing changes on a small scale, measuring their effects, and refining them before broader implementation. This approach is pragmatic and rapid, but it has been criticized for producing changes that are not rigorously evaluated. A change that appears to work in one hospital may not work in another, and the Model for Improvement does not always provide the evidence needed to know whether an observed improvement is actually due to the intervention.
More rigorous approaches draw on the traditions of clinical epidemiology and biostatistics. Interrupted time series analysis, stepped-wedge cluster randomized trials, and other quasi-experimental designs are used to evaluate improvement interventions with greater confidence. The tension between these two traditions—the pragmatic, rapid-cycle approach favored by improvement practitioners and the rigorous, slow approach favored by academic evaluators—is a persistent feature of the field. Proponents of each accuse the other of failing to serve patients: the rigorous approach produces knowledge too slowly to help anyone, while the pragmatic approach produces changes that may not actually work.
Quality and safety are not purely technical problems. They are shaped by the culture of healthcare organizations—the shared values, beliefs, and norms that influence how clinicians behave. A hospital can have excellent protocols and still have poor safety if its culture discourages speaking up, punishes error, or treats safety as someone else's responsibility.
The concept of safety culture has become central to the field. It refers to the extent to which an organization's members share a commitment to safety, believe that safety is a priority, and feel able to raise concerns without fear of retribution. Surveys such as the Hospital Survey on Patient Safety Culture attempt to measure this construct, and the results are used to guide improvement efforts. Research suggests that a positive safety culture is associated with fewer adverse events, although the direction of causality is not always clear—good safety may create a positive culture, rather than the reverse.
Related to culture is the problem of professional hierarchy. Healthcare has traditionally been organized around the authority of physicians, and this hierarchy can inhibit communication. A nurse who notices a potential problem may be reluctant to challenge a physician's decision, and a junior physician may hesitate to question a senior colleague. The field has therefore emphasized the importance of teamwork and communication, developing tools such as structured handoff protocols, briefings before procedures, and "speak up" campaigns that encourage all team members to voice concerns. These interventions are often described as flattening the hierarchy, though their success in doing so is variable.
Quality and safety are not only matters of professional practice; they are also matters of public policy. Governments and regulatory bodies have developed a range of mechanisms to hold healthcare organizations accountable for the quality of their care. These include licensing and accreditation (external inspections against standards), public reporting of performance data, pay-for-performance programs that tie reimbursement to quality measures, and, in some countries, legal liability for preventable harm.
The relationship between these external pressures and internal improvement efforts is complex. On one hand, regulation can drive attention to quality and provide resources for improvement. On the other hand, it can create a compliance mentality, in which organizations focus on meeting the letter of the requirements rather than improving the substance of care. Public reporting can lead to improvements in the reported measures, but it can also lead to avoidance of high-risk patients or to "cherry-picking" of easy cases. Pay-for-performance has been shown to produce modest improvements in some settings, but it can also create perverse incentives, such as focusing on what is rewarded at the expense of what is not.
The field has also grappled with the question of whether quality problems should be addressed through professional self-regulation or through external oversight. The traditional model, in which the medical profession polices its own members, has been criticized as ineffective and self-protective. The alternative model, in which government or independent bodies set standards and enforce them, has been criticized as bureaucratic and insensitive to local context. Most healthcare systems use a combination of both, but the balance varies considerably across countries.
The field today is characterized by several ongoing debates and emerging concerns. One is the question of measurement burden: the proliferation of quality indicators has created significant administrative work for clinicians, and there is growing concern that this burden is itself a threat to quality, as it diverts time and attention from patient care. Some have called for "measurement that matters"—fewer, more meaningful measures that focus on outcomes rather than processes.
Another is the challenge of health equity. Quality and safety research has traditionally focused on average performance, but there is growing recognition that disparities in care—by race, ethnicity, income, geography, and other factors—are themselves quality problems. A healthcare system that provides excellent care to some patients and poor care to others is not a high-quality system, even if its average performance looks good. This has led to calls for quality measures to be stratified by patient characteristics and for improvement efforts to explicitly address equity.
A third concern is the application of quality and safety methods to new settings and new technologies. As care moves out of hospitals and into ambulatory clinics, homes, and digital platforms, the field must adapt its methods to these new contexts. The rise of artificial intelligence in healthcare raises new questions about how to ensure the safety of algorithmic decision-making and how to monitor the performance of AI systems over time.
Finally, the field continues to struggle with the problem of spread and sustainability. Many improvement projects succeed in one unit or one hospital, but the changes do not persist or spread to other settings. Understanding how to scale up successful interventions, and how to maintain improvements once the initial enthusiasm has faded, remains an open challenge.
The field of quality and safety is thus best understood not as a settled body of knowledge but as an ongoing effort to confront a persistent problem: healthcare does not reliably deliver the care that patients need and deserve. Its methods—measurement, systems analysis, improvement cycles, cultural change, and regulation—are tools for addressing this problem, and its history is a record of how those tools have been developed, tested, and refined. The work is far from complete, and the field's central questions remain as urgent as ever.