Confirmation and evidence is the branch of the philosophy of science that asks what it means for a piece of information to support a scientific hypothesis, and how that support should be measured, compared, and rationally applied. It sits at the intersection of logic, probability, and the actual practice of science. The field does not primarily ask whether a particular scientific claim is true; it asks what makes a claim warranted, how much warrant it has, and how new data should change that warrant. Its central questions include: When does a body of evidence confirm a hypothesis? Can two hypotheses be equally compatible with the same data yet be confirmed to different degrees? What is the difference between a prediction that succeeds and a hypothesis that merely accommodates already-known facts? And how should scientists weigh evidence when hypotheses are complex, uncertain, or competing?
The stakes are practical as well as theoretical. Scientific conclusions—about climate change, the safety of a drug, the existence of a particle—are only as strong as the evidence behind them. Confirmation theory tries to make explicit the standards that scientists often apply implicitly, and to test whether those standards are coherent. It also feeds into debates about scientific realism, underdetermination, and the rationality of theory choice.
The modern field grows out of a long tradition of thinking about induction—the inference from observed particulars to general laws. Aristotle and later scholastic thinkers treated induction as a simple enumeration: if many swans are white and none are black, conclude that all swans are white. Francis Bacon in the early seventeenth century refined this into a method of eliminative induction, in which one systematically rules out hypotheses by collecting negative instances. But the philosophical problem that defines the field was stated most sharply by David Hume in the eighteenth century. Hume argued that no amount of past observation can logically guarantee that the future will resemble the past. The sun has risen every day, but there is no contradiction in supposing it will not rise tomorrow. Any attempt to justify induction by appealing to past success of induction would itself be circular. This "problem of induction" does not say that induction is useless; it says that induction cannot be given a deductive justification. The entire subsequent history of confirmation theory can be read as a series of attempts to say what induction can be, given that it cannot be deduction.
For much of the nineteenth century, the dominant response was a form of inductivism associated with William Whewell and John Stuart Mill. Whewell emphasized that hypotheses are not merely collected from data but are "superinduced" by the mind, and that a good hypothesis must explain the data and also predict new phenomena. Mill, in his System of Logic, codified methods of experimental inquiry—agreement, difference, concomitant variation—that remain a useful informal toolkit. But neither offered a formal account of how much support a hypothesis receives. That task fell to the twentieth century.
The first systematic, formal approach to confirmation emerged within logical empiricism, the dominant philosophy of science in the first half of the twentieth century. Rudolf Carnap and Carl Hempel, among others, sought to reconstruct confirmation as a purely logical relation between a hypothesis and an evidence statement, independent of psychology, sociology, or the actual history of science. The goal was to specify, in advance, a function that would assign a degree of confirmation to any hypothesis given any evidence.
Hempel first clarified a basic distinction that still structures the field: the difference between qualitative confirmation (whether evidence supports a hypothesis at all), comparative confirmation (whether one hypothesis is better supported than another), and quantitative confirmation (how much support, on a numerical scale). He also identified a set of intuitive adequacy conditions that any confirmation relation should satisfy. One of these, the "consequence condition," said that if evidence confirms a hypothesis, it should also confirm any logical consequence of that hypothesis. This led to the famous "raven paradox": the hypothesis "all ravens are black" is logically equivalent to "all non-black things are non-ravens." A white shoe is a non-black non-raven, so by the consequence condition, observing a white shoe confirms that all ravens are black. This seems absurd, and the paradox forced philosophers to ask whether confirmation is really a purely logical relation or whether it depends on background knowledge about the relative sizes of the classes involved.
Carnap attempted a full quantitative theory using inductive logic. He defined a continuum of confirmation functions, each assigning a probability to a hypothesis given evidence, based on the structure of a formal language. The key idea was that the degree of confirmation is a logical probability: a measure of the extent to which the evidence entails the hypothesis, not in the deductive sense but in a weaker, inductive sense. Carnap's system was elegant but faced severe problems. It required choosing a language with fixed predicates, and the results depended heavily on that choice. It also could not handle universal generalizations over infinite domains without assigning them zero probability, which would make all scientific laws unconfirmable. By the 1960s, Carnap's programme had largely stalled, though its influence on later Bayesian approaches was profound.
Running alongside the logical empiricist programme was a simpler, older idea: the hypothetico-deductive (H-D) model. On this view, a hypothesis is confirmed when it entails an observation statement that is then verified. If the prediction fails, the hypothesis is falsified; if it succeeds, the hypothesis is confirmed, though not proven. This model has an intuitive appeal and matches much of scientific practice, especially in physics and astronomy. It was defended by Karl Popper, though with an important twist.
Popper rejected the very idea of confirmation as a positive support relation. He argued that science does not proceed by accumulating confirmations but by making bold conjectures and then attempting to refute them. A hypothesis that survives severe testing is not "confirmed" in the sense of being made more probable; it is merely "corroborated"—it has shown itself to be the best available candidate so far. Popper's falsificationism was a response to the problem of induction: since induction cannot be justified, science should not rely on it. Instead, science should use deduction, because a single counterexample logically refutes a universal statement. The asymmetry between verification and falsification—no number of white swans proves "all swans are white," but one black swan disproves it—became the cornerstone of his philosophy.
Popper's view had enormous influence on scientific self-understanding, but it faced serious difficulties. First, as Pierre Duhem had already argued in the early twentieth century, a hypothesis is rarely tested in isolation. When a prediction fails, the fault may lie not in the hypothesis under test but in auxiliary assumptions about instruments, background conditions, or other theories. This is the "Duhem–Quine problem": falsification is not decisive because one can always save a hypothesis by adjusting auxiliary assumptions. Second, Popper's criterion of corroboration was itself comparative and informal; he never provided a measure of how severely a hypothesis had been tested. Third, many philosophers argued that Popper had simply renamed the problem of induction. A corroborated hypothesis is one that has survived past tests, but why should that make it a good guide to the future? Popper's answer—that we have no better option—was pragmatic but did not solve the underlying epistemic question.
The H-D model, in its simpler form, also faced the problem of "underdetermination of theory by evidence." If a hypothesis entails an observation, then so does the hypothesis conjoined with any arbitrary extra claim. The evidence confirms the conjunction just as well, which suggests that confirmation cannot be a simple relation between a single hypothesis and an observation. Moreover, the H-D model gives no account of how to compare two hypotheses that both entail the same evidence. This is where Bayesian approaches entered the picture.
The most influential contemporary framework for confirmation is Bayesianism. Its central claim is that degrees of belief should obey the axioms of probability, and that evidence updates those degrees by conditionalization. If you have a prior probability for a hypothesis, P(H), and you observe evidence E, your posterior probability is P(H|E) = P(E|H)P(H)/P(E), by Bayes' theorem. The term P(E|H) is the likelihood—the probability of the evidence given the hypothesis. The ratio P(E|H)/P(E) is the "Bayes factor," which measures how much the evidence shifts your belief.
Bayesianism solves several problems that plagued earlier approaches. It gives a natural account of comparative confirmation: hypothesis H1 is better confirmed by E than H2 if P(H1|E) > P(H2|E). It handles the raven paradox by noting that the likelihood of observing a white shoe is nearly the same whether or not all ravens are black, so the evidence has negligible confirmatory power. It also handles the Duhem–Quine problem: when a prediction fails, the posterior probability of the hypothesis drops, but so does the posterior of each auxiliary assumption, and the drop is distributed according to prior probabilities and likelihoods. There is no single decisive falsification, but there is a rational redistribution of confidence.
Bayesianism is not a single doctrine but a family of approaches. "Subjective Bayesians" allow priors to vary freely among individuals, constrained only by coherence with the probability axioms. "Objective Bayesians" argue for particular choices of priors, often based on symmetry or maximum entropy. "Frequentist" statisticians, by contrast, reject the very idea of assigning probabilities to hypotheses, since hypotheses are not random events with long-run frequencies. The debate between Bayesian and frequentist methods is one of the liveliest in the field, with implications for how scientific results are reported and interpreted.
The most serious internal problem for Bayesianism is the "problem of old evidence." If a hypothesis was designed to fit already-known facts, those facts have probability 1 in your current belief state, so they cannot raise the hypothesis's probability. Yet scientists often cite old evidence as supporting a theory—for example, Newton's theory explained the already-known orbit of the Moon. Several responses have been proposed: one can consider what your probability would have been had you not known the evidence, or one can treat the evidence as supporting the theory in the sense of increasing its "explanatory power" rather than its probability. None of these responses is fully satisfactory, and the problem remains an active research area.
Another challenge is the "problem of priors." In many scientific contexts, there is no obvious way to assign a prior probability to a hypothesis, especially a novel one. Subjective Bayesians reply that priors are personal and that the likelihoods will eventually dominate as evidence accumulates. But this is only true in the long run, and for finite data, different priors can lead to very different posteriors. This has led to a search for "robust" Bayesian methods that give conclusions insensitive to the choice of prior, and to a broader debate about whether Bayesianism can be objective enough for scientific practice.
A related but distinct approach focuses on likelihoods rather than posterior probabilities. The likelihood principle, associated with statisticians like R. A. Fisher and philosophers like Ian Hacking and A. W. F. Edwards, holds that all the evidence relevant to a hypothesis is contained in the likelihood function—the probability of the observed data under that hypothesis. On this view, evidence E supports hypothesis H1 over H2 if and only if P(E|H1) > P(E|H2). The ratio of likelihoods, the "likelihood ratio," measures the strength of the evidence.
This approach avoids the problem of priors entirely, since it never assigns probabilities to hypotheses. It also gives a natural account of why a successful prediction is more impressive than an accommodation: if H1 predicted E before E was known, while H2 was constructed after the fact to fit E, then P(E|H2) may be high, but the likelihood ratio may still favor H1 if H2 is more flexible or has more free parameters. The likelihood approach thus captures an intuition that the H-D model could not formalize.
However, the likelihood principle has its own limits. It can only compare hypotheses that are already specified; it cannot tell you how much confidence to place in a single hypothesis. It also requires that the likelihoods be well-defined, which is not always the case for complex scientific theories. And it has been criticized for ignoring the role of prior plausibility: a hypothesis with a very high likelihood may still be absurd if it was contrived to fit the data. Likelihoodists reply that prior plausibility is a matter of background theory, not evidence, and that the likelihood ratio is the correct measure of evidential support.
In 1954, Nelson Goodman posed a challenge that cut deeper than the old problem of induction. He invented a predicate "grue," which applies to things that are green before a certain time and blue after it. The hypothesis "all emeralds are grue" is just as well supported by past observations as "all emeralds are green," since all observed emeralds have been green and hence grue. Yet no one would project "grue" into the future. Goodman's point was that induction requires a distinction between "projectible" predicates like "green" and "non-projectible" ones like "grue," and that this distinction cannot be drawn from the data alone. It depends on which predicates are "entrenched" in our language and practice.
Goodman's riddle showed that confirmation is not purely a matter of logical form. Two hypotheses with identical logical structure and identical evidence can differ in confirmability because of the predicates they use. This undermined the logical empiricist hope of a purely formal confirmation theory. It also connected confirmation to the philosophy of language and to the social history of science: what counts as a natural kind or a projectible predicate is partly a matter of convention and practice. The riddle remains unresolved in the sense that no fully general criterion of projectibility has been accepted, though many philosophers have argued that the distinction is pragmatic rather than logical.
The current field is characterized by a plurality of approaches rather than a single dominant paradigm. Bayesianism is the most widely discussed framework, but it is not monolithic, and it coexists with likelihoodism, with formal learning theory, and with more naturalistic approaches that study how scientists actually reason.
Formal learning theory, developed by philosophers like Kevin Kelly and Clark Glymour, treats scientific inquiry as a problem of reliable belief revision. It asks: under what conditions can a method of hypothesis selection be guaranteed to converge to the truth in the limit, as evidence accumulates? This approach uses tools from computability theory and topology, and it has produced results about when underdetermination is unavoidable and when it can be overcome by strategic choices of hypotheses. It is less concerned with degrees of belief than with the reliability of methods, and it offers a different answer to the problem of induction: induction is justified not because it is deductively valid, but because some methods are demonstrably more reliable than others in the long run.
Another important strand is the "new experimentalism," associated with philosophers like Ian Hacking and Nancy Cartwright. This approach downplays the role of grand theories and emphasizes the local, material practices of experimentation. On this view, evidence is often generated by instruments whose reliability is established by calibration, not by theory. The confirmation of a hypothesis may depend less on its logical relation to data and more on the robustness of the experimental procedures that produced the data. This naturalistic turn has brought confirmation theory closer to the history and sociology of science, though it has also been criticized for abandoning the normative questions that define the field.
There is also a growing literature on "evidence amalgamation" and "model selection." In many sciences, especially biology, economics, and climate science, researchers combine evidence from multiple sources—observational studies, experiments, simulations—and choose among models with different numbers of parameters. The Akaike information criterion and the Bayesian information criterion are statistical tools for model selection that balance fit against complexity. Philosophers have asked whether these tools embody a coherent theory of evidence, and whether the preference for simpler models can be justified on evidential grounds or is merely a pragmatic convenience. This connects confirmation theory to the philosophy of statistics and to debates about scientific realism.
Despite the diversity of approaches, the field is unified by a set of enduring questions. The problem of induction remains the background condition for all work in the area: no approach has shown that induction is deductively valid, and none is likely to. What the field offers instead is a set of increasingly sophisticated ways to say what induction is—a matter of probability, of likelihood, of reliability, or of practice.
The distinction between prediction and accommodation remains central. A hypothesis that predicts new phenomena seems to earn more credit than one that merely fits existing data, and Bayesian and likelihoodist accounts both try to explain why. The Duhem–Quine problem remains a permanent feature of the landscape: no hypothesis is tested alone, and the rationality of theory choice must therefore be holistic. And the question of what makes a predicate or a hypothesis "projectible" remains open, connecting confirmation to the deepest issues in the philosophy of language and mind.
The field also continues to grapple with the tension between formal rigor and scientific practice. Formal models—Bayesian, likelihoodist, or learning-theoretic—offer clarity and precision, but they often idealize away the messiness of real science: vague hypotheses, unreliable instruments, social pressures, and changing background theories. Naturalistic approaches capture that messiness but risk losing the normative force that makes confirmation theory a philosophical discipline rather than a branch of sociology. The most productive work in the field tends to combine both: formal models that are tested against case studies, and case studies that are used to refine formal models.
Confirmation and evidence is thus not a settled body of doctrine but a living research programme. Its central insight is that the relationship between evidence and hypothesis is neither arbitrary nor purely deductive. It is a relationship that can be analyzed, formalized, and criticized, and that analysis has real consequences for how science is done and how its results are interpreted. The field's history shows a steady movement away from the hope of a single, universal logic of confirmation and toward a more pluralistic picture, in which different formal tools serve different purposes, and in which the rationality of science is understood as a complex achievement rather than a simple rule.