Information science is the field that studies how human beings create, seek, organize, use, and share recorded knowledge and data. At its core, it asks a deceptively simple question: given that the world contains vastly more information than any person can absorb, how do we ensure that the right information reaches the right person at the right time, in a form they can actually use? The field is not primarily about building technology, though technology is central to its practice. It is about the relationship between information artifacts—books, databases, web pages, datasets—and the people who need them. Its practitioners design the systems, standards, and methods that make information findable, trustworthy, and usable, and they study how those systems succeed or fail in real human contexts.
The discipline is often confused with computer science, library science, or data science, and the confusion is understandable because information science draws on all three. The distinction lies in its focus. Computer science asks how to compute; information science asks how to represent, store, retrieve, and evaluate information for human use. Library science is a precursor and a sibling, focused on the institutions that curate physical and digital collections; information science generalizes those concerns to all information systems. Data science focuses on extracting patterns from data; information science focuses on the entire lifecycle of information, including the messy human decisions about what counts as information in the first place. The field's central commitment is that information is not a neutral resource. It is shaped by the people who produce it, the systems that organize it, and the users who interpret it, and each of those shaping forces is a legitimate object of study.
Information science is organized around a cluster of enduring problems rather than a single unified theory. The most fundamental is representation: how do you describe an information object so that someone who needs it can find it? A book can be described by its author, title, and subject; a dataset by its variables and collection method; a photograph by its visual content and provenance. Every representation is a lossy compression of the original object, and the field studies what to preserve, what to discard, and how to make the representation useful across different contexts. This problem has generated the field's most durable intellectual contributions, from classification schemes to metadata standards to the algorithms that power modern search engines.
The second problem is retrieval: given a representation of a need, how do you locate the information that satisfies it? This is not merely a technical problem of matching query terms to document terms. People rarely know exactly what they want, and they often cannot articulate their need in the language the system uses. The classic formulation, from the mid-twentieth century, distinguishes between a user's actual information need, the query they express, and the system's interpretation of that query. The gap between these three is the central difficulty of retrieval, and it explains why search is so much harder than it appears.
The third problem is organization: how do you arrange a collection so that related items are near each other, both conceptually and physically? Classification systems, taxonomies, and thesauri are the traditional answers, but the problem has been transformed by the web, where there is no single collection and no central authority. The field studies how formal organizational schemes, informal tagging, and algorithmic clustering each solve parts of the problem, and how they fail in different ways.
The fourth problem is evaluation: how do you know whether an information system is actually working? This requires defining what "working" means, which turns out to be deeply contested. Is a search engine good because it returns relevant results, because it returns them quickly, because it helps the user learn something they did not know they needed, or because it does not mislead them? The field has developed rigorous experimental methods for measuring some of these qualities, particularly relevance, but the harder questions about whether information systems genuinely improve human understanding remain open.
The fifth problem is human information behavior: what do people actually do when they need information, and why do they do it? This is the field's social-scientific wing, studying how people browse, ask colleagues, abandon searches, trust sources, and share findings. It has consistently shown that people are not rational information-seekers in the way that early system designers assumed. They satisfic—they take the first adequate answer rather than the best one—and they rely heavily on other people, not just on formal systems. Understanding this behavior is essential for designing systems that people will actually use.
Information science emerged as a distinct field in the mid-twentieth century, but its roots reach back to earlier efforts to manage the explosion of recorded knowledge. The nineteenth century saw the rise of modern librarianship, with its systematic classification schemes and cataloging rules, and the beginnings of scientific documentation—the practice of creating indexes, abstracts, and bibliographies for specialized research communities. These were responses to a real crisis: the number of scientific journals and books was growing faster than any individual could track, and scholars needed help finding what had already been published.
The field's modern identity crystallized after World War II, when the scale of scientific and military information became overwhelming. The term "information science" came into use in the 1950s and 1960s, and the field's early practitioners were often scientists and engineers who were frustrated with the limitations of traditional library methods. They believed that the problem of information overload could be solved by applying scientific methods and, increasingly, computers. This period produced the field's foundational concepts: the distinction between data, information, and knowledge; the idea of relevance as a measurable property of a retrieved document; and the first mathematical models of retrieval.
The 1960s and 1970s were the era of large-scale experimental retrieval research, funded heavily by government agencies. Researchers built test collections of documents and queries, ran different retrieval algorithms against them, and measured which performed best. This tradition, known as the Cranfield paradigm after the British experiments that established it, remains the dominant method for evaluating search systems. It has been enormously influential, but it has also been criticized for reducing information retrieval to a laboratory exercise that ignores the messy realities of real users.
The 1980s brought a shift toward the user. Researchers in library and information science began studying how people actually search for information, rather than how they ought to. This produced a rich body of qualitative research on information seeking, and it also produced a critique of the field's technological optimism. The argument was that information systems were being designed for idealized users who did not exist, and that the field needed to start from the user's perspective rather than the system's.
The 1990s transformed the field again with the rise of the web. The web was not designed by information scientists, but it posed exactly the problems the field had been studying for decades: how to organize, retrieve, and evaluate information at a scale no one had imagined. The field responded by developing web-specific methods, and it also absorbed insights from computer science, particularly from the emerging field of machine learning. The result was a hybrid discipline that is comfortable with both the formal rigor of algorithmic evaluation and the interpretive complexity of human behavior.
The field is best understood not as a single tradition but as a set of approaches that coexist, overlap, and sometimes conflict. The most important division is between the system-centered and user-centered traditions. The system-centered approach, which dominated the field's early decades, treats information retrieval as an engineering problem. Its goal is to build systems that match queries to documents as accurately as possible, and its method is controlled experimentation. It assumes that relevance is a property of the document-query pair, and that the system's job is to rank documents by their likely relevance. This approach has produced the field's most rigorous methods and its most impressive practical achievements, including the algorithms that power modern search engines. Its limitation is that it treats the user as a source of queries rather than as a person with a context, a history, and a purpose.
The user-centered approach, which gained prominence in the 1980s, starts from the opposite assumption: that relevance is not a property of documents but a judgment made by a person in a particular situation. Its methods are qualitative—interviews, observation, diary studies—and its goal is to understand the full context of information seeking, including the emotional, social, and cognitive factors that shape it. This approach has produced rich descriptions of how different communities seek and use information, and it has influenced system design by showing that users need support for browsing, exploration, and serendipity, not just for precise querying. Its limitation is that its findings are often difficult to translate into concrete design recommendations, and its methods are less rigorous by the standards of experimental science.
A third approach, sometimes called the cognitive tradition, tries to bridge the gap between system and user. It treats information seeking as a process of knowledge construction, in which the user's mental model of their problem changes as they encounter new information. The system's job is not simply to return relevant documents but to support the user's evolving understanding. This approach has influenced the design of interactive search systems and has produced influential models of the search process, but it has been criticized for being too abstract and for not producing testable predictions.
A fourth approach is the sociotechnical or critical tradition, which examines information systems as social and political artifacts. It asks who benefits from a particular way of organizing information, whose knowledge is included and whose is excluded, and how information systems reinforce or challenge existing power structures. This tradition has grown in importance with the recognition that search engines, recommendation algorithms, and databases are not neutral tools but are shaped by the values of their creators and the data they are trained on. It has produced important critiques of algorithmic bias and has pushed the field to consider the ethical dimensions of its work.
These approaches are not sequential stages in a linear progression. They coexist in the field today, and many researchers combine them. A modern information scientist might use the experimental methods of the system-centered tradition to evaluate a new retrieval algorithm, the qualitative methods of the user-centered tradition to understand why users struggle with it, and the critical lens of the sociotechnical tradition to examine what the algorithm's training data excludes. The field's strength lies in this methodological pluralism, even though it sometimes makes the field feel fragmented.
The present landscape of information science is shaped by several durable features. The first is the continuing centrality of the web and of large-scale digital information systems. Search engines, social media platforms, and recommendation systems are now the primary means by which most people access information, and they are the field's most visible objects of study. The questions the field asks about these systems are the same questions it has always asked—How should information be represented? How should it be retrieved? How do we know if the system is working?—but the answers are complicated by scale, opacity, and commercial interest.
The second durable feature is the rise of data as a category of information. The field has always dealt with data, but the term now carries new weight, referring to the massive datasets collected by governments, corporations, and scientific projects. Information science contributes to the study of data through its work on metadata, data curation, and data quality, and it asks questions that data science often ignores: Who owns the data? Who can access it? How do we preserve it for future use? How do we ensure that it is not misinterpreted?
The third durable feature is the field's engagement with the social and ethical dimensions of information. The recognition that information systems are not neutral has moved from the margins to the center of the field. Researchers study algorithmic bias, misinformation, filter bubbles, and the digital divide, and they work on methods for auditing and improving information systems. This work is often interdisciplinary, drawing on sociology, political science, and philosophy, but it is grounded in the field's core concern with the relationship between information and human well-being.
The fourth durable feature is the field's institutional home. Information science is taught in university departments that often also house library science, and many of its practitioners work in libraries, archives, and museums. These institutions have been transformed by digital technology, but they remain important sites for the field's research and practice. The field's connection to libraries gives it a practical orientation and a commitment to values—access, preservation, intellectual freedom—that are not always shared by the commercial technology sector.
The field's boundaries remain porous, and that is a source of both strength and tension. Information science overlaps with computer science, communication studies, management, and the digital humanities, and it is not always clear where one field ends and another begins. Some information scientists worry that the field lacks a distinct identity; others argue that its identity lies precisely in its willingness to cross boundaries and to ask questions that no single discipline can answer. What is not in dispute is the importance of the questions. As the volume of information continues to grow, and as the systems that mediate it become more powerful and more opaque, the need for a field that studies how information and people interact has never been greater.