Knowledge organization is the field of information science concerned with how human knowledge is described, structured, classified, and represented so that it can be reliably found, understood, and used. It is not the study of how people think or learn—that belongs to cognitive science—but rather the study of the systems, standards, and conceptual schemes through which recorded knowledge is arranged for retrieval. At its center is a persistent tension: the world of ideas is fluid, overlapping, and constantly changing, while the systems we build to manage it require stability, boundaries, and fixed categories.
The fundamental problem of knowledge organization is that documents, data, and other information objects do not come with their own labels. A book about the history of medicine, a dataset of clinical trials, and a photograph of a hospital ward might all be relevant to a single inquiry, but they share no inherent property that makes them co-located. Someone must decide what each item is about, how that subject relates to other subjects, and how the item should be represented in a system that allows others to find it later. Every act of cataloging, indexing, and classification is therefore an act of interpretation, not merely a clerical task.
This interpretive dimension raises the field's central questions. What is a subject, and can it be identified objectively? Should an item be described by its topic, its form, its intended audience, or its function? How should a system accommodate new knowledge that does not fit existing categories? Who decides which categories matter, and whose knowledge is privileged or excluded in that decision? These questions are not merely theoretical; they shape what any library catalog, database, or search engine can retrieve, and therefore what any researcher, student, or citizen can know.
Long before information science existed as a discipline, libraries and scholars developed practices for arranging knowledge. The great library of Alexandria in the Hellenistic period organized scrolls by broad subject categories, and medieval European libraries often arranged books by the faculties of the university—theology, law, medicine, and arts. In the Islamic world, bibliographers such as Ibn al-Nadim in the tenth century produced classified catalogs of known writings. These were practical schemes, built for specific collections, not general theories of knowledge.
The philosophical tradition of classifying knowledge itself—rather than books—provided a different foundation. Aristotle's division of the sciences, Francis Bacon's seventeenth-century map of human understanding (memory, imagination, and reason), and the Enlightenment encyclopedists' attempts to chart all human knowledge all influenced later library classifications. These schemes were not knowledge organization systems in the modern sense; they were philosophical arguments about the structure of knowledge, and their connection to the later field is one of inheritance rather than direct practice. When nineteenth-century librarians began building general classification systems, they borrowed the idea that knowledge could be mapped, but they adapted it to the practical demands of shelving and retrieving physical objects.
Knowledge organization became a distinct professional and scholarly concern in the late nineteenth century, when the growth of published material made older, local arrangements untenable. The two most influential figures of this period were Melvil Dewey and Paul Otlet. Dewey's Decimal Classification, first published in 1876, divided all knowledge into ten main classes, each subdivided by ten, and so on. It was designed for the practical arrangement of books on shelves and became the dominant system in American public and school libraries. Otlet, working in Belgium, pursued a more ambitious vision. With Henri La Fontaine, he created the Universal Decimal Classification, an expansion of Dewey's scheme that used punctuation and symbols to express complex, multi-faceted subjects. Otlet also imagined a "mundaneum"—a universal repository of recorded knowledge—and wrote extensively about the documentation of all human intellectual output. His work anticipated many concerns of later information science, though it was largely neglected for decades and only rediscovered in the late twentieth century.
The early twentieth century also saw the rise of subject cataloging as a distinct practice. The Library of Congress Subject Headings, developed from the 1890s onward, provided a controlled vocabulary for describing what books were about, independent of where they were shelved. This separation of classification (where an item sits) from subject description (what an item is about) became a foundational distinction. Classification organizes knowledge into a hierarchical structure; subject headings provide a flexible, alphabetical access point. Both are forms of knowledge organization, but they serve different functions and rest on different assumptions.
The most significant theoretical development in the field came from the Indian librarian S. R. Ranganathan, whose work from the 1930s onward challenged the dominance of enumerative classification. Enumerative systems, like Dewey's, attempt to list all possible subjects in advance. Ranganathan argued that this was impossible: knowledge grows, and any fixed list will eventually fail. His Colon Classification introduced the idea of faceted analysis, in which a subject is not a single category but a combination of fundamental categories—personality, matter, energy, space, and time—that can be combined according to rules. A book about the treatment of tuberculosis in rural India, for example, would be analyzed into its facets and synthesized into a notation expressing each component.
Faceted classification was not merely a technical improvement; it represented a different philosophy of knowledge. Enumerative systems treat subjects as pre-existing entities to be discovered and placed. Faceted systems treat subjects as constructions, built from more basic conceptual building blocks according to the needs of the user. This shift had profound implications. It made classification more flexible and more expressive, but it also made it more complex and more dependent on the judgment of the classifier. Ranganathan's work remained marginal in practice for many years—the Colon Classification was never widely adopted—but his theoretical writings, especially the Prolegomena to Library Classification, became a touchstone for later theorists. The faceted approach was later taken up and refined by the British Classification Research Group in the 1950s and 1960s, and it now underpins many modern information architectures, including the design of databases and websites, even where the original notation has been abandoned.
Beginning in the 1970s and 1980s, a group of researchers, many associated with the Danish librarian and information scientist Birger Hjørland, argued that knowledge organization had paid too much attention to the structure of documents and too little to the people who use them. This "cognitive" or "domain-analytic" approach insisted that the meaning of a subject is not fixed by the text itself but is determined by the discourse community that produces and uses it. A book about "depression" means something different in psychiatry, in economics, and in meteorology. A classification system that ignores these differences will fail its users, because it will lump together items that belong to different worlds of practice.
This perspective reframed the central questions of the field. Instead of asking "What is the true structure of knowledge?" it asked "What structures serve the purposes of particular communities?" It also drew attention to the politics of classification. Systems like the Dewey Decimal Classification and the Library of Congress Subject Headings were created in specific cultural contexts and carry the assumptions of those contexts. The historian of classification Hope Olson and others documented how these systems marginalized women, non-Western peoples, and queer identities, either by omitting them, by placing them under headings that implied deviance, or by subordinating them to dominant categories. This critical work did not reject classification as such; it argued that classification is always a form of power and that responsible knowledge organization requires awareness of that power.
The cognitive turn also connected knowledge organization to the broader field of information retrieval. Researchers began to study how users actually search for information and how their search terms relate to the controlled vocabularies of indexing systems. This work revealed a persistent gap between the language of documents and the language of users, and it motivated efforts to build systems that could bridge that gap through thesauri, semantic networks, and, later, automated methods.
The rise of digital computing transformed knowledge organization in ways that are still being assimilated. Early online catalogs and databases simply automated existing manual practices, but the web created a new environment in which the traditional gatekeepers of knowledge organization—librarians and indexers—were bypassed. Anyone could publish, and anyone could describe their own content. The result was a proliferation of metadata, much of it inconsistent, incomplete, or misleading.
The response from the knowledge organization community was twofold. One strand, associated with the semantic web movement, sought to create formal, machine-readable structures for describing knowledge. The Resource Description Framework (RDF) and the Web Ontology Language (OWL) allow the expression of statements about resources—"this document is about X," "X is a subclass of Y"—in a form that computers can process. Ontologies, in this context, are formal specifications of a shared conceptualization: they define the classes, properties, and relationships that are relevant to a domain. This approach inherits the ambitions of Otlet and Ranganathan, but it replaces human-readable classification with logic-based formalization. Its promise is that machines, not just humans, can reason over organized knowledge. Its limitation is that building and maintaining ontologies is difficult and expensive, and the formal rigor required can be brittle in the face of the messy, evolving nature of real knowledge.
The other strand, more pragmatic, embraced the statistical and algorithmic methods of modern information retrieval. Search engines like Google do not classify the web into subjects; they index the text of pages and use ranking algorithms to determine relevance. From the perspective of traditional knowledge organization, this looks like a retreat from the field's core mission. But many researchers argue that the two approaches are complementary rather than opposed. Controlled vocabularies and classification structures can improve search by providing synonyms, broader and narrower terms, and disambiguation. Conversely, statistical methods can reveal patterns of usage that suggest how a classification should be revised. The contemporary field is characterized less by a single dominant paradigm than by a productive, sometimes uneasy, coexistence of these approaches.
Contemporary knowledge organization is a broad and pluralistic field. Its traditional heart remains in libraries, where catalogers and metadata specialists apply standards like the Resource Description and Access (RDA) and the Dewey Decimal or Library of Congress classifications. But the same skills are now applied in museums, archives, scientific data repositories, and corporate information systems. The rise of linked data has pushed libraries to expose their catalogs as machine-readable graphs, connecting their records to those of other institutions and to external datasets. This has revived interest in the theoretical foundations of the field, since linked data requires explicit decisions about what entities exist, how they are named, and how they relate.
The field's enduring questions remain unresolved. Is there a stable structure to knowledge, or is all organization contingent on purpose and perspective? Can a classification be both universal and just, or does universality inevitably impose the categories of the powerful? How should systems accommodate change without losing the stability that makes retrieval possible? These questions have no final answers, and the field's vitality lies in the ongoing attempt to address them. What distinguishes knowledge organization from mere information management is precisely this theoretical self-awareness: the recognition that every system embodies choices, and that those choices have consequences for what can be known and by whom.