Biochemistry is the study of the chemical processes within and relating to living organisms. It sits at the intersection of biology and chemistry, seeking to explain life's phenomena—from the growth of a single cell to the function of a complex organ—through the structure, behavior, and interactions of biological molecules. The central question of biochemistry is deceptively simple: how do the molecules that constitute living systems assemble and operate to produce the properties we associate with being alive? This question encompasses the energy transformations that power life, the storage and expression of genetic information, the catalysis of metabolic reactions, and the intricate communication systems that coordinate cellular activities.
The field is historically and conceptually woven from two distinct traditions that have now merged into a unified discipline. The first, physiological or biological chemistry, grew out of 19th-century medicine and physiology. Its focus was on the chemistry of living material in the context of the whole organism: digestion, respiration, metabolism, and the composition of tissues. Researchers in this tradition asked how the food an animal eats becomes the substance of its body and the energy for its activity. They were concerned with pathways and processes, tracking the fate of molecules through the body. This tradition is the direct ancestor of modern metabolism research, enzymology, and much of nutrition science.
The second tradition, which became known as molecular biology, emerged in the mid-20th century from genetics, microbiology, and biophysics. Its focus was narrower and more structural: what is the physical nature of the gene, and how is information stored and expressed? This tradition asked how a linear sequence of chemical units in a molecule of DNA could specify the three-dimensional structure of a protein and, ultimately, the traits of an organism. The elucidation of the double-helical structure of DNA in 1953 is the emblematic achievement of this tradition, and its powerful explanatory framework—the central dogma, which describes the flow of genetic information from DNA to RNA to protein—became a cornerstone of all modern biology. The emphasis of this tradition is on information and its translation into molecular structure.
These two traditions were not isolated. The chemical description of metabolism built the stage for molecular biology by providing the tools for biochemical analysis—enzyme purification, spectroscopy, and chromatography—and the conceptual framework of macromolecular structure and function. Molecular biology, in turn, answered the deepest question of biological chemistry: how enzymes and structural proteins, so precisely tailored for their roles, come to be. The convergence of these traditions in the late 20th century created the modern field, where the study of a protein's function in a metabolic pathway is inseparable from an understanding of its gene, its synthesis, and its regulation. This unified discipline uses a shared core of methods and concepts, and its unifying theme is the relationship between macromolecular structure and biological function.
At its core, biochemistry is the chemistry of carbon-based molecules, but the reactivity of carbon is organized in living systems by a relatively small set of universal molecular themes. The principal classes of biomolecules are the amino acids and proteins, the nucleotides and nucleic acids, the carbohydrates, and the lipids. Each class has a characteristic chemistry that dictates its biological role.
Proteins are the most versatile biomolecules. They are linear polymers of amino acids, and the precise sequence of amino acids in a protein determines how the chain folds into a complex three-dimensional structure. This structure, in turn, creates precisely shaped pockets and surfaces that confer function. Proteins act as enzymes that accelerate biochemical reactions with extraordinary specificity, as structural components that give cells and tissues their shape, as transporters moving molecules across membranes, and as signaling molecules that transmit information between cells. Understanding how sequence dictates folding—and how misfolding can lead to disease—remains a central problem in the field.
Nucleic acids, the DNA and RNA, are the informational molecules of the cell. DNA stores the genetic blueprint as a sequence of four nucleotide bases—adenine, thymine, cytosine, and guanine—arranged in a double helix. The sequence of these bases carries the instructions for synthesizing proteins. RNA is the intermediate messenger, a single-stranded copy of the DNA sequence that is used as a template for protein synthesis. The flow of information from DNA to RNA to protein is the core of gene expression, and its regulation is how a single genome can give rise to hundreds of distinct cell types in a complex organism.
Carbohydrates serve two primary functions: as energy storage and transport molecules and as structural components. Simple sugars like glucose are the primary fuel for most cells, and their oxidation releases energy that is captured in a usable form. Carbohydrates can be polymerized into more complex structures, such as glycogen in animals and cellulose in plants, which serve as storage and structural materials, respectively. Lipids are a diverse group of water-insoluble molecules. They form the core of cellular membranes, creating the selectively permeable barriers that define cells and organelles. Lipids also serve as energy stores, as signaling molecules, and as pigments.
These molecular classes do not operate independently. The life of a cell is a dense web of interactions: proteins synthesize and degrade other proteins, lipids and carbohydrates are metabolized by enzymes, and nucleic acids encode the instructions for all of them. The complexity that emerges from this web of interactions is the central subject of the field.
The chemical reactions of life are collectively called metabolism, and they are organized into coordinated networks that achieve two overarching goals: to extract energy from the environment and to synthesize the molecules needed for cellular function. A foundational concept in this area is the role of adenosine triphosphate (ATP), which acts as the universal energy currency of the cell. The hydrolysis of ATP releases energy that can be coupled to energy-requiring reactions, such as the synthesis of macromolecules or the mechanical work of muscle contraction. Metabolism is divided into catabolism—the breakdown of complex molecules to release energy and smaller building blocks—and anabolism, the synthesis of complex molecules using energy from ATP and the reducing power of other coenzymes.
The study of metabolism has largely been a story of mapping pathways. The central pathways—glycolysis, the citric acid cycle, and oxidative phosphorylation—process carbohydrates to extract the maximum energy in the form of ATP. Gluconeogenesis, the pentose phosphate pathway, and fatty acid metabolism provide alternative routes for synthesis and breakdown. These pathways are not static but are highly regulated. Enzymes at key steps of a pathway are activated or inhibited by the concentrations of substrates, products, and allosteric effectors, allowing the network to respond dynamically to the cell's needs. The discipline of metabolic biochemistry has evolved from cataloging these pathways to understanding their control: what determines the flux through a pathway, and how are competing pathways balanced?
An essential dimension of metabolism is its thermodynamics. Life exists far from equilibrium, maintaining a highly ordered state by constantly dissipating energy from the environment. The study of bioenergetics examines how cells capture energy from sunlight or from the oxidation of chemical fuels and couple it to the synthesis of ATP. The chemiosmotic theory, which describes how a proton gradient across a membrane drives ATP synthesis, is a central explanatory framework here, uniting our understanding of respiration, photosynthesis, and many transport processes.
The second major branch of biochemistry concerns the storage, transmission, and expression of genetic information. The structure of DNA, its faithful replication, and its repair are fundamental to heredity. The processes of transcription, by which DNA is copied into messenger RNA, and translation, by which RNA directs the assembly of a protein from amino acids, are the core of gene expression.
This informational perspective quickly folds back into chemistry. Transcription and translation are carried out by large molecular machines—RNA polymerase and the ribosome—which are themselves complexes of proteins and nucleic acids. Their activity is controlled by regulatory proteins that bind to specific DNA sequences and either promote or block transcription. A deeper level of regulation occurs post-transcriptionally, where messenger RNA is processed, modified, and can be selectively degraded. Post-translational modifications of proteins, such as phosphorylation or methylation, can rapidly alter an enzyme's activity or its location within the cell.
The regulatory networks that control gene expression are staggeringly complex, and their study is a central focus of contemporary biochemistry. Epigenetic modifications—chemical changes to DNA or to the histone proteins it wraps around—alter gene expression without changing the underlying sequence, providing a molecular mechanism for cellular memory and for the stable inheritance of different cell fates. The field asks how a cell integrates multiple signals to make a decision, how signaling cascades amplify a weak external stimulus into a robust internal response, and how errors in these processes lead to diseases such as cancer.
Between the realms of metabolism and gene expression lies the task of understanding how the three-dimensional structure of a macromolecule gives rise to its function. This is the structural branch of biochemistry, deeply connected to the traditions of biophysics. The folding of a protein into its native structure is a spontaneous process driven by the hydrophobic effect, hydrogen bonding, and other noncovalent interactions, but it is often assisted in the cell by molecular chaperones. The structure determines the protein's function: a catalytic site with precisely positioned amino acids can stabilize a transition state, speeding a reaction by many orders of magnitude; a binding pocket complementary in shape and charge to a ligand can select it from a sea of similar molecules with exquisite specificity.
The methods of structural biochemistry—X-ray crystallography, nuclear magnetic resonance spectroscopy, and, increasingly, cryo-electron microscopy—allow researchers to determine atomic-resolution structures of proteins, nucleic acids, and their complexes. These methods have revealed not only static structures but also the dynamic conformational changes that are central to protein function. Many proteins act as molecular machines, undergoing large-scale movements that couple the energy of ATP hydrolysis to mechanical work or to the directional transport of other molecules.
The emergence of computational protein structure prediction has marked a qualitative shift in this area. Deep learning methods now allow the prediction of a protein's three-dimensional structure from its amino acid sequence with an accuracy that rivals experimental determination for many cases. This has transformed the field, enabling the modeling of entire proteomes and opening new avenues for drug design and for understanding the structural basis of disease mutations. Yet, the approach is not a replacement for experimental study; it provides hypotheses that need experimental validation, and it remains challenged by the prediction of protein dynamics, interactions, and the effects of post-translational modifications.
Contemporary biochemistry increasingly faces the challenge of synthesis: how to integrate the detailed knowledge of individual molecules and pathways into an understanding of the behavior of whole cells and organisms. This integrative impulse goes by several names—systems biology, functional genomics, and quantitative biology—but its core methods and concepts are firmly biochemical.
The advent of high-throughput technologies, such as DNA sequencing, mass spectrometry-based proteomics, and automated metabolite profiling, allows the measurement of thousands of molecular species simultaneously. These omics approaches generate vast datasets, and their interpretation requires computational tools and statistical models. The field has moved from asking "What is this protein doing?" to "How does this entire network of proteins and metabolites respond to a perturbation?" This shift of focus has revealed that biological systems are often robust, with redundant pathways and feedback loops that maintain function in the face of genetic or environmental perturbations. It has also highlighted the importance of dynamic behavior: concentrations, modifications, and interactions change over time, and understanding these dynamics is essential to understanding function.
The promise of this integrative biochemistry is a deeper understanding of disease. Many diseases, from diabetes to cancer, are not caused by a single molecular defect but by the disruption of complex networks. A biochemical understanding that spans from the structure of an individual enzyme to the behavior of an entire metabolic network provides the basis for rational drug design, for the development of biomarkers for early diagnosis, and for a predictive understanding of how an individual patient might respond to a given therapy. The field thus progresses in two directions simultaneously: toward greater molecular detail, as new techniques resolve ever finer structures and interactions, and toward greater systemic scope, as computational models begin to simulate more of the complexity of life. The task of biochemistry is to keep these two directions connected, ensuring that the detail informs the system and the system gives meaning to the detail.