Functional genomics is the subfield of genomics that aims to understand the relationship between an organism's genome—its complete set of DNA sequences—and the dynamic expression of that information into observable traits. Where structural genomics focuses on determining the sequence and physical arrangement of genes, functional genomics asks what those genes actually do, when and where they are turned on or off, and how their products interact to produce a living cell or organism. The field is defined less by a single technique than by a shared commitment to moving from static sequence data to a dynamic, systems-level picture of biological function.
The foundational problem of functional genomics is that a genome sequence alone does not explain an organism. Knowing the order of nucleotides in DNA tells you the potential repertoire of genes, but not which genes are active in a particular cell type, under particular conditions, or at a particular developmental stage. The central questions of the field therefore cluster around gene expression and its consequences: Which genes are transcribed into RNA, and at what levels? How does that transcription vary across tissues, time, or environmental perturbations? Which RNA molecules are translated into proteins, and what post-translational modifications alter their activity? How do proteins and other molecules physically interact to form the regulatory and metabolic networks that sustain life?
The stakes are both fundamental and practical. Functionally, the genome is not a static blueprint but a highly regulated system that responds to internal and external signals. Misregulation of gene expression underlies many diseases, from cancer—where oncogenes become inappropriately active and tumor suppressors silenced—to developmental disorders caused by mutations in regulatory regions rather than in protein-coding sequences. Functional genomics provides tools to identify which genes matter for a given phenotype, to map the regulatory circuits that control them, and to predict how perturbations—genetic, pharmacological, or environmental—will propagate through the system. This makes the field central to drug target discovery, personalized medicine, and the basic understanding of how genotype gives rise to phenotype.
Functional genomics emerged in the late 1990s as a direct response to the successes and limitations of genome sequencing. The Human Genome Project and parallel efforts in model organisms produced complete reference sequences, but these immediately raised the question of what to do with them. Early sequencing efforts had already revealed that the number of genes in a genome did not correlate with organismal complexity—humans have roughly the same number of protein-coding genes as much simpler organisms—suggesting that regulation, splicing, and post-transcriptional control were at least as important as gene count.
The term "functional genomics" was coined to describe a new kind of biology that would systematically assign functions to the thousands of uncharacterized genes that sequencing had revealed. The field's early development was driven by two technological revolutions. The first was the microarray, which allowed researchers to measure the expression levels of thousands of genes simultaneously by hybridizing labeled RNA to arrays of DNA probes. For the first time, one could ask which genes were up- or down-regulated across entire transcriptomes under different conditions. The second was high-throughput sequencing itself, which, when applied to RNA (RNA-seq), replaced microarrays with a more sensitive and comprehensive method for quantifying transcripts.
A third major input came from classical genetics, but at genome scale. Instead of studying one mutant at a time, functional genomics borrowed the logic of genetic screens and scaled them up. Techniques such as RNA interference (RNAi) and later CRISPR-Cas9 gene editing allowed researchers to systematically knock down or knock out every gene in a genome, one at a time or in pools, and observe the phenotypic consequences. This transformed functional genomics from a descriptive discipline—measuring what genes do under natural conditions—into an interventionist one—perturbing genes to learn what they are required for.
Functional genomics is not organized around a single paradigm but rather around a set of complementary approaches that address different aspects of gene function. These approaches are best understood as a division of labor, with each answering a distinct type of question and the results feeding into integrated models.
The most mature and widely used approach is transcriptomics, which aims to catalog and quantify all RNA molecules in a cell or tissue. The central tool is RNA-seq, which converts RNA to complementary DNA, sequences it in high throughput, and aligns the resulting reads to a reference genome to determine which genes are expressed and at what levels. Transcriptomics reveals the "expression profile" of a cell—its active gene set at a moment in time. This is the entry point for most functional genomics studies because it is relatively straightforward and provides a global readout of cellular state.
Transcriptomics has several important variants. Differential expression analysis compares two or more conditions to identify genes whose expression changes, for example between healthy and diseased tissue. Single-cell RNA-seq extends this to individual cells, revealing the heterogeneity of cell types within a tissue and allowing researchers to reconstruct developmental trajectories or identify rare cell populations. A key limitation is that RNA levels do not always correlate with protein levels, because of post-transcriptional regulation, so transcriptomics alone cannot fully describe function.
A second major approach asks how gene expression is controlled. Epigenomics maps the chemical modifications to DNA and histones—the proteins around which DNA is wound—that influence whether genes are accessible to the transcriptional machinery. Techniques such as chromatin immunoprecipitation followed by sequencing (ChIP-seq) identify where specific transcription factors or histone modifications are located across the genome. ATAC-seq (assay for transposase-accessible chromatin) maps regions of open chromatin, indicating where regulatory proteins can bind. These methods reveal the regulatory landscape: the promoters, enhancers, and silencers that control when and where genes are expressed.
Regulatory genomics also includes the study of non-coding RNAs, such as microRNAs and long non-coding RNAs, which modulate gene expression post-transcriptionally. The relationship between epigenomics and transcriptomics is intimate: epigenetic marks are often the cause of expression differences, but they are also dynamic and responsive to environmental signals, so the causal direction is not always clear. This approach has been particularly important for understanding development, where cell fate decisions are driven by changes in chromatin state, and cancer, where epigenetic dysregulation is common.
A third approach moves beyond RNA to the proteins themselves. Proteomics aims to identify and quantify the full complement of proteins in a cell, including their post-translational modifications, which often determine their activity, localization, and interactions. Mass spectrometry is the central technology, allowing researchers to identify thousands of proteins in a single experiment. Interactomics goes one step further, mapping the physical interactions between proteins—the "interactome"—using methods such as yeast two-hybrid screens or affinity purification followed by mass spectrometry.
These approaches address a different layer of function than transcriptomics. A gene may be transcribed but not translated, or translated but rapidly degraded. Proteins may be inactive until modified, or may function only when bound to partners. Proteomics and interactomics therefore provide a more direct readout of cellular activity, but they are technically more challenging and historically less comprehensive than transcriptomics. The relationship between the layers is not one-to-one: the correlation between mRNA and protein abundance is often modest, and the field has devoted considerable effort to understanding the sources of this discrepancy.
A fourth approach is distinguished not by what it measures but by how it intervenes. Functional genetic screens systematically perturb genes and observe the phenotypic consequences. The modern workhorse is CRISPR-Cas9, which can introduce targeted mutations into essentially any gene. In a typical screen, a library of guide RNAs targeting thousands of genes is introduced into a cell population; cells are then subjected to a selective condition—such as a drug, a pathogen, or a nutrient deprivation—and the guide RNAs that are depleted or enriched reveal which genes are required for survival or resistance.
This approach provides causal information that measurement-based methods cannot. Transcriptomics can show that a gene is upregulated in disease, but only a perturbation experiment can show whether that upregulation is causally important or merely correlated. Screens can be performed in cell lines, in primary cells, or in vivo, and can be combined with other approaches—for example, screening for genes that, when knocked out, alter the expression of a reporter gene, thereby linking perturbation to regulatory function. The main limitation is that screens are typically binary or quantitative readouts of a single phenotype, and they may miss genes with redundant functions or those required only under specific conditions.
The final major approach is integrative: combining data from multiple layers—genome sequence, transcriptome, epigenome, proteome, and perturbation screens—to build computational models of gene regulatory networks and cellular pathways. This is often called systems biology, and it is the natural endpoint of functional genomics. The goal is not just to list parts but to understand how they work together. For example, one might integrate ChIP-seq data for a transcription factor with RNA-seq data after its knockdown to infer its target genes, then add protein interaction data to place those targets in a broader network.
These integrative approaches face significant challenges. Data from different technologies have different noise characteristics, biases, and scales. The relationship between layers is complex and often non-linear. Computational models must therefore make simplifying assumptions, and their predictions require experimental validation. Nevertheless, the field has made substantial progress in building predictive models of cellular behavior, particularly in model organisms such as yeast and in well-characterized mammalian cell lines.
The current landscape of functional genomics is characterized by several durable trends. First, the field is increasingly single-cell. Technologies that measure the transcriptome, epigenome, or proteome of individual cells have revealed that cellular heterogeneity is the rule rather than the exception, and that population-level averages can obscure important biology. Single-cell methods are now being combined—for example, measuring RNA and chromatin accessibility in the same cell—to build multi-modal atlases of developing tissues and tumors.
Second, the field is becoming more causal and more systematic. CRISPR screens have moved from cell lines to primary cells and in vivo models, and are being combined with single-cell readouts to link perturbations to molecular phenotypes at scale. The development of base editing and prime editing has expanded the types of genetic changes that can be introduced, allowing more subtle questions about the effects of specific mutations.
Third, the field is increasingly integrated with human genetics. Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with disease, but most of these variants lie in non-coding regions and their functional consequences are unknown. Functional genomics provides the tools to interpret these variants: by mapping regulatory elements, measuring allele-specific expression, and perturbing candidate regions, researchers can connect statistical associations to molecular mechanisms. This has given rise to a subfield sometimes called "functional genomics of disease" or "variant-to-function" studies.
Fourth, the field is becoming more quantitative and more computational. The cost of sequencing has fallen dramatically, but the cost of analyzing the resulting data has not. Machine learning methods are increasingly used to predict the functional consequences of genetic variants, to integrate heterogeneous data types, and to build models of gene regulation. These models are powerful but require careful validation, and the field has developed rigorous standards for benchmarking and reproducibility.
Finally, the field retains a fundamental tension between comprehensiveness and depth. Genome-wide methods provide breadth but often at the cost of resolution or accuracy. A researcher may identify hundreds of candidate genes from a screen, but validating each one requires focused, low-throughput experiments. The field's progress has therefore been iterative: genome-wide discovery followed by targeted validation, with the validated results feeding back into improved models and better genome-wide tools.
Functional genomics is best understood not as a single method or theory but as a research programme organized around a central commitment: that the function of a genome must be studied as a dynamic, integrated system, not as a static list of parts. Its approaches are complementary, each revealing a different layer of the complex relationship between genotype and phenotype. The field's enduring contribution has been to transform genomics from a descriptive science of sequences into an experimental science of function, and in doing so, to provide the conceptual and technical tools for understanding how genomes actually work.