Single cell genomics is the study of genomes, gene expression, and other molecular features one cell at a time. Rather than extracting DNA or RNA from a bulk tissue sample and measuring an average across millions of cells, researchers isolate individual cells, amplify their genetic material, and profile it with high-throughput sequencing or related technologies. The field’s central promise is resolution: it reveals the heterogeneity that bulk measurements obscure, allowing scientists to see distinct cell types, developmental states, rare mutants, and dynamic regulatory programs as they actually exist in tissues, tumors, and microbial communities.
Before single cell methods matured, almost all genomic analysis was performed on bulk samples. A typical experiment would grind up a piece of tissue, extract DNA or RNA from the pooled cells, and sequence the mixture. The resulting data represented an average over the entire population. For many questions this was sufficient, but it concealed fundamental biological facts. A tumor, for example, is not a uniform mass of identical cancer cells; it contains subclones with different mutations, immune cells infiltrating the tissue, and stromal cells supporting it. A bulk RNA-seq measurement of a tumor reports the mean expression across all these populations, which may not correspond to any actual cell in the sample. Similarly, a developing embryo contains cells that are diverging into distinct lineages; averaging their transcriptomes erases the very signals that define those lineages.
The conceptual shift of single cell genomics is to treat the individual cell as the fundamental unit of analysis. This is not merely a technical improvement but a change in the kinds of questions that can be asked. Instead of asking, “What is the average expression of gene X in this tissue?”, one can ask, “How many cells express gene X, and at what levels?” Instead of asking whether a mutation is present in a tumor, one can ask which cells carry it and which do not. This shift has made it possible to discover new cell types, trace developmental trajectories, and measure the stochastic variability that is intrinsic to gene expression.
The field rests on a set of molecular and computational techniques that were developed over roughly two decades. The core challenge is that a single cell contains only a tiny amount of nucleic acid—about 10 picograms of RNA, for example—far below what sequencing instruments require. The material must therefore be amplified before sequencing, and the amplification must be faithful enough to preserve the original information.
For DNA, the key technique is whole-genome amplification (WGA). Multiple displacement amplification, which uses a strand-displacing polymerase and random primers, is a common approach, but it introduces errors and uneven coverage. More recent methods use a combination of fragmentation, ligation, and PCR to amplify the genome more uniformly. Even with improvements, single cell DNA sequencing has higher error rates and more missing data than bulk sequencing, and researchers must account for these artifacts when calling mutations.
For RNA, the standard approach is to convert the cell’s mRNA into complementary DNA (cDNA), then amplify that cDNA. Early methods used PCR, which biases against long transcripts and can distort abundance measurements. The introduction of in vitro transcription and, later, of unique molecular identifiers (UMIs)—short random barcodes attached to each cDNA molecule before amplification—allowed researchers to correct for amplification bias by counting molecules rather than reads. The development of microfluidic platforms, which isolate individual cells in nanoliter droplets, made it possible to profile thousands or tens of thousands of cells in a single experiment, a dramatic scale increase over earlier plate-based methods that processed dozens or hundreds of cells.
The computational side of the field is equally essential. A single cell RNA-seq experiment produces a matrix of counts: rows are genes, columns are cells, and each entry is the number of transcripts detected for that gene in that cell. This matrix is sparse—most genes are not detected in most cells—and noisy, because the capture and amplification steps lose a large fraction of the original mRNA. Analyzing these data requires specialized statistical methods for normalization, dimensionality reduction, clustering, and trajectory inference. The field has developed its own algorithmic toolkit, including methods for embedding cells in low-dimensional space, identifying cell types by unsupervised clustering, and ordering cells along continuous processes such as differentiation.
Single cell genomics is not organized around rival schools in the way that, say, evolutionary biology has been. It is a methodologically driven field, and its internal structure is better described by the type of molecule being measured and the biological question being asked. Nevertheless, several distinct approaches have emerged, each with its own assumptions, strengths, and limitations.
Single cell RNA sequencing (scRNA-seq) is the most widely used approach. It measures the transcriptome of individual cells, providing a snapshot of which genes are active at the moment of lysis. The central assumption is that the mRNA content of a cell reflects its identity and state. This assumption is powerful but imperfect: mRNA levels do not always correspond to protein levels, and the act of dissociating a tissue into single cells can itself alter gene expression. scRNA-seq is used to catalog cell types in tissues, to identify rare populations, and to study how gene expression changes during development or disease. Its main limitation is that it measures only RNA, not the genome or the epigenetic state, and it destroys the cell in the process.
Single cell DNA sequencing addresses a different question: what mutations are present in each cell? This approach is used primarily in cancer research, where it can reveal the clonal architecture of a tumor—which mutations are shared by all cancer cells and which are present in only a subset. It can also be used to study somatic mosaicism in normal tissues, where mutations accumulate with age. The technical challenges are greater than for RNA, because the genome is much larger than the transcriptome and amplification errors can be mistaken for real mutations. Single cell DNA sequencing is therefore less mature than scRNA-seq, but it provides information that RNA cannot: the actual genetic alterations that drive disease.
Single cell epigenomics measures chemical modifications to DNA or chromatin that regulate gene expression without changing the underlying sequence. The most common targets are DNA methylation and chromatin accessibility. DNA methylation is a stable mark that is often associated with gene silencing; single cell methylation sequencing can reveal how epigenetic states vary across cells and how they change during differentiation. Chromatin accessibility, measured by methods such as ATAC-seq (assay for transposase-accessible chromatin), identifies regions of the genome that are open and potentially active. Single cell ATAC-seq can be used to infer regulatory elements and transcription factor activity in individual cells. These approaches are technically demanding and produce sparser data than scRNA-seq, but they add a crucial layer of information about the regulatory logic of the cell.
Multi-omics approaches attempt to measure several molecular layers from the same cell. For example, a method might capture both the transcriptome and the accessible chromatin from a single nucleus, or both the genome and the transcriptome from a single cell. The motivation is that these layers are deeply interconnected: a mutation can affect gene expression, and an epigenetic change can alter the response to a signal. Measuring them together allows researchers to link genotype to phenotype within individual cells. The cost is increased complexity and lower throughput; multi-omic methods typically profile fewer cells than single-layer approaches.
Spatial transcriptomics, while not strictly single cell, is closely related and often used in conjunction with single cell methods. It measures gene expression in tissue sections while retaining spatial information, so that each measurement is associated with a location in the tissue. Some spatial methods achieve single-cell resolution; others measure small groups of cells. The relationship to single cell genomics is complementary: scRNA-seq provides high-resolution molecular data without spatial context, while spatial methods provide the tissue architecture that single cell data lack. Combining the two, often through computational integration, allows researchers to map cell types back onto their native positions in the tissue.
The field’s methods are in service of several enduring biological questions. The most fundamental is the identification and characterization of cell types. For decades, cell types were defined by morphology, location, and a handful of marker genes. Single cell genomics has revealed that the number of distinct cell states is far larger than previously appreciated, and that the boundaries between types are often fuzzy. A major ongoing effort is to build comprehensive cell atlases—catalogs of all cell types in an organ or organism—using single cell data. These atlases are not just lists; they provide the molecular definitions that allow researchers to compare cell types across species, developmental stages, and disease states.
A second central question concerns development and differentiation. How does a single fertilized egg give rise to hundreds of distinct cell types? Single cell genomics allows researchers to capture cells at many points along a developmental trajectory and to infer the sequence of gene expression changes that occur as cells commit to a fate. This has led to the concept of the “trajectory” or “pseudotime”—an ordering of cells along a continuous process, inferred from their transcriptomes. Trajectory inference is a powerful tool, but it is important to recognize that it is an inference, not a direct observation; the ordering is a computational reconstruction from static snapshots.
A third question concerns cellular variability itself. Even within a seemingly homogeneous population, cells differ in their gene expression. Some of this variability is due to the stochastic nature of transcription—the random on-and-off switching of genes—and some is due to subtle differences in the cellular environment. Single cell genomics has made it possible to quantify this noise and to ask whether it is biologically meaningful or merely technical artifact. In some cases, variability is functional: it allows a population of cells to respond flexibly to changing conditions, or it underlies the decision of a cell to differentiate or remain stem-like.
A fourth question, particularly relevant to disease, is how mutations and gene expression changes are distributed across cells in a tissue. In cancer, this has led to the study of intratumor heterogeneity: the coexistence of multiple genetically distinct subclones within a single tumor. Single cell DNA and RNA sequencing have shown that this heterogeneity is common and that it has clinical implications, because different subclones may respond differently to therapy. In the brain, single cell genomics has been used to study the cellular composition of Alzheimer’s disease and other neurodegenerative conditions, revealing which cell types are most affected.
The origins of single cell genomics lie in earlier efforts to study individual cells using PCR and microarrays. In the 1990s and early 2000s, researchers developed methods to amplify the RNA of single cells and to measure the expression of a few dozen or a few hundred genes. These early studies were technically heroic but limited in scale. The field took its modern form with the advent of next-generation sequencing, which made it possible to measure the entire transcriptome of a single cell, and with the development of microfluidic devices that could process many cells in parallel.
A key milestone was the publication of the first single cell RNA-seq studies in 2009, which demonstrated that it was possible to sequence the full transcriptome of individual cells. Over the following years, the throughput of the technology increased by orders of magnitude. The introduction of droplet-based methods in 2015 made it possible to profile tens of thousands of cells in a single experiment, transforming the field from a specialized technique into a standard tool. The cost per cell fell correspondingly, and the methods became commercially available, allowing laboratories without deep microfluidics expertise to adopt them.
The development of the field has been driven by a close interplay between experimental and computational innovation. Each new experimental method has required new statistical tools to analyze its output, and each computational advance has suggested new experimental designs. The field has also been shaped by large collaborative projects, such as the Human Cell Atlas, which aim to create comprehensive reference maps of all human cells. These projects have pushed the technology toward higher throughput, lower cost, and better data quality, and they have established standards for data sharing and analysis.
The present state of single cell genomics is one of rapid expansion and consolidation. The technology is now widely used across biology and medicine, and it has become a standard tool in many laboratories. Commercial platforms offer turnkey solutions for scRNA-seq, and the computational tools for analyzing the resulting data are mature and well documented. The field has moved from proof-of-concept to routine application, and the focus has shifted from developing the basic methods to applying them to specific biological questions.
At the same time, important limitations remain. The most fundamental is that single cell methods destroy the cell, so they provide only a snapshot of its state at the moment of lysis. They cannot follow the same cell over time, although lineage tracing methods—which introduce heritable barcodes into cells and read them out later—can partially overcome this limitation. A second limitation is the loss of spatial context: dissociating a tissue into single cells destroys the architecture that is often crucial for understanding function. Spatial methods address this, but they are less mature and have lower resolution. A third limitation is the sparsity and noise of the data. Even the best methods capture only a fraction of the mRNA in a cell, and distinguishing true biological variability from technical dropout remains a challenge.
The field is also grappling with questions of scale and interpretation. As experiments profile millions of cells, the computational burden becomes substantial, and the statistical methods for comparing datasets across experiments and laboratories are still evolving. The interpretation of cell types is itself contested: there is no universally accepted definition of a cell type, and different clustering methods can produce different partitions of the same data. These are active areas of methodological research, not settled problems.
Single cell genomics has not replaced bulk genomics; rather, it has added a new dimension to genomic analysis. Bulk methods remain useful for many questions, particularly when the population is known to be homogeneous or when the goal is to measure an average effect. Single cell methods are most powerful when heterogeneity is the object of study. The two approaches are often combined: bulk sequencing provides depth and coverage, while single cell sequencing provides resolution and diversity. The field’s future likely lies in integrating these layers—genome, transcriptome, epigenome, and spatial position—into a unified picture of cellular function, one cell at a time.