RNA sequencing experiments depend on collecting enough sequencing data to reliably measure the RNA molecules present in a sample. Sequencing too little can cause researchers to miss important molecules, while sequencing more than necessary increases costs without providing equivalent gains in useful information.
Researchers from the University of Toronto have developed a modeling framework called NB-Lib to help determine how much sequencing is needed for experiments that use unique molecular identifiers, or UMIs.
Why sequencing depth matters
In many single-cell and spatial RNA sequencing experiments, researchers attach a short sequence called a UMI to individual RNA molecules. These molecular barcodes help distinguish original RNA molecules from copies generated during amplification.
Researchers commonly use measurements such as reads per cell and sequencing saturation to decide whether an RNA sequencing library has been sequenced deeply enough.
Sequencing saturation describes how often additional sequencing reads represent molecules that have already been detected. A highly saturated library might appear to suggest that most available molecules have already been recovered.
However, the researchers found that saturation alone does not provide a complete picture of molecular recovery.
Sequencing saturation and molecular recovery are not the same
One reason is amplification heterogeneity. During library preparation, RNA-derived molecules are amplified to produce enough material for sequencing, but individual molecules may not be amplified equally.
Some molecules can produce many sequencing reads, while others produce relatively few. As a result, two libraries with similar sequencing saturation can contain different numbers of recoverable molecules.
This distinction is important when deciding whether additional sequencing is worthwhile. What researchers ultimately want to know is how many additional original molecules, and therefore how much additional biological information, can be recovered by sequencing more deeply.
Modeling RNA sequencing libraries with NB-Lib
To address this problem, the researchers developed NB-Lib, a statistical modeling framework that estimates both amplification heterogeneity and library complexity.
Library complexity describes the number and diversity of unique molecules available for sequencing. By considering complexity together with differences in amplification, NB-Lib models the relationship among sequencing depth, sequencing saturation, and molecular recovery.
The researchers evaluated the framework using 150 single-cell and spatial transcriptomic datasets representing different experimental platforms.
NB-Lib accurately reconstructed sequencing saturation curves and predicted how much sequencing would be required to achieve different levels of molecular recovery.
Amplification heterogeneity controls the relationship between saturation and recovery
(A) ZT-NB probability mass functions of reads per molecule for a fixed saturation (70%). Curves correspond to different values of the dispersion/amplification heterogeneity parameter as indicated by the legend. (B) Recovery versus saturation expressed in percentages for a fixed library size . Each curve corresponds to a different amplification heterogeneity value () as indicated by the legend. Points indicate 30% and 50% saturation values. The special case where r = 1 is indicated with a dashed line, where the NB-Lib function reduces to a Michaelis-Menten form and recovery is equal to saturation. (C) Predicted amplification heterogeneity () from the ZT-NB model across the full cohort of samples. Samples are grouped by dataset and colored by library type as indicated in the legend. Vertical lines denote dataset boundaries and the horizontal lines indicate .
Predicting sequencing needs from pilot experiments
A particularly useful feature of NB-Lib is its ability to make predictions from shallow pilot sequencing.
Instead of deeply sequencing an entire library before determining whether the chosen depth was appropriate, researchers can generate a smaller amount of sequencing data and use NB-Lib to estimate how additional sequencing is likely to affect molecular recovery.
This could help researchers plan sequencing depth based on the characteristics of an individual library rather than relying primarily on general recommendations such as a fixed number of reads per cell.
The analysis also showed that differences in amplification heterogeneity and library complexity can explain substantial variation in sequencing efficiency and cost among samples and transcriptomic technologies.
Molecular recovery affects detection of low-abundance genes
Sequencing depth becomes particularly important when researchers are interested in genes expressed at low levels.
The researchers found that molecular recovery directly influenced the reproducibility of low-abundance gene detection. When too few molecules are recovered, weakly expressed genes may be detected inconsistently across measurements.
This means that selecting an appropriate sequencing depth can affect more than cost and overall data quantity. It can influence whether biologically important but uncommon transcripts are reliably observed.
Improving RNA sequencing experiment planning
NB-Lib provides a more detailed way to evaluate sequencing requirements by connecting sequencing depth with the molecules researchers actually recover from their libraries.
Rather than assuming that sequencing saturation alone indicates whether an experiment has been sequenced deeply enough, the framework accounts for library complexity and amplification differences that can alter molecular recovery.
For single-cell and spatial RNA sequencing experiments, this approach could help researchers make more informed decisions about sequencing depth, improve detection of low-abundance genes, and allocate sequencing resources more efficiently.
Availability – The scdepth package is available at https://github.com/gwlab-ca/scdepth.
Wilson G, N Kalimuthu S, Yeung J. (2026) Sequencing saturation does not uniquely determine molecular recovery in UMI transcriptomics. Bioinformatics 42(8): btag593. [article]
RNA sequencing experiments depend on collecting enough sequencing data to reliably measure the RNA molecules present in a sample. Sequencing too little can cause researchers to miss important molecules, while sequencing more than necessary increases costs without providing equivalent gains in useful information.
Researchers from the University of Toronto have developed a modeling framework called NB-Lib to help determine how much sequencing is needed for experiments that use unique molecular identifiers, or UMIs.
Why sequencing depth matters
In many single-cell and spatial RNA sequencing experiments, researchers attach a short sequence called a UMI to individual RNA molecules. These molecular barcodes help distinguish original RNA molecules from copies generated during amplification.
Researchers commonly use measurements such as reads per cell and sequencing saturation to decide whether an RNA sequencing library has been sequenced deeply enough.
Sequencing saturation describes how often additional sequencing reads represent molecules that have already been detected. A highly saturated library might appear to suggest that most available molecules have already been recovered.
However, the researchers found that saturation alone does not provide a complete picture of molecular recovery.
Sequencing saturation and molecular recovery are not the same
One reason is amplification heterogeneity. During library preparation, RNA-derived molecules are amplified to produce enough material for sequencing, but individual molecules may not be amplified equally.
Some molecules can produce many sequencing reads, while others produce relatively few. As a result, two libraries with similar sequencing saturation can contain different numbers of recoverable molecules.
This distinction is important when deciding whether additional sequencing is worthwhile. What researchers ultimately want to know is how many additional original molecules, and therefore how much additional biological information, can be recovered by sequencing more deeply.
Modeling RNA sequencing libraries with NB-Lib
To address this problem, the researchers developed NB-Lib, a statistical modeling framework that estimates both amplification heterogeneity and library complexity.
Library complexity describes the number and diversity of unique molecules available for sequencing. By considering complexity together with differences in amplification, NB-Lib models the relationship among sequencing depth, sequencing saturation, and molecular recovery.
The researchers evaluated the framework using 150 single-cell and spatial transcriptomic datasets representing different experimental platforms.
NB-Lib accurately reconstructed sequencing saturation curves and predicted how much sequencing would be required to achieve different levels of molecular recovery.
Amplification heterogeneity controls the relationship between saturation and recovery
(A) ZT-NB probability mass functions of reads per molecule for a fixed saturation (70%). Curves correspond to different values of the dispersion/amplification heterogeneity parameter as indicated by the legend. (B) Recovery versus saturation expressed in percentages for a fixed library size . Each curve corresponds to a different amplification heterogeneity value () as indicated by the legend. Points indicate 30% and 50% saturation values. The special case where r = 1 is indicated with a dashed line, where the NB-Lib function reduces to a Michaelis-Menten form and recovery is equal to saturation. (C) Predicted amplification heterogeneity () from the ZT-NB model across the full cohort of samples. Samples are grouped by dataset and colored by library type as indicated in the legend. Vertical lines denote dataset boundaries and the horizontal lines indicate .
Predicting sequencing needs from pilot experiments
A particularly useful feature of NB-Lib is its ability to make predictions from shallow pilot sequencing.
Instead of deeply sequencing an entire library before determining whether the chosen depth was appropriate, researchers can generate a smaller amount of sequencing data and use NB-Lib to estimate how additional sequencing is likely to affect molecular recovery.
This could help researchers plan sequencing depth based on the characteristics of an individual library rather than relying primarily on general recommendations such as a fixed number of reads per cell.
The analysis also showed that differences in amplification heterogeneity and library complexity can explain substantial variation in sequencing efficiency and cost among samples and transcriptomic technologies.
Molecular recovery affects detection of low-abundance genes
Sequencing depth becomes particularly important when researchers are interested in genes expressed at low levels.
The researchers found that molecular recovery directly influenced the reproducibility of low-abundance gene detection. When too few molecules are recovered, weakly expressed genes may be detected inconsistently across measurements.
This means that selecting an appropriate sequencing depth can affect more than cost and overall data quantity. It can influence whether biologically important but uncommon transcripts are reliably observed.
Improving RNA sequencing experiment planning
NB-Lib provides a more detailed way to evaluate sequencing requirements by connecting sequencing depth with the molecules researchers actually recover from their libraries.
Rather than assuming that sequencing saturation alone indicates whether an experiment has been sequenced deeply enough, the framework accounts for library complexity and amplification differences that can alter molecular recovery.
For single-cell and spatial RNA sequencing experiments, this approach could help researchers make more informed decisions about sequencing depth, improve detection of low-abundance genes, and allocate sequencing resources more efficiently.
Availability – The scdepth package is available at https://github.com/gwlab-ca/scdepth.
Wilson G, N Kalimuthu S, Yeung J. (2026) Sequencing saturation does not uniquely determine molecular recovery in UMI transcriptomics. Bioinformatics 42(8): btag593. [article]












Stay Connected