High-throughput RNA sequencing is a powerful way to study how genes are turned on or off in cells, but it comes with a challenge. Not all differences detected in the data reflect real biological signals, some are caused by technical noise, such as variation in sample handling, sequencing machines, or other experimental factors. If this unwanted variation is not properly managed, it can hide true signals or create false ones.
Researchers at Hasselt University, Belgium have explored strategies for removing unwanted variation in pseudobulked single-cell RNA sequencing data. Pseudobulking involves combining the RNA data from cells of the same type to improve the statistical power of analyses. The team evaluated three widely used methods called RUV2, RUVIII, and RUV4, which aim to separate real biological differences from technical noise. They also introduced a new approach, RUVIII PBPS, designed to handle situations where standard methods struggle, such as when there are too few technical replicates or missing control genes.
Schematic overview of the process to generate PBPS with six single-cell samples,
two biological variables, and one known source of unwanted variation
The first biological covariate, factor of interest, has two levels, and the second biological covariate, cell type, has two levels; therefore . There are two batches, therefore . The first biological subgroup is composed by the squared blue cells; it has seven cells from three homogeneous single-cell samples, therefore is the average number of cells per single-cell sample within the set . Of those 7 cells, four were processed in batch 1, and form the set : Two in light blue from sample 1, and two in a darker shade from sample 2. We randomly sampled two cells, times; some pseudosamples will have cells from different samples, while others will have cells from only one sample. The cells are then pseudobulked and added to the matrix.
Their findings show that removing unwanted variation per cell type using RUV2 or RUVIII effectively extracts technical noise factors and keeps false discoveries under control, even when there are confounding factors. The novel RUVIII PBPS method succeeds in controlling false discoveries in more challenging scenarios, making it a robust tool for analyzing pseudobulk single-cell RNA sequencing data. These strategies help researchers confidently interpret gene expression changes, leading to more accurate biological conclusions.
Availability – The R package to create PBPS is available at https://github.com/esprietol/ruvPBPS
Prieto León S, De Troyer E, Geys H, Van den Berge K, Thas O, (2025) Removal of unwanted variation in pseudobulk analysis of single-cell RNA sequencing data and the leveraging of pseudoreplicates. NAR Genomics and Bioinformatics 7(4): lqaf179. [article]
High-throughput RNA sequencing is a powerful way to study how genes are turned on or off in cells, but it comes with a challenge. Not all differences detected in the data reflect real biological signals, some are caused by technical noise, such as variation in sample handling, sequencing machines, or other experimental factors. If this unwanted variation is not properly managed, it can hide true signals or create false ones.
Researchers at Hasselt University, Belgium have explored strategies for removing unwanted variation in pseudobulked single-cell RNA sequencing data. Pseudobulking involves combining the RNA data from cells of the same type to improve the statistical power of analyses. The team evaluated three widely used methods called RUV2, RUVIII, and RUV4, which aim to separate real biological differences from technical noise. They also introduced a new approach, RUVIII PBPS, designed to handle situations where standard methods struggle, such as when there are too few technical replicates or missing control genes.
Schematic overview of the process to generate PBPS with six single-cell samples,
two biological variables, and one known source of unwanted variation
The first biological covariate, factor of interest, has two levels, and the second biological covariate, cell type, has two levels; therefore . There are two batches, therefore . The first biological subgroup is composed by the squared blue cells; it has seven cells from three homogeneous single-cell samples, therefore is the average number of cells per single-cell sample within the set . Of those 7 cells, four were processed in batch 1, and form the set : Two in light blue from sample 1, and two in a darker shade from sample 2. We randomly sampled two cells, times; some pseudosamples will have cells from different samples, while others will have cells from only one sample. The cells are then pseudobulked and added to the matrix.
Their findings show that removing unwanted variation per cell type using RUV2 or RUVIII effectively extracts technical noise factors and keeps false discoveries under control, even when there are confounding factors. The novel RUVIII PBPS method succeeds in controlling false discoveries in more challenging scenarios, making it a robust tool for analyzing pseudobulk single-cell RNA sequencing data. These strategies help researchers confidently interpret gene expression changes, leading to more accurate biological conclusions.
Availability – The R package to create PBPS is available at https://github.com/esprietol/ruvPBPS
Prieto León S, De Troyer E, Geys H, Van den Berge K, Thas O, (2025) Removal of unwanted variation in pseudobulk analysis of single-cell RNA sequencing data and the leveraging of pseudoreplicates. NAR Genomics and Bioinformatics 7(4): lqaf179. [article]












Stay Connected