Single-cell RNA sequencing allows researchers to examine gene activity in individual cells, making it possible to study the many different cell types within a tissue. However, single-cell experiments can be expensive and technically demanding, especially when large numbers of samples are involved.
An alternative approach is cellular deconvolution, a computational method that estimates the proportions of different cell types in a tissue using bulk RNA sequencing data. Bulk RNA sequencing measures gene activity across all cells in a sample at once, so the resulting signal is a mixture of many cell types.
Researchers at the Department of Computer Science at Wayne State University developed a new deconvolution method called DeOPUS, short for Deconvolution via Optimized Power-transformed Unmixing with Shrinkage.
Estimating cell types from mixed RNA signals
Imagine a tissue sample containing immune cells, epithelial cells and connective tissue cells. Bulk RNA sequencing measures RNA from all of these cells together.
Because different cell types have characteristic gene expression patterns, computational methods can compare the mixed bulk RNA sequencing signal with reference profiles from known cell types. From this information, they estimate what proportion of the sample came from each type of cell.
The challenge is that RNA sequencing data can vary enormously in expression level. Some genes are highly expressed, while others are detected at much lower levels. Technical noise and outliers can also interfere with estimates.
DeOPUS was designed to reduce the influence of these problems.
How DeOPUS works
DeOPUS uses a statistical approach called a hierarchical shrinkage transformation to stabilize variation in gene expression data.
The method combines several techniques, including variance-stabilizing transformations, adaptive weighting and rank-based quantile normalization. Together, these steps reduce the influence of unusually high expression values and technical noise while preserving useful biological signals.
DeOPUS takes two main inputs, a bulk RNA sequencing expression profile and a reference gene expression profile for known cell types, often derived from single-cell RNA sequencing. It then estimates the proportion of each cell type within the bulk sample.
Overview of the DeOPUS deconvolution pipeline
Input data: A bulk RNA-seq expression matrix (genes by samples) and a reference cell-type expression matrix (genes by cell types) derived from scRNA-seq. HST: Expression values undergo preprocessing and depth normalization, variance-stabilizing power transformation using global shrinkage, rank-based quantile normalization, and gene-specific weighted blending, smoothing, and scaling to stabilize variance and mitigate outlier effects. Optimization: The loss between the transformed bulk profile and the transformed reconstructed profile is minimized independently for each sample using the L-BFGS-B algorithm with non-negativity constraints. Output: Estimated cell-type proportions (sum to one) for each bulk sample.
Testing DeOPUS across many tissues
The researchers compared DeOPUS with eight existing cellular deconvolution methods across 122 tissues representing 12 organ systems.
Across these datasets, DeOPUS achieved mean Pearson and Spearman correlations of 0.82 and 0.77, respectively, while also producing the lowest mean squared error among the methods tested.
The method also maintained strong performance as tissue complexity increased, including samples containing more than 30 different cell types.
Validation with real biological samples
The researchers performed additional testing using 18 bulk datasets in which the actual cell-type proportions had been measured experimentally.
DeOPUS again produced the highest average correlations among the evaluated methods and was the only method that produced positive correlations across all 18 datasets. It also showed the highest accuracy for identifying the most abundant cell type within a sample.
These results suggest that stabilizing variation in RNA sequencing data can improve the reliability of cellular deconvolution across a wide range of tissues.
Expanding the value of bulk RNA sequencing
Single-cell RNA sequencing provides detailed information about cellular diversity, but large collections of existing biological and clinical samples have already been analyzed using bulk RNA sequencing.
Methods such as DeOPUS offer a way to extract additional information from these datasets by estimating which cell types contributed to the overall gene expression profile.
This could be particularly useful for investigating complex tissues such as tumors, where differences in immune cells, cancer cells and surrounding tissue can influence disease behavior.
Availability – DeOPUS is available as an open-source R package at https://github.com/tinnlab/DeOPUS.
Nguyen H, Nguyen K, Bya P, Alafif T, Quan TT, Nguyen T. (2026) DeOPUS: cellular deconvolution via optimized power-transformed unmixing with shrinkage. Briefings in Bioinformatics 27(5): bbag524. [article]
Single-cell RNA sequencing allows researchers to examine gene activity in individual cells, making it possible to study the many different cell types within a tissue. However, single-cell experiments can be expensive and technically demanding, especially when large numbers of samples are involved.
An alternative approach is cellular deconvolution, a computational method that estimates the proportions of different cell types in a tissue using bulk RNA sequencing data. Bulk RNA sequencing measures gene activity across all cells in a sample at once, so the resulting signal is a mixture of many cell types.
Researchers at the Department of Computer Science at Wayne State University developed a new deconvolution method called DeOPUS, short for Deconvolution via Optimized Power-transformed Unmixing with Shrinkage.
Estimating cell types from mixed RNA signals
Imagine a tissue sample containing immune cells, epithelial cells and connective tissue cells. Bulk RNA sequencing measures RNA from all of these cells together.
Because different cell types have characteristic gene expression patterns, computational methods can compare the mixed bulk RNA sequencing signal with reference profiles from known cell types. From this information, they estimate what proportion of the sample came from each type of cell.
The challenge is that RNA sequencing data can vary enormously in expression level. Some genes are highly expressed, while others are detected at much lower levels. Technical noise and outliers can also interfere with estimates.
DeOPUS was designed to reduce the influence of these problems.
How DeOPUS works
DeOPUS uses a statistical approach called a hierarchical shrinkage transformation to stabilize variation in gene expression data.
The method combines several techniques, including variance-stabilizing transformations, adaptive weighting and rank-based quantile normalization. Together, these steps reduce the influence of unusually high expression values and technical noise while preserving useful biological signals.
DeOPUS takes two main inputs, a bulk RNA sequencing expression profile and a reference gene expression profile for known cell types, often derived from single-cell RNA sequencing. It then estimates the proportion of each cell type within the bulk sample.
Overview of the DeOPUS deconvolution pipeline
Input data: A bulk RNA-seq expression matrix (genes by samples) and a reference cell-type expression matrix (genes by cell types) derived from scRNA-seq. HST: Expression values undergo preprocessing and depth normalization, variance-stabilizing power transformation using global shrinkage, rank-based quantile normalization, and gene-specific weighted blending, smoothing, and scaling to stabilize variance and mitigate outlier effects. Optimization: The loss between the transformed bulk profile and the transformed reconstructed profile is minimized independently for each sample using the L-BFGS-B algorithm with non-negativity constraints. Output: Estimated cell-type proportions (sum to one) for each bulk sample.
Testing DeOPUS across many tissues
The researchers compared DeOPUS with eight existing cellular deconvolution methods across 122 tissues representing 12 organ systems.
Across these datasets, DeOPUS achieved mean Pearson and Spearman correlations of 0.82 and 0.77, respectively, while also producing the lowest mean squared error among the methods tested.
The method also maintained strong performance as tissue complexity increased, including samples containing more than 30 different cell types.
Validation with real biological samples
The researchers performed additional testing using 18 bulk datasets in which the actual cell-type proportions had been measured experimentally.
DeOPUS again produced the highest average correlations among the evaluated methods and was the only method that produced positive correlations across all 18 datasets. It also showed the highest accuracy for identifying the most abundant cell type within a sample.
These results suggest that stabilizing variation in RNA sequencing data can improve the reliability of cellular deconvolution across a wide range of tissues.
Expanding the value of bulk RNA sequencing
Single-cell RNA sequencing provides detailed information about cellular diversity, but large collections of existing biological and clinical samples have already been analyzed using bulk RNA sequencing.
Methods such as DeOPUS offer a way to extract additional information from these datasets by estimating which cell types contributed to the overall gene expression profile.
This could be particularly useful for investigating complex tissues such as tumors, where differences in immune cells, cancer cells and surrounding tissue can influence disease behavior.
Availability – DeOPUS is available as an open-source R package at https://github.com/tinnlab/DeOPUS.
Nguyen H, Nguyen K, Bya P, Alafif T, Quan TT, Nguyen T. (2026) DeOPUS: cellular deconvolution via optimized power-transformed unmixing with shrinkage. Briefings in Bioinformatics 27(5): bbag524. [article]












Stay Connected