Single cell technologies have made it possible to look at gene expression in individual cells, revealing differences that were once hidden in averaged data. However, one major challenge in single cell RNA sequencing is the presence of large numbers of zeros in the data. Some of these zeros reflect true biological absence, meaning the gene is not active in that cell. Others occur because of technical limitations that cause the sequencing system to miss transcripts that are actually present. These non biological zeros can mislead researchers and distort the patterns they try to uncover.
To address this problem, researchers at the China University of Geosciences developed D3Impute, a new computational framework designed to tell the difference between real and artificial zeros. It introduces three major components. The first is a normalization step that adjusts the data while keeping true biological variation intact. The second is a dual network discriminator that uses bulk RNA sequencing as a reference point to distinguish non biological zeros from true ones. The third is a density guided imputation engine that fills in missing values while preserving the structure of cell neighborhoods.
Workflow of D3Impute framework
(A) Distribution-aware preprocessing denoiser: Quality control and denoising of raw scRNA-seq (Xraw−seq) and bulk RNA-seq (Xraw−bulk) matrices. (B) Dropout-aware discriminator: Construction of interaction networks and generation of imputation index matrix I1. (C) Density-guided imputation engine: SNN graph-based neighbor identification and matrix updating . (D) Downstream analysis: Clustering, trajectory inference, and differential expression analysis.
The team tested D3Impute against twelve leading imputation tools using six different datasets. In every case, D3Impute improved key downstream analyses, including how well cells could be clustered, how developmental trajectories were inferred, and how differential expression was detected. They also explored how the tool performs under different data qualities and provided guidance to help users get the best results. By improving the accuracy of single cell datasets, D3Impute strengthens the biological insights researchers can gain from their experiments.
Availability – The MATLAB and Python package D3Impute is freely available at https://github.com/zhuyuan-cug/D3Impute.
Huang S, Jiang L, Yi M, Zhu Y. (2025) D3Impute, dropout aware discrimination, distribution aware modeling, and density guide imputation for scRNA seq data. PLoS Computational Biology 21(12): e1013744. [article]
Single cell technologies have made it possible to look at gene expression in individual cells, revealing differences that were once hidden in averaged data. However, one major challenge in single cell RNA sequencing is the presence of large numbers of zeros in the data. Some of these zeros reflect true biological absence, meaning the gene is not active in that cell. Others occur because of technical limitations that cause the sequencing system to miss transcripts that are actually present. These non biological zeros can mislead researchers and distort the patterns they try to uncover.
To address this problem, researchers at the China University of Geosciences developed D3Impute, a new computational framework designed to tell the difference between real and artificial zeros. It introduces three major components. The first is a normalization step that adjusts the data while keeping true biological variation intact. The second is a dual network discriminator that uses bulk RNA sequencing as a reference point to distinguish non biological zeros from true ones. The third is a density guided imputation engine that fills in missing values while preserving the structure of cell neighborhoods.
Workflow of D3Impute framework
(A) Distribution-aware preprocessing denoiser: Quality control and denoising of raw scRNA-seq (Xraw−seq) and bulk RNA-seq (Xraw−bulk) matrices. (B) Dropout-aware discriminator: Construction of interaction networks and generation of imputation index matrix I1. (C) Density-guided imputation engine: SNN graph-based neighbor identification and matrix updating . (D) Downstream analysis: Clustering, trajectory inference, and differential expression analysis.
The team tested D3Impute against twelve leading imputation tools using six different datasets. In every case, D3Impute improved key downstream analyses, including how well cells could be clustered, how developmental trajectories were inferred, and how differential expression was detected. They also explored how the tool performs under different data qualities and provided guidance to help users get the best results. By improving the accuracy of single cell datasets, D3Impute strengthens the biological insights researchers can gain from their experiments.
Availability – The MATLAB and Python package D3Impute is freely available at https://github.com/zhuyuan-cug/D3Impute.
Huang S, Jiang L, Yi M, Zhu Y. (2025) D3Impute, dropout aware discrimination, distribution aware modeling, and density guide imputation for scRNA seq data. PLoS Computational Biology 21(12): e1013744. [article]











Stay Connected