RNA sequencing has become one of the most widely used tools for studying gene expression. Researchers use it to identify disease mechanisms, discover biomarkers, and understand how genes are regulated. However, analyzing RNA sequencing data is often more complicated than it appears.

One major challenge is the presence of hidden confounding factors. These are sources of variation that influence gene expression measurements but are not directly measured or recorded. Examples include subtle technical differences between experiments, variations in sample processing, or interactions between biological and technical factors.

When these hidden influences are not properly addressed, they can distort results and make it difficult to distinguish genuine biological signals from experimental noise.

Researchers from Sun Yat-Sen University have developed a new deep learning framework called UnSupervised Adversarial Deconfounding AutoEncoder, or USADAE, to tackle this problem.

Traditional approaches for identifying hidden confounders include methods such as Surrogate Variable Analysis (SVA), Probabilistic Estimation of Expression Residuals (PEER), and Remove Unwanted Variation (RUVSeq). While these tools have been useful, they are generally based on linear assumptions. Real biological datasets often contain complex nonlinear relationships that these methods may not fully capture.

Sketch of USADAE framework

Sketch of USADAE framework.

USADAE employs dual encoders to decompose RNA-seq data into biologically relevant (⁠⁠) and confounder-associated (⁠⁠) latent representations, concatenated for decoder-driven reconstruction. The autoencoder operates along the gene-expression feature dimension, reducing the dimensionality of expression profiles while preserving sample-wise relationships. An adversarial discriminator (⁠⁠) enforces statistical independence between  and  by distinguishing true  from shuffled pairs, while a weak discriminator (⁠⁠) guides the biological encoder to retain label-discriminative features (e.g. case/control). Through adversarial optimization, USADAE explicitly isolates hidden covariates from observed signals.

USADAE uses deep learning and adversarial learning strategies to separate biological signals from hidden confounding signals. The model creates two distinct internal representations of the data. One captures the biological information researchers want to study, while the other captures unwanted sources of variation.

By separating these components, USADAE can improve downstream analyses such as differential gene expression studies and expression quantitative trait locus, or eQTL, analysis.

The researchers evaluated the method using both simulated datasets and real-world RNA sequencing data. In simulations, USADAE consistently outperformed existing approaches in identifying hidden covariates while preserving meaningful biological information.

The team also applied the framework to real datasets from cancer genomics and eQTL studies. Across these diverse applications, the model demonstrated strong performance and robustness, suggesting it may be useful for many different types of transcriptomics research.

As RNA sequencing datasets continue to grow in size and complexity, methods that can accurately separate biological signals from technical noise are becoming increasingly important. Deep learning approaches such as USADAE may help researchers generate more reliable results, improve reproducibility, and uncover biological insights that might otherwise remain hidden.

Availability – The code for the USADAE algorithm is available at https://github.com/chenxuya/USADAE.

Chen X, Guo L, Chen Y, Mo D, Liu X. (2026) USADAE: a deep learning approach to disentangle hidden covariates in RNA-seq data. Briefings in Bioinformatics 27(3). [article]

RNA sequencing has become one of the most widely used tools for studying gene expression. Researchers use it to identify disease mechanisms, discover biomarkers, and understand how genes are regulated. However, analyzing RNA sequencing data is often more complicated than it appears.

One major challenge is the presence of hidden confounding factors. These are sources of variation that influence gene expression measurements but are not directly measured or recorded. Examples include subtle technical differences between experiments, variations in sample processing, or interactions between biological and technical factors.

When these hidden influences are not properly addressed, they can distort results and make it difficult to distinguish genuine biological signals from experimental noise.

Researchers from Sun Yat-Sen University have developed a new deep learning framework called UnSupervised Adversarial Deconfounding AutoEncoder, or USADAE, to tackle this problem.

Traditional approaches for identifying hidden confounders include methods such as Surrogate Variable Analysis (SVA), Probabilistic Estimation of Expression Residuals (PEER), and Remove Unwanted Variation (RUVSeq). While these tools have been useful, they are generally based on linear assumptions. Real biological datasets often contain complex nonlinear relationships that these methods may not fully capture.

Sketch of USADAE framework

Sketch of USADAE framework.

USADAE employs dual encoders to decompose RNA-seq data into biologically relevant (⁠⁠) and confounder-associated (⁠⁠) latent representations, concatenated for decoder-driven reconstruction. The autoencoder operates along the gene-expression feature dimension, reducing the dimensionality of expression profiles while preserving sample-wise relationships. An adversarial discriminator (⁠⁠) enforces statistical independence between  and  by distinguishing true  from shuffled pairs, while a weak discriminator (⁠⁠) guides the biological encoder to retain label-discriminative features (e.g. case/control). Through adversarial optimization, USADAE explicitly isolates hidden covariates from observed signals.

USADAE uses deep learning and adversarial learning strategies to separate biological signals from hidden confounding signals. The model creates two distinct internal representations of the data. One captures the biological information researchers want to study, while the other captures unwanted sources of variation.

By separating these components, USADAE can improve downstream analyses such as differential gene expression studies and expression quantitative trait locus, or eQTL, analysis.

The researchers evaluated the method using both simulated datasets and real-world RNA sequencing data. In simulations, USADAE consistently outperformed existing approaches in identifying hidden covariates while preserving meaningful biological information.

The team also applied the framework to real datasets from cancer genomics and eQTL studies. Across these diverse applications, the model demonstrated strong performance and robustness, suggesting it may be useful for many different types of transcriptomics research.

As RNA sequencing datasets continue to grow in size and complexity, methods that can accurately separate biological signals from technical noise are becoming increasingly important. Deep learning approaches such as USADAE may help researchers generate more reliable results, improve reproducibility, and uncover biological insights that might otherwise remain hidden.

Availability – The code for the USADAE algorithm is available at https://github.com/chenxuya/USADAE.

Chen X, Guo L, Chen Y, Mo D, Liu X. (2026) USADAE: a deep learning approach to disentangle hidden covariates in RNA-seq data. Briefings in Bioinformatics 27(3). [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services