
Gene expression studies often focus on measuring how much of a gene is being produced in different cells or tissues. However, genes can also be regulated after transcription through processes that affect how RNA molecules are processed. One important mechanism is alternative polyadenylation, or APA, which can change the length and structure of messenger RNA molecules and influence how genes function.
APA has been linked to numerous biological processes and diseases, including cancer. As interest in APA continues to grow, researchers have developed specialized sequencing methods known as APA-seq that focus on identifying polyadenylation sites across the genome. While these methods provide valuable information about RNA processing, analyzing gene expression changes from APA-seq data has remained challenging.
Researchers from Shanghai Jiao Tong University in China have developed a new computational method called APAdeg to address this problem. The tool was specifically designed to identify differentially expressed genes from APA-seq datasets more accurately than conventional RNA sequencing analysis methods.
Traditional RNA sequencing approaches for differential gene expression analysis were created for whole transcriptome data and do not fully account for the unique characteristics of APA-seq datasets. Most existing APA analysis tools focus primarily on identifying polyadenylation sites and comparing their usage between samples. They generally do not provide robust methods for detecting genes whose overall expression levels differ between conditions.
APAdeg was developed to bridge this gap. The method combines information about a gene’s total sequencing reads with read counts from individual polyadenylation sites within the gene. By integrating both types of information into a statistical framework, APAdeg can more accurately determine whether a gene is differentially expressed.
The researchers evaluated APAdeg using both simulated datasets and real APA-seq datasets. Across multiple tests, APAdeg consistently outperformed conventional RNA sequencing based approaches in identifying differentially expressed genes. This suggests that methods specifically designed for APA-seq data can provide more reliable biological insights than repurposing tools originally developed for other sequencing technologies.
To demonstrate the practical value of the approach, the team applied APAdeg to APA-seq datasets from multiple cancer types. Interestingly, they found that only a relatively small fraction of differentially expressed genes showed significant changes in the length of their 3′ untranslated regions. Likewise, only a small proportion displayed substantial changes in intronic APA usage.
These findings suggest that changes in gene expression and changes in APA patterns do not always occur together. Some genes may exhibit altered expression without major shifts in polyadenylation site usage, while others may undergo APA changes with little effect on overall expression levels. Understanding these distinctions can provide a more complete picture of gene regulation in health and disease.
The researchers have made APAdeg available as an R software package, allowing other scientists to incorporate the method into their own studies. This accessibility could help accelerate research into APA biology and improve the interpretation of APA-seq datasets across many fields of biomedical science.
As sequencing technologies continue to evolve, specialized analysis tools will become increasingly important. APAdeg represents an important advance for researchers investigating alternative polyadenylation, providing a more accurate way to identify differentially expressed genes and better understand how RNA processing contributes to disease and normal cellular function.
Sarker B, Zhou T, Deng X, Zhang L, Zhao X, Huang W, Yi H, Jiang H, Xu C. (2026) APAdeg enhances differentially expressed gene inference by leveraging site-specific signals in APA-seq data. Briefings in Bioinformatics 27(3): bbag295. [article]

Gene expression studies often focus on measuring how much of a gene is being produced in different cells or tissues. However, genes can also be regulated after transcription through processes that affect how RNA molecules are processed. One important mechanism is alternative polyadenylation, or APA, which can change the length and structure of messenger RNA molecules and influence how genes function.
APA has been linked to numerous biological processes and diseases, including cancer. As interest in APA continues to grow, researchers have developed specialized sequencing methods known as APA-seq that focus on identifying polyadenylation sites across the genome. While these methods provide valuable information about RNA processing, analyzing gene expression changes from APA-seq data has remained challenging.
Researchers from Shanghai Jiao Tong University in China have developed a new computational method called APAdeg to address this problem. The tool was specifically designed to identify differentially expressed genes from APA-seq datasets more accurately than conventional RNA sequencing analysis methods.
Traditional RNA sequencing approaches for differential gene expression analysis were created for whole transcriptome data and do not fully account for the unique characteristics of APA-seq datasets. Most existing APA analysis tools focus primarily on identifying polyadenylation sites and comparing their usage between samples. They generally do not provide robust methods for detecting genes whose overall expression levels differ between conditions.
APAdeg was developed to bridge this gap. The method combines information about a gene’s total sequencing reads with read counts from individual polyadenylation sites within the gene. By integrating both types of information into a statistical framework, APAdeg can more accurately determine whether a gene is differentially expressed.
The researchers evaluated APAdeg using both simulated datasets and real APA-seq datasets. Across multiple tests, APAdeg consistently outperformed conventional RNA sequencing based approaches in identifying differentially expressed genes. This suggests that methods specifically designed for APA-seq data can provide more reliable biological insights than repurposing tools originally developed for other sequencing technologies.
To demonstrate the practical value of the approach, the team applied APAdeg to APA-seq datasets from multiple cancer types. Interestingly, they found that only a relatively small fraction of differentially expressed genes showed significant changes in the length of their 3′ untranslated regions. Likewise, only a small proportion displayed substantial changes in intronic APA usage.
These findings suggest that changes in gene expression and changes in APA patterns do not always occur together. Some genes may exhibit altered expression without major shifts in polyadenylation site usage, while others may undergo APA changes with little effect on overall expression levels. Understanding these distinctions can provide a more complete picture of gene regulation in health and disease.
The researchers have made APAdeg available as an R software package, allowing other scientists to incorporate the method into their own studies. This accessibility could help accelerate research into APA biology and improve the interpretation of APA-seq datasets across many fields of biomedical science.
As sequencing technologies continue to evolve, specialized analysis tools will become increasingly important. APAdeg represents an important advance for researchers investigating alternative polyadenylation, providing a more accurate way to identify differentially expressed genes and better understand how RNA processing contributes to disease and normal cellular function.
Sarker B, Zhou T, Deng X, Zhang L, Zhao X, Huang W, Yi H, Jiang H, Xu C. (2026) APAdeg enhances differentially expressed gene inference by leveraging site-specific signals in APA-seq data. Briefings in Bioinformatics 27(3): bbag295. [article]












Stay Connected