A single gene can produce multiple RNA molecules with different structures. One important process behind this diversity is alternative polyadenylation, or APA, in which cells choose different locations near the end of an RNA transcript where the RNA is cut and given a poly(A) tail.

APA affects more than 70 percent of human genes and can influence RNA stability, transport and translation. Researchers at the Department of Genetics and Genome Sciences at the University of Connecticut School of Medicine developed a computational framework called APALORD to study this process using long-read RNA sequencing.

Using long-read RNA sequencing to study APA

Traditional short-read RNA sequencing breaks RNA into small fragments, which then have to be computationally reconstructed. Long-read RNA sequencing can capture much longer stretches of individual RNA molecules, making it easier to identify where transcripts end.

APALORD, short for Alternative Polyadenylation Analysis of LOng-ReaDs, compares RNA cleavage sites between biological conditions to identify changes in polyadenylation patterns.

The researchers applied APALORD to direct RNA sequencing data from human embryonic stem cells and neurons derived from them. The method identified both known and previously unannotated polyadenylation sites with high positional accuracy.

APALORD workflow

Fig. 1: APALORD workflow.

LR RNA-seq reads are first aligned to the reference genome (with genome sequence and gene annotation as reference) using minimap2, generating BAM files. Aligned reads are then assigned to unique host genes using IsoQuant. APALORD takes these reads assignment results as input; alternatively, BAM files can be loaded directly into APALORD, where reads are assigned to genes with a built-in function based on Bambu. APALORD subsequently identifies PASs from the data and quantifies PAS usage (PAU) within each gene. The framework enables multiple levels of APA analysis, including PAS-level PAU change, gene-level APA changes, and proximal/distal PAU, along with transcriptome-wide and gene-specific visualization.

Neurons show longer RNA transcripts

As stem cells differentiated into neurons, the researchers observed a broad shift toward longer RNA transcripts. Neurons more frequently used polyadenylation sites located farther downstream, resulting in longer transcript ends.

These sites also tended to contain stronger polyadenylation signals, including the canonical AAUAAA sequence, upstream UGUA motifs and downstream GU or U-rich regions.

The researchers also found that stronger polyadenylation sites were associated with higher gene expression and with genes containing fewer alternative sites.

Similar patterns across species

APALORD was also applied to long-read RNA sequencing data from Drosophila embryos. The analysis expanded the number of known polyadenylation sites in fruit flies and identified regulatory features shared with human cells.

The researchers also detected a previously unrecognized group of low-abundance transcripts that did not map directly to identified polyadenylation sites but still showed patterns associated with APA regulation.

Improving analysis of RNA processing

Alternative polyadenylation adds another layer of regulation to gene expression by allowing cells to produce RNA transcripts with different endings.

APALORD provides a way to study these differences at high resolution using long-read RNA sequencing. By identifying known and novel polyadenylation sites and comparing their use across cell types and species, the framework may help researchers better understand how RNA processing contributes to development and cell specialization.

Availability – Software is available at GitHub [https://github.com/markandtwin/APALORD]

Zhang Z, Glatt-Deeley H, Zou L, Song D, Miura P. (2026) Leveraging long read RNA-seq to decipher neuronal regulation of alternative polyadenylation. Nature Communications 17(1): 10171. [article]

A single gene can produce multiple RNA molecules with different structures. One important process behind this diversity is alternative polyadenylation, or APA, in which cells choose different locations near the end of an RNA transcript where the RNA is cut and given a poly(A) tail.

APA affects more than 70 percent of human genes and can influence RNA stability, transport and translation. Researchers at the Department of Genetics and Genome Sciences at the University of Connecticut School of Medicine developed a computational framework called APALORD to study this process using long-read RNA sequencing.

Using long-read RNA sequencing to study APA

Traditional short-read RNA sequencing breaks RNA into small fragments, which then have to be computationally reconstructed. Long-read RNA sequencing can capture much longer stretches of individual RNA molecules, making it easier to identify where transcripts end.

APALORD, short for Alternative Polyadenylation Analysis of LOng-ReaDs, compares RNA cleavage sites between biological conditions to identify changes in polyadenylation patterns.

The researchers applied APALORD to direct RNA sequencing data from human embryonic stem cells and neurons derived from them. The method identified both known and previously unannotated polyadenylation sites with high positional accuracy.

APALORD workflow

Fig. 1: APALORD workflow.

LR RNA-seq reads are first aligned to the reference genome (with genome sequence and gene annotation as reference) using minimap2, generating BAM files. Aligned reads are then assigned to unique host genes using IsoQuant. APALORD takes these reads assignment results as input; alternatively, BAM files can be loaded directly into APALORD, where reads are assigned to genes with a built-in function based on Bambu. APALORD subsequently identifies PASs from the data and quantifies PAS usage (PAU) within each gene. The framework enables multiple levels of APA analysis, including PAS-level PAU change, gene-level APA changes, and proximal/distal PAU, along with transcriptome-wide and gene-specific visualization.

Neurons show longer RNA transcripts

As stem cells differentiated into neurons, the researchers observed a broad shift toward longer RNA transcripts. Neurons more frequently used polyadenylation sites located farther downstream, resulting in longer transcript ends.

These sites also tended to contain stronger polyadenylation signals, including the canonical AAUAAA sequence, upstream UGUA motifs and downstream GU or U-rich regions.

The researchers also found that stronger polyadenylation sites were associated with higher gene expression and with genes containing fewer alternative sites.

Similar patterns across species

APALORD was also applied to long-read RNA sequencing data from Drosophila embryos. The analysis expanded the number of known polyadenylation sites in fruit flies and identified regulatory features shared with human cells.

The researchers also detected a previously unrecognized group of low-abundance transcripts that did not map directly to identified polyadenylation sites but still showed patterns associated with APA regulation.

Improving analysis of RNA processing

Alternative polyadenylation adds another layer of regulation to gene expression by allowing cells to produce RNA transcripts with different endings.

APALORD provides a way to study these differences at high resolution using long-read RNA sequencing. By identifying known and novel polyadenylation sites and comparing their use across cell types and species, the framework may help researchers better understand how RNA processing contributes to development and cell specialization.

Availability – Software is available at GitHub [https://github.com/markandtwin/APALORD]

Zhang Z, Glatt-Deeley H, Zou L, Song D, Miura P. (2026) Leveraging long read RNA-seq to decipher neuronal regulation of alternative polyadenylation. Nature Communications 17(1): 10171. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services