Long-read RNA sequencing has given researchers a powerful way to examine the full range of RNA molecules produced by cells. Technologies from Pacific Biosciences, or PacBio, and Oxford Nanopore Technologies, or ONT, can sequence much longer stretches of RNA than traditional short-read methods. This makes it easier to identify complete RNA transcripts and understand how a single gene can produce multiple RNA variants.
Researchers at the Washington University School of Medicine have developed a new computational pipeline called NextLongIso to simplify the analysis of long-read RNA sequencing data.e The framework brings several types of transcript analysis together within a single reproducible workflow.
Overview and application of the NextLongIso pipeline
(a) Schematic representation of the NextLongIso workflow. The pipeline processes long-read RNA-seq data through alignment and quality control, followed by integrated downstream analyses across multiple regulatory layers, including transcript discovery and quantification, isoform switching, alternative splicing, alternative promoter usage, alternative polyadenylation and transposable element (TE)-associated exonization. Representative output files generated by each analysis module are also shown. (b) Multi-layer analysis of PacBio long-read RNA-seq data from GM12878 and H1-ESC cell lines. A representative genome browser view illustrates regulatory differences between two cell lines, including isoform switching, alternative splicing, alternative promoter usage, alternative transcription end sites and TE-associated exonization events.
One of the major advantages of long-read RNA sequencing is its ability to capture full-length transcripts. Genes are often more complicated than a simple one-gene, one-transcript relationship. A single gene can produce several RNA isoforms depending on how different sections of the RNA are combined and where transcription begins or ends. These differences can affect the proteins that cells produce and how genes function.
Analyzing all of these variations can be challenging. Researchers often need to use several separate software tools to reconstruct transcripts, measure their abundance, identify alternative splicing, and investigate other forms of RNA regulation. Moving data between these programs can require extensive formatting and processing, making analyses more complicated and potentially harder to reproduce.
NextLongIso addresses this problem by combining these analytical steps into a unified Nextflow pipeline. Nextflow is a workflow management system designed to make computational analyses reproducible and scalable, particularly when researchers are processing large biological datasets.
Rather than stopping after identifying transcripts, NextLongIso allows researchers to examine several layers of RNA regulation. The pipeline can analyze alternative splicing, where different combinations of RNA segments produce different transcripts from the same gene. It can also identify isoform switching, which occurs when cells change which version of a transcript they predominantly produce.
NextLongIso also examines where transcription begins and ends. Alternative promoter analysis can reveal when cells use different starting points to produce RNA from the same gene, while alternative polyadenylation analysis identifies differences in transcript ending sites. Both processes can influence RNA stability, localization, and protein production.
Another feature of the pipeline is its ability to investigate transcription associated with transposable elements and other repetitive sequences. These regions make up a substantial portion of the genome and can influence gene regulation, but they can be difficult to analyze using conventional sequencing approaches.
Importantly, NextLongIso supports data generated using both PacBio and Oxford Nanopore sequencing platforms. This gives researchers greater flexibility to apply the same analytical framework across different long-read technologies.
By integrating transcript discovery with multiple downstream analyses, NextLongIso helps researchers move from simply cataloging RNA transcripts toward understanding how those transcripts are regulated. As long-read RNA sequencing datasets continue to increase in size and complexity, unified pipelines such as NextLongIso could make it easier to investigate alternative splicing, transcript isoforms, promoter usage, polyadenylation, and other important aspects of gene regulation.
Availability – NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.
Tan J, Wu Y, Sun Y. (2026) NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis. Bioinformatics 42(8): btag518. [article]
Long-read RNA sequencing has given researchers a powerful way to examine the full range of RNA molecules produced by cells. Technologies from Pacific Biosciences, or PacBio, and Oxford Nanopore Technologies, or ONT, can sequence much longer stretches of RNA than traditional short-read methods. This makes it easier to identify complete RNA transcripts and understand how a single gene can produce multiple RNA variants.
Researchers at the Washington University School of Medicine have developed a new computational pipeline called NextLongIso to simplify the analysis of long-read RNA sequencing data.e The framework brings several types of transcript analysis together within a single reproducible workflow.
Overview and application of the NextLongIso pipeline
(a) Schematic representation of the NextLongIso workflow. The pipeline processes long-read RNA-seq data through alignment and quality control, followed by integrated downstream analyses across multiple regulatory layers, including transcript discovery and quantification, isoform switching, alternative splicing, alternative promoter usage, alternative polyadenylation and transposable element (TE)-associated exonization. Representative output files generated by each analysis module are also shown. (b) Multi-layer analysis of PacBio long-read RNA-seq data from GM12878 and H1-ESC cell lines. A representative genome browser view illustrates regulatory differences between two cell lines, including isoform switching, alternative splicing, alternative promoter usage, alternative transcription end sites and TE-associated exonization events.
One of the major advantages of long-read RNA sequencing is its ability to capture full-length transcripts. Genes are often more complicated than a simple one-gene, one-transcript relationship. A single gene can produce several RNA isoforms depending on how different sections of the RNA are combined and where transcription begins or ends. These differences can affect the proteins that cells produce and how genes function.
Analyzing all of these variations can be challenging. Researchers often need to use several separate software tools to reconstruct transcripts, measure their abundance, identify alternative splicing, and investigate other forms of RNA regulation. Moving data between these programs can require extensive formatting and processing, making analyses more complicated and potentially harder to reproduce.
NextLongIso addresses this problem by combining these analytical steps into a unified Nextflow pipeline. Nextflow is a workflow management system designed to make computational analyses reproducible and scalable, particularly when researchers are processing large biological datasets.
Rather than stopping after identifying transcripts, NextLongIso allows researchers to examine several layers of RNA regulation. The pipeline can analyze alternative splicing, where different combinations of RNA segments produce different transcripts from the same gene. It can also identify isoform switching, which occurs when cells change which version of a transcript they predominantly produce.
NextLongIso also examines where transcription begins and ends. Alternative promoter analysis can reveal when cells use different starting points to produce RNA from the same gene, while alternative polyadenylation analysis identifies differences in transcript ending sites. Both processes can influence RNA stability, localization, and protein production.
Another feature of the pipeline is its ability to investigate transcription associated with transposable elements and other repetitive sequences. These regions make up a substantial portion of the genome and can influence gene regulation, but they can be difficult to analyze using conventional sequencing approaches.
Importantly, NextLongIso supports data generated using both PacBio and Oxford Nanopore sequencing platforms. This gives researchers greater flexibility to apply the same analytical framework across different long-read technologies.
By integrating transcript discovery with multiple downstream analyses, NextLongIso helps researchers move from simply cataloging RNA transcripts toward understanding how those transcripts are regulated. As long-read RNA sequencing datasets continue to increase in size and complexity, unified pipelines such as NextLongIso could make it easier to investigate alternative splicing, transcript isoforms, promoter usage, polyadenylation, and other important aspects of gene regulation.
Availability – NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.
Tan J, Wu Y, Sun Y. (2026) NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis. Bioinformatics 42(8): btag518. [article]












Stay Connected