Understanding how genes are expressed is not just about which genes are turned on or off, but also how different versions of those genes are produced. A team led by researcher at the Dana-Farber Cancer Institute have developed a new tool called longcallR to better analyze long-read RNA sequencing data.

Long-read RNA sequencing is especially useful because it captures full-length RNA molecules, allowing scientists to see complete transcript structures. However, analyzing this type of data has been challenging due to limited tools. LongcallR addresses this by combining several important analyses into one platform, including identifying genetic variants, determining how those variants are inherited together, and measuring how each version of a gene is expressed.

Overview of the longcallR algorithm

Fig. 1: Overview of the longcallR algorithm.

a, The longcallR-nn module for RNA SNP calling. It constructs a seven-channel image representing a 41-bp flanking region around each candidate SNP (left). This image is processed by a ResNet-50 convolutional neural network, which outputs two classifications (right): zygosity (0/0, 0/1, 1/1, 1/2 and A-to-I RNA editing) and genotype (combinations of reference alleles (A, C, G, T) including alternate alleles (A, C, G, T)). b, The longcallR-phase module for SNP filtering and haplotype phasing. It extracts read alleles at candidate heterozygous SNP sites (left) and then iteratively refines read haplotags (H1 or H2) and SNP haplotypes to minimize conflicts between read alleles and haplotypes. The phased reads (green, H1; blue, H2) are then used to distinguish true SNPs, sequencing errors and A-to-I RNA editing events (right). c, Allele-specific analysis using longcallR. Phased reads are assigned to the closest gene based on the longest exonic matching length and read junctions are extracted from sequencing data. PS1 and PS2 denote two distinct phase blocks (left). For ASE analysis (top right), the module quantifies haplotype-specific read counts and applies a beta-binomial test. For ASJ analysis (bottom right), the module separately counts junction inclusion and exclusion reads for each haplotype and uses Fisher’s exact test to assess haplotype-specific splicing differences. The Benjamin–Hochberg procedure is applied for controlling the false-positive rate.

When applied to over 200 human samples, the tool uncovered a large number of allele-specific splicing events. This means that different copies of the same gene can be processed in distinct ways, leading to different RNA products. On average, each sample showed dozens of these events, many of which had not been previously documented.

One notable finding was that nearly half of these splicing events involved previously unannotated junctions. This suggests that current gene annotations are incomplete and that RNA sequencing continues to reveal new layers of biological complexity.

By linking genetic variation directly to RNA structure and expression, long-read RNA sequencing provides a more detailed picture of gene regulation. Tools like longcallR make it easier for researchers to uncover how genetic differences influence RNA processing, which has important implications for understanding disease mechanisms and developing targeted therapies.

Availability – LongcallR is available via GitHub at https://github.com/huangnengCSU/longcallR.

Huang N, Li H. (2026) SNP calling, haplotype phasing and allele-specific analysis with long RNA-seq reads. Nature Methods (Epub ahead of print)[article]

Understanding how genes are expressed is not just about which genes are turned on or off, but also how different versions of those genes are produced. A team led by researcher at the Dana-Farber Cancer Institute have developed a new tool called longcallR to better analyze long-read RNA sequencing data.

Long-read RNA sequencing is especially useful because it captures full-length RNA molecules, allowing scientists to see complete transcript structures. However, analyzing this type of data has been challenging due to limited tools. LongcallR addresses this by combining several important analyses into one platform, including identifying genetic variants, determining how those variants are inherited together, and measuring how each version of a gene is expressed.

Overview of the longcallR algorithm

Fig. 1: Overview of the longcallR algorithm.

a, The longcallR-nn module for RNA SNP calling. It constructs a seven-channel image representing a 41-bp flanking region around each candidate SNP (left). This image is processed by a ResNet-50 convolutional neural network, which outputs two classifications (right): zygosity (0/0, 0/1, 1/1, 1/2 and A-to-I RNA editing) and genotype (combinations of reference alleles (A, C, G, T) including alternate alleles (A, C, G, T)). b, The longcallR-phase module for SNP filtering and haplotype phasing. It extracts read alleles at candidate heterozygous SNP sites (left) and then iteratively refines read haplotags (H1 or H2) and SNP haplotypes to minimize conflicts between read alleles and haplotypes. The phased reads (green, H1; blue, H2) are then used to distinguish true SNPs, sequencing errors and A-to-I RNA editing events (right). c, Allele-specific analysis using longcallR. Phased reads are assigned to the closest gene based on the longest exonic matching length and read junctions are extracted from sequencing data. PS1 and PS2 denote two distinct phase blocks (left). For ASE analysis (top right), the module quantifies haplotype-specific read counts and applies a beta-binomial test. For ASJ analysis (bottom right), the module separately counts junction inclusion and exclusion reads for each haplotype and uses Fisher’s exact test to assess haplotype-specific splicing differences. The Benjamin–Hochberg procedure is applied for controlling the false-positive rate.

When applied to over 200 human samples, the tool uncovered a large number of allele-specific splicing events. This means that different copies of the same gene can be processed in distinct ways, leading to different RNA products. On average, each sample showed dozens of these events, many of which had not been previously documented.

One notable finding was that nearly half of these splicing events involved previously unannotated junctions. This suggests that current gene annotations are incomplete and that RNA sequencing continues to reveal new layers of biological complexity.

By linking genetic variation directly to RNA structure and expression, long-read RNA sequencing provides a more detailed picture of gene regulation. Tools like longcallR make it easier for researchers to uncover how genetic differences influence RNA processing, which has important implications for understanding disease mechanisms and developing targeted therapies.

Availability – LongcallR is available via GitHub at https://github.com/huangnengCSU/longcallR.

Huang N, Li H. (2026) SNP calling, haplotype phasing and allele-specific analysis with long RNA-seq reads. Nature Methods (Epub ahead of print)[article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services