RNA sequencing allows researchers to measure which genes are active in a cell or tissue and how strongly they are expressed. However, determining whether different copies, or alleles, of the same gene are expressed at different levels can be more complicated. Errors introduced while sequencing reads are mapped to a reference genome can make one allele appear more or less active than it actually is.
A new analysis led by researchers at the National Institute of Biology, Slovenia, examines how long-read RNA sequencing can help reduce these problems.
The challenge of measuring allelic imbalance
Most organisms carry multiple versions of many genes. In diploid organisms such as humans, one copy generally comes from each parent. These two copies may not always be expressed at the same level, a phenomenon known as allelic imbalance or allele-specific expression.
Accurately detecting these differences can provide important information about how genetic variation influences gene regulation. The problem is that apparent differences between alleles do not always reflect biology. They can also result from the way sequencing reads are assigned to a reference genome.
Mapping bias occurs when reads from one allele align more easily to the reference genome than reads from another allele. Genome assembly and annotation errors can introduce additional problems. Together, these issues can generate both false positives, where researchers detect an allelic difference that is not actually present, and false negatives, where a genuine difference goes undetected.
Using long reads to improve allele assignment
The researchers examined whether long-read RNA sequencing combined with relatively straightforward quality-control procedures could improve allele-specific expression measurements.
Unlike short-read sequencing, which breaks RNA-derived sequences into relatively small fragments, long-read technologies can sequence much larger portions of individual RNA molecules. The additional sequence information can make it easier to determine where a read originated and which allele it represents.
The team evaluated the approach using data from four very different species: the fruit fly Drosophila melanogaster, potato Solanum tuberosum, Sumatran orangutan Pongo abelii, and humans. These organisms provided a useful range of genomic complexity, including diploid, highly heterozygous, and polyploid genomes.
Personalized genomes can reduce mapping bias
One of the team’s main recommendations is to map RNA sequencing reads against a personalized genome whenever possible.
A standard reference genome represents only one version of a species’ genome and does not contain every genetic variant found in an individual sample. Reads containing variants that differ from the reference may therefore map less efficiently, potentially favoring the allele that more closely resembles the reference sequence.
A personalized reference incorporates variants from the individual being analyzed. The researchers found that this approach can increase the number of reads that can be confidently assigned to specific alleles and reduce bias in allele counts.
The reference used affects equivalence groups. Tracking multimapping results in equivalent allele and genotype counts regardless of whether mapping is in parallel or competitive.
(a) Haploid reference genomes and mapping bias. Reads containing the reference allele may be more likely to align than those containing alternative alleles. (b) Haplotype-phased personalized reference genomes mitigate mapping bias. RNA-seq data can be mapped either (c) in parallel to each haplotype or (d) to a combined genotype reference. The proportion of default uniquely mapped reads (black) versus multimapping reads (gray) differs in these two strategies. Tracking the multimapping and equivalently mapping reads separates gene-specific and multigenic reads. Long-read RNA-seq is expected to have less ambiguity and overall fewer multimapping reads compared to short-read RNA-seq. (e–g) Case studies for the two mapping strategies applied to (e) diploid D. melanogaster using a SNP-updated personalized reference, (f) diploid Pongo abelii with a phased assembly for the individual assayed, and (g) tetraploid Solanum tuberosum with a phased assembly for the cultivar assayed. (h) Barplot of the number of reads in each category with concordant mapping locations in parallel and competitive mapping for S. tuberosum. Reads mapping to only one allele are assigned to mapping category “unique_haplotype[1,2,3,4]”, reads mapping equally well to multiple alleles of a gene are assigned as “multimapping_same_gene”, these are the reads which will be discordant without tracking, and reads mapping equally well to multiple alleles and genes are classified as “multimapping_multi_genes” (i) Heatmap showing number of reads with different mapping location in parallel and competitive mapping for S. tuberosum.
Multimapping reads require careful attention
Another important source of uncertainty comes from multimapping reads, sequences that can align to more than one location in the genome.
Instead of simply ignoring this issue, the researchers recommend tracking multimapping reads and adjusting mapping parameters to determine how different settings affect allele and gene expression measurements. Careful evaluation of these reads can help prevent mapping decisions from producing misleading expression differences.
This is particularly important in genomes containing duplicated genes, repetitive regions, or multiple similar chromosome copies.
Extreme allele bias can signal genome problems
The researchers also recommend investigating cases in which RNA sequencing appears to show an extremely large difference between alleles.
While such differences can be biologically real, unusually strong allelic imbalance can also indicate problems with the underlying genome assembly or gene annotation. Examining these extreme cases can therefore serve as a useful quality-control step.
Rather than automatically interpreting a highly imbalanced result as an important biological discovery, researchers can first determine whether technical factors provide a better explanation.
Improving confidence in allele-specific RNA sequencing
The findings demonstrate that long-read RNA sequencing alone does not eliminate every source of error in allele-specific expression analysis. How sequencing reads are mapped, how the reference genome is constructed, and how ambiguous reads are handled can all influence the final results.
Importantly, the researchers show that reducing these biases does not necessarily require an overly complicated analysis. Personalized reference genomes, monitoring of multimapping reads, careful selection of mapping parameters, and investigation of extreme allele biases provide practical quality-control measures that can improve the reliability of allele and gene expression measurements.
As long-read RNA sequencing becomes more widely used, these approaches could help researchers distinguish genuine biological differences between alleles from artifacts introduced during data analysis.
Nolte N, Petek M, Angulo Lara P, Mulroney L, Nicassio F, Marroni F, McIntyre L. (2026) The promise of long-read RNA-seq: reducing bias in analyses of allele imbalance NAR Genomics and Bioinformatics 8(3): lqag071. [article]
RNA sequencing allows researchers to measure which genes are active in a cell or tissue and how strongly they are expressed. However, determining whether different copies, or alleles, of the same gene are expressed at different levels can be more complicated. Errors introduced while sequencing reads are mapped to a reference genome can make one allele appear more or less active than it actually is.
A new analysis led by researchers at the National Institute of Biology, Slovenia, examines how long-read RNA sequencing can help reduce these problems.
The challenge of measuring allelic imbalance
Most organisms carry multiple versions of many genes. In diploid organisms such as humans, one copy generally comes from each parent. These two copies may not always be expressed at the same level, a phenomenon known as allelic imbalance or allele-specific expression.
Accurately detecting these differences can provide important information about how genetic variation influences gene regulation. The problem is that apparent differences between alleles do not always reflect biology. They can also result from the way sequencing reads are assigned to a reference genome.
Mapping bias occurs when reads from one allele align more easily to the reference genome than reads from another allele. Genome assembly and annotation errors can introduce additional problems. Together, these issues can generate both false positives, where researchers detect an allelic difference that is not actually present, and false negatives, where a genuine difference goes undetected.
Using long reads to improve allele assignment
The researchers examined whether long-read RNA sequencing combined with relatively straightforward quality-control procedures could improve allele-specific expression measurements.
Unlike short-read sequencing, which breaks RNA-derived sequences into relatively small fragments, long-read technologies can sequence much larger portions of individual RNA molecules. The additional sequence information can make it easier to determine where a read originated and which allele it represents.
The team evaluated the approach using data from four very different species: the fruit fly Drosophila melanogaster, potato Solanum tuberosum, Sumatran orangutan Pongo abelii, and humans. These organisms provided a useful range of genomic complexity, including diploid, highly heterozygous, and polyploid genomes.
Personalized genomes can reduce mapping bias
One of the team’s main recommendations is to map RNA sequencing reads against a personalized genome whenever possible.
A standard reference genome represents only one version of a species’ genome and does not contain every genetic variant found in an individual sample. Reads containing variants that differ from the reference may therefore map less efficiently, potentially favoring the allele that more closely resembles the reference sequence.
A personalized reference incorporates variants from the individual being analyzed. The researchers found that this approach can increase the number of reads that can be confidently assigned to specific alleles and reduce bias in allele counts.
The reference used affects equivalence groups. Tracking multimapping results in equivalent allele and genotype counts regardless of whether mapping is in parallel or competitive.
(a) Haploid reference genomes and mapping bias. Reads containing the reference allele may be more likely to align than those containing alternative alleles. (b) Haplotype-phased personalized reference genomes mitigate mapping bias. RNA-seq data can be mapped either (c) in parallel to each haplotype or (d) to a combined genotype reference. The proportion of default uniquely mapped reads (black) versus multimapping reads (gray) differs in these two strategies. Tracking the multimapping and equivalently mapping reads separates gene-specific and multigenic reads. Long-read RNA-seq is expected to have less ambiguity and overall fewer multimapping reads compared to short-read RNA-seq. (e–g) Case studies for the two mapping strategies applied to (e) diploid D. melanogaster using a SNP-updated personalized reference, (f) diploid Pongo abelii with a phased assembly for the individual assayed, and (g) tetraploid Solanum tuberosum with a phased assembly for the cultivar assayed. (h) Barplot of the number of reads in each category with concordant mapping locations in parallel and competitive mapping for S. tuberosum. Reads mapping to only one allele are assigned to mapping category “unique_haplotype[1,2,3,4]”, reads mapping equally well to multiple alleles of a gene are assigned as “multimapping_same_gene”, these are the reads which will be discordant without tracking, and reads mapping equally well to multiple alleles and genes are classified as “multimapping_multi_genes” (i) Heatmap showing number of reads with different mapping location in parallel and competitive mapping for S. tuberosum.
Multimapping reads require careful attention
Another important source of uncertainty comes from multimapping reads, sequences that can align to more than one location in the genome.
Instead of simply ignoring this issue, the researchers recommend tracking multimapping reads and adjusting mapping parameters to determine how different settings affect allele and gene expression measurements. Careful evaluation of these reads can help prevent mapping decisions from producing misleading expression differences.
This is particularly important in genomes containing duplicated genes, repetitive regions, or multiple similar chromosome copies.
Extreme allele bias can signal genome problems
The researchers also recommend investigating cases in which RNA sequencing appears to show an extremely large difference between alleles.
While such differences can be biologically real, unusually strong allelic imbalance can also indicate problems with the underlying genome assembly or gene annotation. Examining these extreme cases can therefore serve as a useful quality-control step.
Rather than automatically interpreting a highly imbalanced result as an important biological discovery, researchers can first determine whether technical factors provide a better explanation.
Improving confidence in allele-specific RNA sequencing
The findings demonstrate that long-read RNA sequencing alone does not eliminate every source of error in allele-specific expression analysis. How sequencing reads are mapped, how the reference genome is constructed, and how ambiguous reads are handled can all influence the final results.
Importantly, the researchers show that reducing these biases does not necessarily require an overly complicated analysis. Personalized reference genomes, monitoring of multimapping reads, careful selection of mapping parameters, and investigation of extreme allele biases provide practical quality-control measures that can improve the reliability of allele and gene expression measurements.
As long-read RNA sequencing becomes more widely used, these approaches could help researchers distinguish genuine biological differences between alleles from artifacts introduced during data analysis.
Nolte N, Petek M, Angulo Lara P, Mulroney L, Nicassio F, Marroni F, McIntyre L. (2026) The promise of long-read RNA-seq: reducing bias in analyses of allele imbalance NAR Genomics and Bioinformatics 8(3): lqag071. [article]












Stay Connected