Understanding how genes vary and which RNA isoforms are produced is key to modern biology and medicine. Long-read RNA sequencing makes this possible by reading entire RNA molecules in one piece, but it also introduces a challenge, the data are noisy, with higher error rates and complex signals from RNA editing and many transcript forms.

Researchers from the University of Hong Kong have developed Clair3-RNA, a deep learning tool designed specifically to call genetic variants from long-read RNA data.

Overview of Clair3-RNA variant-calling workflow

(a) The figure contrasts RNA sequencing techniques, specifically highlighting the mechanisms of direct RNA sequencing (dRNA) and complementary DNA (cDNA) sequencing with reverse transcription. In dRNA sequencing, Inosine editing is sequenced as is, while in cDNA sequencing, it is converted to Guanine nucleotide. (b) The illustration of training tensor generation outlines the process of deriving training tensors from DNA and RNA alignments, incorporating labels obtained from GIAB truths. Variants located outside the callable region are omitted, and the zygosity switch is applied, whereby heterozygous variants displaying allelic fractions indicative of homozygous variant or homozygous reference in RNA are reclassified. Read subsampling is employed for data augmentation, and pileup tensors are produced for every candidate site, along with flanking positions. (c) The illustration of tensor generation in inference shows the process of creating tensors in inference. Coverage normalization is adopted for exceeding coverage in inference. (d) The figure illustrates 18 pileup features for each position in the forward and reverse strands. More details are provided in the Supplementary methods − Description of RNA pileup input features section. (e) In this illustration of the Clair3-RNA model architecture, the pileup network includes a two-layer LSTMs with two dense layers. The output supports two probabilistic tasks for classification: 21-genotype and zygosity. Variants identified by the model are tagged by the REDIportal database and recorded in VCF files.

Clair3-RNA builds on earlier Clair methods that were successful for DNA sequencing. The team adapted the approach for RNA by handling uneven coverage across transcripts, improving training datasets, detecting RNA editing sites, and using haplotype phasing to better separate variants coming from different chromosome copies. These improvements help the model distinguish real genetic changes from sequencing errors.

One of the strengths of Clair3-RNA is its flexibility. It works with data from multiple platforms, including PacBio Iso-Seq and MAS-Seq, as well as Oxford Nanopore Technologies complementary DNA and direct RNA sequencing. This makes it useful for many labs using different long-read technologies.

When tested on benchmark samples, Clair3-RNA reached high accuracy for identifying single nucleotide variants. With moderate to high read coverage, the method achieved F1-scores in the mid to high 90 percent range, and performance improved even further after phasing. Importantly, it also performed well at finding RNA editing events, which are chemical changes to RNA that do not exist in the DNA but can affect how genes function.

For researchers interested in full-length transcripts, allele-specific expression, or disease-related variants, tools like Clair3-RNA can make long-read RNA sequencing far more informative. Because the software is open source, the community can use it, test it, and continue to improve it for future applications.

Availability – Clair3-RNA is open-source at (https://github.com/HKU-BAL/Clair3-RNA).

Luo R, Cheng H, Wong K C, Au K F, Lam T W. (2025) Clair3-RNA, a deep learning-based variant caller tailored for long-read RNA sequencing. Bioinformatics 41(1): 1–10. [article]

Understanding how genes vary and which RNA isoforms are produced is key to modern biology and medicine. Long-read RNA sequencing makes this possible by reading entire RNA molecules in one piece, but it also introduces a challenge, the data are noisy, with higher error rates and complex signals from RNA editing and many transcript forms.

Researchers from the University of Hong Kong have developed Clair3-RNA, a deep learning tool designed specifically to call genetic variants from long-read RNA data.

Overview of Clair3-RNA variant-calling workflow

(a) The figure contrasts RNA sequencing techniques, specifically highlighting the mechanisms of direct RNA sequencing (dRNA) and complementary DNA (cDNA) sequencing with reverse transcription. In dRNA sequencing, Inosine editing is sequenced as is, while in cDNA sequencing, it is converted to Guanine nucleotide. (b) The illustration of training tensor generation outlines the process of deriving training tensors from DNA and RNA alignments, incorporating labels obtained from GIAB truths. Variants located outside the callable region are omitted, and the zygosity switch is applied, whereby heterozygous variants displaying allelic fractions indicative of homozygous variant or homozygous reference in RNA are reclassified. Read subsampling is employed for data augmentation, and pileup tensors are produced for every candidate site, along with flanking positions. (c) The illustration of tensor generation in inference shows the process of creating tensors in inference. Coverage normalization is adopted for exceeding coverage in inference. (d) The figure illustrates 18 pileup features for each position in the forward and reverse strands. More details are provided in the Supplementary methods − Description of RNA pileup input features section. (e) In this illustration of the Clair3-RNA model architecture, the pileup network includes a two-layer LSTMs with two dense layers. The output supports two probabilistic tasks for classification: 21-genotype and zygosity. Variants identified by the model are tagged by the REDIportal database and recorded in VCF files.

Clair3-RNA builds on earlier Clair methods that were successful for DNA sequencing. The team adapted the approach for RNA by handling uneven coverage across transcripts, improving training datasets, detecting RNA editing sites, and using haplotype phasing to better separate variants coming from different chromosome copies. These improvements help the model distinguish real genetic changes from sequencing errors.

One of the strengths of Clair3-RNA is its flexibility. It works with data from multiple platforms, including PacBio Iso-Seq and MAS-Seq, as well as Oxford Nanopore Technologies complementary DNA and direct RNA sequencing. This makes it useful for many labs using different long-read technologies.

When tested on benchmark samples, Clair3-RNA reached high accuracy for identifying single nucleotide variants. With moderate to high read coverage, the method achieved F1-scores in the mid to high 90 percent range, and performance improved even further after phasing. Importantly, it also performed well at finding RNA editing events, which are chemical changes to RNA that do not exist in the DNA but can affect how genes function.

For researchers interested in full-length transcripts, allele-specific expression, or disease-related variants, tools like Clair3-RNA can make long-read RNA sequencing far more informative. Because the software is open source, the community can use it, test it, and continue to improve it for future applications.

Availability – Clair3-RNA is open-source at (https://github.com/HKU-BAL/Clair3-RNA).

Luo R, Cheng H, Wong K C, Au K F, Lam T W. (2025) Clair3-RNA, a deep learning-based variant caller tailored for long-read RNA sequencing. Bioinformatics 41(1): 1–10. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services