Researchers from the Faculty of Engineering at The University of Hong Kong (HKU) have developed two innovative deep-learning algorithms, ClairS-TO and Clair3-RNA, that significantly advance genetic mutation detection in cancer diagnostics and RNA-based genomic studies.

The pioneering research team, led by Professor Ruibang Luo from the School of Computing and Data Science, Faculty of Engineering, has unveiled two groundbreaking deep-learning algorithms—ClairS-TO and Clair3-RNA—set to revolutionise genetic analysis in both clinical and research settings. Leveraging long-read sequencing technologies, these tools significantly improve the accuracy of detecting genetic mutations in complex samples, opening new horizons for precision medicine and genomic discovery. Both research articles have been published in Nature Communications.

Long-read sequencing technologies capture continuous stretches of DNA and RNA, providing detailed insights into genetic information. However, interpreting this data, especially identifying mutations in challenging conditions, has remained a hurdle. The two new algorithms aim to overcome these obstacles, making genomic analysis faster, more accurate, and more accessible.

ClairS-TO addresses a critical challenge in cancer diagnostics: analysing tumour DNA without needing matched healthy tissue samples. Standard methods require both tumour and normal samples for comparison, which are not always available. Using a sophisticated dual-network approach—one to confirm genuine mutations and another to reject errors— ClairS-TO eliminates this requirement. This breakthrough allows for cost-effective, reliable tumour analysis even when sample material is limited, broadening access to precise cancer diagnostics.

Overview of ClairS-TO somatic variant calling and model training workflow

Fig. 1

a The three steps in the somatic variant calling workflow of ClairS-TO. In Step 1, the pileup inputs of variant candidates in the tumor are fed into the affirmative neural network (AFF) and the negational neural network (NEG) to obtain the probabilities of the candidate being a somatic variant and not being a somatic variant, respectively. In Step 2, a joint posterior probability  P(y|(PAFF(y|x),1−PNEG(¬y|x)))

“>

 is calculated using the output probabilities from both networks for inference.  P(y)

“>

 is derived from the same samples used for training AFF and NEG. In Step 3, multiple post-processing strategies are applied to tag non-somatic variants. b The training workflow of ClairS-TO. ClairS-TO provides two types of pre-trained models: a model trained exclusively on synthetic tumor samples (ClairS-TO SS) and a model initially trained on synthetic tumor samples and then augmented with real tumor samples (ClairS-TO SSRS). c The data synthesis workflow of ClairS-TO. For example, using 90-fold coverage of GIAB HG002 (Sample A, with GIAB-known truth germline variants as GA) and 80-fold coverage of GIAB HG001 (Sample B, with GIAB-known truth germline variants as GB) from ONT WGS alignments as sources for synthetic tumor and normal, the alignments were split into smaller non-overlapping chunks with an average of 4-fold coverage. These smaller chunks were used to simulate synthetic tumors with varying coverages and different VAFs. For variants in the synthetic tumor (T), “Somatic” refers to GIAB truth germline variants present in A but not in B (T ∩ (GA−GB)); “Germline” refers to GIAB truth germline variants present in both A and B (T ∩ GA ∩ GB); and “Artifact” refers to variants not present in the GIAB truth germline variants of either A or B (T-GA-GB). When using Sample B as the tumor and Sample A as the normal, the definitions remain identical except for switching the subscripts. d The VAF distributions of synthetic somatic SNVs at simulated tumor purities of 100, 75, 50, and 25%, using either HG002 or HG001 as the tumor source. e The VAF distributions of real somatic SNVs in the four real cancer cell lines (i.e., HCC1937, HCC1954, H1437, and H2009).

Meanwhile, Clair3-RNA marks the world’s first deep-learning-based small variant caller specifically tailored for long-read RNA sequencing. RNA editing and technical sequencing errors can easily confuse the identification of true genetic variants. Clair3-RNA employs advanced deep learning techniques to accurately distinguish real mutations from biological noise and editing, enabling researchers and clinicians to simultaneously analyse gene expression and mutations with exceptional accuracy.

Overview of Clair3-RNA variant-calling workflow

Fig. 1

a Mechanism of cDNA and dRNA sequencing: The figure contrasts RNA sequencing techniques, specifically highlighting the mechanisms of direct RNA sequencing (dRNA) and complementary DNA (cDNA) sequencing with reverse transcription. In dRNA sequencing, Inosine editing is sequenced as is, while in cDNA sequencing, it is converted to Guanine nucleotide. b Training tensor generation: The illustration of training tensor generation outlines the process of deriving training tensors from DNA and RNA alignments, incorporating labels obtained from GIAB truths. Variants located outside the callable region are omitted, and the zygosity flipping is applied, whereby heterozygous variants displaying allelic fractions indicative of homozygous variant or homozygous reference in RNA are reclassified. Read subsampling is employed for data augmentation, and pileup tensors are produced for every candidate site, along with flanking positions. c Tensor generation in Inference: The illustration of tensor generation in inference shows the process of creating tensors in inference. Coverage normalization is adopted for exceeding coverage in inference. d Pileup features: The figure illustrates 18 pileup features for each position in the forward and reverse strands. e Model architecture: In this illustration of the Clair3-RNA model architecture, the pileup network includes a two-layer LSTMs with two dense layers. The output supports two probabilistic tasks for classification: 21-genotype and zygosity. Variants identified by the model are tagged by the REDIportal database and recorded in VCF files.

These algorithms are the latest additions to the renowned Clair series, a suite of artificial intelligence (AI)-driven genomic tools developed by Professor Luo’s team. The series, including the industry-standard Clair3, has become a cornerstone in the field of computational biology. Known for their speed, accuracy, and robustness, these open-source algorithms have amassed over 400,000 downloads. They are widely adopted by leading research institutes and sequencing companies globally, setting the benchmark for processing third-generation sequencing data.

Professor Ruibang Luo commented, “ClairS-TO and Clair3-RNA, along with other algorithms in the Clair series, have established a solid foundation for deep-learning-driven genetic mutation discovery, and accelerated the adoption of precision medicine and clinical genomics.”

These advances represent a significant leap toward more accessible, accurate, and comprehensive genetic analysis. They hold the potential to improve cancer diagnosis, enable personalised medicine, and accelerate genomic research—delivering tangible benefits to patients and scientists around the world.

Availability

SourceThe University of Hong Kong

Chen L, Zheng Z, Su J, Yu X, Wong AOK, Zhang J, Lee YL, Luo R. (2025) ClairS-TO a deep-learning method for long-read tumor-only somatic small variant calling. Nature Communications 16(1):9630. [article]

Zheng Z, Yu X, Chen L, Lee YL, Xin C, Wong AOK, Jain M, Kesharwani RK, Sedlazeck FJ, Luo R. (2025) Clair3-RNA a deep learning-based small variant caller for long-read RNA sequencing data. Nature Communications 16(1):11553. [article]

Researchers from the Faculty of Engineering at The University of Hong Kong (HKU) have developed two innovative deep-learning algorithms, ClairS-TO and Clair3-RNA, that significantly advance genetic mutation detection in cancer diagnostics and RNA-based genomic studies.

The pioneering research team, led by Professor Ruibang Luo from the School of Computing and Data Science, Faculty of Engineering, has unveiled two groundbreaking deep-learning algorithms—ClairS-TO and Clair3-RNA—set to revolutionise genetic analysis in both clinical and research settings. Leveraging long-read sequencing technologies, these tools significantly improve the accuracy of detecting genetic mutations in complex samples, opening new horizons for precision medicine and genomic discovery. Both research articles have been published in Nature Communications.

Long-read sequencing technologies capture continuous stretches of DNA and RNA, providing detailed insights into genetic information. However, interpreting this data, especially identifying mutations in challenging conditions, has remained a hurdle. The two new algorithms aim to overcome these obstacles, making genomic analysis faster, more accurate, and more accessible.

ClairS-TO addresses a critical challenge in cancer diagnostics: analysing tumour DNA without needing matched healthy tissue samples. Standard methods require both tumour and normal samples for comparison, which are not always available. Using a sophisticated dual-network approach—one to confirm genuine mutations and another to reject errors— ClairS-TO eliminates this requirement. This breakthrough allows for cost-effective, reliable tumour analysis even when sample material is limited, broadening access to precise cancer diagnostics.

Overview of ClairS-TO somatic variant calling and model training workflow

Fig. 1

a The three steps in the somatic variant calling workflow of ClairS-TO. In Step 1, the pileup inputs of variant candidates in the tumor are fed into the affirmative neural network (AFF) and the negational neural network (NEG) to obtain the probabilities of the candidate being a somatic variant and not being a somatic variant, respectively. In Step 2, a joint posterior probability  P(y|(PAFF(y|x),1−PNEG(¬y|x)))

“>

 is calculated using the output probabilities from both networks for inference.  P(y)

“>

 is derived from the same samples used for training AFF and NEG. In Step 3, multiple post-processing strategies are applied to tag non-somatic variants. b The training workflow of ClairS-TO. ClairS-TO provides two types of pre-trained models: a model trained exclusively on synthetic tumor samples (ClairS-TO SS) and a model initially trained on synthetic tumor samples and then augmented with real tumor samples (ClairS-TO SSRS). c The data synthesis workflow of ClairS-TO. For example, using 90-fold coverage of GIAB HG002 (Sample A, with GIAB-known truth germline variants as GA) and 80-fold coverage of GIAB HG001 (Sample B, with GIAB-known truth germline variants as GB) from ONT WGS alignments as sources for synthetic tumor and normal, the alignments were split into smaller non-overlapping chunks with an average of 4-fold coverage. These smaller chunks were used to simulate synthetic tumors with varying coverages and different VAFs. For variants in the synthetic tumor (T), “Somatic” refers to GIAB truth germline variants present in A but not in B (T ∩ (GA−GB)); “Germline” refers to GIAB truth germline variants present in both A and B (T ∩ GA ∩ GB); and “Artifact” refers to variants not present in the GIAB truth germline variants of either A or B (T-GA-GB). When using Sample B as the tumor and Sample A as the normal, the definitions remain identical except for switching the subscripts. d The VAF distributions of synthetic somatic SNVs at simulated tumor purities of 100, 75, 50, and 25%, using either HG002 or HG001 as the tumor source. e The VAF distributions of real somatic SNVs in the four real cancer cell lines (i.e., HCC1937, HCC1954, H1437, and H2009).

Meanwhile, Clair3-RNA marks the world’s first deep-learning-based small variant caller specifically tailored for long-read RNA sequencing. RNA editing and technical sequencing errors can easily confuse the identification of true genetic variants. Clair3-RNA employs advanced deep learning techniques to accurately distinguish real mutations from biological noise and editing, enabling researchers and clinicians to simultaneously analyse gene expression and mutations with exceptional accuracy.

Overview of Clair3-RNA variant-calling workflow

Fig. 1

a Mechanism of cDNA and dRNA sequencing: The figure contrasts RNA sequencing techniques, specifically highlighting the mechanisms of direct RNA sequencing (dRNA) and complementary DNA (cDNA) sequencing with reverse transcription. In dRNA sequencing, Inosine editing is sequenced as is, while in cDNA sequencing, it is converted to Guanine nucleotide. b Training tensor generation: The illustration of training tensor generation outlines the process of deriving training tensors from DNA and RNA alignments, incorporating labels obtained from GIAB truths. Variants located outside the callable region are omitted, and the zygosity flipping is applied, whereby heterozygous variants displaying allelic fractions indicative of homozygous variant or homozygous reference in RNA are reclassified. Read subsampling is employed for data augmentation, and pileup tensors are produced for every candidate site, along with flanking positions. c Tensor generation in Inference: The illustration of tensor generation in inference shows the process of creating tensors in inference. Coverage normalization is adopted for exceeding coverage in inference. d Pileup features: The figure illustrates 18 pileup features for each position in the forward and reverse strands. e Model architecture: In this illustration of the Clair3-RNA model architecture, the pileup network includes a two-layer LSTMs with two dense layers. The output supports two probabilistic tasks for classification: 21-genotype and zygosity. Variants identified by the model are tagged by the REDIportal database and recorded in VCF files.

These algorithms are the latest additions to the renowned Clair series, a suite of artificial intelligence (AI)-driven genomic tools developed by Professor Luo’s team. The series, including the industry-standard Clair3, has become a cornerstone in the field of computational biology. Known for their speed, accuracy, and robustness, these open-source algorithms have amassed over 400,000 downloads. They are widely adopted by leading research institutes and sequencing companies globally, setting the benchmark for processing third-generation sequencing data.

Professor Ruibang Luo commented, “ClairS-TO and Clair3-RNA, along with other algorithms in the Clair series, have established a solid foundation for deep-learning-driven genetic mutation discovery, and accelerated the adoption of precision medicine and clinical genomics.”

These advances represent a significant leap toward more accessible, accurate, and comprehensive genetic analysis. They hold the potential to improve cancer diagnosis, enable personalised medicine, and accelerate genomic research—delivering tangible benefits to patients and scientists around the world.

Availability

SourceThe University of Hong Kong

Chen L, Zheng Z, Su J, Yu X, Wong AOK, Zhang J, Lee YL, Luo R. (2025) ClairS-TO a deep-learning method for long-read tumor-only somatic small variant calling. Nature Communications 16(1):9630. [article]

Zheng Z, Yu X, Chen L, Lee YL, Xin C, Wong AOK, Jain M, Kesharwani RK, Sedlazeck FJ, Luo R. (2025) Clair3-RNA a deep learning-based small variant caller for long-read RNA sequencing data. Nature Communications 16(1):11553. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services