Single-cell RNA sequencing has transformed how researchers study complex tissues by revealing differences in gene expression from one cell to the next. But understanding variation goes beyond expression alone. Cells can differ in the genetic variants they express, including single-nucleotide variants, which can reflect mutations, allele-specific expression, or regulatory processes such as imprinting and transcriptional bursting.

Researchers at George Washington University have developed a new software tool designed to make these patterns easier to explore. The tool, called scSNViz, focuses on visualizing and analyzing expressed genetic variants directly from single-cell RNA sequencing data.

(a) scSNViz workflow: scSNViz accepts either raw or processed gene–cell expression values as its first input, and a list of SNVs with cellular barcodes as its second input. It is recommended that the list of SNVs be processed through SCReadCounts, facilitating quantitative visualization based on the number of reference and variant read counts, expressed as color gradients. SCReadCounts-processed data also enables the distinction between SNV loci with solely reference read counts, solely variant read counts, and no read counts, thus allowing for the differentiation between cells with monoallelic reference expression and monoallelic variant expression. The software tools developed by our team and focused on variant analysis from single cells are shown in red. (b) Sets of SNVs: Overall statistical metrics, accompanied by cell type classification by scType (offered as an option in scSNViz). To enhance visualization, users can select from multiple customizable color designs, including mono- and bi-chromatic gradients. The visualization is exemplified on the publicly accessible sample SAMN12799270, isolated from neuroblastoma tumor tissue; the list of SNVs used is shown in Table 2, available as supplementary data at Bioinformatics online. (c) scSNViz UMAP visualization of individual SNVs with likely germline origin in sample SAMN13012147 (cholangiocarcinoma primary tumor). For each SNV in the submitted list, VAF_RNA, N_VAR, and N_REF are visualized to evaluate allele-specific expression. Distinct patterns differentiate heterozygous from homozygous SNVs, with homozygous variants showing exclusive expression of the variant allele—enabling clear discrimination based on allelic configuration. (d) scSNViz UMAP visualization of individual SNV 17:48895725 C > T in ATP5MC1, corresponding to the previously reported somatic mutation COSV63506293, across three samples from different tumors: prostate cancer, non-small cell lung carcinoma and cholangiocarcinoma. In all samples, similar bi-allelic expression patterns were observed. (e) scSNViz UMAP visualization of individual SNV Y:2865219_C > T in RPS4Y1, with likely RNA origin, across three samples from different tissues: normal fetal adrenal, cholangiocarcinoma and prostate cancer. In all samples, similar expression patterns were observed, with low N_VAR and VAF_RNA values. (f) scSNViz UMAP visualization of individual SNVs from heterozygous loci exhibiting random monoallelic expression at the single-cell level. These patterns are characterized by a similar number of cells expressing either the variant or the reference allele, typically supported by low read counts (often fewer than 10). Such expression profiles are consistent with both X-chromosome inactivation and transcriptional bursting.

(a) scSNViz workflow: scSNViz accepts either raw or processed gene–cell expression values as its first input, and a list of SNVs with cellular barcodes as its second input. (b) Sets of SNVs: Overall statistical metrics, accompanied by cell type classification by scType (offered as an option in scSNViz). (c) scSNViz UMAP visualization of individual SNVs with likely germline origin in sample SAMN13012147 (cholangiocarcinoma primary tumor). (d) scSNViz UMAP visualization of individual SNV 17:48895725 C > T in ATP5MC1, corresponding to the previously reported somatic mutation COSV63506293, across three samples from different tumors: prostate cancer, non-small cell lung carcinoma and cholangiocarcinoma. (e) scSNViz UMAP visualization of individual SNV Y:2865219_C > T in RPS4Y1, with likely RNA origin, across three samples from different tissues: normal fetal adrenal, cholangiocarcinoma and prostate cancer. (f) scSNViz UMAP visualization of individual SNVs from heterozygous loci exhibiting random monoallelic expression at the single-cell level. 

scSNViz works by identifying single-nucleotide variants that are present in RNA sequencing reads and tracking how often each variant is expressed in individual cells. It calculates variant allele fractions, which show the balance between different alleles in each cell, and helps researchers see how these variants are distributed across cell types, clusters, or developmental trajectories.

One of the strengths of scSNViz is its visualization capability. The software can display variant expression in two or three dimensions, making it easier to spot cell populations that share similar genetic signatures. This is particularly useful for studying processes such as lineage relationships, clonal expansion, or how mutations emerge and persist in diseases like cancer.

The tool is designed to integrate smoothly with widely used single-cell analysis frameworks, including Seurat for clustering, Slingshot for trajectory analysis, and CopyKat for copy number profiling. This allows researchers to combine genetic variation data with gene expression, cell identity, and inferred lineage information in a single workflow.

Importantly, scSNViz is built with accessibility in mind. It includes clear documentation and example workflows, making advanced variant analysis approachable even for users with limited bioinformatics experience. By lowering the barrier to exploring expressed genetic variation, the tool opens new opportunities to study how genetic diversity shapes cellular behavior.

Overall, scSNViz adds an important layer to single-cell analysis, helping researchers move from gene expression patterns to a deeper understanding of how genetic variation is expressed and regulated at the level of individual cells.

Availability – scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz

Martinez S, Sharma T, Johnson L, Kim A, Ballesteros Prieto V, Arestakesyan H, Harish S, Dias J, Goldfrank J, Edwards N, Horvath A. (2026) scSNViz visualization and analysis of cell-specific expressed SNVs. Bioinformatics 42(2): btag023. [article]

Single-cell RNA sequencing has transformed how researchers study complex tissues by revealing differences in gene expression from one cell to the next. But understanding variation goes beyond expression alone. Cells can differ in the genetic variants they express, including single-nucleotide variants, which can reflect mutations, allele-specific expression, or regulatory processes such as imprinting and transcriptional bursting.

Researchers at George Washington University have developed a new software tool designed to make these patterns easier to explore. The tool, called scSNViz, focuses on visualizing and analyzing expressed genetic variants directly from single-cell RNA sequencing data.

(a) scSNViz workflow: scSNViz accepts either raw or processed gene–cell expression values as its first input, and a list of SNVs with cellular barcodes as its second input. It is recommended that the list of SNVs be processed through SCReadCounts, facilitating quantitative visualization based on the number of reference and variant read counts, expressed as color gradients. SCReadCounts-processed data also enables the distinction between SNV loci with solely reference read counts, solely variant read counts, and no read counts, thus allowing for the differentiation between cells with monoallelic reference expression and monoallelic variant expression. The software tools developed by our team and focused on variant analysis from single cells are shown in red. (b) Sets of SNVs: Overall statistical metrics, accompanied by cell type classification by scType (offered as an option in scSNViz). To enhance visualization, users can select from multiple customizable color designs, including mono- and bi-chromatic gradients. The visualization is exemplified on the publicly accessible sample SAMN12799270, isolated from neuroblastoma tumor tissue; the list of SNVs used is shown in Table 2, available as supplementary data at Bioinformatics online. (c) scSNViz UMAP visualization of individual SNVs with likely germline origin in sample SAMN13012147 (cholangiocarcinoma primary tumor). For each SNV in the submitted list, VAF_RNA, N_VAR, and N_REF are visualized to evaluate allele-specific expression. Distinct patterns differentiate heterozygous from homozygous SNVs, with homozygous variants showing exclusive expression of the variant allele—enabling clear discrimination based on allelic configuration. (d) scSNViz UMAP visualization of individual SNV 17:48895725 C > T in ATP5MC1, corresponding to the previously reported somatic mutation COSV63506293, across three samples from different tumors: prostate cancer, non-small cell lung carcinoma and cholangiocarcinoma. In all samples, similar bi-allelic expression patterns were observed. (e) scSNViz UMAP visualization of individual SNV Y:2865219_C > T in RPS4Y1, with likely RNA origin, across three samples from different tissues: normal fetal adrenal, cholangiocarcinoma and prostate cancer. In all samples, similar expression patterns were observed, with low N_VAR and VAF_RNA values. (f) scSNViz UMAP visualization of individual SNVs from heterozygous loci exhibiting random monoallelic expression at the single-cell level. These patterns are characterized by a similar number of cells expressing either the variant or the reference allele, typically supported by low read counts (often fewer than 10). Such expression profiles are consistent with both X-chromosome inactivation and transcriptional bursting.

(a) scSNViz workflow: scSNViz accepts either raw or processed gene–cell expression values as its first input, and a list of SNVs with cellular barcodes as its second input. (b) Sets of SNVs: Overall statistical metrics, accompanied by cell type classification by scType (offered as an option in scSNViz). (c) scSNViz UMAP visualization of individual SNVs with likely germline origin in sample SAMN13012147 (cholangiocarcinoma primary tumor). (d) scSNViz UMAP visualization of individual SNV 17:48895725 C > T in ATP5MC1, corresponding to the previously reported somatic mutation COSV63506293, across three samples from different tumors: prostate cancer, non-small cell lung carcinoma and cholangiocarcinoma. (e) scSNViz UMAP visualization of individual SNV Y:2865219_C > T in RPS4Y1, with likely RNA origin, across three samples from different tissues: normal fetal adrenal, cholangiocarcinoma and prostate cancer. (f) scSNViz UMAP visualization of individual SNVs from heterozygous loci exhibiting random monoallelic expression at the single-cell level. 

scSNViz works by identifying single-nucleotide variants that are present in RNA sequencing reads and tracking how often each variant is expressed in individual cells. It calculates variant allele fractions, which show the balance between different alleles in each cell, and helps researchers see how these variants are distributed across cell types, clusters, or developmental trajectories.

One of the strengths of scSNViz is its visualization capability. The software can display variant expression in two or three dimensions, making it easier to spot cell populations that share similar genetic signatures. This is particularly useful for studying processes such as lineage relationships, clonal expansion, or how mutations emerge and persist in diseases like cancer.

The tool is designed to integrate smoothly with widely used single-cell analysis frameworks, including Seurat for clustering, Slingshot for trajectory analysis, and CopyKat for copy number profiling. This allows researchers to combine genetic variation data with gene expression, cell identity, and inferred lineage information in a single workflow.

Importantly, scSNViz is built with accessibility in mind. It includes clear documentation and example workflows, making advanced variant analysis approachable even for users with limited bioinformatics experience. By lowering the barrier to exploring expressed genetic variation, the tool opens new opportunities to study how genetic diversity shapes cellular behavior.

Overall, scSNViz adds an important layer to single-cell analysis, helping researchers move from gene expression patterns to a deeper understanding of how genetic variation is expressed and regulated at the level of individual cells.

Availability – scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz

Martinez S, Sharma T, Johnson L, Kim A, Ballesteros Prieto V, Arestakesyan H, Harish S, Dias J, Goldfrank J, Edwards N, Horvath A. (2026) scSNViz visualization and analysis of cell-specific expressed SNVs. Bioinformatics 42(2): btag023. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services