Single-cell RNA sequencing, or scRNA-seq, allows researchers to measure gene activity in thousands or even millions of individual cells. Instead of averaging gene expression across an entire tissue sample, researchers can examine how different cell types and even individual cells behave.

This detailed view of cellular activity has made single-cell RNA sequencing an increasingly important tool for biomarker discovery. Biomarkers are measurable biological characteristics that can help researchers detect disease, predict how a disease may progress, or determine how a patient might respond to treatment.

Researchers at the School of Computer Science and Engineering at UNSW Sydney in Australia examine how machine learning is increasingly being combined with single-cell RNA sequencing to identify potential biomarkers.

General workflow diagram illustrating scRNA-seq data processing and machine learning for biomarker discovery, divided into three sections: (a) feature selection from train and test sets to identify candidate biomarkers, (b) classification model training and evaluation using selected biomarkers, and (c) downstream assessment via survival, enrichment, differential expression, and literature-based analyses.

(a) Single cell sequencing data comprised of healthy and disease samples is first split into training and test sets. The training set is utilized as input for feature selection to identify potential biomarkers. Identified biomarkers then undergo further evaluation through two approaches: (b) The assessment of classification performance through classification metrics and feature importance score, and (c) The assessment of biological relevance through downstream analysis.

Finding biomarkers within complex cell populations

Traditional RNA sequencing generally measures gene expression across a mixture of cells. This approach can identify important differences between samples, but signals from relatively rare cell populations may be obscured by more abundant cells.

Single-cell RNA sequencing addresses this problem by measuring RNA in individual cells. Researchers can identify different cell populations and examine which genes are active within each population.

This is particularly valuable in diseases such as cancer, where cells from the same tumor may behave very differently. Some cells may respond to therapy, while others may possess molecular characteristics associated with treatment resistance or disease progression.

The large amount of information produced by scRNA-seq, however, creates another challenge. Researchers may need to evaluate the expression of thousands of genes across thousands of cells to determine which molecular patterns are most informative.

Machine learning provides another way to search these complex datasets.

Moving beyond individual gene comparisons

One of the most common methods for identifying potential biomarkers from RNA sequencing data is differential gene expression analysis.

Researchers compare groups, such as healthy and diseased cells, and identify genes whose expression differs significantly between them. This approach has produced many useful discoveries, but it commonly evaluates genes individually.

Biological systems are rarely controlled by a single gene. Groups of genes often work together within pathways and regulatory networks.

Machine learning methods can evaluate multiple genes simultaneously and search for combinations of gene-expression patterns that distinguish one biological condition from another. This multivariable approach may identify useful biomarkers that would be less apparent when genes are examined independently.

Training computers to recognize disease-associated patterns

Many machine learning approaches used for single-cell biomarker discovery rely on supervised learning.

In supervised learning, researchers provide an algorithm with examples whose classifications are already known. For example, the algorithm might receive gene-expression profiles from cells obtained from patients with a particular disease and from healthy individuals.

The algorithm attempts to identify patterns that distinguish the groups.

Genes whose expression strongly contributes to that classification may then become potential biomarker candidates.

Researchers can divide their data into training and testing groups. The machine learning model learns from the training data and is then evaluated using data it has not previously seen. This helps determine whether the identified gene patterns can reliably classify new samples rather than simply describing the original dataset.

Biomarkers can be identified at different biological levels

An important feature of single-cell RNA sequencing is that researchers can approach biomarker discovery from more than one level.

Machine learning can classify individual cells based on their gene-expression profiles. This may help identify molecular signatures associated with a particular disease-related cell population.

Researchers can also analyze data at the patient level. Gene-expression information from many cells can be combined to search for molecular patterns capable of distinguishing patients with different diseases, outcomes, or responses to treatment.

The researchers describe both cell-level and patient-level approaches, highlighting how the appropriate strategy depends on the biological question being investigated.

Selecting the genes that matter most

One of the major challenges with single-cell RNA sequencing is the enormous number of potential variables.

A dataset might contain expression measurements for thousands of genes, but only a relatively small number may provide useful information for distinguishing disease from healthy samples.

Machine learning workflows therefore commonly include a process known as feature selection.

In this context, the features are genes and their expression levels. Feature-selection methods attempt to identify the genes that provide the most useful information for classification.

Reducing thousands of genes to a smaller group of informative candidates can make machine learning models easier to interpret while also providing researchers with potential biomarkers for further investigation.

Prediction alone is not enough

A machine learning model that accurately classifies cells or patients does not automatically prove that the genes it identifies are biologically meaningful biomarkers.

Researchers still need to investigate what those genes do.

Candidate biomarkers can be examined to determine whether they participate in disease-related pathways, correlate with patient outcomes, show consistent expression differences, or have previously been associated with the disease.

Independent validation is also important. A biomarker that performs well in one dataset may not necessarily perform equally well in another population.

The researchers emphasize that machine learning performance and biological relevance both need to be considered when evaluating potential biomarkers.

A rapidly developing field still needs common standards

Machine learning approaches for single-cell biomarker discovery remain highly diverse.

Researchers are using different algorithms, feature-selection strategies, classification metrics, and levels of analysis. This diversity encourages experimentation, but it also makes it difficult to compare results between research groups.

Another challenge is accessibility. Many machine learning approaches require substantial programming and computational expertise, creating a barrier for researchers whose primary training is in molecular biology or medicine.

The authors suggest that standardized benchmarking methods and more accessible software tools will be important for increasing the reproducibility and adoption of these approaches.

Combining single-cell biology with machine learning

Single-cell RNA sequencing provides an unusually detailed picture of gene activity by revealing molecular differences between individual cells. Machine learning offers methods for searching that enormous amount of information for combinations of genes that may be associated with disease.

Together, these technologies could help researchers move beyond searching for individual differentially expressed genes and instead identify more complex molecular signatures associated with specific cells, patients, or biological conditions.

The challenge now is determining which computational methods provide reliable and reproducible biomarkers. As machine learning methods become better standardized and easier to use, their combination with single-cell RNA sequencing could become an increasingly useful component of biomarker discovery and precision medicine.

Dewa G, Munier CML, Ballouz S, Louie R. (2026) Machine learning approaches for biomarker discovery using single-cell RNA sequencing. Frontiers in Bioinformatics 6: 1767362. [article]

Single-cell RNA sequencing, or scRNA-seq, allows researchers to measure gene activity in thousands or even millions of individual cells. Instead of averaging gene expression across an entire tissue sample, researchers can examine how different cell types and even individual cells behave.

This detailed view of cellular activity has made single-cell RNA sequencing an increasingly important tool for biomarker discovery. Biomarkers are measurable biological characteristics that can help researchers detect disease, predict how a disease may progress, or determine how a patient might respond to treatment.

Researchers at the School of Computer Science and Engineering at UNSW Sydney in Australia examine how machine learning is increasingly being combined with single-cell RNA sequencing to identify potential biomarkers.

General workflow diagram illustrating scRNA-seq data processing and machine learning for biomarker discovery, divided into three sections: (a) feature selection from train and test sets to identify candidate biomarkers, (b) classification model training and evaluation using selected biomarkers, and (c) downstream assessment via survival, enrichment, differential expression, and literature-based analyses.

(a) Single cell sequencing data comprised of healthy and disease samples is first split into training and test sets. The training set is utilized as input for feature selection to identify potential biomarkers. Identified biomarkers then undergo further evaluation through two approaches: (b) The assessment of classification performance through classification metrics and feature importance score, and (c) The assessment of biological relevance through downstream analysis.

Finding biomarkers within complex cell populations

Traditional RNA sequencing generally measures gene expression across a mixture of cells. This approach can identify important differences between samples, but signals from relatively rare cell populations may be obscured by more abundant cells.

Single-cell RNA sequencing addresses this problem by measuring RNA in individual cells. Researchers can identify different cell populations and examine which genes are active within each population.

This is particularly valuable in diseases such as cancer, where cells from the same tumor may behave very differently. Some cells may respond to therapy, while others may possess molecular characteristics associated with treatment resistance or disease progression.

The large amount of information produced by scRNA-seq, however, creates another challenge. Researchers may need to evaluate the expression of thousands of genes across thousands of cells to determine which molecular patterns are most informative.

Machine learning provides another way to search these complex datasets.

Moving beyond individual gene comparisons

One of the most common methods for identifying potential biomarkers from RNA sequencing data is differential gene expression analysis.

Researchers compare groups, such as healthy and diseased cells, and identify genes whose expression differs significantly between them. This approach has produced many useful discoveries, but it commonly evaluates genes individually.

Biological systems are rarely controlled by a single gene. Groups of genes often work together within pathways and regulatory networks.

Machine learning methods can evaluate multiple genes simultaneously and search for combinations of gene-expression patterns that distinguish one biological condition from another. This multivariable approach may identify useful biomarkers that would be less apparent when genes are examined independently.

Training computers to recognize disease-associated patterns

Many machine learning approaches used for single-cell biomarker discovery rely on supervised learning.

In supervised learning, researchers provide an algorithm with examples whose classifications are already known. For example, the algorithm might receive gene-expression profiles from cells obtained from patients with a particular disease and from healthy individuals.

The algorithm attempts to identify patterns that distinguish the groups.

Genes whose expression strongly contributes to that classification may then become potential biomarker candidates.

Researchers can divide their data into training and testing groups. The machine learning model learns from the training data and is then evaluated using data it has not previously seen. This helps determine whether the identified gene patterns can reliably classify new samples rather than simply describing the original dataset.

Biomarkers can be identified at different biological levels

An important feature of single-cell RNA sequencing is that researchers can approach biomarker discovery from more than one level.

Machine learning can classify individual cells based on their gene-expression profiles. This may help identify molecular signatures associated with a particular disease-related cell population.

Researchers can also analyze data at the patient level. Gene-expression information from many cells can be combined to search for molecular patterns capable of distinguishing patients with different diseases, outcomes, or responses to treatment.

The researchers describe both cell-level and patient-level approaches, highlighting how the appropriate strategy depends on the biological question being investigated.

Selecting the genes that matter most

One of the major challenges with single-cell RNA sequencing is the enormous number of potential variables.

A dataset might contain expression measurements for thousands of genes, but only a relatively small number may provide useful information for distinguishing disease from healthy samples.

Machine learning workflows therefore commonly include a process known as feature selection.

In this context, the features are genes and their expression levels. Feature-selection methods attempt to identify the genes that provide the most useful information for classification.

Reducing thousands of genes to a smaller group of informative candidates can make machine learning models easier to interpret while also providing researchers with potential biomarkers for further investigation.

Prediction alone is not enough

A machine learning model that accurately classifies cells or patients does not automatically prove that the genes it identifies are biologically meaningful biomarkers.

Researchers still need to investigate what those genes do.

Candidate biomarkers can be examined to determine whether they participate in disease-related pathways, correlate with patient outcomes, show consistent expression differences, or have previously been associated with the disease.

Independent validation is also important. A biomarker that performs well in one dataset may not necessarily perform equally well in another population.

The researchers emphasize that machine learning performance and biological relevance both need to be considered when evaluating potential biomarkers.

A rapidly developing field still needs common standards

Machine learning approaches for single-cell biomarker discovery remain highly diverse.

Researchers are using different algorithms, feature-selection strategies, classification metrics, and levels of analysis. This diversity encourages experimentation, but it also makes it difficult to compare results between research groups.

Another challenge is accessibility. Many machine learning approaches require substantial programming and computational expertise, creating a barrier for researchers whose primary training is in molecular biology or medicine.

The authors suggest that standardized benchmarking methods and more accessible software tools will be important for increasing the reproducibility and adoption of these approaches.

Combining single-cell biology with machine learning

Single-cell RNA sequencing provides an unusually detailed picture of gene activity by revealing molecular differences between individual cells. Machine learning offers methods for searching that enormous amount of information for combinations of genes that may be associated with disease.

Together, these technologies could help researchers move beyond searching for individual differentially expressed genes and instead identify more complex molecular signatures associated with specific cells, patients, or biological conditions.

The challenge now is determining which computational methods provide reliable and reproducible biomarkers. As machine learning methods become better standardized and easier to use, their combination with single-cell RNA sequencing could become an increasingly useful component of biomarker discovery and precision medicine.

Dewa G, Munier CML, Ballouz S, Louie R. (2026) Machine learning approaches for biomarker discovery using single-cell RNA sequencing. Frontiers in Bioinformatics 6: 1767362. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services