Single-cell experiments have transformed biology by allowing researchers to measure gene activity in thousands to millions of individual cells. Techniques such as RNA sequencing reveal how cells differ from one another, but combining data from different experiments remains difficult. Technical noise, batch effects, and the sheer complexity of high-dimensional data can obscure real biological signals and make it hard to compare results across studies, technologies, or even species.

Researchers at the University of California, San Francisco have developed a new computational framework called CONCORD to tackle these challenges in a unified way. Instead of treating batch correction, denoising, and dimensionality reduction as separate steps, CONCORD addresses all three at once using a single self-supervised learning model. This integrated approach simplifies analysis while improving the biological clarity of the results.

CONCORD minibatch sampling enables high-resolution, batch-effect-mitigated representation learning of single-cell data

Fig. 1

a, Schematic of hypothetical cell-state landscapes and corresponding low-dimensional representations that capture key structural features. b, Overview of the CONCORD framework, which replaces the conventional minibatch sampler with a joint hard-negative and dataset-aware sampling scheme, enabling integrated, high-resolution representation learning with a minimalist contrastive model. c, Uniform versus hard-negative sampling in a simulated four-state dataset. Heat maps show simulated expression and latent space, accompanied by density curves with black lines indicating the distribution of cells in an example minibatch under each scheme. Resulting UMAP embeddings are shown. d, Contrastive learning on a single dataset using the conventional uniform sampler, which draws cells uniformly from the entire dataset to form minibatches. e, Standard contrastive learning mixes cells from different datasets within minibatches, amplifying batch effects in the resulting latent embedding. f, CONCORD mitigates batch effects by predominantly contrasting cells within each dataset and randomly shuffling minibatches during training.

At the heart of CONCORD is a clever sampling strategy that learns from the structure of the data itself. By being aware of which dataset each cell comes from, the model corrects batch effects while still preserving meaningful biological differences. At the same time, it uses contrastive learning to sharpen distinctions between cell states, even when those differences are subtle. Notably, CONCORD achieves strong performance using a very simple neural network, avoiding the need for deep or complex architectures.

When applied to diverse single-cell datasets, including RNA sequencing data from different platforms and species, CONCORD produces clean and biologically meaningful representations of cellular identity. These representations capture gene coexpression programs, trace developmental lineages, and preserve both fine-scale and large-scale relationships among cells. As a result, researchers can build high-resolution cell atlases that are easier to interpret and compare across experiments.

By providing a general-purpose framework for integrating and analyzing single-cell data, CONCORD helps researchers move closer to a coherent view of how cell states are organized and how they change over time. This capability is essential for understanding development, disease progression, and the dynamic behavior of cells across biological systems.

Availability – CONCORD is available from GitHub (https://github.com/Gartner-Lab/Concord) under the MIT License.

Zhu Q, Jiang Z, Zuckerman B, Weinberger L, Thomson M, Gartner Z J. (2026) Revealing a coherent cell-state landscape across single-cell datasets with CONCORD. Nature Biotechnology [Epub ahead of print]. [article]

Single-cell experiments have transformed biology by allowing researchers to measure gene activity in thousands to millions of individual cells. Techniques such as RNA sequencing reveal how cells differ from one another, but combining data from different experiments remains difficult. Technical noise, batch effects, and the sheer complexity of high-dimensional data can obscure real biological signals and make it hard to compare results across studies, technologies, or even species.

Researchers at the University of California, San Francisco have developed a new computational framework called CONCORD to tackle these challenges in a unified way. Instead of treating batch correction, denoising, and dimensionality reduction as separate steps, CONCORD addresses all three at once using a single self-supervised learning model. This integrated approach simplifies analysis while improving the biological clarity of the results.

CONCORD minibatch sampling enables high-resolution, batch-effect-mitigated representation learning of single-cell data

Fig. 1

a, Schematic of hypothetical cell-state landscapes and corresponding low-dimensional representations that capture key structural features. b, Overview of the CONCORD framework, which replaces the conventional minibatch sampler with a joint hard-negative and dataset-aware sampling scheme, enabling integrated, high-resolution representation learning with a minimalist contrastive model. c, Uniform versus hard-negative sampling in a simulated four-state dataset. Heat maps show simulated expression and latent space, accompanied by density curves with black lines indicating the distribution of cells in an example minibatch under each scheme. Resulting UMAP embeddings are shown. d, Contrastive learning on a single dataset using the conventional uniform sampler, which draws cells uniformly from the entire dataset to form minibatches. e, Standard contrastive learning mixes cells from different datasets within minibatches, amplifying batch effects in the resulting latent embedding. f, CONCORD mitigates batch effects by predominantly contrasting cells within each dataset and randomly shuffling minibatches during training.

At the heart of CONCORD is a clever sampling strategy that learns from the structure of the data itself. By being aware of which dataset each cell comes from, the model corrects batch effects while still preserving meaningful biological differences. At the same time, it uses contrastive learning to sharpen distinctions between cell states, even when those differences are subtle. Notably, CONCORD achieves strong performance using a very simple neural network, avoiding the need for deep or complex architectures.

When applied to diverse single-cell datasets, including RNA sequencing data from different platforms and species, CONCORD produces clean and biologically meaningful representations of cellular identity. These representations capture gene coexpression programs, trace developmental lineages, and preserve both fine-scale and large-scale relationships among cells. As a result, researchers can build high-resolution cell atlases that are easier to interpret and compare across experiments.

By providing a general-purpose framework for integrating and analyzing single-cell data, CONCORD helps researchers move closer to a coherent view of how cell states are organized and how they change over time. This capability is essential for understanding development, disease progression, and the dynamic behavior of cells across biological systems.

Availability – CONCORD is available from GitHub (https://github.com/Gartner-Lab/Concord) under the MIT License.

Zhu Q, Jiang Z, Zuckerman B, Weinberger L, Thomson M, Gartner Z J. (2026) Revealing a coherent cell-state landscape across single-cell datasets with CONCORD. Nature Biotechnology [Epub ahead of print]. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services