
Genes are controlled by regulatory DNA sequences that determine when and where genes are turned on. Understanding how these DNA sequences influence gene expression is one of the major goals of modern genomics.
Over the past several years, researchers have developed computational models that can predict gene expression directly from DNA sequence. These models have helped scientists better understand gene regulation and how noncoding genetic variants contribute to disease. However, most existing models rely on bulk RNA data collected from mixed populations of cells. Because of this, they often miss the unique regulatory patterns found in individual cell types.
Researchers from Genentech have developed a new model called Decima that addresses this limitation by incorporating large-scale single-cell and single-nucleus RNA sequencing data.
The Decima model and its evaluation
a, Schematic of Decima. Single-cell datasets were combined, filtered and aggregated into a pseudobulk gene expression matrix. Each row of this matrix represents a unique combination of cell type, tissue, disease and study. Each column represents a single gene. The model takes the DNA sequence surrounding a gene and predicts the corresponding column of the matrix. b, Schematic showing evaluation of the trained Decima model on 1,811 test-set genes. c, Histogram of Pearson correlation coefficients between the measured and predicted expression vectors for each pseudobulk (row of the matrix in b), over 1,811 test-set genes. d, Histogram of Pearson correlation coefficients between the measured and predicted per-gene expression vectors (columns of the matrix in b) for each of the 1,811 test-set genes. e, Scatter plot showing the measured and predicted expression (log(CPM + 1)) of FABP1 in all pseudobulks. f, Box plots showing the measured and predicted expression of FABP1 in pseudobulks representing enterocytes, hepatocytes, other gut cells, other liver cells and all remaining pseudobulks. g, Scatter plot showing the measured and predicted expression of DNAH6 in all pseudobulks. h, Box plots showing the expression of DNAH6 in pseudobulks representing ependymal cells, choroid plexus cells, ciliated cells, other central nervous system (CNS) cells, other lung cells and all remaining pseudobulks. i, Scatter plot showing the measured and predicted expression of SPI1 in all pseudobulks. j, Box plots showing the expression of SPI1 in pseudobulks representing monocytes, macrophages, microglia, other blood cells, other CNS cells and all remaining pseudobulks. In all box plots, the center represents the median, lower and upper hinges correspond to the first and third quartiles, whiskers extend to 1.5 × interquartile range (IQR) and remaining points are plotted individually. The number of pseudobulks is indicated in parentheses.Â
Decima was trained using RNA sequencing data from more than 22 million individual cells. This allowed the model to learn how surrounding DNA sequences influence gene expression in highly specific cell types and cellular states.
One of the most important features of Decima is its ability to predict the expression of genes it has never previously seen. The model can also identify regulatory mechanisms that control cell-type-specific gene expression and detect how these patterns change during disease.
The researchers demonstrated that Decima can predict how noncoding genetic variants affect gene expression in particular cell types. This is important because many disease-associated mutations occur outside protein-coding regions of the genome, making them difficult to interpret using traditional approaches.
Another interesting capability of Decima is the ability to design regulatory DNA elements tailored for specific biological contexts. This could eventually support synthetic biology applications and more precise gene therapies.
By combining large-scale RNA sequencing datasets with advanced machine learning, Decima provides a more detailed view of gene regulation across diverse cell types and disease states. Models like this may improve understanding of complex diseases and help researchers identify new therapeutic targets.
Availability – Decima is available via GitHub at https://github.com/Genentech/decima.
Lal A, Karollus A, Gunsalus L, Garfield D, Nair S, Tseng AM, Gordon MG, Blischak J, Van De Geijn B, Bhangale T, Collier JL, Diamant N, Biancalani T, Corrada Bravo H, Scalia G, Eraslan G. (2026) Decoding sequence determinants of gene expression in diverse cellular and disease states. Nature Methods [Epub ahead of print]. [article]

Genes are controlled by regulatory DNA sequences that determine when and where genes are turned on. Understanding how these DNA sequences influence gene expression is one of the major goals of modern genomics.
Over the past several years, researchers have developed computational models that can predict gene expression directly from DNA sequence. These models have helped scientists better understand gene regulation and how noncoding genetic variants contribute to disease. However, most existing models rely on bulk RNA data collected from mixed populations of cells. Because of this, they often miss the unique regulatory patterns found in individual cell types.
Researchers from Genentech have developed a new model called Decima that addresses this limitation by incorporating large-scale single-cell and single-nucleus RNA sequencing data.
The Decima model and its evaluation
a, Schematic of Decima. Single-cell datasets were combined, filtered and aggregated into a pseudobulk gene expression matrix. Each row of this matrix represents a unique combination of cell type, tissue, disease and study. Each column represents a single gene. The model takes the DNA sequence surrounding a gene and predicts the corresponding column of the matrix. b, Schematic showing evaluation of the trained Decima model on 1,811 test-set genes. c, Histogram of Pearson correlation coefficients between the measured and predicted expression vectors for each pseudobulk (row of the matrix in b), over 1,811 test-set genes. d, Histogram of Pearson correlation coefficients between the measured and predicted per-gene expression vectors (columns of the matrix in b) for each of the 1,811 test-set genes. e, Scatter plot showing the measured and predicted expression (log(CPM + 1)) of FABP1 in all pseudobulks. f, Box plots showing the measured and predicted expression of FABP1 in pseudobulks representing enterocytes, hepatocytes, other gut cells, other liver cells and all remaining pseudobulks. g, Scatter plot showing the measured and predicted expression of DNAH6 in all pseudobulks. h, Box plots showing the expression of DNAH6 in pseudobulks representing ependymal cells, choroid plexus cells, ciliated cells, other central nervous system (CNS) cells, other lung cells and all remaining pseudobulks. i, Scatter plot showing the measured and predicted expression of SPI1 in all pseudobulks. j, Box plots showing the expression of SPI1 in pseudobulks representing monocytes, macrophages, microglia, other blood cells, other CNS cells and all remaining pseudobulks. In all box plots, the center represents the median, lower and upper hinges correspond to the first and third quartiles, whiskers extend to 1.5 × interquartile range (IQR) and remaining points are plotted individually. The number of pseudobulks is indicated in parentheses.Â
Decima was trained using RNA sequencing data from more than 22 million individual cells. This allowed the model to learn how surrounding DNA sequences influence gene expression in highly specific cell types and cellular states.
One of the most important features of Decima is its ability to predict the expression of genes it has never previously seen. The model can also identify regulatory mechanisms that control cell-type-specific gene expression and detect how these patterns change during disease.
The researchers demonstrated that Decima can predict how noncoding genetic variants affect gene expression in particular cell types. This is important because many disease-associated mutations occur outside protein-coding regions of the genome, making them difficult to interpret using traditional approaches.
Another interesting capability of Decima is the ability to design regulatory DNA elements tailored for specific biological contexts. This could eventually support synthetic biology applications and more precise gene therapies.
By combining large-scale RNA sequencing datasets with advanced machine learning, Decima provides a more detailed view of gene regulation across diverse cell types and disease states. Models like this may improve understanding of complex diseases and help researchers identify new therapeutic targets.
Availability – Decima is available via GitHub at https://github.com/Genentech/decima.
Lal A, Karollus A, Gunsalus L, Garfield D, Nair S, Tseng AM, Gordon MG, Blischak J, Van De Geijn B, Bhangale T, Collier JL, Diamant N, Biancalani T, Corrada Bravo H, Scalia G, Eraslan G. (2026) Decoding sequence determinants of gene expression in diverse cellular and disease states. Nature Methods [Epub ahead of print]. [article]












Stay Connected