Computational biologists at Carnegie Mellon University have devised an algorithm to rapidly sort through mountains of gene expression data to find unexpected phenomena that might merit further study. What’s more, the algorithm then re-examines its own output, looking for mistakes it has made and then correcting them.
This work by Carl Kingsford, a professor in CMU’s Computational Biology Department, and Cong Ma, a Ph.D. student in computational biology, is the first attempt at automating the search for these anomalies in gene expression inferred by RNA sequencing, or RNA-seq, the leading method for inferring the activity level of genes.
As they report today in the journal Cell Systems, the researchers already have detected 88 anomalies — unexpectedly high or low levels of expression of regions within genes — in two widely used RNA-seq libraries that are both common and not previously known.
“We don’t yet know why we’re seeing those 88 weird patterns,” Kingsford said, noting that they could be a subject of further investigation.
Though an organism’s genetic makeup is static, the activity level, or expression, of genes varies greatly over time. Gene expression analysis has thus become a major tool for biological research, as well as for diagnosing and monitoring cancers.
Anomalies can be important clues for researchers, but until now finding them has been a painstaking, manual process, sometimes called “sequence gazing.” Finding one anomaly might require examining 200,000 transcript sequences — sequences of RNA that encode information from the gene’s DNA, Kingsford said. Most researchers therefore zero in on regions of genes that they think are important, largely ignoring the vast majority of potential anomalies.
The algorithm developed by Ma and Kingsford automates the search for anomalies, enabling researchers to consider all of the transcript sequences, not just those regions where they expect to see anomalies. This technology could uncover many new phenomena, such as the 88 previously unknown common anomalies found in the multi-tissue RNA-seq libraries.

But Ma noted that identifying anomalies is often not clear cut. Some RNA-seq “reads,” for instance, are common to multiple genes and transcripts and sometimes get mapped to the wrong one. If that occurs, a genetic region might appear more or less active than expected. So the algorithm re-examines any anomalies it detects and sees if they disappear when the RNA-seq reads are redistributed between the genes.
“By correcting anomalies when possible, we reduce the number of falsely predicted instances of differential expression,” Ma said.
Ma C, Kingsford C. (2019) Detecting, Categorizing, and Correcting Coverage Anomalies of RNA-Seq Quantification. Cell Syst [Epub ahead of print]. [article]
Source – Carnegie Mellon University
Computational biologists at Carnegie Mellon University have devised an algorithm to rapidly sort through mountains of gene expression data to find unexpected phenomena that might merit further study. What’s more, the algorithm then re-examines its own output, looking for mistakes it has made and then correcting them.
This work by Carl Kingsford, a professor in CMU’s Computational Biology Department, and Cong Ma, a Ph.D. student in computational biology, is the first attempt at automating the search for these anomalies in gene expression inferred by RNA sequencing, or RNA-seq, the leading method for inferring the activity level of genes.
As they report today in the journal Cell Systems, the researchers already have detected 88 anomalies — unexpectedly high or low levels of expression of regions within genes — in two widely used RNA-seq libraries that are both common and not previously known.
Though an organism’s genetic makeup is static, the activity level, or expression, of genes varies greatly over time. Gene expression analysis has thus become a major tool for biological research, as well as for diagnosing and monitoring cancers.
Anomalies can be important clues for researchers, but until now finding them has been a painstaking, manual process, sometimes called “sequence gazing.” Finding one anomaly might require examining 200,000 transcript sequences — sequences of RNA that encode information from the gene’s DNA, Kingsford said. Most researchers therefore zero in on regions of genes that they think are important, largely ignoring the vast majority of potential anomalies.
The algorithm developed by Ma and Kingsford automates the search for anomalies, enabling researchers to consider all of the transcript sequences, not just those regions where they expect to see anomalies. This technology could uncover many new phenomena, such as the 88 previously unknown common anomalies found in the multi-tissue RNA-seq libraries.
But Ma noted that identifying anomalies is often not clear cut. Some RNA-seq “reads,” for instance, are common to multiple genes and transcripts and sometimes get mapped to the wrong one. If that occurs, a genetic region might appear more or less active than expected. So the algorithm re-examines any anomalies it detects and sees if they disappear when the RNA-seq reads are redistributed between the genes.
Ma C, Kingsford C. (2019) Detecting, Categorizing, and Correcting Coverage Anomalies of RNA-Seq Quantification. Cell Syst [Epub ahead of print]. [article]
Source – Carnegie Mellon University
Related Posts
RNA sequencing reveals functional chimeric mRNAs in mammalian immunity
Atlas of the brain’s striatum could guide researchers to new drug treatments
Immune cells offer insights on billion-dollar virus
A functionally integrated cross-tissue alternative splicing program during short-term calorie restriction
Dietary oxidized plant sterol shifts macrophage state to fuel aortic inflammation
Unlocking the past – new method helps gain insights into old tissue
Novel AI model trained on RNA-Seq data accurately detects key gene mutations and predicts biomarkers across 32 cancer types
Transcriptomic aging clock reveals age-related molecular patterns in opioid dependence
RNA sequencing helps predict stem cell transplant benefit in pediatric AML
Protein ‘switch’ determines whether liposarcoma cells will become aggressive
Precursor tRNAs sense temperature changes: heat stress-induced capped pre-tRNAs suppress protein synthesis
Ketamine increases neuroplasticity in female mice but not in males
Somatic mutations linked to vascular damage in progeria
Scientists map dormant cancer cells’ hideouts, opening new targets for treatment
Soluble signals released by neighboring cells direct how the human kidney is built
Genetics influence how cancer arises – and how it evolves
RNA-based testing uncovers extraordinary diversity in mutations driving lung cancer
Study offers new insights into why ex-smokers remain at elevated risk of lung disease
Learning the grammar of gene regulation
New findings could transform new treatment for rare brain tumor astroblastoma
Computational biologists at Carnegie Mellon University have devised an algorithm to rapidly sort through mountains of gene expression data to find unexpected phenomena that might merit further study. What’s more, the algorithm then re-examines its own output, looking for mistakes it has made and then correcting them.
This work by Carl Kingsford, a professor in CMU’s Computational Biology Department, and Cong Ma, a Ph.D. student in computational biology, is the first attempt at automating the search for these anomalies in gene expression inferred by RNA sequencing, or RNA-seq, the leading method for inferring the activity level of genes.
As they report today in the journal Cell Systems, the researchers already have detected 88 anomalies — unexpectedly high or low levels of expression of regions within genes — in two widely used RNA-seq libraries that are both common and not previously known.
Though an organism’s genetic makeup is static, the activity level, or expression, of genes varies greatly over time. Gene expression analysis has thus become a major tool for biological research, as well as for diagnosing and monitoring cancers.
Anomalies can be important clues for researchers, but until now finding them has been a painstaking, manual process, sometimes called “sequence gazing.” Finding one anomaly might require examining 200,000 transcript sequences — sequences of RNA that encode information from the gene’s DNA, Kingsford said. Most researchers therefore zero in on regions of genes that they think are important, largely ignoring the vast majority of potential anomalies.
The algorithm developed by Ma and Kingsford automates the search for anomalies, enabling researchers to consider all of the transcript sequences, not just those regions where they expect to see anomalies. This technology could uncover many new phenomena, such as the 88 previously unknown common anomalies found in the multi-tissue RNA-seq libraries.
But Ma noted that identifying anomalies is often not clear cut. Some RNA-seq “reads,” for instance, are common to multiple genes and transcripts and sometimes get mapped to the wrong one. If that occurs, a genetic region might appear more or less active than expected. So the algorithm re-examines any anomalies it detects and sees if they disappear when the RNA-seq reads are redistributed between the genes.
Ma C, Kingsford C. (2019) Detecting, Categorizing, and Correcting Coverage Anomalies of RNA-Seq Quantification. Cell Syst [Epub ahead of print]. [article]
Source – Carnegie Mellon University
Related Posts
RNA sequencing reveals functional chimeric mRNAs in mammalian immunity
Atlas of the brain’s striatum could guide researchers to new drug treatments
Immune cells offer insights on billion-dollar virus
A functionally integrated cross-tissue alternative splicing program during short-term calorie restriction
Dietary oxidized plant sterol shifts macrophage state to fuel aortic inflammation
Unlocking the past – new method helps gain insights into old tissue
Novel AI model trained on RNA-Seq data accurately detects key gene mutations and predicts biomarkers across 32 cancer types
Transcriptomic aging clock reveals age-related molecular patterns in opioid dependence
RNA sequencing helps predict stem cell transplant benefit in pediatric AML
Protein ‘switch’ determines whether liposarcoma cells will become aggressive
Precursor tRNAs sense temperature changes: heat stress-induced capped pre-tRNAs suppress protein synthesis
Ketamine increases neuroplasticity in female mice but not in males
Somatic mutations linked to vascular damage in progeria
Scientists map dormant cancer cells’ hideouts, opening new targets for treatment
Soluble signals released by neighboring cells direct how the human kidney is built
Genetics influence how cancer arises – and how it evolves
RNA-based testing uncovers extraordinary diversity in mutations driving lung cancer
Study offers new insights into why ex-smokers remain at elevated risk of lung disease
Learning the grammar of gene regulation
New findings could transform new treatment for rare brain tumor astroblastoma
Stay Connected
Submit a Post to the Blog
Recent Posts
Subscribe to the RNA-Seq Blog
RNA-Seq Products & Services