Plasma cell-free RNA (cfRNA) is one of the most exciting frontiers in liquid biopsy, promising a real-time, systemic view of human health. However, as many researchers have discovered, the road from discovery to clinical translation is paved with technical hurdles, from low input concentrations to extreme fragmentation. A new preprint from the team at Flomics Biotech — “Systematic cross-study assessment of RNA-Seq experimental workflows for plasma cell-free transcriptome profiling” — provides the most comprehensive map to date of this complex landscape.

The Scale: 15 Studies, One Uniform Pipeline

To move past the anecdotal evidence that often plagues the field, the authors curated a massive collection of 2,166 sequenced samples across 15 independent studies. Crucially, they bypassed the variability of non-standardized bioinformatics by processing every sample through a uniform, best-practice pipeline (nf-core/rnaseq). This allowed them to isolate the real-world impact of experimental workflows across global laboratories.

Key Finding: The “Technical Floor”

The study reveals that transcriptomic variation in cfRNA is overwhelmingly driven by technical factors rather than biology. Using Variance Partition Analysis (VPA), the researchers demonstrated that the Broad Protocol Category (BPC) and technical metrics like library diversity are the primary drivers of variation, while donor phenotype explains a negligible fraction.

The technical noise is so profound that variation across cfRNA datasets is larger than that observed in a wide variety of structured human tissue transcriptomes. The authors highlight two “master regulators” of this noise:

  1. Genomic DNA (gDNA) Contamination: Tracked via the Fraction of Spliced Reads (FSR), gDNA is a major pitfall that can completely mask the true RNA signal, especially in protocols lacking DNase treatment.
  2. Library Diversity (LD): The study introduces NG80 (the number of genes accounting for 80% of reads) as a robust, intuitive metric for LD. They show that cfRNA libraries are often dominated by a few abundant, uninformative species like srpRNAs.

Can We “Fix” It Post-Hoc?

A critical takeaway is the non-recoverability of the biological signal. The researchers used linear modelling to regress out major technical confounders. While this massively reduced the primary technical variance (with PC1+PC2 dropping from roughly 60% to 18%), biological clustering was not restored. This suggests that technical noise creates an embedded “floor” that resists simple retrospective correction, making prospective, experimental standardization the only viable path forward.

The Confounding Trap: When Technical Noise Masquerades as Biology

Perhaps most concerning for those in the biomarker space is the high degree of confounding between donor phenotypes and technical variables. In several analyzed datasets, the researchers found that donor status (e.g. healthy vs. disease) was inextricably linked to pre-analytical factors like the collection center or centrifugation protocol. Using Cramér’s V score to quantify these associations, the study reveals several “red flags” in the current literature, where observed “biomarkers” might actually reflect differences in sample handling rather than the disease itself.

Actionable Recommendations

The preprint concludes with evidence-based guidelines for choosing the right workflow:

  • Exome Capture (EB): Best for protein-coding biomarker discovery, as it consistently produces diverse, gDNA-free libraries.
  • Whole RNA-Seq, Random-primed (WRR): Essential for exploratory or microbial signatures, but requires rigorous DNase treatment and affinity-based depletion of uninformative transcripts.

This work serves as a much-needed benchmark, transforming our understanding of cfRNA-Seq from a collection of fragmented studies into a rigorous, quantitative framework for future clinical diagnostics.

Read the full preprint here: https://www.biorxiv.org/content/10.1101/2025.07.10.664092

Lead contact: Julien Lagarde, PhD (email / LinkedIn)

Plasma cell-free RNA (cfRNA) is one of the most exciting frontiers in liquid biopsy, promising a real-time, systemic view of human health. However, as many researchers have discovered, the road from discovery to clinical translation is paved with technical hurdles, from low input concentrations to extreme fragmentation. A new preprint from the team at Flomics Biotech — “Systematic cross-study assessment of RNA-Seq experimental workflows for plasma cell-free transcriptome profiling” — provides the most comprehensive map to date of this complex landscape.

The Scale: 15 Studies, One Uniform Pipeline

To move past the anecdotal evidence that often plagues the field, the authors curated a massive collection of 2,166 sequenced samples across 15 independent studies. Crucially, they bypassed the variability of non-standardized bioinformatics by processing every sample through a uniform, best-practice pipeline (nf-core/rnaseq). This allowed them to isolate the real-world impact of experimental workflows across global laboratories.

Key Finding: The “Technical Floor”

The study reveals that transcriptomic variation in cfRNA is overwhelmingly driven by technical factors rather than biology. Using Variance Partition Analysis (VPA), the researchers demonstrated that the Broad Protocol Category (BPC) and technical metrics like library diversity are the primary drivers of variation, while donor phenotype explains a negligible fraction.

The technical noise is so profound that variation across cfRNA datasets is larger than that observed in a wide variety of structured human tissue transcriptomes. The authors highlight two “master regulators” of this noise:

  1. Genomic DNA (gDNA) Contamination: Tracked via the Fraction of Spliced Reads (FSR), gDNA is a major pitfall that can completely mask the true RNA signal, especially in protocols lacking DNase treatment.
  2. Library Diversity (LD): The study introduces NG80 (the number of genes accounting for 80% of reads) as a robust, intuitive metric for LD. They show that cfRNA libraries are often dominated by a few abundant, uninformative species like srpRNAs.

Can We “Fix” It Post-Hoc?

A critical takeaway is the non-recoverability of the biological signal. The researchers used linear modelling to regress out major technical confounders. While this massively reduced the primary technical variance (with PC1+PC2 dropping from roughly 60% to 18%), biological clustering was not restored. This suggests that technical noise creates an embedded “floor” that resists simple retrospective correction, making prospective, experimental standardization the only viable path forward.

The Confounding Trap: When Technical Noise Masquerades as Biology

Perhaps most concerning for those in the biomarker space is the high degree of confounding between donor phenotypes and technical variables. In several analyzed datasets, the researchers found that donor status (e.g. healthy vs. disease) was inextricably linked to pre-analytical factors like the collection center or centrifugation protocol. Using Cramér’s V score to quantify these associations, the study reveals several “red flags” in the current literature, where observed “biomarkers” might actually reflect differences in sample handling rather than the disease itself.

Actionable Recommendations

The preprint concludes with evidence-based guidelines for choosing the right workflow:

  • Exome Capture (EB): Best for protein-coding biomarker discovery, as it consistently produces diverse, gDNA-free libraries.
  • Whole RNA-Seq, Random-primed (WRR): Essential for exploratory or microbial signatures, but requires rigorous DNase treatment and affinity-based depletion of uninformative transcripts.

This work serves as a much-needed benchmark, transforming our understanding of cfRNA-Seq from a collection of fragmented studies into a rigorous, quantitative framework for future clinical diagnostics.

Read the full preprint here: https://www.biorxiv.org/content/10.1101/2025.07.10.664092

Lead contact: Julien Lagarde, PhD (email / LinkedIn)

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services