With the rapid growth of RNA sequencing (RNA-seq) technologies, researchers now have access to vast amounts of short-read sequence data, enabling them to study gene expression in greater detail. However, analyzing this data to construct an accurate transcriptome – the complete set of RNA molecules in a cell – can be a challenging and time-consuming task.

RNA-seq data typically requires careful processing before it can be assembled into a transcriptome. This includes several essential steps such as error correction, trimming adapters, and removing unwanted sequences (known as chimera filtration). Additionally, data often needs to be transferred between different programs, which can involve manual reformatting and restructuring – a process that can slow down the overall workflow and impact the accuracy of the assembly.

Researchers at the University of Illinois at Chicago have developed Semblans, a new computational tool designed to address these challenges. It streamlines the entire transcriptome assembly process, from raw RNA-seq data to fully annotated coding sequences. By automating the quality control, reconstitution, and post-processing steps, Semblans reduces the need for manual intervention, making the assembly process much more efficient.

A graphical workflow of the pipeline

Through the integration of several external packages and the leveraging of C++ data streaming performance, Semblans streamlines the necessary pre-processing, quality control, assembly, and post-assembly steps, allowing a hands-off assembly process without loss to versatility. The reference proteome has been omitted for simplicity, but is utilized by Diamond during the BLASTX / BLASTP steps of postprocess.

When tested against other methods, Semblans demonstrated its ability to produce higher quality transcriptomes in 98 out of 101 short-read runs, showcasing its effectiveness and reliability. This makes it a powerful tool for researchers looking to assemble accurate transcriptomes quickly and consistently.

In summary, Semblans offers a user-friendly solution that simplifies the RNA-seq data processing pipeline, ensuring high-quality transcriptome assemblies with minimal manual effort. With this tool, researchers can focus more on analyzing gene expression and less on the complex technical aspects of data assembly, ultimately accelerating the pace of scientific discovery.

Availability – Source code, documentation, and compiled binaries are hosted under the GNU General Public License at https://github.com/gladshire/Semblans

Woodcock-Girard MD, Bretz EC, Robertson HM, Ramanauskas K, Hampton-Marcell JT, Walker JF. (2025) Semblans: Automated assembly and processing of RNA-Seq data. Bioinformatics [Epub ahead of print]. [article]

With the rapid growth of RNA sequencing (RNA-seq) technologies, researchers now have access to vast amounts of short-read sequence data, enabling them to study gene expression in greater detail. However, analyzing this data to construct an accurate transcriptome – the complete set of RNA molecules in a cell – can be a challenging and time-consuming task.

RNA-seq data typically requires careful processing before it can be assembled into a transcriptome. This includes several essential steps such as error correction, trimming adapters, and removing unwanted sequences (known as chimera filtration). Additionally, data often needs to be transferred between different programs, which can involve manual reformatting and restructuring – a process that can slow down the overall workflow and impact the accuracy of the assembly.

Researchers at the University of Illinois at Chicago have developed Semblans, a new computational tool designed to address these challenges. It streamlines the entire transcriptome assembly process, from raw RNA-seq data to fully annotated coding sequences. By automating the quality control, reconstitution, and post-processing steps, Semblans reduces the need for manual intervention, making the assembly process much more efficient.

A graphical workflow of the pipeline

Through the integration of several external packages and the leveraging of C++ data streaming performance, Semblans streamlines the necessary pre-processing, quality control, assembly, and post-assembly steps, allowing a hands-off assembly process without loss to versatility. The reference proteome has been omitted for simplicity, but is utilized by Diamond during the BLASTX / BLASTP steps of postprocess.

When tested against other methods, Semblans demonstrated its ability to produce higher quality transcriptomes in 98 out of 101 short-read runs, showcasing its effectiveness and reliability. This makes it a powerful tool for researchers looking to assemble accurate transcriptomes quickly and consistently.

In summary, Semblans offers a user-friendly solution that simplifies the RNA-seq data processing pipeline, ensuring high-quality transcriptome assemblies with minimal manual effort. With this tool, researchers can focus more on analyzing gene expression and less on the complex technical aspects of data assembly, ultimately accelerating the pace of scientific discovery.

Availability – Source code, documentation, and compiled binaries are hosted under the GNU General Public License at https://github.com/gladshire/Semblans

Woodcock-Girard MD, Bretz EC, Robertson HM, Ramanauskas K, Hampton-Marcell JT, Walker JF. (2025) Semblans: Automated assembly and processing of RNA-Seq data. Bioinformatics [Epub ahead of print]. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services