Genes can produce more than one RNA transcript, allowing a single gene to generate different versions of RNA that may ultimately produce different proteins. Accurately identifying these transcript variants, known as isoforms, is important for understanding how genes function and how their activity changes across tissues and environmental conditions.

A team led by researchers at The James Hutton Institute in Scotland has developed improved reference transcript datasets for potato and tomato. The new resources, called PotatoRTD and TomatoRTD, combine long-read Iso-Seq with short-read RNA sequencing to provide a more complete view of the transcripts produced by these important crop species.

Why transcript annotations matter

RNA sequencing experiments rely heavily on transcriptome annotations. These annotations describe where transcripts begin and end, which sections of a gene are included, and where RNA molecules are spliced together.

If an annotation is incomplete, sequencing reads may be assigned incorrectly or important transcript variants may be missed entirely. This can affect measurements of gene expression and make it difficult to accurately identify alternative splicing or changes in transcript usage.

Potato and tomato are examples of species where existing annotations do not capture the full diversity of RNA transcripts. Many splice junctions and transcript isoforms have been missing from previous reference datasets.

Combining short-read and long-read RNA sequencing

To improve these annotations, the researchers analyzed RNA from a variety of potato and tomato tissues and environmental conditions.

They combined two complementary sequencing approaches. Illumina RNA sequencing generated large numbers of short reads with deep coverage, while PacBio Iso-Seq produced long reads capable of capturing complete RNA transcripts from beginning to end.

Short-read RNA sequencing is useful for measuring gene expression with high depth, but reconstructing complete transcripts from many short fragments can be difficult. Long-read Iso-Seq can directly capture full-length transcript structures, making it easier to identify transcript start sites, end sites, and alternative splice patterns.

Using both approaches allowed the researchers to take advantage of the strengths of each technology.

Revealing more transcript isoforms

One of the main goals was to improve coverage of transcript isoforms.

A gene can produce multiple RNA molecules by using different transcription start sites, ending transcription at different locations, or joining exons together in different combinations. These processes include alternative transcription initiation, alternative polyadenylation, and alternative splicing.

Different isoforms can have different biological functions, and some may appear only in certain tissues or under particular environmental conditions.

By sequencing RNA from several tissues and stress conditions, the researchers were able to capture a wider range of this transcript diversity than would be possible from a single tissue or growth condition.

Fig. 6

Nine transcript isoforms captured in PotatoRTD for a well characterized alternative spliced gene WRKY33.

Improving RNA sequencing analysis

More complete transcript annotations can directly improve downstream RNA sequencing analysis.

Gene-expression tools such as Salmon and Kallisto depend on reference transcript models to determine which transcripts sequencing reads originated from. If transcripts are missing or incorrectly assembled, estimates of transcript abundance can also be inaccurate.

The new PotatoRTD and TomatoRTD resources provide more accurate splice junctions, transcript boundaries, and isoform structures. This should allow researchers to measure individual transcripts with greater confidence and detect biological changes that may have been overlooked using older annotations.

These improvements are particularly important when researchers want to examine alternative splicing or isoform switching, where different versions of a transcript change in abundance under different biological conditions.

Supporting crop biology and biotechnology

Potato and tomato are economically important crops, and improved transcript annotations could support research in areas ranging from plant development to disease resistance and environmental stress responses.

Accurate transcript structures are also valuable for biotechnology applications. Technologies such as CRISPR gene editing and PCR depend on knowing the exact locations of exons, introns, untranslated regions, and transcript boundaries. Incorrect annotations can lead researchers to target the wrong region or misinterpret experimental results.

The researchers have also made PotatoRTD and TomatoRTD available through genome browsers, allowing scientists to visually explore individual genes and their transcript isoforms.

By combining long-read Iso-Seq and short-read RNA sequencing, these new reference datasets provide a more detailed picture of gene expression in potato and tomato. Better transcript annotations should help researchers measure RNA more accurately, investigate alternative splicing, and uncover biological differences that were previously hidden by incomplete reference resources.

Availabilityhttps://github.com/wyguo/potatoRTD_tomatoRTD

Guo W, Milne L, Lopez-Gomollon S, Kaur A, Hein I, Baulcombe D, Zhang R. (2026) PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery. Scientific Data 13: 1222. [article]

Genes can produce more than one RNA transcript, allowing a single gene to generate different versions of RNA that may ultimately produce different proteins. Accurately identifying these transcript variants, known as isoforms, is important for understanding how genes function and how their activity changes across tissues and environmental conditions.

A team led by researchers at The James Hutton Institute in Scotland has developed improved reference transcript datasets for potato and tomato. The new resources, called PotatoRTD and TomatoRTD, combine long-read Iso-Seq with short-read RNA sequencing to provide a more complete view of the transcripts produced by these important crop species.

Why transcript annotations matter

RNA sequencing experiments rely heavily on transcriptome annotations. These annotations describe where transcripts begin and end, which sections of a gene are included, and where RNA molecules are spliced together.

If an annotation is incomplete, sequencing reads may be assigned incorrectly or important transcript variants may be missed entirely. This can affect measurements of gene expression and make it difficult to accurately identify alternative splicing or changes in transcript usage.

Potato and tomato are examples of species where existing annotations do not capture the full diversity of RNA transcripts. Many splice junctions and transcript isoforms have been missing from previous reference datasets.

Combining short-read and long-read RNA sequencing

To improve these annotations, the researchers analyzed RNA from a variety of potato and tomato tissues and environmental conditions.

They combined two complementary sequencing approaches. Illumina RNA sequencing generated large numbers of short reads with deep coverage, while PacBio Iso-Seq produced long reads capable of capturing complete RNA transcripts from beginning to end.

Short-read RNA sequencing is useful for measuring gene expression with high depth, but reconstructing complete transcripts from many short fragments can be difficult. Long-read Iso-Seq can directly capture full-length transcript structures, making it easier to identify transcript start sites, end sites, and alternative splice patterns.

Using both approaches allowed the researchers to take advantage of the strengths of each technology.

Revealing more transcript isoforms

One of the main goals was to improve coverage of transcript isoforms.

A gene can produce multiple RNA molecules by using different transcription start sites, ending transcription at different locations, or joining exons together in different combinations. These processes include alternative transcription initiation, alternative polyadenylation, and alternative splicing.

Different isoforms can have different biological functions, and some may appear only in certain tissues or under particular environmental conditions.

By sequencing RNA from several tissues and stress conditions, the researchers were able to capture a wider range of this transcript diversity than would be possible from a single tissue or growth condition.

Fig. 6

Nine transcript isoforms captured in PotatoRTD for a well characterized alternative spliced gene WRKY33.

Improving RNA sequencing analysis

More complete transcript annotations can directly improve downstream RNA sequencing analysis.

Gene-expression tools such as Salmon and Kallisto depend on reference transcript models to determine which transcripts sequencing reads originated from. If transcripts are missing or incorrectly assembled, estimates of transcript abundance can also be inaccurate.

The new PotatoRTD and TomatoRTD resources provide more accurate splice junctions, transcript boundaries, and isoform structures. This should allow researchers to measure individual transcripts with greater confidence and detect biological changes that may have been overlooked using older annotations.

These improvements are particularly important when researchers want to examine alternative splicing or isoform switching, where different versions of a transcript change in abundance under different biological conditions.

Supporting crop biology and biotechnology

Potato and tomato are economically important crops, and improved transcript annotations could support research in areas ranging from plant development to disease resistance and environmental stress responses.

Accurate transcript structures are also valuable for biotechnology applications. Technologies such as CRISPR gene editing and PCR depend on knowing the exact locations of exons, introns, untranslated regions, and transcript boundaries. Incorrect annotations can lead researchers to target the wrong region or misinterpret experimental results.

The researchers have also made PotatoRTD and TomatoRTD available through genome browsers, allowing scientists to visually explore individual genes and their transcript isoforms.

By combining long-read Iso-Seq and short-read RNA sequencing, these new reference datasets provide a more detailed picture of gene expression in potato and tomato. Better transcript annotations should help researchers measure RNA more accurately, investigate alternative splicing, and uncover biological differences that were previously hidden by incomplete reference resources.

Availabilityhttps://github.com/wyguo/potatoRTD_tomatoRTD

Guo W, Milne L, Lopez-Gomollon S, Kaur A, Hein I, Baulcombe D, Zhang R. (2026) PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery. Scientific Data 13: 1222. [article]

Submit a Post to the Blog

SUBMIT CONTENT

Subscribe to the RNA-Seq Blog

RNA-Seq Products & Services