0%

Long read sequencing

A key and recent development, especially in the context of sequencing methods, primarily geared towards metabarcoding, metagenomics, and metatranscriptomics is long read sequencing. While both PacBio and Oxford Nanopore Technologies launched their sequencers a decade ago, they have in recent times become more affordable, efficient and nearly error-free (Figure 5).

While a majority of the metabarcoding community still utilises short-reads starting from 75 bp upto 500 bp in length and focussing on specific hypervariable regions (V1-V9) of the 16S rRNA gene, long reads now enable the sequencing of the entire 16S rRNA gene. This has led to improved taxonomic resolutions, increased sample multiplexing and confident identification of most genera and species. 

The image illustrates comparison between PacBio SMRT sequencing (left) and Oxford Nanopore Technology (ONT) sequencing (right). The PacBio SMRT sequencing process includes a circular DNA template with polymerase, flow cells with immobilised DNA, and ZMW (Zero Mode Waveguide) wells, where fluorescently labelled nucleotides are added and recorded. The readout is shown as coloured dots representing fluorescence intensity over time. ONT sequencing shows a DNA template with motor protein and adapters, a nanopore within a synthetic membrane, and sequencing through electric current as the DNA strand is passed through the nanopore. The readout represents a current change, translating into DNA sequence data.
Figure 5 Schematic outlining the long read sequencing approaches used by (a) PacBio and (b) Oxford Nanopore Technology (image reproduced from Oehler et al, 2023)

These advancements have furthered the upper limit of sequencing capacities, where nearly a full Mbp of genome can be sequenced in a single run. This has additional knock-on effects (Kovaka et al. 2023). For example, high accuracy of sequencing with long reads has improved the precision of calling single nucleotide variants and insertions or deletions (indels). This is limited to none GC-content or PCR-amplification bias, whilst ensuring high contiguity of the sequenced regions, due to little to no fragmentation steps involved in the library preparation.

The overall cycle efficiencies are improved when compared to short-read sequencing, where the sequence error rates increase with length of the reads. Most importantly, additional molecular characteristics such as DNA methylation can be detected, allowing one to perform simultaneous epigenetic characterisation based on diverse methylation types. Interestingly, this methodology was highlighted as the ‘Method of the year: 2022’ by the Nature Journal Publications group through its flagship journal for methodologies (Marx 2023). This has led to the use of long read sequencing for the ‘All of Us’ initiative aimed at sequencing over one million Americans from diverse ethnic backgrounds for improving personalised medical care (Mahmoud 2024). 

Long read sequencing, thus offers various solutions, including high-quality, cutting edge ways for testing hypotheses about community structure and functioning, especially from complex environmental samples. This space is expected to be ever-improving and constantly being updated, so it is advised that the reader refer to the latest literature to understand the nuances of using long reads for their own research questions and hypotheses.

After reviewing the advances in biodiversity exploration, why not take a short quiz to check what you’ve learned?