Low-Pass Long-Read vs. Low-Pass Short-Read Sequencing: What Changes at the Same Depth?
Low-pass sequencing has earned its place in genomics. Sequence each sample at a fraction of the depth used for a reference-quality genome, rely on imputation to fill the gaps, and a population study that once required a genotyping array or a restriction-digest library becomes a whole-genome study.
Most of that work has been done with short reads, and most study designs carry an assumption that rarely gets tested: at low pass, depth is the variable that matters. Pick a target coverage, price it out, and treat the read technology as interchangeable.
We tested that assumption directly. The answer changes how low-pass studies should be designed, especially in complex crop genomes.
What Low-Pass Sequencing Is Designed to Do
Low-pass sequencing, also called low-coverage whole-genome sequencing, sequences each sample at an average depth typically below 5x. Short read sequencing at that depth cannot cover every single sample. Sparse reads from many related individuals, combined with a reference panel and statistical imputation, can reconstruct genotypes across the genome at a fraction of the cost of deep sequencing.
That logic holds for any read length. What differs between technologies is how much of each read can actually be used and how much of the sample’s genome is covered.
The Comparison: Same Population, Same Depth, Different Read Length
In a study published in G3: Genes|Genomes|Genetics, our team compared long-read low-pass (LRLP) sequencing on PacBio HiFi with short-read low-pass (SRLP) sequencing in a peanut MAGIC breeding population. Peanut is an allotetraploid with two closely related subgenomes, which makes it a demanding test for any genotyping method.
Both approaches were sequenced to nearly the same average depth: 1.63x for long reads and 1.68x for short reads, a difference that was not statistically significant. With depth held constant, differences in the results come from the reads themselves.
Finding 1: Equal Depth, More Than Three Times the Genome Covered
At matched depth, long-read samples covered an average of 55% of the genome. Short-read samples covered 17.3%. In gene space, the gap was wider: 58% for long reads versus 11% for short reads.
Nominal depth is calculated from the total bases sequenced. It says nothing about how many of those bases land somewhere useful. Short reads that cannot be placed confidently are filtered out before genotyping, and in a polyploid genome, that filter removes a large share of the data you paid for.
Finding 2: Alignment Confidence Decides What Survives
Every aligned read receives a mapping quality score that reflects how confident the aligner is that the read belongs where it was placed. Reads below a quality threshold are usually discarded, because a misplaced read produces false variant calls.
Averaged across the peanut population, long reads aligned with 92.9% confidence. Short reads aligned with 54.9% confidence.
The reason is structural. Peanut's A and B subgenomes share long stretches of nearly identical sequence. A 150 base pair read drawn from one of those stretches can match both subgenomes equally well, so the aligner cannot choose between them. A HiFi read spanning many kilobases almost always reaches a distinguishing difference, which anchors it to the correct subgenome.
The same mechanism applies to any genome with recent duplications, large repeat families, or closely related gene copies. Polyploid crops make the effect dramatic, but diploid genomes experience it too.
Finding 3: A Resistance Variant Short Reads Could Not See
Coverage percentages matter because of what lives in the regions that go uncovered. In this population, long-read low-pass sequencing identified 16 lines carrying an insertion associated with resistance to tomato spotted wilt virus, one of the most damaging diseases in peanut production. Short reads failed to align within the inserted segment.
For a breeding program, that is the difference between tracking a resistance source and losing it without knowing it was there. The short-read data did not report the variant as absent. It had nothing to say about it at all, and in a marker dataset, silence is easy to mistake for a negative result.
Why Depth Alone Is the Wrong Planning Unit
The practical lesson is that nominal depth overstates what short reads deliver in complex genomes. A more useful planning concept is usable depth: the depth that remains after reads that cannot be placed confidently are removed, measured across the regions your research question depends on.
Two studies at the same nominal depth can have very different usable depth. When your traits of interest sit in repetitive regions, duplicated gene families, or homeologous chromosome segments, read length determines whether those regions are genotyped at all.
Before committing to a low-pass design, three questions will tell you more than the coverage number:
What fraction of reads is expected to align confidently in my species, and in the regions where my traits are located?
Which variant classes does my design need to detect: SNPs only, or also insertions, deletions, and larger structural variants?
Do I have an imputation reference panel and if so, does it contain the haplotypes and variants that matter in my population?
When Short-Read Low-Pass Is Still a Good Fit
Short-read low-pass sequencing remains a strong choice under the right conditions. Human populations with large, well-phased reference panels, diploid species with simple and well-assembled genomes, and studies whose questions sit entirely within easily mapped regions can all be served well by short reads at low depth.
The per-sample price of short-read low-pass is also lower. The more informative comparison is what each approach returns per sample: how much of the genome, how much of the gene space, and which variants are visible when the data comes back. In the peanut population, the long-read data answered questions that adding short-read depth could not reach, because the limiting factor was alignment rather than budget.
The Bottom Line
Low-pass study designs usually treat depth as the key variable, but read length matters as much.
At matched depth of about 1.6x in a peanut breeding population, long-read low-pass (LRLP) sequencing covered 55% of the genome versus 17.3% for short-read low-pass, and 58% of gene space versus 11%.
Alignment confidence averaged 92.9% for long reads and 54.9% for short reads, because long reads span the near-identical sequence that confuses short-read aligners.
LRLP identified a disease resistance insertion in 16 lines that short reads could not align to.
Plan low-pass studies around usable depth in the regions your traits depend on, not nominal depth alone.
Source: Lee K, et al. Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population. G3: Genes|Genomes|Genetics, jkag196. https://doi.org/10.1093/g3journal/jkag196
Veil Genomics offers the first commercially available long-read low-pass sequencing service on PacBio HiFi. Planning a low-pass study in a complex genome? We can help you estimate usable depth for your species before you commit a budget.
Talk to a scientist or request a quote.