How Long-Read Low-Pass Sequencing Works in a Pangenome Study

Every genomics study begins with a reference. A researcher sequences a population and aligns the reads to a reference genome to find where the individual samples differ. The variants discovered are measured relative to that reference.

This approach has served genomics well. It has also introduced a systematic limitation the field is now fully reckoning with: the single reference genome captures the biology of one individual. Everything else in the species that differs from that individual gets measured as deviation. Variation that is common in the species but absent from the reference genome becomes invisible.

The pangenome is the solution to this problem. Long-read low-pass sequencing is a key part of what makes population-scale pangenome analysis practical and more accurate.

What a Pangenome is and Why it Matters

A pangenome is a graph-based representation of the full genetic diversity present across a species — or a defined population within a species. Instead of aligning reads to a single linear reference, reads are aligned to a graph that contains all known haplotypes, structural variants, and sequences present in the population.

The pangenome concept emerged from the observation that a single reference genome is not representative of species-level diversity. In many crop species, the reference genome captures only 60 to 80 percent of the total gene content present across cultivars, landraces, and wild relatives. The remaining 20 to 40 percent — what researchers call presence/absence variation — is not biologically trivial. Disease resistance genes, adaptation loci, and domestication-related sequences are disproportionately represented in this absent fraction.

A pangenome graph encodes all of this variation explicitly. When reads are aligned to the graph, they find the path that best represents their sequence — including sequences absent from the linear reference. Variants are discovered against the full diversity of the species, not just against one individual.

The Reference Bias Problem in Standard Sequencing

When a short read cannot map to any path in a linear reference, it is discarded. When it maps ambiguously to two similar sequences, it is discarded or assigned probabilistically. Both cases introduce reference bias: systematic underdetection of variants in regions that differ structurally from the reference.

Reference bias is concentrated in:

  • Genomic regions carrying structural variants relative to the reference

  • Genes present in the study population but absent from the reference

  • Highly repetitive loci where short reads are ambiguously mappable

  • Subgenome-specific regions in polyploid species where reads from one subgenome may falsely align to the homeologous chromosome

Long reads solve the reference bias problem through their alignment advantage. A read spanning 13 to 25 kilobases crosses repetitive elements entirely, mapping uniquely where short reads cannot. When aligned to a pangenome graph, long reads find the correct haplotype path through structurally complex regions that would confuse short reads.

What LRLP Adds to Pangenome Analysis

Variant Discovery Across the Full Spectrum

LRLP sequencing aligned to a pangenome graph produces variant calls across the full size range of mutations — from single nucleotide polymorphisms to large structural variants — in one sequencing run. No separate assay is needed for SVs. No restricted genomic windows define what is detectable.

In a peanut population study using long-read low-pass sequencing aligned to a pangenome graph reference, researchers detected approximately 2 million total variants — an exponential increase over alignment to a single linear reference at the same sequencing depth. The increase comes primarily from SVs present in the study population but absent from the linear reference, and from sequences in the graph that have no equivalent in the linear reference at all.

Lee et al. 2025, bioRxiv preprint. Not yet peer-reviewed.

Haplotype Resolution in Complex Genomes

A pangenome graph contains all known haplotypes for a genomic region. Long reads are long enough to be assigned to the correct haplotype path based on multiple flanking variants captured in a single read — haplotype phasing without parental samples or statistical imputation.

In polyploid species, haplotype assignment in a pangenome graph is substantially more accurate with long reads than with short reads. Short reads are often mappable to multiple paths and must be assigned probabilistically. Long reads span the haplotype-discriminating variants directly.

Population-Scale SV Catalogs

A population sequenced with LRLP and aligned to a pangenome graph produces a structural variant catalog: the set of all SVs present in the population, with frequency estimates for each. This catalog can be used for SV-aware GWAS, for imputing SVs in downstream short-read datasets, and for building the next iteration of the pangenome as more individuals are added.

Building an SV catalog at population scale was previously a task requiring deep whole-genome sequencing of each individual — economically prohibitive for most research programs. LRLP makes it tractable by distributing sequencing cost across large cohorts via multiplexing.

Pangenome Analysis in Practice

Crop Improvement and Breeding

For plant breeding programs, a pangenome built from diverse germplasm — elite lines, landraces, and wild relatives — captures the full variant landscape relevant to trait improvement. Breeders can identify presence/absence variants linked to disease resistance, abiotic stress tolerance, or yield components that would be invisible in a single reference alignment. Marker-assisted selection using pangenome-informed marker panels is more complete because it selects for the variants that matter, not just the variants representable in the reference.

Non-Model Organism Research

Species without high-quality reference genomes benefit most from pangenome approaches. For non-model organisms, a population-sequenced pangenome can substitute for a complete reference assembly. This is particularly relevant for biodiversity and conservation research, where the target organisms may have no reference genome at all and where the scientific value is in characterizing intraspecific genetic diversity across populations.

Human Health and Biobank Studies

In human genomics, the limitations of the GRCh38 reference genome are well-documented. The reference captures European-ancestry haplotypes at higher frequency than other global populations. Variants common in African, Asian, or Indigenous populations but absent from the reference are systematically underdetected. Pangenome references built from diverse human populations address this bias directly, and long-read sequencing data is the primary input for building and aligning to them with high accuracy.

When Linear References Are Still Fine

Pangenome analysis is computationally more complex than linear reference alignment. A single linear reference is sufficient for:

  • Model organisms with well-assembled references and study questions fully contained in the mappable fraction of the genome

  • Studies focused on common, well-characterized variants in low-complexity genomic regions

  • Initial or exploratory analyses where computational simplicity is prioritized over completeness

The pangenome approach adds the most value when research questions require structural variant detection, when the study population is diverse relative to the reference, when the target organism lacks a representative reference, or when the research aim is to characterize the full complement of variation in a population.

The Bottom Line

  • A pangenome is a graph-based representation of all genetic variation in a population, replacing a single linear reference with a structure that contains all known haplotypes and sequences.

  • Linear reference genomes introduce systematic bias against variants absent from the reference, structural variants, and sequences in complex or repetitive regions.

  • Long reads align to pangenome graphs with higher accuracy than short reads, enabling reliable variant discovery across full genomic complexity.

  • LRLP sequencing aligned to a pangenome graph detects substantially more variants than linear reference alignment — across all variant classes.

  • Applications include crop breeding from diverse germplasm, non-model organism characterization, biobank sequencing in diverse human populations, and population-scale SV catalog construction.

  • The pangenome approach adds the most value when the research question requires complete variant detection across diverse, structurally complex populations.

Working on a study that requires complete variant detection across a diverse population? We can help design the sequencing and analysis approach.

Talk to a scientist or request a quote.

Related Reading

Next
Next

What is a Structural Variant and Why Does it Matter in Genomics Research?