What is a Structural Variant and Why Does it Matter in Genomics Research?
Genomics research has a variant problem. Not a shortage of them — an awareness problem. The field has spent decades optimizing for the variants it can easily detect while systematically undercounting the ones that are harder to see.
Structural variants (SVs) fall into the second category. They are large. They are common. They are disproportionately concentrated in genomically complex regions. And most standard sequencing approaches miss a significant fraction of them.
This is not a minor technical footnote. Structural variants are implicated in cancer, rare disease, complex trait heritability, crop yield, and disease resistance. Missing them is not neutral. It has consequences for the conclusions researchers draw and the markers breeders select.
What is a Structural Variant?
A structural variant (SV) is any DNA mutation larger than 50 base pairs. This size threshold distinguishes SVs from single nucleotide polymorphisms (SNPs), which affect a single base, and small insertions and deletions (indels), which affect 2 to 50 base pairs.
The SV category encompasses several distinct mutation classes:
Insertions: Additions of 50 or more base pairs of DNA sequence. Can range from a few hundred base pairs to tens of kilobases.
Deletions: Loss of 50 or more base pairs. One of the most common SV types detected in population studies.
Inversions: A segment of DNA flipped to the reverse orientation. Can affect gene regulation without changing gene content.
Translocations: DNA from one chromosomal region that has moved to a different location, either within the same chromosome or between chromosomes.
Copy Number Variants (CNVs): Regions where an individual carries more or fewer copies of a DNA segment than the reference genome. Relevant in dosage-sensitive genes and agronomic traits.
Mobile Element Insertions: Transposable elements that have inserted into new locations. Common in plant genomes and implicated in adaptation and phenotypic diversity.
Together, these variant classes represent a large fraction of the total sequence difference between any two individuals of the same species. In humans, SVs account for more base pairs of genomic difference than SNPs do, even though there are far fewer SVs than SNPs in absolute count. In plant genomes — particularly polyploid crops — SV density is even higher.
Why Structural Variants Have Been Underdetected
The core reason is read length. Short-read sequencing platforms produce reads of 150 to 300 base pairs. Most structural variants are larger than the reads themselves. This creates a fundamental detection problem: you cannot span a variant with a read that is shorter than the variant.
When short reads encounter a structural variant, several failure modes occur:
Reads in the variant region fail to map to the reference genome uniquely. They get discarded.
Reads that map to flanking regions provide a signal that something has changed, but cannot define the variant's boundaries or sequence content with precision.
Complex SVs involving rearrangements, inversions, or mobile element insertions are often collapsed, split, or misclassified in short-read variant call files.
Long reads solve this problem directly. A 13 to 25 kilobase HiFi read can span most structural variants entirely — from one unaltered flank through the variant to the other. The variant sequence is captured in full. Breakpoints can be defined at base-pair resolution.
What This Looks Like in Practice
The gap in SV detection between long-read and short-read approaches is not theoretical. Researchers studying peanut (Arachis hypogea), an allotetraploid with a 2.5 gigabase genome, compared long-read low-pass (LRLP) and short-read low-pass sequencing across 130 lines.
At approximately 1.6x depth, LRLP detected 51.6 times more structural variants than short-read low-pass at equivalent depth. This is not a marginal improvement in sensitivity. It represents a qualitative difference in the data: short-read low-pass was functionally blind to the vast majority of structural variation present in this population.
Lee et al. 2025, bioRxiv preprint. Not yet peer-reviewed.
The practical consequence: if you are running a GWAS or QTL mapping study in a crop species using short-read sequencing, structural variants linked to your trait of interest are likely absent from your marker set. The loci they drive will show up as unexplained variance in your model - what the field calls missing heritability.
Structural Variants and Missing Heritability
Missing heritability describes the gap between the heritability estimated from pedigree studies and the heritability explained by detected genomic variants. For complex traits, such as yield, disease resistance, height, polygenic disease risk, that gap has been substantial across decades of GWAS.
Structural variants are one of the major contributors to that gap, for several reasons:
SVs are systematically absent from SNP arrays and underrepresented in short-read sequencing data.
SVs often affect gene dosage through copy number variation, which has large phenotypic effects not captured by SNP markers.
SVs disrupt regulatory elements, gene order, and linkage disequilibrium blocks in ways that SNP-level analysis cannot fully reconstruct.
In polyploid species, SVs involving subgenome-specific deletions or rearrangements are particularly difficult to detect with short reads that cannot distinguish homeologous chromosomes.
SV-aware GWAS — association studies that include structural variants alongside SNPs — have identified loci missed by SNP-only approaches in multiple species and trait categories. The variants were there. They were just invisible to the methods being used.
What Detecting Structural Variants Changes for Research
For Plant Breeders
Disease resistance in many crops is conferred by structural variants, not point mutations. An NBS-LRR resistance gene cluster expanded through tandem duplications is a structural variant. A deletion that removes a susceptibility gene is a structural variant. Marker-assisted selection that cannot detect these variants cannot reliably track the traits they control.
Long-read sequencing at the population level enables breeders to build SV-inclusive marker panels. Selection accuracy for complex, structurally-driven traits improves substantially when the underlying variation is visible in the data.
For Human Health Researchers
Structural variants cause a significant fraction of rare genetic disease. CNVs and complex rearrangements involved in neurodevelopmental disorders, cancer predisposition, and immunological conditions are detectable only with long reads at the population scale. Repeat expansions — a subtype of structural variant — are the causative mutation in diseases including Huntington's disease, fragile X syndrome, myotonic dystrophy, and ALS-linked C9orf72 expansion. Long reads can span them, size them, and detect their epigenetic status simultaneously.
For Biodiversity Researchers
Non-model organisms rarely have the curated reference panels that make SNP-based imputation reliable. Structural variants in these genomes are even less well-characterized than in model species. Long-read sequencing is often the only method capable of generating SV catalogs for species where the reference genome is incomplete or absent.
Why Low Coverage is Sufficient for SV Detection
A common assumption is that structural variant detection requires deep coverage. This is true for short reads. For long reads, it is not.
Long reads detect structural variants through spanning: a single read that crosses the variant boundary provides definitive evidence of the variant's presence and structure. At 2x coverage per haplotype most regions of a diploid genome will be covered by at least one read. Structural variants large enough to be biologically significant are large enough to be reliably spanned by long reads that cover the region.
This is fundamentally different from short-read SV detection, which requires accumulating many discordant read pairs across a variant boundary before a statistically confident call can be made. Long reads do not need that statistical accumulation. A single spanning read makes the call.
The Bottom Line
Structural variants are DNA mutations larger than 50 base pairs — insertions, deletions, inversions, translocations, copy number variants, and mobile element insertions.
SVs account for more total sequence difference between individuals than SNPs do, even though they are fewer in number.
Short-read sequencing systematically underdetects SVs because reads are shorter than most SVs.
Long reads span structural variants entirely, enabling base-pair-resolution breakpoint detection at low coverage.
SV detection is directly relevant to missing heritability, disease resistance in crops, rare disease diagnosis, and biodiversity characterization.
Low-pass long-read sequencing detects SVs reliably at 2x coverage per haplotype — a depth that is economically viable at population scale.
If your current sequencing approach is missing structural variants, we can help you design a study that captures them.
Talk to a scientist or request a quote.
Related Reading