Is Genotyping-by-Sequencing Still Worth It? What Plant Breeders Need to Know
Genotyping-by-sequencing was a genuinely important development. When Elshire et al. published the method in 2011, it solved a real problem: how to generate thousands of genetic markers across a population of plants quickly and cheaply, without the expense or inflexibility of a custom SNP array.
More than a decade later, it is worth asking the question directly: is GBS still the right tool, or are researchers using it because it is familiar rather than because it is optimal?
The answer depends on your research question. But for a growing number of applications in plant breeding and crop genomics, the honest answer is that GBS has been outpaced.
What GBS Does and Why It Worked
Genotyping-by-sequencing reduces genome complexity by digesting DNA with restriction enzymes and sequencing only the regions adjacent to restriction cut sites. This produces a manageable number of reads from a predictable subset of the genome, enabling sequencing of large populations at low cost.
The method was a breakthrough for crop species that lacked affordable genotyping arrays and had large, complex genomes that made whole-genome sequencing expensive. GBS gave breeders access to thousands of markers across diverse germplasm for the cost of a sequencing run. That value proposition was real in 2011.
The Structural Limitations of GBS
Restriction Site Dependence
GBS generates markers only at restriction enzyme cut sites. Where those sites fall in the genome is determined by the genome sequence, not by the researcher's scientific priorities. In repetitive or complex genomic regions, restriction sites are often absent or produce fragments that are too large or too small to sequence efficiently.
The result is that GBS marker panels are not genomically representative. They sample the accessible, low-complexity regions of the genome at the expense of the biologically important complex regions — exactly the regions where structural variants, resistance gene clusters, and adaptive loci tend to concentrate.
Locus Dropout
Not every restriction site produces a sequenceable fragment in every sample. Partial digestion, methylation variation, and small insertions or deletions near cut sites all cause specific loci to drop out in specific samples. The result is a marker matrix with a high rate of missing data.
Missing data in GBS datasets is not random. It is correlated with biological variation — the same loci that vary structurally between accessions are most likely to show differential cut-site presence. Missing data is not a nuisance; it is a source of systematic bias in downstream analyses.
Reference Genome Dependency
GBS data is only as good as the reference genome it is aligned to. For species with fragmented or low-quality reference assemblies — which includes a substantial fraction of the crops and wild relatives studied in plant breeding research — GBS alignment is unreliable. Reads falling in unassembled regions produce no marker calls. Reads misaligning to collapsed repeats produce false-positive variant calls.
No Structural Variant Detection
GBS does not detect structural variants. The method samples small genomic windows adjacent to restriction sites. A deletion, inversion, or mobile element insertion outside those windows is invisible. A large insertion at or near a restriction site disrupts the cut-site context and causes locus dropout — meaning the variant causes the marker to disappear rather than appearing as an SV call.
For plant breeders working on traits with a structural genetic basis — disease resistance gene clusters, domestication loci, yield-related CNVs — GBS is not measuring the relevant genetic variation. It is measuring a proxy that is increasingly known to be incomplete.
When GBS Still Makes Sense
GBS is not without legitimate current applications. It remains a reasonable choice when:
Sample sizes are very large and budget constraints are severe, making per-sample cost the dominant constraint.
The research question is fully addressable with a sparse, restriction-site-defined marker panel.
You are maintaining an existing GBS dataset and need new samples to be compatible with historical markers.
The species has a small, diploid genome with low repetitiveness and a well-assembled reference, where GBS dropout is minimized.
Outside these conditions, the case for GBS weakens considerably.
What Has Changed Since GBS Was Introduced
Three things have shifted the calculation substantially since 2011.
First, sequencing costs have dropped dramatically. The per-sample cost of low-coverage whole-genome sequencing has converged toward GBS in many species. The cost advantage that justified GBS's methodological tradeoffs is smaller than it was.
Second, long-read sequencing has become scalable. When GBS was introduced, long-read sequencing required one flowcell per sample and cost thousands of dollars per genome. Multiplexed long-read low-pass sequencing changes that economics entirely. Up to 96 samples can be run in a single PacBio Revio cell, distributing the fixed cost across a population.
Third, the research questions have changed. The field has moved from simple marker discovery to population-scale GWAS, SV-inclusive pangenome analysis, and haplotype-resolved trait mapping. GBS was not designed for these applications.
How Long-Read Low-Pass Sequencing Compares
Long-read low-pass (LRLP) sequencing addresses the core limitations of GBS directly.
Whole-genome coverage: LRLP sequences across the entire genome without restriction site dependence. No bias toward low-complexity regions.
No locus dropout: Coverage is distributed stochastically across all genomic regions. There is no mechanism equivalent to restriction-site dropout.
Structural variant detection: Long reads span SVs directly. GBS cannot detect them at all.
Reference independence at the read level: Long reads map with higher confidence to complex and repetitive genomic regions where GBS reads are lost.
Haplotype phasing: Long reads physically phase SNPs, indels, and SVs into haplotype blocks. GBS produces unphased genotype calls.
In direct comparison studies in plant species, LRLP delivers substantially more usable genomic coverage than short-read low-pass approaches at equivalent depth. In peanut — an allotetraploid with a 2.5 gigabase genome — LRLP covered 55% of the genome and 58% of gene space at 1.63x average depth, compared to 17% and 11% for short-read low-pass at comparable depth. The comparison to GBS, which only samples a fraction of the genome to begin with, would be larger still.
Lee et al. 2025. Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population. G3: Genes|Genomes|Genetics, jkag196. https://doi.org/10.1093/g3journal/jkag196
A Framework for Evaluating Whether to Switch
If you are currently running GBS projects and considering whether to migrate, ask:
What fraction of my GBS markers are in biologically relevant regions versus restriction-site-accessible but low-complexity regions?
Are my traits of interest in genomic regions where GBS has known dropout? For resistance gene clusters, centromeric regions, and highly repetitive loci, the answer is frequently yes.
Am I seeing unexplained variance in my GWAS that might reflect missing structural variation? If heritability estimates fall well below pedigree-based estimates, undetected SVs are one explanation.
What would it cost to run the same population with LRLP? At current pricing with multiplexing, the per-sample cost differential is narrower than most researchers expect when cost per insight is the metric.
The Bottom Line
GBS was a breakthrough method that enabled large-scale plant breeding genomics at low cost.
Its structural limitations — restriction site dependence, locus dropout, reference bias, no SV detection — have become more consequential as research questions have grown more complex.
The cost advantage of GBS over whole-genome sequencing approaches has narrowed substantially.
Long-read low-pass sequencing addresses GBS's core limitations: full genome coverage, structural variant detection, haplotype phasing, and reliable mapping in complex genomic regions.
The decision to migrate from GBS to LRLP should be made on the basis of research question fit, trait biology, and a cost-per-insight comparison — not per-sample price alone.
Evaluating whether to transition from GBS or SNP arrays to whole-genome sequencing? We offer study design consultation as part of every engagement.
Talk to a scientist or request a quote at veilgenomics.com/contact.
Related Reading