# Reference-Based Variant Calling vs. Assembly-Based Variant Detection: Which Approach Is Right for Your Study?


## Key Takeaways

- Reference-based variant calling excels at high-accuracy detection of single nucleotide polymorphisms (SNPs) and small insertions/deletions (indels) by aligning sequencing reads to a known reference genome. This method is computationally less demanding and ideal for large population studies or clinical screening where a high-quality reference is available and structural variants are secondary.
- Assembly-based variant detection reconstructs the sample genome de novo from sequencing reads, making it superior for identifying structural variants, novel insertions, and presence-absence variation, especially in species lacking a reference genome or in complex pangenome construction. This approach requires significantly more computational resources but offers higher sensitivity for sequence divergence from a reference.
- Long-read sequencing technologies are crucial for both approaches, particularly for assembly-based methods and for improving structural variant detection in reference-based workflows by spanning repetitive regions and variant breakpoints. However, reference bias can still lead to spurious variant calls in reference-based calling with long reads when structural differences exist between the sample and reference.
- Combining reference-based calling for SNPs/indels with assembly-based methods for structural variants offers a comprehensive approach, leveraging the strengths of each. Alternatively, assembling unmapped reads from a reference-based workflow can capture novel sequences absent from the reference without the full computational cost of whole-genome assembly.
- Computational cost is a major differentiator: reference-based calling is generally more accessible with standard servers, while de novo assembly, especially for large or repetitive genomes, necessitates high-performance computing clusters due to substantial memory and processing time requirements.
- Phased assemblies, where maternal and paternal chromosome copies are resolved, significantly improve variant genotyping accuracy, particularly in polyploid organisms, and are essential for detailed structural variant analysis and pangenome construction.

---

Researchers studying genetic variation face a fundamental decision at the start of any sequencing project: whether to align reads to an existing reference genome or to assemble reads into new sequences before looking for variants. Reference-based variant calling maps short or long reads against a known reference sequence and identifies differences at each position. Assembly-based variant detection builds contiguous sequences from the reads themselves, then compares those assembled sequences against a reference or against each other to find structural differences. Each approach detects single nucleotide polymorphisms (SNPs), small insertions and deletions (indels), and larger structural variants with different sensitivity and accuracy profiles. The choice affects also which variants you find but also how much computing power you need, how you interpret results, and how confident you can be in your conclusions. This article compares the two approaches across the full range of variant types, explains when to use each method or combine them, and provides practical criteria for matching the approach to your study goals, genome complexity, and available computational resources.

## At a Glance: Choosing Between Reference-Based and Assembly-Based Approaches

The table below summarizes the key differences between reference-based variant calling and assembly-based variant detection across the dimensions that matter most for study design.

| Decision Factor | Reference-Based Variant Calling | Assembly-Based Variant Detection |
|---|---|---|
| Primary variant types detected | SNPs and short indels with high accuracy, structural variants only with specialized tools | Structural variants, novel insertions, and presence-absence variation, SNPs and indels with lower per-base accuracy |
| Reference genome requirement | Requires a high-quality reference sequence from the same or closely related species | No reference required for assembly, but a reference is needed for comparative analysis after assembly |
| Computational cost | Lower, read mapping and variant calling are less memory-intensive than assembly | Higher, de novo assembly requires substantial memory and processing time, especially for large or repetitive genomes |
| Sensitivity to novel sequence | Low, reads from regions absent in the reference often fail to map or map incorrectly | High, novel sequence is assembled directly from reads and can be identified |
| Accuracy in repetitive regions | Prone to spurious variant calls when reads map to multiple locations | Improved resolution of repetitive regions with long reads, but assembly errors remain possible |
| Best use case | Large population studies, clinical variant screening, SNP-based association studies | Pangenome construction, structural variant discovery, species without a reference genome |

## Understanding the Two Approaches

### Reference-Based Variant Calling: Principles and Workflow

Reference-based variant calling begins with sequencing reads that are aligned to a reference genome. The reference serves as a coordinate system against which every read is positioned. After alignment, the variant caller examines each genomic position, counts the bases observed in the reads, and applies statistical models to determine whether the observed differences from the reference represent true genetic variants or sequencing errors.

The standard workflow follows a defined sequence of steps. Raw sequencing reads first undergo quality control to remove adapters and low-quality bases. Cleaned reads are then aligned to the reference genome using a read mapper. The resulting alignment file contains information about where each read maps, how well it maps, and any mismatches or gaps relative to the reference. Variant calling software then analyzes the alignment to identify positions where a significant proportion of reads disagree with the reference. Finally, the candidate variants are filtered and annotated to assess their potential functional impact.

The accuracy of reference-based calling depends heavily on the quality of the reference genome and the mapping step. When reads map uniquely to a single location, variant calls at that position are generally reliable. Problems arise when reads map to multiple locations, which happens frequently in repetitive regions, or when the sample genome contains sequences that are absent from the reference. In these cases, reads may map to the wrong location or fail to map entirely, producing false variant calls or missing true variants.

### Assembly-Based Variant Detection: Principles and Workflow

Assembly-based variant detection takes a different path. Instead of using a reference as a scaffold, the assembler attempts to reconstruct the sample genome directly from the sequencing reads. The assembled contigs or scaffolds represent the sample's own sequence, independent of any reference. Once assembly is complete, the assembled sequences can be compared against a reference genome to identify differences, or multiple assemblies can be compared against each other to find variation within a population.

The assembly workflow also follows a defined sequence. Reads undergo quality control and error correction before assembly. The assembler then builds contigs by finding overlapping reads and extending them into longer sequences. For genomes with complex repeat structures, additional steps such as scaffolding and gap filling may be needed. After assembly, the resulting sequences are evaluated for completeness and accuracy using metrics such as contig N50, which describes the length of the contig at which half the assembled sequence is contained in contigs of that length or longer.

Assembly-based approaches excel at detecting structural variants because they do not depend on reads mapping to a reference. A large insertion that is present in the sample but absent from the reference will be assembled into the sample's contigs, where it can be identified by comparing the assembly to the reference. Similarly, deletions, inversions, and duplications can be detected by aligning assembled contigs to the reference and identifying regions where the alignment breaks down or rearranges.

### The Role of Sequencing Technology in Both Approaches

The choice between reference-based and assembly-based methods is closely tied to sequencing technology. Short-read sequencing platforms produce reads of 150 to 300 base pairs with high per-base accuracy. These reads are well suited to reference-based variant calling for SNPs and small indels, but they struggle with structural variant detection because structural variants often span regions longer than individual reads. Short reads also create difficulties in repetitive regions, where multiple genomic locations share similar sequences.

Long-read sequencing platforms produce reads of thousands to tens of thousands of base pairs. These reads can span entire repetitive regions and structural variant breakpoints, making them valuable for both reference-based structural variant calling and de novo assembly. Long reads improve assembly quality because they provide the long-range information needed to resolve repeats and produce contiguous assemblies. However, long reads historically had higher error rates than short reads, although recent platforms have improved accuracy substantially.

A benchmark study using accurate long reads across human and plant genomes found that variant calling accuracy is strongly influenced by genome complexity, particularly repeat content. The study also identified a critical mechanism affecting variant discovery: structural variations between the reference and sample genomes, especially those containing repetitive elements, can induce spurious read mapping. This effect is likely exacerbated by the length and accuracy of long reads, leading to false variant calls that constitute a distinct and more dominant source of error than allelic dosage uncertainty in polyploids. This finding has direct implications for study design, as it suggests that reference-based calling with long reads can produce false variants in regions where the sample genome differs structurally from the reference.

## Structural Variant Detection: Where the Approaches Diverge

### Why Structural Variants Require Special Consideration

Structural variants include deletions, insertions, duplications, inversions, and translocations that typically involve sequences longer than 50 base pairs. These variants are an underexplored source of genetic diversity compared to SNPs and small indels, largely because they are more difficult to detect with standard short-read approaches. Structural variants can have profound effects on phenotype because they can remove entire genes, duplicate regulatory regions, or disrupt gene structure.

The difficulty in detecting structural variants with reference-based methods stems from the nature of read mapping. When a read spans a deletion breakpoint, part of the read maps to one side of the deletion and part maps to the other side, creating a discordant alignment that can be detected by specialized tools. Insertions are harder to detect because reads from the inserted sequence may not map to the reference at all. In repetitive regions, the problem becomes more severe because reads may map to multiple locations, making it difficult to determine whether an apparent structural variant is real or an artifact of ambiguous mapping.

Assembly-based methods avoid many of these problems because they do not require reads to map to a reference. The sample's own sequence is assembled directly, so novel insertions and complex rearrangements are represented in the assembly. Comparing the assembly to a reference reveals structural differences that would be missed by read mapping alone.

### Evidence from Pangenome Studies

The value of assembly-based approaches for structural variant detection is demonstrated by recent pangenome research. A study of dairy cattle constructed a breed-specific pangenome graph using 40 phased haploid assemblies from 20 Holstein cows. This pangenome graph outperformed both assembly-based and read-based long-read variant callers and far exceeded short-read approaches, identifying over 10,000 additional structural variants per sample. The study also found that structural variant detection and genotyping improved significantly when using graphs built from phased assemblies within the same breed, compared to graphs built across breeds or from fewer or unphased assemblies.

This evidence has practical implications for researchers working with any species. A single reference genome does not capture the full genetic diversity of a species. Pangenome approaches that incorporate multiple assemblies can reveal structural variants that are invisible to reference-based methods. For species with existing pangenome resources, researchers can leverage these tools to improve structural variant detection. For species without such resources, building a pangenome may be a worthwhile investment if structural variants are a primary study focus.

### The Problem of Reference Bias

Reference bias occurs when reads from a sample that differs from the reference genome fail to map or map incorrectly, leading to systematic errors in variant detection. This bias affects both SNP calling and structural variant detection, but it is more severe for structural variants because they involve larger sequence differences.

A study of disease resistance genes in sunflower illustrated this problem. The researchers performed de novo assembly of reads that did not map to the reference genome and identified 198 disease resistance genes that were absent from the reference. These genes would have been completely missed by reference-based variant calling because their sequences were not present in the reference. The study also identified SNPs, short indels, and large deletions within disease resistance genes, demonstrating that a combination of approaches provides a more complete picture of genetic variation.

For researchers studying traits influenced by genes that vary in presence or absence across individuals, reference-based methods alone will provide an incomplete view. Assembly-based approaches, or at minimum the assembly of unmapped reads, are necessary to capture this type of variation.

## Practical Workflow: Implementing Each Approach

### Reference-Based Variant Calling Workflow

The reference-based workflow begins with obtaining a reference genome and sequencing reads. The reference should be from the same species and ideally from a closely related individual or a high-quality representative of the species. The sequencing reads should be generated with sufficient depth to support confident variant calls, typically 30x or higher for diploid genomes.

Quality control is the first analysis step. Raw reads are examined for adapter contamination, low-quality bases, and sequencing artifacts. Tools for this purpose are available through platforms such as the Galaxy Training Network, which provides accessible workflow training and analysis tutorials for quality control and downstream steps. After cleaning, reads are aligned to the reference using a read mapper appropriate for the sequencing technology. Short reads are typically aligned with splice-aware or gapped aligners, while long reads require aligners designed for their error profiles.

Variant calling follows alignment. Multiple algorithms are available, each with different strengths. A client-server interface for large-scale NGS analysis described in a recent study supports three distinct variant calling algorithms: GATK, DeepVariant, and VarScan. This flexibility allows researchers to choose an algorithm suited to their data and study goals, or to run multiple callers and compare results. The study also noted that such interfaces make powerful computational resources more accessible to researchers who are not comfortable with command-line interfaces.

After variant calling, filtering is essential. Raw variant calls include false positives from sequencing errors, mapping errors, and alignment artifacts. Filtering criteria typically include read depth, mapping quality, allele balance, and strand bias. The specific thresholds depend on the sequencing platform, the variant caller, and the study design.

### Assembly-Based Variant Detection Workflow

The assembly workflow begins with the same quality control steps as reference-based calling. Read quality is critical for assembly because errors in reads propagate into assembly errors. Error correction may be performed before assembly, particularly for long reads with higher error rates.

The assembly step itself requires substantial computational resources. The assembler builds contigs from reads, and the quality of the assembly depends on read length, read depth, genome complexity, and the assembler's algorithms. For complex genomes with high repeat content, assembly may require specialized assemblers designed to handle repeats. After assembly, the contigs are typically scaffolded and polished to improve accuracy.

Once a high-quality assembly is obtained, variant detection proceeds by comparing the assembly to a reference genome. This comparison can be done through whole-genome alignment, which identifies regions of conserved sequence and rearrangements. Structural variants are identified as differences in the alignment structure, such as large gaps, inversions, or translocations. SNPs and small indels can also be identified from the alignment, though with lower accuracy than reference-based methods because assembly errors can be mistaken for variants.

### Combining Both Approaches

Many studies benefit from combining reference-based and assembly-based methods. A common strategy is to use reference-based calling for SNPs and small indels, which are detected with high accuracy, and assembly-based methods for structural variants, which require the long-range information provided by assembly. This combined approach captures the strengths of each method while mitigating their weaknesses.

The sunflower disease resistance study provides an example of this combined strategy. The researchers used reference-based methods to identify SNPs, short indels, and large deletions within known disease resistance genes. They then performed de novo assembly of reads that did not map to the reference to identify disease resistance genes absent from the reference genome. This two-pronged approach revealed a more complete picture of disease resistance gene diversity than either method alone.

For researchers with limited computational resources, a staged approach may be practical. Start with reference-based variant calling to identify SNPs and small indels. Then assemble unmapped reads, which are reads that failed to map to the reference, to identify novel sequence present in the sample but absent from the reference. This approach captures some of the benefits of assembly without the full computational cost of whole-genome assembly.

## Computational Considerations and Resource Planning

### Hardware and Software Requirements

Reference-based variant calling is computationally modest compared to assembly. Read mapping and variant calling for a single human genome can be completed on a server with 16 to 32 gigabytes of RAM and multiple processor cores. The analysis typically takes hours instead of days. Software installation and workflow management can be handled through platforms such as Bioconductor, which provides official package and workflow documentation for reproducible genomic analysis, or through community pipeline standards such as nf-core, which offers documented workflows for reproducible analysis.

Assembly-based variant detection requires substantially more resources. De novo assembly of a complex genome can require hundreds of gigabytes of RAM and days of processing time. The exact requirements depend on genome size, repeat content, read length, and read depth. For large or highly repetitive genomes, assembly may require access to high-performance computing clusters.

The computational barrier is a real consideration for many researchers. A study describing AgrOmicSo noted that large-scale NGS data analysis requires substantial computational power, often necessitating high-performance computing environments, and that command-line interfaces for these resources create a significant barrier for many researchers. The study developed a client-server interface to make powerful computational resources more accessible through a graphical user interface. Researchers without access to such tools may need to invest time in learning command-line skills, with training available through resources such as The Carpentries, which offers foundational computing and data lessons.

### Cost Considerations

The cost of a variant detection study includes sequencing costs, computational costs, and personnel time. Sequencing costs are determined by the platform, read length, and depth. Long-read sequencing is generally more expensive per base than short-read sequencing, but the cost difference has narrowed in recent years. Computational costs include hardware, cloud computing, or institutional high-performance computing allocations.

For reference-based variant calling, the main cost is sequencing depth. Higher depth improves variant calling accuracy but increases sequencing cost. For assembly-based approaches, the cost includes both sequencing and computation. Long reads are often preferred for assembly because they produce better assemblies, but they cost more than short reads. The computational cost of assembly can also be significant, particularly for large genomes.

A practical approach to cost management is to match the method to the study question. If the study focuses on SNPs for association analysis, reference-based calling with short reads is cost-effective. If the study focuses on structural variants, the additional cost of long reads and assembly may be justified. If the study requires both, a combined approach using short reads for SNPs and long reads for structural variants may be the most cost-effective strategy.

## Quality Assessment and Validation

### Metrics for Reference-Based Variant Calling

Quality assessment for reference-based variant calling focuses on the accuracy and completeness of the variant calls. Key metrics include transition-to-transversion ratio, which is expected to be around 2.0 for whole-genome data in mammals, heterozygous-to-homozygous ratio, which should reflect the sample's ploidy and population history, and the number of variants in known problematic regions, such as segmental duplications and low-complexity sequences.

Validation of variant calls can be performed through several methods. Sanger sequencing of selected variants provides independent confirmation but is impractical for large numbers of variants. Genotyping arrays can validate SNPs that are represented on the array. For structural variants, PCR amplification across breakpoints can confirm deletions and insertions. The choice of validation method depends on the number of variants to validate and the available resources.

### Metrics for Assembly Quality

Assembly quality is assessed before variant detection because errors in the assembly will produce false variant calls. Key metrics include contig N50, which measures assembly contiguity, the number of contigs, which should be low for a good assembly, and the total assembled length, which should approximate the expected genome size. Completeness can be assessed by checking for the presence of conserved single-copy genes, which should be present in a complete assembly.

Assembly accuracy is more difficult to assess than contiguity. One approach is to align reads back to the assembly and check for consistent coverage. Regions with unusually high or low coverage may indicate assembly errors. Another approach is to compare the assembly to a closely related reference genome, though this comparison is complicated by genuine differences between the species or individuals.

### The Importance of Phasing

Phasing refers to the assignment of variants to the maternal and paternal copies of each chromosome. Phasing is important for accurate genotyping, particularly in polyploid organisms where multiple allele copies exist. A study of long-read variant calling in diploid and polyploid genomes found that genotyping accuracy decreases with increasing ploidy due to allelic dosage uncertainty. This challenge is separate from initial variant discovery and highlights the difficulty of assigning correct allele counts in polyploids even with high sequencing depth.

Phased assemblies provide a solution to this problem. The dairy cattle pangenome study used 40 phased haploid assemblies from 20 cows, meaning each cow contributed two separate assemblies representing its two chromosome copies. This phasing allowed the researchers to distinguish variants on each chromosome copy and improved structural variant detection and genotyping. For researchers working with polyploid species or with samples where phasing is important, phased assembly may be worth the additional cost and complexity.

## Common Failure Patterns and How to Avoid Them

### Reference-Based Calling Failures

The most common failure pattern in reference-based variant calling is the production of false variant calls in repetitive regions. When reads map to multiple locations, the variant caller may interpret mapping artifacts as genuine variants. This problem is particularly severe in genomes with high repeat content, such as plant genomes with large transposable element populations. The benchmark study of long-read variant calling found that genome complexity, particularly repeat content, strongly influences overall variant calling accuracy.

Another common failure is the complete absence of variant calls in regions where the sample differs substantially from the reference. If the sample has a large insertion that is absent from the reference, reads from the insertion will not map and will be discarded. The variant caller will not report any variant at that location because there is no evidence of a difference. This failure mode is silent, meaning the researcher is unaware that variation has been missed.

A third failure pattern involves spurious variant calls induced by structural variation between the reference and sample. The long-read benchmark study identified this as a distinct and more dominant source of error than allelic dosage uncertainty. When structural variants contain repetitive elements, they can induce spurious read mapping that leads to false variant calls. This finding underscores the need for bias-aware mapping strategies that account for structural differences between the reference and sample.

### Assembly-Based Detection Failures

Assembly-based detection has its own failure patterns. The most common is incomplete assembly, where the assembler fails to reconstruct entire chromosomes or large genomic regions. Incomplete assembly leads to missing variants in the unassembled regions. This problem is more severe in repetitive genomes, where the assembler may be unable to resolve repeat structures.

Assembly errors are another failure pattern. The assembler may join sequences incorrectly, creating chimeric contigs that do not exist in the sample genome. These errors produce false structural variants when the assembly is compared to a reference. Assembly errors are more common with short reads, which lack the long-range information needed to resolve complex regions.

A third failure pattern is the misidentification of assembly errors as genuine variants. When comparing an assembly to a reference, small differences may represent assembly errors instead of true genetic variants. Distinguishing between the two requires careful examination of the evidence, including read depth and the consistency of the difference across multiple reads.

### Strategies to Reduce Failures

Several strategies can reduce the frequency of these failure patterns. For reference-based calling, using a reference from the same population or breed reduces reference bias. The dairy cattle study found that within-breed pangenome graphs improved structural variant detection relative to graphs built across breeds. For species with multiple available references, selecting the most closely related reference can improve accuracy.

For assembly-based detection, using long reads improves assembly quality and reduces assembly errors. The benchmark study of long-read variant calling used accurate long reads and still found challenges in complex genomes, suggesting that even the best available technology has limitations. Combining multiple assembly approaches or using assembly polishing can further improve quality.

For both approaches, validating a subset of variants with an independent method provides confidence in the overall results. Validation is particularly important for structural variants, which are more prone to false positives than SNPs and small indels.

## Records and Documentation for Reproducibility

### What to Record

Reproducibility requires detailed documentation of every step in the analysis. For reference-based variant calling, records should include the reference genome version and source, the read mapper and version, the variant caller and version, the filtering thresholds, and the parameters used at each step. For assembly-based detection, records should include the assembler and version, the assembly parameters, the polishing steps, and the comparison method used to identify variants.

The sequencing data themselves must be documented, including the platform, read length, depth, and quality metrics. This information is essential for interpreting variant calls and for comparing results across studies.

### Where to Store Records

Analysis workflows should be stored in version-controlled repositories to track changes over time. Platforms such as Bioconductor provide official documentation for reproducible genomic analysis workflows, while nf-core offers community pipeline standards that emphasize reproducibility. The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility through documented analysis tutorials.

Raw sequencing data should be deposited in public databases such as those maintained by the National Center for Biotechnology Information, which provides official descriptions of its databases, search systems, sequence resources, and analysis services. Processed data, including variant calls and assemblies, should also be deposited to allow other researchers to reproduce or extend the analysis.

### Documentation Standards

Documentation should include both the commands used and the rationale for key decisions. For example, the choice of variant caller should be documented with an explanation of why that caller was selected. The choice of assembly parameters should be documented with an explanation of how they were determined. This level of documentation allows other researchers to understand beyond what was done but why it was done.

The European Bioinformatics Institute provides training on bioinformatics data resources and practical analysis education, which can help researchers develop documentation skills. Training resources such as The Carpentries offer foundational computing and data lessons that include best practices for reproducible analysis.

## Interpretation Limits and Reporting Guidelines

### What Variant Calls Do and Do Not Tell You

Variant calls identify positions where the sample genome differs from the reference or from other samples. They do not directly tell you the functional consequences of those differences. A SNP in a coding region may be synonymous, nonsynonymous, or nonsense, and the functional impact depends on the specific change and the gene context. A structural variant may remove a gene, duplicate a gene, or disrupt regulatory elements, but the phenotypic consequence depends on many factors beyond the variant itself.

Variant calls also depend on the reference genome. A variant call is always relative to the reference, so a sample that is identical to the reference at a position will have no variant call at that position. If the reference is not representative of the population being studied, many variants will be missed or misclassified.

### Reporting Variant Calls

Reports should include the number of variants by type, the distribution of variants across the genome, and the quality metrics used to filter the calls. For structural variants, reports should include the size distribution and the genomic locations of the variants. Reports should also describe the limitations of the analysis, including the reference genome used, the sequencing platform, and the expected false positive and false negative rates.

When reporting variants, it is important to distinguish between variants that have been validated and those that have not. Unvalidated variants should be clearly labeled as such, and the validation method should be described for validated variants.

### Professional Escalation Criteria

Some findings warrant escalation to specialists or additional analysis. If variant calls suggest a large number of structural variants in a region of known biological importance, such as a disease resistance gene cluster or a quantitative trait locus, additional validation may be warranted. If assembly quality metrics are poor, the assembly should not be used for variant detection until the quality issues are resolved. If the number of variants detected is unexpectedly high or low, the analysis should be reviewed for technical artifacts.

For researchers working with species that have limited genomic resources, the absence of a high-quality reference genome may require assembly-based approaches or the construction of a reference before variant calling can proceed. In these cases, consultation with bioinformatics specialists may be appropriate.

## Safety and Regulatory Context

### Data Management and Privacy

Genomic data from humans and some animal species are subject to privacy and data protection regulations. Researchers must ensure that their data management practices comply with applicable regulations and institutional policies. Public databases such as those maintained by the National Center for Biotechnology Information provide guidance on data deposition and access, including controlled access for sensitive data.

For non-human species, data sharing is generally less restricted, but researchers should still consider the implications of sharing genomic data, particularly for commercially valuable traits or endangered species.

### Ethical Considerations

Variant detection studies should consider the ethical implications of their findings. For agricultural species, variant information may be used for breeding decisions that affect animal welfare or genetic diversity. For human studies, variant information may have implications for health and disease risk. Researchers should consider these implications when designing studies and reporting results.

### Compliance with Platform Policies

Researchers using public analysis platforms should comply with the platforms' policies and guidelines. The Galaxy Training Network provides accessible workflow training and analysis tutorials, and its documentation includes guidance on responsible use of the platform. Similarly, nf-core documentation describes community pipeline standards and usage guidelines that promote reproducible and responsible analysis.

## Frequently Asked Questions

### What is the main difference between reference-based variant calling and assembly-based variant detection?

Reference-based variant calling aligns sequencing reads to an existing reference genome and identifies positions where the reads differ from the reference. Assembly-based variant detection builds the sample genome from the reads themselves, then compares the assembled sequence to a reference or to other assemblies. The main difference is that reference-based methods depend on the reference being similar to the sample, while assembly-based methods reconstruct the sample's own sequence and can identify variation that is absent from the reference.

### Which approach is better for detecting SNPs and small indels?

Reference-based variant calling is generally better for SNPs and small indels because read mapping provides high per-base accuracy and the statistical models used by variant callers are well calibrated for these variant types. Assembly-based methods can detect SNPs and small indels, but assembly errors can be mistaken for genuine variants, reducing accuracy. For studies focused primarily on SNPs and small indels, reference-based calling is the recommended approach.

### When should I use assembly-based variant detection?

Assembly-based variant detection is recommended when structural variants are a primary study focus, when the species lacks a high-quality reference genome, or when the sample is expected to contain substantial sequence that is absent from the reference. Assembly-based methods are also valuable for pangenome studies that aim to capture the full genetic diversity of a species. The dairy cattle pangenome study demonstrated that assembly-based approaches identify over 10,000 additional structural variants per sample compared to short-read approaches.

### Can I combine both approaches in a single study?

Yes, combining both approaches is often the best strategy. A common design uses reference-based calling for SNPs and small indels, which are detected with high accuracy, and assembly-based methods for structural variants, which require the long-range information provided by assembly. The sunflower disease resistance study used this combined approach, identifying known disease resistance genes with reference-based methods and discovering novel genes absent from the reference through assembly of unmapped reads.

### How much computational power do I need for each approach?

Reference-based variant calling requires modest computational resources, typically a server with 16 to 32 gigabytes of RAM and multiple processor cores. Assembly-based variant detection requires substantially more resources, potentially hundreds of gigabytes of RAM and days of processing time for complex genomes. Researchers without access to high-performance computing may benefit from user-friendly interfaces such as AgrOmicSo, which provides a graphical interface to a server-side analysis engine.

### What causes false variant calls in reference-based methods?

False variant calls in reference-based methods are most commonly caused by reads mapping to multiple locations in repetitive regions, by structural variation between the reference and sample that induces spurious read mapping, and by sequencing errors that are not filtered out. A benchmark study of long-read variant calling found that structural variations between the reference and sample, particularly those containing repetitive elements, can induce spurious read mapping and constitute a dominant source of error.

### How do I validate structural variants detected by assembly?

Structural variants detected by assembly can be validated through several methods. PCR amplification across breakpoints can confirm deletions and insertions. Read depth analysis can confirm copy number changes. Comparison of the variant across multiple samples can identify recurrent variants that are more likely to be genuine. For variants in regions of known biological importance, additional validation is recommended before drawing conclusions.

### What is reference bias and why does it matter?

Reference bias occurs when reads from a sample that differs from the reference genome fail to map or map incorrectly, leading to systematic errors in variant detection. This bias affects both SNP calling and structural variant detection, but it is more severe for structural variants because they involve larger sequence differences. The sunflower disease resistance study found 198 disease resistance genes that were absent from the reference genome, demonstrating that reference-based methods alone would miss this variation entirely.

## Related Bioinformatics Guides

- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)
- [Hybrid Genome Assembly: Combining Short and Long Reads for Better Results](/knowledge/bioinformatics/hybrid-genome-assembly-combining-short-and-long-reads-for-better-results)
- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)
- [Metagenomics vs Metabarcoding: Choosing the Right Approach for Your Study](/knowledge/bioinformatics/metagenomics-vs-metabarcoding-choosing-the-right-approach-for-your-study)
- [Transcriptome Assembly Without a Reference Genome](/knowledge/bioinformatics/transcriptome-assembly-without-a-reference-genome)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Phased-assembly-driven pangenome graphs for structural variant genotyping and complex trait mapping in dairy cattle.](https://doi.org/10.1038/s41467-026-68807-4). 2026.
- [AgrOmicSo: A client-server interface for accessible large-scale analysis of next-generation sequencing data.](https://doi.org/10.1371/journal.pone.0348571). 2026.
- [Benchmarking long-read variant calling in diploid and polyploid genomes: insights from human and plants.](https://doi.org/10.1186/s12864-025-12259-5). 2026.
- [Expanding the disease-resistance gene repertoire for sunflower breeding through pan-genomic insights](https://doi.org/10.21203/rs.3.rs-9558005/v1). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.