Detecting Structural Variants with Reference-Guided Assembly: A Tutorial Using Assemblytics and SVIM-asm
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Structural variant (SV) detection from long-read sequencing necessitates a workflow involving read alignment to a reference, assembly construction, and subsequent alignment of the assembly to the reference to identify differences >50 bp. Tools like Assemblytics and SVIM-asm offer complementary approaches for SV calling from these alignments.
- Reference-guided assembly mitigates reference bias by incorporating sample-specific sequences during assembly, unlike traditional single-reference genotyping which can miss population-specific genetic variation crucial for crop improvement and understanding phenotypic diversity.
- Long-read sequencing (PacBio HiFi, Oxford Nanopore) is critical for SV detection due to its ability to span repetitive regions and variant breakpoints, which short-read sequencing cannot reliably resolve, thereby enabling more comprehensive pangenome construction.
- Assembly quality is paramount; metrics like N50, L50, and BUSCO completeness must be assessed to minimize false positive SV calls arising from assembly errors or fragmentation.
- Minimap2 is a primary tool for aligning assembled contigs to the reference genome, producing PAF files essential for downstream SV callers like Assemblytics, with specific presets (e.g.,
asm5) chosen based on assembly accuracy. - Comparing SV calls from multiple tools (e.g., Assemblytics and SVIM-asm) and performing manual inspection using genome browsers like IGV are crucial quality control steps to validate high-confidence variants and filter out false positives.
Structural variant detection from long-read sequencing data requires a workflow that converts raw reads into a genome assembly, aligns that assembly to a reference genome, and then identifies differences larger than 50 base pairs. Reference-guided assembly provides the computational framework for this process, and tools such as Assemblytics and SVIM-asm offer distinct approaches to variant calling from the resulting alignments. This tutorial walks through the complete protocol, from input data preparation through output interpretation, with attention to quality controls, common failure modes, and the biological meaning of each variant class.
The practical problem addressed here is straightforward. A researcher has generated long-read sequencing data from an organism of interest and wants to know how its genome differs structurally from a reference genome. The differences may include deletions, insertions, inversions, duplications, and more complex rearrangements. Short-read resequencing approaches capture single-nucleotide polymorphisms and small insertions and deletions but struggle with structural variants, which leads to a loss of heritability in genome-wide association studies. Long-read sequencing has improved pangenome construction for diverse eukaryotic species, including humans, crops, and other organisms of ecological and economic importance, addressing this issue to some extent. The workflow described here converts long-read data into a reference-guided assembly and then applies two complementary variant callers to extract structural variant information.
At a Glance
The table below summarizes the key decisions in a reference-guided structural variant detection workflow. Each row represents a workflow stage with the primary tool options, the input required, and the output produced.
| Workflow Stage | Primary Tool Options | Required Input | Output Produced |
|---|---|---|---|
| Read alignment to reference | Minimap2, Winnowmap2 | Long reads (HiFi, ONT), reference genome FASTA | Aligned reads in BAM or PAF format |
| Assembly construction | Hifiasm, Flye, HiFiCCL | Long reads, optional reference guidance | Contig-level assembly in FASTA format |
| Assembly-to-reference alignment | Minimap2, Nucmer | Assembled contigs, reference genome FASTA | Alignment file in PAF or delta format |
| Variant calling from assembly alignment | Assemblytics, SVIM-asm | Alignment file, reference genome FASTA | Structural variant calls in VCF or BED format |
| Variant filtering and interpretation | Bcftools, custom scripts | Raw variant calls, annotation files | Filtered variant set with functional annotations |
The choice of assembly strategy depends on read coverage and computational resources. Sufficient-coverage high-fidelity data for population genomics is often prohibitively expensive, limiting its use in large-scale populations and broader eukaryotic species and creating an urgent need for robust low-coverage assemblies. Current assemblers underperform in such conditions, which has motivated the development of reference-guided assembly frameworks specifically designed for low-coverage high-fidelity reads. These frameworks use a reference-guided, chromosome-by-chromosome assembly approach that improves low-coverage assembly performance of existing assemblers and outperforms the state-of-the-art assemblers on human and plant datasets.
Scope and Reader Context
This tutorial assumes the reader has a working knowledge of the Linux command line, basic familiarity with FASTA and BAM file formats, and access to a computing environment with at least 16 gigabytes of RAM. The workflow applies to diploid and polyploid genomes, though polyploid genomes require additional care during variant interpretation. The methods described here work with PacBio HiFi reads, Oxford Nanopore Technologies reads, and assembled contigs from any source. The primary output is a set of structural variant calls that can be used for downstream analysis such as population genetics, association studies, or functional annotation.
The reader should understand that reference-guided assembly differs from de novo assembly in an important way. In reference-guided assembly, the reference genome serves as a scaffold for ordering and orienting contigs, which can improve the contiguity of the final assembly. RaGOO, a reference-guided contig ordering and orienting tool, leverages the speed and sensitivity of Minimap2 to accurately achieve chromosome-scale assemblies in minutes. After the pseudomolecules are constructed, RaGOO identifies structural variants, including those spanning sequencing gaps. This approach has been used to accurately order and orient three de novo tomato genome assemblies and to perform a pan-genome analysis of 103 Arabidopsis thaliana accessions by examining the structural variants detected in the newly assembled pseudomolecules.
Core Principles of Structural Variant Detection
Structural variants are genomic alterations that affect segments of DNA larger than 50 base pairs. They include deletions, insertions, inversions, duplications, and translocations. The biological consequences of structural variants can be substantial. In Brassica napus, a recently evolved, globally important allopolyploid crop, species-wide genome structural variation plays a role in intraspecific and ecogeographical diversification. Whole-genome long-read DNA sequencing and reference-guided genome assemblies for 94 accessions, including winter-type, spring-type, and East Asian oilseed, along with kale forms and swedes or rutabagas, revealed pangenome-wide patterns for insertions, deletions, inversions, and large chromosomal deletions and duplications. These patterns reflect evolutionary diversification across morphotypes and ecotypes.
The detection of structural variants from reference-guided assemblies relies on a fundamental principle. When an assembled contig is aligned to a reference genome, differences in the alignment structure reveal the presence of variants. A deletion in the sample genome appears as a gap in the alignment where the reference has sequence but the assembly does not. An insertion appears as sequence in the assembly that does not align to the reference. An inversion appears as a reversal in the alignment orientation. Duplications appear as regions where the assembly has extra copies of a reference segment.
The accuracy of structural variant detection depends on the quality of the assembly and the alignment. Assembly errors can create false structural variant calls, and alignment errors can miss true variants. The workflow described here includes multiple quality checkpoints to minimize both types of errors.
The Role of Long-Read Sequencing
Long-read sequencing technologies produce reads that span repetitive regions and structural variant breakpoints, which makes them well suited for structural variant detection. Short-read sequencing cannot reliably resolve these regions because the reads are too short to span the variant or its surrounding repetitive context. The advent of the pangenome era has unraveled previously unknown genetic variation existing within diverse crop plants, including rice. This untapped genetic variation accounts for a major portion of phenotypic variation existing in crop plants. However, the use of conventional single reference-guided genotyping often fails to capture a large portion of this genetic variation, leading to a reference bias. This makes it difficult to identify and utilize novel population or cultivar-specific genes for crop improvement.
Long-read sequencing has improved pangenome construction for diverse eukaryotic species, including humans, crops, and other organisms of ecological and economic importance. The cost of sufficient-coverage high-fidelity data for population genomics is often prohibitively expensive, limiting its use in large-scale populations and broader eukaryotic species. This cost constraint has created an urgent need for robust low-coverage assemblies, and reference-guided assembly approaches have been developed to address this need.
Reference Bias and Its Mitigation
Reference bias occurs when a sample genome is compared to a reference genome that does not represent the sample's population. Sequences that are present in the sample but absent from the reference are missed by conventional analysis approaches. The Rice Pangenome Genotyping Array was developed to address this problem by harboring probes assaying 80K single-nucleotide polymorphisms and presence-absence variants spanning the entire 3K rice pangenome. A genome-wide association study conducted using this array detected a total of 42 loci, including previously known as well as novel genomic loci regulating grain size and weight traits in rice. Eight of these identified trait-associated loci could not be detected with conventional single reference genome-based genome-wide association studies.
Reference-guided assembly mitigates reference bias by using the reference as a guide instead of as the sole source of variant information. The assembly process incorporates reads that do not align to the reference, which preserves sample-specific sequence. The subsequent alignment of the assembly to the reference then identifies both variants that are present in the reference and variants that are absent from it.
Input Data Preparation
The structural variant detection workflow begins with input data preparation. The quality of the input data directly affects the quality of the assembly and the accuracy of the variant calls. This section describes the steps required to prepare long-read data and reference genome files for the workflow.
Long-Read Data Requirements
Long-read data can come from PacBio HiFi sequencing or Oxford Nanopore Technologies sequencing. HiFi reads have high accuracy, typically above 99 percent, and lengths ranging from 10 to 25 kilobases. Nanopore reads can be longer but have lower per-read accuracy. Both types of reads can be used for structural variant detection, but the assembly parameters and quality expectations differ.
The coverage requirement depends on the assembly strategy. For de novo assembly with hifiasm, 30-fold coverage of HiFi reads is typically recommended. For reference-guided assembly approaches designed for low-coverage data, such as HiFiCCL, coverage of approximately 5-fold has been demonstrated to produce assemblies that excel in detecting large germline structural variants, minimize inter-chromosome mis-scaffolding, and improve the detection of specific germline and tumor somatic structural variants based on the pangenome graph.
Before beginning the workflow, check the read data for quality. Use FastQC or similar tools to assess read length distributions, quality scores, and adapter contamination. Remove adapters and low-quality bases with tools such as porechop for Nanopore data or HiFiAdapterFilt for PacBio data. The NCBI provides official descriptions of sequence resources and analysis services that can help with data quality assessment and submission.
Reference Genome Selection
The choice of reference genome affects the sensitivity and specificity of structural variant detection. The reference should be closely related to the sample genome. A distantly related reference introduces alignment ambiguity and increases the number of false variant calls. For well-studied organisms, a high-quality reference genome is usually available from NCBI or Ensembl. For less-studied organisms, the reference may be a draft assembly with lower contiguity.
Check the reference genome for the following properties before use. The assembly should be chromosome-level or at least scaffold-level. The sequence should be free of vector contamination and ambiguous bases. The annotation should be current if functional interpretation of variants is planned. The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that can help identify appropriate reference genomes.
Directory Structure and File Organization
Organize the workflow in a clear directory structure to ensure reproducibility. A suggested structure is shown below.
project/
reads/
sample1.fastq.gz
sample2.fastq.gz
reference/
reference.fasta
reference.fasta.fai
assembly/
sample1.asm.fasta
sample2.asm.fasta
alignment/
sample1.asm.paf
sample2.asm.paf
variants/
sample1.assemblytics.vcf
sample1.svim.vcf
logs/
assembly.log
alignment.log
variant_calling.log
Record the versions of all software tools used in the workflow. This information is essential for reproducibility and for troubleshooting when results differ between runs. The nf-core documentation provides community pipeline standards, usage, configuration, and reproducible workflow context that can inform the organization of analysis projects.
Assembly Construction
The assembly step converts raw long reads into contiguous sequences that represent the sample genome. The choice of assembler and assembly strategy depends on the read type, coverage, and computational resources available.
De Novo Assembly with Hifiasm
Hifiasm is a haplotype-resolved assembler designed for HiFi reads. It produces primary and alternate assemblies that represent the two haplotypes of a diploid genome. The primary assembly is typically used for structural variant detection because it represents the consensus of the two haplotypes.
Run hifiasm with the following command structure.
hifiasm -o sample.asm -t 32 reads.fastq.gz
The output includes the primary assembly in the file sample.asm.bp.p_ctg.fa and the alternate assembly in sample.asm.bp.a_ctg.fa. The primary assembly is the default input for downstream structural variant detection.
For low-coverage data, standard assemblers underperform. HiFiCCL, a reference-guided, chromosome-by-chromosome assembly approach, has been shown to improve low-coverage assembly performance of existing assemblers and outperform the state-of-the-art assemblers on human and plant datasets. Tested on 45 human datasets at approximately 5-fold coverage, HiFiCCL combined with hifiasm reduces the length of misassembled contigs relative to hifiasm by an average of 21.19 percent and up to 38.58 percent.
Reference-Guided Assembly Approaches
Reference-guided assembly uses the reference genome to guide the assembly process. This approach is particularly useful for low-coverage data where de novo assembly produces fragmented results. The reference provides a framework for ordering and orienting contigs, which improves the contiguity of the final assembly.
RaGOO is a reference-guided contig ordering and orienting tool that leverages the speed and sensitivity of Minimap2 to accurately achieve chromosome-scale assemblies in minutes. After the pseudomolecules are constructed, RaGOO identifies structural variants, including those spanning sequencing gaps. This tool has been used to accurately order and orient three de novo tomato genome assemblies, including the widely used M82 reference cultivar, and to perform a pan-genome analysis of 103 Arabidopsis thaliana accessions.
The choice between de novo and reference-guided assembly depends on the research question. De novo assembly preserves all sample-specific sequence without reference bias but may produce fragmented assemblies at low coverage. Reference-guided assembly improves contiguity but may introduce reference bias if the reference is distantly related to the sample.
Assembly Quality Assessment
Assembly quality assessment is a critical checkpoint before structural variant detection. A poor-quality assembly produces false structural variant calls that are difficult to distinguish from true variants.
Use QUAST to assess assembly contiguity metrics such as N50, L50, and total assembly size. Compare these metrics to the expected genome size for the organism. A substantially smaller assembly may indicate missing sequence, while a substantially larger assembly may indicate contamination or assembly errors.
Use BUSCO to assess the completeness of the assembly by searching for conserved single-copy orthologs. A complete assembly should contain most BUSCO genes in single copy. Missing BUSCO genes may indicate assembly gaps, while duplicated BUSCO genes may indicate haplotype duplication or assembly errors.
The Galaxy Training Network provides accessible workflow training, analysis tutorials, and reproducibility context that can help with assembly quality assessment. The Carpentries lessons provide foundational computing, data, shell, Git, and programming training context that is useful for researchers new to command-line analysis.
Alignment of Assembly to Reference
The alignment step maps the assembled contigs to the reference genome. The alignment file serves as the input for both Assemblytics and SVIM-asm. The quality of the alignment directly affects the accuracy of the structural variant calls.
Minimap2 Alignment
Minimap2 is a fast and sensitive aligner that works well for long sequences. It can align assembled contigs to a reference genome in PAF format, which is the input format for Assemblytics.
Run Minimap2 with the following command structure.
minimap2 -x asm5 reference.fasta assembly.fasta > assembly.paf
The asm5 preset is appropriate for aligning assemblies with high accuracy, such as HiFi assemblies. For assemblies with lower accuracy, such as Nanopore assemblies, the asm10 or asm20 presets may be more appropriate. These presets adjust the alignment parameters to tolerate higher error rates.
The PAF output format contains the following columns: query name, query length, query start, query end, strand, target name, target length, target start, target end, and mapping quality. The alignment information in these columns is used by Assemblytics to identify structural variants.
Alignment Quality Checks
Check the alignment file for the following properties before proceeding to variant calling. The alignment rate should be high, typically above 90 percent of the assembly length. The mapping quality distribution should show a peak at high values. The alignment coverage should be uniform across the reference, with no large gaps that might indicate missing sequence or misassembly.
Use the PAF file to calculate alignment statistics. The number of aligned bases divided by the total assembly length gives the alignment rate. The number of primary alignments divided by the total number of alignments gives the fraction of unique alignments. Low values for either metric indicate alignment problems that should be investigated before variant calling.
Handling Multi-Contig Assemblies
Assemblies often contain multiple contigs, especially when the assembly is fragmented or when the sample is polyploid. The alignment step treats each contig independently, which means that a structural variant spanning a contig boundary will not be detected. This limitation should be considered when interpreting the results.
For chromosome-level assemblies, the contigs correspond to chromosomes, and the alignment is straightforward. For fragmented assemblies, the contigs may correspond to smaller genomic segments, and structural variants that span contig boundaries will be missed. Reference-guided scaffolding with tools such as RaGOO can improve the contiguity of the assembly before alignment, which reduces this problem.
Structural Variant Calling with Assemblytics
Assemblytics is a tool that identifies structural variants from the alignment of an assembly to a reference genome. It uses the PAF alignment file produced by Minimap2 and the reference genome sequence to identify deletions, insertions, inversions, and duplications.
Running Assemblytics
Assemblytics is available as a command-line tool and as a web server. The command-line version is appropriate for batch processing and for integration into automated workflows.
Run Assemblytics with the following command structure.
Assemblytics assembly.paf reference.fasta 500 10000 output_prefix
The first numeric argument is the minimum variant size, and the second is the maximum variant size. Variants smaller than the minimum size are ignored, and variants larger than the maximum size are reported separately. The default minimum size is 50 base pairs, and the default maximum size is 10,000 base pairs. Adjust these values based on the research question. For detection of large structural variants, increase the maximum size to 100,000 base pairs or more.
The output includes a VCF file with the structural variant calls and a set of BED files with the variant coordinates. The VCF file is the primary output for downstream analysis.
Understanding Assemblytics Output
Assemblytics identifies structural variants by analyzing the alignment structure. A deletion appears as a region where the reference has sequence but the assembly does not. An insertion appears as a region where the assembly has sequence but the reference does not. A tandem duplication appears as a region where the assembly has extra copies of a reference segment. An inversion appears as a region where the assembly aligns in the reverse orientation.
The VCF output contains the following information for each variant: chromosome, position, reference allele, alternate allele, variant type, and genotype. The variant type is encoded in the INFO field as DEL, INS, DUP, or INV. The size of the variant is encoded in the INFO field as SVLEN.
The Assemblytics output also includes a set of summary files that describe the distribution of variant sizes and types. These files are useful for quality assessment and for comparing multiple samples.
Interpreting Assemblytics Results
The interpretation of Assemblytics results requires an understanding of the biological context. A deletion in a gene may cause loss of function, while an insertion in a regulatory region may affect gene expression. The functional impact of a structural variant depends on its location and size.
For crop improvement applications, structural variants can be associated with phenotypic variation. In rice, the advent of the pangenome era has unraveled previously unknown genetic variation existing within diverse crop plants. This untapped genetic variation accounts for a major portion of phenotypic variation existing in crop plants. The use of conventional single reference-guided genotyping often fails to capture a large portion of this genetic variation, leading to a reference bias.
For evolutionary studies, structural variants can reveal patterns of diversification. In Brassica napus, collective structural variants are unevenly distributed and biased toward subgenome A, with asymmetrical selection pattern favoring subgenome C. Selection signatures for inversions exhibit no subgenome asymmetry, however, increased selection signal strength and frequency are detected in paracentric chromosome regions, highlighting evolutionary significance.
Structural Variant Calling with SVIM-asm
SVIM-asm is a tool that identifies structural variants from the alignment of an assembly to a reference genome. It uses a different algorithm than Assemblytics and provides complementary information. Running both tools and comparing the results can improve the confidence of structural variant calls.
Running SVIM-asm
SVIM-asm is available as a command-line tool. It requires the alignment file in PAF format and the reference genome sequence.
Run SVIM-asm with the following command structure.
svim-asm diploid sample_prefix assembly.paf reference.fasta
The diploid mode is appropriate for diploid samples. For haploid samples, use the haploid mode. The output includes a VCF file with the structural variant calls and a set of BAM files with the supporting alignments.
Understanding SVIM-asm Output
SVIM-asm identifies structural variants by analyzing the alignment structure in a manner similar to Assemblytics but with different algorithmic details. The output includes deletions, insertions, inversions, and duplications. The VCF file contains the variant coordinates, types, sizes, and supporting evidence.
The SVIM-asm output also includes a set of quality metrics for each variant. These metrics describe the number of supporting reads, the consistency of the breakpoints, and the confidence of the call. Use these metrics to filter low-quality calls.
Comparing Assemblytics and SVIM-asm Results
Assemblytics and SVIM-asm use different algorithms to identify structural variants, so the results will not be identical. Some variants will be detected by both tools, some by only one tool, and some by neither. The overlap between the two tools provides a measure of confidence.
Use the following approach to compare the results. Convert both VCF files to BED format. Use BEDTools to find the intersection of the two variant sets. Variants that are detected by both tools with consistent breakpoints and sizes are high-confidence calls. Variants that are detected by only one tool should be examined manually to determine whether they are true variants or false positives.
The comparison of results from multiple tools is a standard practice in bioinformatics. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation that can help with the comparison and downstream analysis of variant calls.
Variant Filtering and Quality Control
The raw variant calls from Assemblytics and SVIM-asm contain false positives that must be filtered before downstream analysis. The filtering strategy depends on the research question and the quality of the assembly.
Filtering Criteria
Apply the following filtering criteria to the raw variant calls. Remove variants that are smaller than the minimum size threshold. Remove variants that are supported by fewer than the minimum number of reads. Remove variants that overlap assembly gaps or low-complexity regions. Remove variants that have inconsistent breakpoints between the two callers.
The specific thresholds depend on the data quality and the research question. For high-quality HiFi assemblies, a minimum of two supporting reads is often sufficient. For lower-quality assemblies, a higher threshold may be needed.
Manual Inspection of Variant Calls
Manual inspection of variant calls is essential for validating the results. Use a genome browser such as IGV to visualize the alignment at each variant locus. Check that the alignment structure is consistent with the variant call. A deletion should appear as a gap in the assembly coverage. An insertion should appear as extra sequence in the assembly. An inversion should appear as a reversal in the alignment orientation.
The manual inspection step is time-consuming but necessary for high-confidence results. The number of variants that require manual inspection depends on the filtering stringency and the quality of the assembly.
Quality Metrics for Variant Calls
Record the following quality metrics for the variant calls. The total number of variants detected by each tool. The number of variants detected by both tools. The size distribution of the variants. The type distribution of the variants. The number of variants that pass the filtering criteria.
These metrics provide a summary of the structural variant landscape of the sample and allow comparison between samples. The metrics also provide a quality check on the workflow. An unusually high number of variants may indicate assembly or alignment problems, while an unusually low number may indicate that the sample is closely related to the reference.
Records and Measurements
Maintaining detailed records of the structural variant detection workflow is essential for reproducibility and for troubleshooting. The records should include the software versions, the parameters used, the input data, and the output files.
Workflow Log
Create a workflow log that records the following information for each step. The date and time of the run. The software version. The command used. The input files. The output files. The runtime and resource usage. Any warnings or errors.
The workflow log serves as a permanent record of the analysis. It allows other researchers to reproduce the results and allows the original researcher to understand what was done when revisiting the analysis at a later time.
Variant Call Summary
Create a variant call summary that records the following information for each sample. The number of deletions, insertions, inversions, and duplications detected. The size distribution of each variant type. The number of variants that overlap genes. The number of variants that overlap regulatory regions.
The variant call summary provides a high-level view of the structural variant landscape of the sample. It is useful for comparing samples and for identifying samples with unusual structural variant profiles.
Reproducibility Considerations
Reproducibility requires that the analysis can be repeated with the same results. The following practices support reproducibility. Record the exact versions of all software tools. Record the exact parameters used for each tool. Record the exact input files. Use a workflow management system such as Snakemake or Nextflow to automate the analysis.
The nf-core documentation provides community pipeline standards, usage, configuration, and reproducible workflow context that can inform the development of reproducible analysis pipelines. The Galaxy Training Network provides accessible workflow training, analysis tutorials, and reproducibility context that is useful for researchers who prefer a graphical interface.
Common Failure Patterns
Several common failure patterns occur in structural variant detection from reference-guided assemblies. Recognizing these patterns allows the researcher to diagnose and correct problems quickly.
Low Alignment Rate
A low alignment rate indicates that a substantial fraction of the assembly does not align to the reference. This can occur when the reference is distantly related to the sample, when the assembly contains contamination, or when the assembly contains sequence from a different organism.
Check the alignment rate by calculating the fraction of the assembly length that is covered by alignments. If the alignment rate is below 90 percent, investigate the unaligned sequence. Use BLAST to search the unaligned sequence against the NCBI nucleotide database to identify the source of the sequence.
Excessive Variant Calls
An unusually high number of structural variant calls may indicate assembly or alignment problems. Assembly errors create false structural variants, and alignment errors can create false breakpoints.
Check the assembly quality metrics. A fragmented assembly with a low N50 will produce more false variants than a contiguous assembly. Check the alignment quality. Low mapping quality scores indicate ambiguous alignments that may produce false variants.
Missing Expected Variants
A failure to detect expected structural variants may indicate that the variant size threshold is too high, that the assembly is missing the variant region, or that the alignment is not sensitive enough.
Check the variant size threshold. If the expected variants are smaller than the minimum size threshold, they will not be detected. Check the assembly coverage of the variant region. If the assembly has a gap at the variant locus, the variant will not be detected.
Inconsistent Results Between Callers
Inconsistent results between Assemblytics and SVIM-asm are expected to some degree, but large discrepancies indicate a problem. The discrepancy may be due to differences in the algorithms, differences in the filtering criteria, or errors in one of the callers.
Compare the variant calls from the two tools. Identify the variants that are detected by only one tool. Examine these variants manually to determine whether they are true variants or false positives.
Limitations of Reference-Guided Assembly
Reference-guided assembly has several limitations that should be considered when interpreting the results. The reference genome introduces bias, the assembly process can miss complex variants, and the alignment step can introduce errors.
Reference Bias
Reference-guided assembly uses the reference genome as a guide, which introduces reference bias. Sequences that are present in the sample but absent from the reference may be underrepresented in the assembly. This bias is particularly problematic for samples that are distantly related to the reference.
The Rice Pangenome Genotyping Array was developed to overcome reference bias by providing a genotyping solution that spans the entire 3K rice pangenome. This array provides a simple, user-friendly and cost-effective solution for rapid pangenome-based genotyping in rice. The genome-wide association study conducted using this array detected a total of 42 loci, including previously known as well as novel genomic loci regulating grain size and weight traits in rice.
Complex Variant Missed
Reference-guided assembly can miss complex structural variants that involve multiple breakpoints or that span assembly gaps. The assembly process may not resolve these regions correctly, and the alignment step may not detect the variant.
The detection of complex variants requires specialized approaches. The study of rearrangement hotspots within segmental duplications in humans used a hierarchical method to fragment segmental duplications into multiple smaller units and combined an end space free pairwise alignment algorithm with a seed and extend approach to detect complex structural rearrangements within the reference-guided assembly of the NA18507 human genome. This approach identified 1,963 rearrangement hotspots within segmental duplications that encompass 166 genes.
Alignment Errors
The alignment of the assembly to the reference can introduce errors, particularly in repetitive regions. The aligner may place the assembly sequence at the wrong location, creating false structural variants or missing true variants.
The use of multiple aligners and the comparison of results can help identify alignment errors. The manual inspection of variant calls in a genome browser is the most reliable method for validating the results.
Welfare and Safety Context
The structural variant detection workflow described here is a computational analysis that does not involve animal subjects or biological safety concerns. However, the workflow may be used in the context of agricultural research, where the results inform breeding decisions and crop improvement strategies.
The use of structural variant information in crop improvement has the potential to accelerate the development of improved varieties. In Brassica napus, the analysis of candidate loci under selection identified regions for collective structural variants and inversions, harboring genes for organ formation, cell division and expansion in swede, and stress responses in East Asian oilseed rape. Large chromosomal duplications and deletions distinguish swede from oilseed rape, particularly in subgenome C, including copy-number variation in flowering-time genes BnFLC.C09 and BnATX2.C08, and cell wall development gene BnCEL2.C08.
The responsible use of structural variant information requires consideration of the ethical and regulatory context. Researchers should ensure that their work complies with applicable regulations and guidelines for genetic analysis and crop improvement. The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that can help with data management and compliance.
Professional Escalation Criteria
The structural variant detection workflow may produce results that require professional escalation. The following situations warrant consultation with a bioinformatics specialist or a domain expert.
Unexpected Variant Patterns
If the structural variant calls show an unexpected pattern, such as an unusually high number of variants in a specific genomic region or an unusual distribution of variant types, consult a specialist. The pattern may indicate a biological phenomenon of interest or a technical artifact.
Discrepant Results Between Samples
If the structural variant calls are discrepant between samples that are expected to be similar, consult a specialist. The discrepancy may indicate a sample quality issue, a contamination issue, or a technical problem.
Results with Major Biological Implications
If the structural variant calls have major biological implications, such as the identification of a variant that affects a trait of economic importance, consult a domain expert. The expert can help interpret the results in the context of the biological system and can guide the design of follow-up experiments.
Technical Problems
If the workflow produces technical problems that cannot be resolved with the troubleshooting steps described here, consult a bioinformatics specialist. The specialist can help diagnose the problem and can recommend alternative approaches.
Frequently Asked Questions
What is the difference between Assemblytics and SVIM-asm?
Assemblytics and SVIM-asm are both tools for detecting structural variants from the alignment of an assembly to a reference genome, but they use different algorithms. Assemblytics analyzes the PAF alignment file to identify deletions, insertions, inversions, and duplications based on the alignment structure. SVIM-asm uses a similar approach but with different algorithmic details and provides additional quality metrics for each variant. Running both tools and comparing the results can improve the confidence of structural variant calls.
What coverage of long reads is needed for structural variant detection?
The coverage requirement depends on the assembly strategy. For de novo assembly with hifiasm, 30-fold coverage of HiFi reads is typically recommended. For reference-guided assembly approaches designed for low-coverage data, coverage of approximately 5-fold has been demonstrated to produce assemblies that excel in detecting large germline structural variants. The choice of coverage should balance the cost of sequencing against the quality of the assembly and the accuracy of the variant calls.
How do I choose the minimum and maximum variant size thresholds?
The minimum and maximum variant size thresholds depend on the research question. The default minimum size is 50 base pairs, which is the standard definition of a structural variant. The default maximum size is 10,000 base pairs, but this can be increased to detect larger variants. For detection of large structural variants, increase the maximum size to 100,000 base pairs or more. The choice of thresholds should be recorded in the workflow log for reproducibility.
What is reference bias and how does it affect structural variant detection?
Reference bias occurs when a sample genome is compared to a reference genome that does not represent the sample population. Sequences that are present in the sample but absent from the reference are missed by conventional analysis approaches. Reference-guided assembly mitigates reference bias by using the reference as a guide instead of as the sole source of variant information. The assembly process incorporates reads that do not align to the reference, which preserves sample-specific sequence.
How do I validate structural variant calls?
Validation of structural variant calls involves multiple steps. First, compare the results from Assemblytics and SVIM-asm and focus on the variants detected by both tools. Second, manually inspect the variant calls in a genome browser to confirm that the alignment structure is consistent with the variant call. Third, if possible, validate a subset of variants using an independent method such as PCR amplification or Southern blotting.
What should I do if the alignment rate is low?
A low alignment rate indicates that a substantial fraction of the assembly does not align to the reference. This can occur when the reference is distantly related to the sample, when the assembly contains contamination, or when the assembly contains sequence from a different organism. Check the alignment rate by calculating the fraction of the assembly length that is covered by alignments. If the alignment rate is below 90 percent, investigate the unaligned sequence using BLAST against the NCBI nucleotide database.
How do I handle polyploid genomes in structural variant detection?
Polyploid genomes require additional care during structural variant detection. The assembly process must correctly represent the multiple haplotypes, and the alignment step must correctly map the haplotypes to the reference. The interpretation of variant calls in polyploid genomes is more complex because a variant may be present in one haplotype but not in another. Consult a specialist with experience in polyploid genomics for guidance on the appropriate approach.
What are the common causes of false positive structural variant calls?
Common causes of false positive structural variant calls include assembly errors, alignment errors, and repetitive regions. Assembly errors create false structural variants by introducing sequence that does not match the true genome. Alignment errors create false breakpoints by placing the assembly sequence at the wrong location. Repetitive regions create ambiguity in the alignment, which can lead to false variant calls. The use of multiple callers and manual inspection can help identify and remove false positives.
Related Bioinformatics Guides
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Hybrid Genome Assembly: Combining Short and Long Reads for Better Results
- Long-Read Metagenome Assembly: Overcoming Challenges with Nanopore and PacBio Data
- Gene Set Enrichment Analysis in R: A Practical Tutorial for Interpreting Omics Data
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Reference-Guided Chromosome-by-Chromosome de novo Assembly at Scale Using Low-Coverage High-Fidelity Long-Reads with HiFiCCL.. Advanced science (Weinheim, Baden-Wurttemberg, Germany), 2026.
- RaGOO: fast and accurate reference-guided scaffolding of draft genomes.. Genome biology, 2019.
- Rice Pangenome Genotyping Array: an efficient genotyping solution for pangenome-based accelerated genetic improvement in rice.. The Plant journal : for cell and molecular biology, 2023.
- Pangenomic structural variant patterns reflect evolutionary diversification in Brassica napus.. Genome biology, 2025.
- Genome-wide signatures of 'rearrangement hotspots' within segmental duplications in humans.. PloS one, 2011.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.