# How to Perform Reference-Based Scaffolding: A Step-by-Step Pipeline with BWA, SAMtools, and RagTag


## Key Takeaways

- Reference-based scaffolding orders and orients draft genome assembly contigs using a closely related, high-quality reference genome as a template, significantly improving contiguity for downstream comparative genomics.
- The pipeline employs BWA-MEM for efficient read mapping to the reference, SAMtools for processing SAM/BAM alignment files into sorted and indexed BAMs, and RagTag for assembly correction and scaffold construction.
- Critical quality checks include a mapping rate exceeding 90% for closely related references, verification of BAM file integrity with `samtools quickcheck`, and assessment of scaffold count reduction and N50 increase.
- Common failure modes include low mapping rates due to divergent references, undetected misjoins requiring sufficient read coverage, and chimeric scaffolds arising from structural rearrangements or repetitive sequences.
- The choice of reference genome is paramount; it must be taxonomically proximate, complete (chromosome-level), and publicly available, with divergence estimated by mapping a read subset to ensure reliable alignment statistics.
- Validation of scaffolded assemblies is crucial, involving alignment-based checks for consistency with the reference, assessment of gap sizes against read evidence, and potentially independent data such as optical or Hi-C maps.

---

Reference-based scaffolding is a computational procedure that orders and orients draft assembly contigs against a closely related reference genome to produce a chromosome-level layout. This article provides a concrete, reproducible command-line protocol using BWA for read mapping, SAMtools for alignment processing, and RagTag for scaffold construction, with explicit parameter choices, quality checks, and troubleshooting guidance for common failure modes including misjoins and chimeric contigs.

The intended reader is a biology student, researcher, or laboratory professional who has already produced a draft genome assembly and needs to improve its contiguity for downstream comparative analysis. The protocol assumes basic familiarity with the Unix shell, FASTA and FASTQ formats, and the concept of sequence alignment. For foundational shell and data skills, structured lessons from [The Carpentries](https://carpentries.org/lessons) provide practical training in command-line navigation, file manipulation, and scripting that will be useful before attempting this pipeline.

## Why Reference-Based Scaffolding Is Needed

De novo assembly from short sequencing reads frequently produces fragmented draft genomes. Repetitive elements longer than the read length create ambiguous graph traversals, and heuristic assembly algorithms can misresolve branches, leaving contigs disconnected even when the underlying sequence is complete. A study of bacterial genome assembly tools noted that high-throughput short-read platforms generate millions of fragmented reads that must be assembled into contigs, and those contigs often require further downstream processing before they are useful for comparative genomics. The same study observed that scaffolding based on paired-end information alone struggles when repetitive elements exceed the read size, while long-read technologies can resolve those repeats but are not universally available to researchers. Reference-based ordering offers a practical middle path: it uses a complete genome from a related organism as a template to arrange draft contigs in the correct linear order.

The value of this approach extends beyond bacteria. Viral metagenomic studies frequently produce highly fragmented contigs from de novo assembly, and recovering complete or near-complete viral genomes from such data is a recognized challenge. A computational pipeline that maps reads to viral reference genomes, assembles the mapped reads, and extends contigs through iterative assembly was able to recover complete viral genomes from ocean metagenomic samples, including two novel viruses not present in existing databases. That work demonstrates that hybrid approaches combining reference-based mapping with de novo assembly can reconstruct genomes that neither method achieves alone.

Reference-based scaffolding is also relevant to transcriptomic analysis. A comprehensive reanalysis of bulk RNA sequencing data from the New York Genome Center ALS Consortium cohort used dual analytical pipelines, one reference-based to model canonical splicing events and one de novo to detect transcript structural novelties. The reference-based approach revealed that the ALS transcriptome is defined primarily by splicing failure instead of changes in gene expression, with aberrant intron retention outnumbering differentially expressed genes by an order of magnitude. This finding illustrates that reference-based methods can detect biological signals that de novo approaches miss, and vice versa, which is why many research projects run both strategies in parallel.

## Core Principles of Reference-Based Scaffolding

### The Reference Genome as a Template

Reference-based scaffolding treats a high-quality genome from a related organism as a linear template. Each draft contig is aligned to the reference, and the alignment coordinates determine the contig's position and orientation in the final scaffold. The underlying assumption is that the draft genome and the reference share sufficient synteny, meaning conserved gene order along the chromosome, for the reference to be a reliable guide. When that assumption holds, scaffolding dramatically improves contiguity and produces an ordered layout that reflects the true chromosome structure.

The choice of reference genome is the single most important decision in this pipeline. The reference should be the closest available relative to the sequenced organism, ideally from the same species or genus. For bacterial projects, the NCBI [GenBank and RefSeq databases](https://www.ncbi.nlm.nih.gov/) provide curated reference genomes for thousands of species, including complete chromosomes with annotated genes. For eukaryotic projects, the same databases host chromosome-level assemblies for model organisms and many non-model species. The [EMBL-EBI training portal](https://www.ebi.ac.uk/training) offers structured learning materials on selecting and using reference data resources, which can help researchers evaluate reference quality before committing to a scaffolding run.

### The Role of Read Mapping

The RagTag scaffolder does not align contigs directly to the reference. Instead, it maps the original sequencing reads to the reference genome, then uses the read alignments to infer the order and orientation of the draft contigs. This indirect approach is more robust than direct contig-to-reference alignment because reads carry information about adjacency that contigs alone do not. When paired-end reads map to the reference with one mate on each of two draft contigs, that read pair provides evidence that the two contigs are adjacent in the genome. The BWA aligner performs this read mapping efficiently, and SAMtools processes the resulting alignments into a format that RagTag can consume.

### The Scaffolding Algorithm

RagTag's correction module uses the read alignments to identify and fix errors in the draft assembly before scaffolding. It detects misjoins, where two unrelated sequences have been incorrectly joined into a single contig, and breaks the contig at the error point. It also identifies small-scale errors such as insertions, deletions, and substitutions relative to the reference, and corrects them in the output. The scaffolding module then orders the corrected contigs along the reference, determines the orientation of each contig, and joins them with gaps of estimated size. The final output is a set of scaffold sequences that follow the reference layout.

## At a Glance

| Pipeline Stage | Primary Tool | Input | Output | Key Quality Check |
| --- | --- | --- | --- | --- |
| Read mapping | BWA-MEM | Draft assembly, raw reads, reference genome | SAM alignment file | Mapping rate above 90 percent for closely related reference |
| Alignment processing | SAMtools | SAM file | Sorted, indexed BAM file | No truncation errors, proper mate pairs present |
| Assembly correction | RagTag correct | Draft assembly, BAM file | Corrected FASTA | Misjoin count reported, corrected contigs align cleanly |
| Scaffold construction | RagTag scaffold | Corrected FASTA, reference genome | Scaffolded FASTA and AGP file | Scaffold count reduced, N50 increased, no chimeric joins |

## Preparing the Input Data

### Draft Assembly Requirements

The draft assembly should be in FASTA format with contig names that are unique and free of special characters. Spaces, pipes, and colons in contig names can cause parsing errors in downstream tools. If the assembly contains duplicate contig names, rename them before starting. The assembly should be the product of a completed de novo assembly run, not an intermediate graph or unitig output. For short-read assemblies, the draft will typically contain hundreds or thousands of contigs. For long-read assemblies, the draft may already be near-complete, and scaffolding will primarily serve to order contigs and close gaps.

### Raw Read Requirements

The pipeline requires the raw sequencing reads that were used to produce the draft assembly. Reads should be in FASTQ format, and paired-end reads should be in two files with matching read names. The read data must correspond to the same DNA sample as the draft assembly. Mixing reads from different samples or sequencing runs will produce incorrect scaffolding results. If the draft assembly was produced from a subset of the available reads, use the same subset for scaffolding to maintain consistency.

### Reference Genome Requirements

The reference genome should be a complete or chromosome-level assembly from a closely related organism. For bacterial projects, a complete genome from the same species is ideal. For eukaryotic projects, a chromosome-level assembly from the same species or a congener is preferred. The reference should be in FASTA format, and each chromosome or scaffold should have a unique name. Download the reference from a trusted source such as the NCBI [GenBank or RefSeq databases](https://www.ncbi.nlm.nih.gov/), and record the accession number and download date in the project notebook.

### Directory Structure and File Organization

Create a project directory with subdirectories for each pipeline stage. A suggested layout is:

```
scaffolding_project/
  reads/
  assembly/
  reference/
  mapping/
  correction/
  scaffolding/
  logs/
```

Store raw reads in `reads/`, the draft assembly in `assembly/`, and the reference genome in `reference/`. Write all intermediate files to the stage-specific directories. Keep a plain-text log file in `logs/` that records every command run, the tool version, and the date. This practice supports reproducibility and makes troubleshooting easier when results are unexpected.

## Step 1: Index the Reference Genome

Before mapping reads, the reference genome must be indexed with BWA. Indexing creates a set of files that allow BWA to rapidly locate seed matches during alignment. Run the following command from the `reference/` directory:

```bash
bwa index reference.fasta
```

This command produces five index files with suffixes `.amb`, `.ann`, `.bwt`, `.pac`, and `.sa`. The indexing step is computationally inexpensive for bacterial genomes and moderate for mammalian genomes. For very large plant or animal genomes, indexing may take several hours and require substantial memory. Check that all five index files are present before proceeding.

The reference FASTA file should contain only nucleotide characters and no duplicate sequence names. If the reference contains ambiguous bases, BWA will still index it, but reads mapping to ambiguous regions will be less informative. For most projects, the standard reference from NCBI is appropriate without additional masking.

## Step 2: Map Reads to the Reference with BWA-MEM

BWA-MEM is the recommended alignment algorithm for reads longer than 70 base pairs, which includes most Illumina and all long-read data. The command is:

```bash
bwa mem -t 8 -R '@RG\tID:sample\tSM:sample\tPL:ILLUMINA' \
  reference/reference.fasta \
  reads/sample_R1.fastq.gz reads/sample_R2.fastq.gz \
  > mapping/sample.sam
```

The `-t 8` flag specifies eight threads for parallel alignment. Adjust this number to match the available CPU cores. The `-R` flag adds a read group header line to the SAM output, which records the sample name and sequencing platform. Read groups are required for downstream tools that use sample metadata and are essential if the project will later involve variant calling or multi-sample analysis.

For single-end reads, omit the second FASTQ file:

```bash
bwa mem -t 8 -R '@RG\tID:sample\tSM:sample\tPL:ILLUMINA' \
  reference/reference.fasta \
  reads/sample_R1.fastq.gz \
  > mapping/sample.sam
```

For long reads from Oxford Nanopore or Pacific Biosciences platforms, BWA-MEM still works but may be slower than platform-specific aligners. The protocol described here is compatible with any aligner that produces SAM or BAM output, so long-read projects can substitute minimap2 or another aligner at this stage.

### Mapping Quality Assessment

After mapping completes, assess the mapping rate before proceeding. A high mapping rate indicates that the reference is closely related to the sequenced sample. Use SAMtools to compute basic alignment statistics:

```bash
samtools flagstat mapping/sample.sam
```

The output reports the total number of reads, the number mapped, and the mapping percentage. For a closely related reference, expect more than 90 percent of reads to map. Lower mapping rates suggest that the reference is too divergent, and scaffolding results will be unreliable. In that case, consider a different reference genome or a de novo scaffolding approach.

## Step 3: Process Alignments with SAMtools

The SAM file produced by BWA must be converted to a sorted, indexed BAM file before RagTag can use it. SAMtools provides the necessary utilities. Run the following commands:

```bash
samtools view -bS mapping/sample.sam > mapping/sample.bam
samtools sort -o mapping/sample.sorted.bam mapping/sample.bam
samtools index mapping/sample.sorted.bam
```

The first command converts the text-based SAM file to the binary BAM format. The second command sorts the alignments by reference coordinate, which is required for efficient random access. The third command creates a `.bai` index file that allows tools to quickly retrieve alignments from specific genomic regions.

### Verifying the BAM File

Before proceeding, verify that the sorted BAM file is complete and correctly formatted. Run:

```bash
samtools quickcheck mapping/sample.sorted.bam
```

This command exits with a zero status if the file is valid and a non-zero status if it is truncated or corrupted. Also run `samtools flagstat` on the sorted BAM to confirm that the mapping statistics match the original SAM file. Discrepancies indicate a processing error that must be resolved before scaffolding.

The sorted BAM file can be large, especially for eukaryotic genomes. Check available disk space before running the sort command, and use the `-T` flag to specify a temporary directory if needed:

```bash
samtools sort -T /tmp/samtools_sort -o mapping/sample.sorted.bam mapping/sample.bam
```

## Step 4: Correct the Assembly with RagTag

RagTag's correction module uses the read alignments to identify and fix assembly errors. Run the correction step before scaffolding so that misjoins and small-scale errors are resolved in the input assembly. The command is:

```bash
ragtag.py correct \
  reference/reference.fasta \
  mapping/sample.sorted.bam \
  -o correction/ \
  -t 8 \
  --debug
```

The first positional argument is the reference FASTA, and the second is the sorted BAM file. The `-o` flag specifies the output directory, and `-t` sets the number of threads. The `--debug` flag produces additional diagnostic output that is useful for troubleshooting but can be omitted for routine runs.

RagTag correct produces a corrected assembly FASTA file in the output directory. The file name follows the pattern `ragtag.correct.fasta`. The correction module also reports statistics about the number of misjoins detected and corrected. Record these numbers in the project log.

### Understanding the Correction Output

The corrected assembly should have the same number of contigs as the input assembly, unless misjoins were detected. When RagTag identifies a misjoin, it breaks the contig at the error point, increasing the contig count. Each break is reported in the log with the contig name and the break position. Review these breakpoints to confirm that they are biologically plausible. A misjoin break at a position where read coverage drops sharply or where the reference alignment switches from one chromosome to another is credible. A break in the middle of a well-covered, syntenic region may indicate an error in the correction algorithm.

The corrected assembly is the input to the scaffolding step. Do not skip the correction step, because scaffolding a misassembled draft will propagate errors into the final scaffold layout.

## Step 5: Scaffold the Corrected Assembly with RagTag

The scaffolding module orders and orients the corrected contigs along the reference. Run:

```bash
ragtag.py scaffold \
  reference/reference.fasta \
  correction/ragtag.correct.fasta \
  -o scaffolding/ \
  -t 8 \
  -C
```

The `-C` flag tells RagTag to keep the corrected assembly as the basis for scaffolding instead of using the original draft. This flag is important when the correction step was run separately, as in this protocol. Without `-C`, RagTag will attempt to correct the assembly again, which is redundant and may produce different results.

The scaffolding output directory contains several files:

- `ragtag.scaffold.fasta`: the scaffolded assembly in FASTA format
- `ragtag.scaffold.agp`: the AGP file describing the order, orientation, and gap sizes of contigs within each scaffold
- `ragtag.scaffold.conf`: a configuration file with the parameters used for the run
- `ragtag.scaffold.stats`: summary statistics for the scaffolding run

### Interpreting the AGP File

The AGP file is the definitive record of the scaffolding layout. Each line describes one component of a scaffold, including the scaffold name, the component's start and end position within the scaffold, the component type (contig or gap), and the contig's orientation. Review the AGP file to confirm that the scaffold layout is consistent with the reference order. Contigs should appear in the same order as their aligned positions in the reference, and orientations should be consistent with the alignment strand.

The AGP file also records gap sizes. RagTag estimates gap sizes from the read coverage and the distance between adjacent contigs in the reference. Gaps that are unexpectedly large, such as gaps spanning more than 10 kilobases in a bacterial genome, may indicate that the reference and the draft genome have structural differences at that locus. Investigate such regions before accepting the scaffold.

## Step 6: Assess Scaffolding Quality

### Assembly Statistics

Compute assembly statistics for the scaffolded output and compare them with the input draft. The key metrics are:

- Number of scaffolds
- N50, the length such that 50 percent of the assembly is in scaffolds of that length or longer
- L50, the number of scaffolds needed to reach half the assembly length
- Total assembly length
- Number of gaps

A successful scaffolding run reduces the scaffold count, increases N50, and maintains total assembly length within a few percent of the draft. Use a tool such as QUAST or a simple script to compute these statistics. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on assembly quality assessment that include practical exercises with these metrics.

### Alignment-Based Quality Checks

Beyond summary statistics, verify that the scaffolded assembly aligns cleanly to the reference. Map the scaffolded contigs to the reference using a whole-genome aligner such as MUMmer or minimap2, and inspect the alignment for large-scale inconsistencies. Each scaffold should align to a single reference chromosome in a consistent orientation. A scaffold that aligns to two different reference chromosomes, or that switches orientation midway, indicates a scaffolding error.

Also check that the gaps in the scaffold are supported by read evidence. The reads that map to the reference region spanning a gap should also map to the draft contigs flanking the gap. If no reads bridge a gap, the gap may be an artifact of the scaffolding algorithm instead of a genuine sequence gap.

### Visual Inspection

For small genomes, visual inspection of the scaffold layout is feasible and recommended. Plot the scaffolded contigs along the reference using a genome browser or a custom script. The [Bioconductor project](https://bioconductor.org/) provides R packages for genomic visualization and analysis that can generate publication-quality figures of scaffold layouts. Visual inspection can reveal ordering errors that summary statistics miss, such as a single contig placed in the wrong orientation or a scaffold that skips a reference region.

## Step 7: Record and Report Results

### Project Records

Maintain a complete record of the scaffolding run in the project log. Include:

- Tool versions for BWA, SAMtools, and RagTag
- Reference genome accession and download date
- Read file names and sequencing platform
- All command lines with parameters
- Mapping rate from `samtools flagstat`
- Correction statistics, including misjoin count
- Scaffolding statistics, including scaffold count and N50
- Any anomalies observed during quality assessment

This record supports reproducibility and provides the information needed to write the methods section of a manuscript or report.

### Reporting in Publications

When reporting reference-based scaffolding in a publication, state the reference genome accession, the versions of all tools, and the key parameters used. Report the assembly statistics before and after scaffolding in a table. Describe any manual curation steps, such as breaking a scaffold at a suspected misjoin or removing a contig that did not align to the reference. The [nf-core documentation](https://nf-co.re/docs) emphasizes the importance of reproducible workflow configuration and version reporting, principles that apply equally to manual pipelines.

## Common Failure Patterns and Troubleshooting

### Low Mapping Rate

A mapping rate below 90 percent for a bacterial genome or below 70 percent for a eukaryotic genome suggests that the reference is too divergent. Check whether a closer reference is available in the NCBI [databases](https://www.ncbi.nlm.nih.gov/). If no closer reference exists, consider whether reference-based scaffolding is appropriate for this project. A divergent reference will produce incorrect contig ordering, and the resulting scaffolds may be worse than the original draft.

### Misjoins Not Detected

RagTag correct identifies misjoins by analyzing read coverage and alignment consistency across contig boundaries. If the correction step reports zero misjoins but the scaffolded assembly has obvious errors, the misjoins may be too small for the algorithm to detect, or the read coverage may be insufficient. Check the read depth across the assembly. Regions with coverage below 10x may not provide enough evidence for reliable correction. Consider whether additional sequencing is needed before scaffolding.

### Chimeric Scaffolds

A chimeric scaffold joins sequences from two different genomic regions, often from different chromosomes. Chimeras arise when the reference has a structural rearrangement relative to the draft genome, or when repetitive sequences cause spurious read alignments. Detect chimeras by aligning the scaffold to the reference and looking for scaffolds that span two reference chromosomes. Break chimeric scaffolds at the junction point and treat the two pieces as separate contigs.

### Excessive Gap Sizes

Gaps in the scaffold represent sequence that is present in the reference but absent from the draft assembly. Large gaps may indicate that the draft assembly is missing a genomic region, such as a ribosomal RNA operon or a prophage, that is present in the reference. Small gaps of a few hundred base pairs are common and usually represent unresolved repeats. If gap sizes are consistently large across the scaffold, the reference may contain sequences that are genuinely absent from the sequenced sample, such as strain-specific mobile elements.

### Orientation Errors

An orientation error places a contig in the reverse orientation relative to the reference. These errors are detected by aligning the scaffold to the reference and looking for contigs whose alignment strand differs from the surrounding contigs. Orientation errors are more common in AT-rich genomes where the forward and reverse strands are difficult to distinguish. Check the read coverage at the contig boundaries, a sharp coverage drop at an orientation switch point supports the orientation error diagnosis.

## Limitations of Reference-Based Scaffolding

### Reference Bias

Reference-based scaffolding imposes the reference genome's structure on the draft assembly. If the sequenced sample has a structural variant, such as an inversion or a large insertion, that is absent from the reference, the scaffolding algorithm will place the affected contigs according to the reference layout, which may be incorrect for the sample. A study of bacterial genotyping noted that reference-based approaches can mask strain-specific information, which is a recognized limitation of this method. Researchers studying strains with known structural variation should interpret scaffold layouts with caution and validate suspicious regions with additional evidence.

### Repetitive Sequence Ambiguity

Repetitive elements longer than the read length create ambiguity in both assembly and scaffolding. Reads from different copies of a repeat map equally well to each copy, so the scaffolding algorithm cannot determine which copy a contig belongs to. This limitation is inherent to short-read data and is not resolved by reference-based scaffolding. Long-read sequencing can resolve repeats, but the protocol described here assumes short-read data. Projects with substantial repetitive content should consider long-read assembly or hybrid approaches.

### Reference Quality Dependence

The scaffolding result is only as good as the reference genome. A reference with assembly errors, such as misjoins or missing regions, will produce incorrect scaffolds. Check the reference assembly statistics before use. A reference with an N50 below the chromosome length for a bacterial genome, or below the chromosome arm length for a eukaryotic genome, may have assembly errors that compromise scaffolding.

### Divergent Genomes

Reference-based scaffolding requires sufficient sequence similarity between the draft and the reference for reads to map reliably. For genomes that are more than 10 to 15 percent divergent at the nucleotide level, read mapping becomes unreliable, and scaffolding results will be poor. In such cases, consider a de novo scaffolding approach using long reads or linked reads, or use a more distant reference only for chromosome-level ordering instead of base-level correction.

## Quality Control and Validation

### Read Coverage Assessment

Before scaffolding, assess the read coverage of the draft assembly. Coverage should be relatively uniform across the genome, with no regions of zero coverage that would indicate assembly gaps. Use SAMtools to compute per-base coverage from the sorted BAM file:

```bash
samtools depth mapping/sample.sorted.bam > mapping/sample.depth
```

Plot the depth profile and inspect it for anomalies. Regions with coverage spikes may be repetitive elements, and regions with coverage drops may be assembly errors or genuine sequence absent from the sample.

### Consistency Between Replicates

If the project includes biological or technical replicates, scaffold each replicate independently and compare the results. Consistent scaffold layouts across replicates increase confidence in the scaffolding. Discrepancies may indicate sample-specific structural variation or technical artifacts. The [Galaxy Training Network](https://training.galaxyproject.org/) provides guidance on designing reproducible analysis workflows that support such comparisons.

### Validation with Independent Data

When possible, validate the scaffold layout with independent data. Optical maps, genetic maps, or Hi-C contact maps provide long-range information that can confirm or refute the scaffold order. For bacterial genomes, PCR amplification across scaffold junctions can confirm that adjacent contigs are genuinely adjacent in the genome. These validation approaches are especially important when the scaffold will be used for downstream analysis such as comparative genomics or variant calling.

## Professional Escalation Criteria

### When to Seek Additional Expertise

Several situations warrant consultation with a bioinformatics specialist or a core facility:

- The draft assembly has an unusually high number of contigs relative to the expected genome size, suggesting a fundamental assembly problem that scaffolding will not resolve.
- The mapping rate is below 50 percent, indicating that the reference is too divergent for reliable scaffolding.
- The scaffolded assembly contains many chimeric scaffolds that cannot be resolved by breaking at obvious junctions.
- The project requires chromosome-level scaffolds for a eukaryotic genome, and the reference-based approach produces only partial chromosome coverage.
- The downstream analysis depends on accurate structural variant detection, and the reference-based scaffold layout may mask or create structural variants.

### Documentation for Escalation

When escalating a problem, provide the specialist with the project log, the draft assembly, the reference accession, the sorted BAM file, and the RagTag output files. Include the specific error messages or quality metrics that prompted the escalation. A clear problem statement, such as "the scaffolded assembly has 15 scaffolds that each align to two reference chromosomes" or "the mapping rate is 62 percent with the closest available reference," helps the specialist diagnose the issue efficiently.

## Workflow Automation and Reproducibility

### Scripting the Pipeline

The commands in this protocol can be combined into a shell script for reproducible execution. Store the script in the project directory and record the script version in the project log. The [nf-core documentation](https://nf-co.re/docs) describes best practices for workflow design, including parameter validation, logging, and error handling, which are applicable to custom scripts as well as community pipelines.

### Containerization

For maximum reproducibility, run the pipeline inside a container such as Docker or Singularity. Containers package the exact tool versions and dependencies, eliminating version-related variability between runs. The [Bioconductor project](https://bioconductor.org/) and the [Galaxy Training Network](https://training.galaxyproject.org/) both provide containerized analysis environments that can be adapted for this pipeline.

### Version Control

Track changes to the pipeline script and the project log using a version control system such as Git. The [Carpentries lessons](https://carpentries.org/lessons) include practical instruction on using Git for version control in research projects. Version control allows the researcher to reproduce any previous analysis state and provides a record of how the analysis evolved.

## Decision Framework for Choosing Between Reference-Based and Alternative Scaffolding Strategies

Before committing to the reference-based scaffolding pipeline described in this article, evaluate whether this approach is appropriate for the specific project. The decision framework below provides a structured method for comparing reference-based scaffolding with alternative strategies, including paired-end scaffolding, long-read assembly, and hybrid approaches. This framework is designed to be applied before the first BWA command is run, because switching strategies mid-pipeline wastes computational time and complicates the project record.

### Step 1: Assess Reference Suitability

The first decision point is whether a suitable reference genome exists. A suitable reference meets three criteria: taxonomic proximity, assembly completeness, and sequence availability. Taxonomic proximity means the reference comes from the same species or a very close congener. For bacteria, a complete genome from the same species is ideal, while for eukaryotes, a chromosome-level assembly from the same species is strongly preferred. Assembly completeness is assessed by checking whether the reference has chromosome-level scaffolds instead of thousands of unplaced contigs. The NCBI [GenBank and RefSeq databases](https://www.ncbi.nlm.nih.gov/) allow filtering by assembly level, so select a reference marked as Complete Genome or Chromosome for the scaffolding project.

Sequence availability refers to whether the reference is publicly accessible and downloadable in FASTA format. If the closest available reference is from a different genus or has an N50 below the expected chromosome length, the reference-based approach will produce unreliable results. In that case, proceed to Step 4 and consider alternative strategies.

### Step 2: Estimate Sequence Divergence

The second decision point is the expected nucleotide divergence between the draft genome and the candidate reference. Divergence can be estimated by mapping a small random subset of reads to the reference and computing the average identity from the alignment. A quick approach is to map 100,000 read pairs with BWA-MEM and inspect the CIGAR strings and MD tags in the SAM output. Reads with high mismatch rates or soft clipping at the ends indicate divergence.

For bacterial genomes, a mapping rate above 90 percent with a median identity above 95 percent supports reference-based scaffolding. For eukaryotic genomes, a mapping rate above 70 percent with a median identity above 90 percent is workable. Below these thresholds, the read alignments become too noisy for reliable misjoin detection and contig ordering. The [EMBL-EBI training portal](https://www.ebi.ac.uk/training) provides practical exercises on interpreting alignment statistics that are directly applicable to this assessment.

### Step 3: Evaluate Repetitive Content

The third decision point is the repetitive content of the genome. Repetitive elements longer than the read length create ambiguity in both assembly and scaffolding. If the draft assembly has an unusually high number of contigs relative to the expected genome size, or if the assembly graph contained many unresolved branches, the genome likely has substantial repetitive content. A bacterial genome with more than 500 contigs from a 50x coverage Illumina run may indicate problematic repeats. For such genomes, reference-based scaffolding will order the unique regions correctly but will place repetitive contigs arbitrarily.

The [Contig-Layout-Authenticator (CLA)](https://pubmed.ncbi.nlm.nih.gov/27248146) study directly addressed this limitation by combining reference-based ordering with paired-end scaffolding. The authors noted that reference-based approaches can mask strain-specific information, while paired-end scaffolding alone struggles when repeats exceed read length. Their pipeline scaffolded reference-sorted contigs using paired reads, which improved assemblies with large repetitive contigs. If the draft genome has substantial repetitive content, consider whether a combined approach that uses both reference ordering and paired-end scaffolding would produce better results than reference-based scaffolding alone.

### Step 4: Compare Alternative Strategies

When the reference is unsuitable, divergent, or the genome is repeat-rich, evaluate the following alternatives:

**Paired-end scaffolding without a reference.** This approach uses mate-pair or paired-end read information to order contigs without a reference template. It avoids reference bias but struggles with repeats longer than the insert size. The CLA study demonstrated that this approach can be combined with reference ordering for improved results.

**Long-read assembly.** Long reads from Oxford Nanopore or Pacific Biosciences platforms can resolve repeats and produce near-complete assemblies without a reference. The [FastViromeExplorer-Novel](https://pubmed.ncbi.nlm.nih.gov/36607772) study showed that iterative assembly with long reads recovered complete viral genomes from metagenomic data, including novel viruses absent from databases. If long-read data are available or can be generated, this approach may eliminate the need for scaffolding entirely.

**Hybrid reference-based and de novo assembly.** The [FastViromeExplorer-Novel](https://pubmed.ncbi.nlm.nih.gov/36607772) pipeline mapped reads to viral reference genomes, assembled the mapped reads de novo, and extended contigs through iterative assembly. This hybrid approach recovered complete genomes that neither reference-based mapping nor de novo assembly achieved alone. A similar strategy can be applied to bacterial or eukaryotic projects by assembling unmapped reads separately and then combining the results with the reference-scaffolded assembly.

**Graph-based assembly without scaffolding.** The [COGRAM](https://doi.org/10.21203/rs.3.rs-9829976/v1) study explored a learning-guided approach that scores edges in a k-mer overlap graph and reconstructs contigs by deterministic path extraction, without paired-end or long-read scaffolding. While this approach is still a proof of concept, it represents an emerging alternative that avoids reference bias entirely.

### Step 5: Document the Decision

Record the decision framework results in the project log before running the pipeline. Include the reference accession, the estimated divergence, the repetitive content assessment, and the rationale for choosing reference-based scaffolding over alternatives. This documentation is essential for the methods section of a manuscript and for troubleshooting if the scaffolding results are unsatisfactory. The [nf-core documentation](https://nf-co.re/docs) emphasizes the importance of documenting workflow decisions and parameter choices for reproducibility, a principle that applies equally to manual pipelines.

### Decision Matrix for Common Scenarios

| Scenario | Recommended Strategy | Rationale |
| --- | --- | --- |
| Same-species reference available, mapping rate above 90 percent | Reference-based scaffolding with RagTag | Closest reference provides reliable ordering and correction |
| Cross-genus reference only, mapping rate below 70 percent | De novo scaffolding with paired-end or long-read data | Divergent reference produces unreliable ordering |
| Bacterial genome with many repetitive contigs | Reference-based scaffolding combined with paired-end scaffolding | CLA study showed combined approach improves repeat-rich assemblies |
| Viral metagenome with fragmented contigs | Hybrid reference-based mapping and iterative de novo assembly | FastViromeExplorer-Novel recovered complete genomes with this approach |
| No suitable reference, long-read data available | Long-read assembly without scaffolding | Long reads resolve repeats and produce near-complete assemblies |
| Eukaryotic genome with chromosome-level reference | Reference-based scaffolding with validation | Chromosome-level reference enables chromosome-scale scaffolds |

### When to Reconsider the Reference-Based Approach Mid-Pipeline

The decision framework should be revisited if the mapping rate from Step 2 of the pipeline is substantially lower than expected. If the mapping rate falls below 70 percent for a bacterial genome or below 50 percent for a eukaryotic genome, stop the pipeline and reassess the reference choice. Continuing with a divergent reference will produce scaffolds that are worse than the original draft because the ordering will reflect the reference structure instead of the sample structure.

Similarly, if the correction step reports an unexpectedly high number of misjoins, such as more than 5 percent of contigs broken, the draft assembly may have fundamental problems that scaffolding will not resolve. In that case, consider whether the draft assembly should be improved with additional sequencing or a different assembler before attempting scaffolding.

### Practical Implementation of the Decision Framework

Create a simple checklist in the project log with the following items:

1. Reference accession and assembly level recorded
2. Mapping rate from a 100,000-read subset computed
3. Median identity from the subset alignment recorded
4. Draft assembly contig count compared with expected genome size
5. Repetitive content assessed from assembly graph or contig count
6. Alternative strategies considered and documented
7. Decision recorded with rationale

This checklist takes approximately 30 minutes to complete and prevents wasted computational time on an unsuitable reference. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on assembly assessment and reference selection that complement this framework with hands-on exercises.

### Relationship to the Main Pipeline

The decision framework is a pre-pipeline step that determines whether the BWA, SAMtools, and RagTag commands described in the main protocol should be executed. It does not replace any step in the pipeline but rather ensures that the pipeline is applied to appropriate data. Projects that pass the decision framework can proceed directly to Step 1 of the main protocol. Projects that fail the framework should pursue one of the alternative strategies listed in the decision matrix.

The framework also informs the interpretation of scaffolding results. If the reference is closely related but the genome is repeat-rich, expect some contigs to be placed arbitrarily and validate those regions with independent evidence. If the reference is moderately divergent, expect more gaps and smaller scaffolds, and consider whether the improvement over the draft justifies the reference bias introduced by scaffolding.

## Frequently Asked Questions

### What is the difference between reference-based scaffolding and reference-guided assembly?

Reference-based scaffolding orders and orients existing draft contigs using a reference genome as a template. The draft assembly is produced independently by de novo assembly, and the reference is used only to arrange the contigs. Reference-guided assembly, by contrast, uses the reference throughout the assembly process, often by mapping reads to the reference and generating consensus sequences directly. Scaffolding preserves the sequence content of the draft assembly, while reference-guided assembly can introduce reference alleles into the consensus.

### Can I scaffold with a reference from a different species?

Scaffolding with a cross-species reference is possible but produces less reliable results. The reads must map to the reference with sufficient specificity for the scaffolding algorithm to infer contig order. As sequence divergence increases, mapping becomes less specific, and the risk of incorrect scaffolding increases. For cross-species scaffolding, expect a lower mapping rate and validate the scaffold layout with independent evidence. A reference from the same genus is usually workable for bacteria, while a reference from the same species is strongly preferred for eukaryotes.

### How many reads do I need for reliable scaffolding?

The read depth required for scaffolding depends on the genome size and complexity. For bacterial genomes, 50x to 100x coverage of paired-end reads is typically sufficient. For eukaryotic genomes, 30x to 50x coverage may be adequate for scaffolding, though higher coverage improves the detection of misjoins and the accuracy of gap size estimates. The key requirement is that reads cover the contig boundaries, so coverage should be uniform across the assembly instead of concentrated in a few regions.

### What does the AGP file contain and why is it important?

The AGP file is a tab-delimited text file that describes the composition of each scaffold. Each line specifies a component of a scaffold, including the scaffold name, the component's position within the scaffold, the component type, and for contig components, the contig name, orientation, and coordinates in the original assembly. The AGP file is the definitive record of the scaffolding layout and is required for submitting scaffolded assemblies to public databases such as NCBI. It also allows researchers to trace any position in a scaffold back to the original contig.

### How do I know if my scaffolded assembly is correct?

No single metric confirms that a scaffolded assembly is correct. The strongest evidence comes from multiple independent sources: high read mapping rates, consistent alignment of scaffolds to the reference, agreement between replicates, and validation with independent long-range data such as optical maps or Hi-C. Summary statistics such as N50 and scaffold count indicate improvement over the draft but do not prove correctness. Manual inspection of the scaffold layout and targeted validation of scaffold junctions are the most reliable quality checks.

### What should I do if RagTag reports no misjoins but my assembly still looks wrong?

If the correction step reports no misjoins but the scaffolded assembly has obvious errors, investigate the read coverage and alignment patterns at the suspected error locations. Low coverage regions may prevent misjoin detection. Check whether the reference has structural differences from the sample, such as inversions or translocations, that would cause the scaffolding algorithm to produce an incorrect layout. Consider running the scaffolding with different parameters or using a different reference genome.

### Can I use this pipeline for eukaryotic genomes?

The pipeline works for eukaryotic genomes, but the computational requirements are higher and the results are less complete. Eukaryotic genomes are larger, so read mapping and sorting take longer and require more memory. Repetitive content is typically higher, which increases the ambiguity in scaffolding. The reference genome must be chromosome-level for the scaffold to produce chromosome-scale output. For eukaryotic projects, expect the scaffolded assembly to contain large scaffolds that cover portions of chromosomes instead of complete chromosomes.

### How do I handle contigs that do not align to the reference?

Contigs that do not align to the reference are left unplaced by the scaffolding algorithm. These contigs may represent sample-specific sequences, such as plasmids, prophages, or genomic islands, that are absent from the reference. They may also be assembly artifacts, such as chimeric contigs or contamination. Examine the coverage and sequence composition of unplaced contigs to determine whether they are biologically meaningful. Unplaced contigs should be reported separately from the scaffolded chromosomes in the final assembly.

## Related Bioinformatics Guides

- [Human Reference Genomes: Navigating and Downloading hg38 and hg19 FASTA Assemblies](/knowledge/bioinformatics/human-reference-genomes-hg38-hg19-assemblies)
- [Binning in Metagenomics: From Contigs to Genomes](/knowledge/bioinformatics/binning-in-metagenomics-from-contigs-to-genomes)
- [Digital Pathology Guidelines: A Reference for Implementation](/knowledge/bioinformatics/digital-pathology-guidelines-a-reference-for-implementation)
- [Metagenomics Pipeline: From Raw Reads to Taxonomic and Functional Profiles](/knowledge/bioinformatics/metagenomics-pipeline-from-raw-reads-to-taxonomic-and-functional-profiles)
- [Oxford Nanopore Sequencing: From Sample to Base Calls](/knowledge/bioinformatics/oxford-nanopore-sequencing-from-sample-to-base-calls)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Transcriptomic profiling uncovers mis-splicing and gene fusions in amyotrophic lateral sclerosis.](https://pubmed.ncbi.nlm.nih.gov/41674618). medRxiv : the preprint server for health sciences, 2026.
- [FastViromeExplorer-Novel: Recovering Draft Genomes of Novel Viruses and Phages in Metagenomic Data.](https://pubmed.ncbi.nlm.nih.gov/36607772). Journal of computational biology : a journal of computational molecular cell biology, 2023.
- [Identification of highly variable sequence fragments in unmapped reads for rapid bacterial genotyping.](https://pubmed.ncbi.nlm.nih.gov/36581824). BMC genomics, 2022.
- [Contig-Layout-Authenticator (CLA): A Combinatorial Approach to Ordering and Scaffolding of Bacterial Contigs for Comparative Genomics and Molecular Epidemiology.](https://pubmed.ncbi.nlm.nih.gov/27248146). PloS one, 2016.
- [COGRAM: Genome Assembly through Graph Neural Networks](https://doi.org/10.21203/rs.3.rs-9829976/v1). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.