Repeat Masking in Genome Assembly: How to Identify and Soft-Mask Repetitive Elements Before Assembly
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Repeat masking is critical for accurate genome assembly, particularly in repeat-rich eukaryotic genomes, by identifying and marking repetitive DNA elements to prevent assemblers from creating chimeric contigs or collapsing distinct genomic regions.
- Repetitive elements, primarily transposable elements and satellite DNAs, disrupt assembly logic by creating ambiguous read placements, leading to significant assembly errors that can be quantified by metrics like BUSCO completeness and contiguity (N50).
- Soft masking, which converts repetitive sequences to lowercase while preserving nucleotide identity, is generally preferred over hard masking (replacing with 'N') as it retains sequence information for downstream analysis and allows some assemblers to resolve local structure.
- A combined approach of de novo repeat identification (e.g., using RepeatModeler) and homology-based masking (e.g., using RepeatMasker against known databases) is standard practice to capture both species-specific and conserved repetitive elements.
- Iterative refinement of repeat masking stringency, guided by metrics such as BUSCO completeness, contiguity, assembly size, and expected repeat copy number, is essential, with a practical framework suggesting a masking matrix and a three-round adjustment process.
- Thorough documentation of repeat masking parameters, tool versions, repeat library contents, and masking percentages is paramount for reproducibility, alongside validation against independent data like genetic or optical maps to confirm assembly accuracy.
Repetitive DNA constitutes the major proportion of nuclear DNA in most eukaryotic genomes, with sequence motifs repeated hundreds or thousands of times across the genome. When assembling a genome from sequencing reads, these repeats create ambiguity in read placement that can collapse distinct genomic regions into a single erroneous contig or produce artificial joins between unrelated loci. Repeat masking is the process of identifying these repetitive elements and marking them so that assembly algorithms handle them appropriately. This article explains how to identify repetitive elements, soft-mask them before assembly, and interpret the results to improve assembly quality in repeat-rich genomes.
The Problem That Repeat Masking Solves
Genome assemblers work by finding overlaps between sequencing reads and extending contigs where the overlap is unambiguous. Repetitive elements break this logic. If a read originates from a transposable element family present at 500 locations in the genome, the assembler cannot determine which of those 500 locations the read came from. The assembler may then merge unrelated regions that share the same repeat, creating a chimeric contig, or it may collapse multiple repeat copies into a single copy, producing an assembly that is shorter than the true genome.
The consequences of unmasked repeats are measurable. In the chromosome-level assembly of the relapsing fever tick Ornithodoros turicata, 33% of the genome was identified as repeats and masked before the final assembly was produced. The resulting reference genome achieved a QV score of 43.58 and a BUSCO completeness score of 97.8%. The authors explicitly attributed the assembly quality to the repeat identification and masking step in their workflow. Similarly, the wine grape Carménère genome project tested approximately one thousand combinations of assembly parameters, sequencing coverage, error correction, and repeat masking methods to optimize contiguity and completeness. The choice of masking strategy was one of the four variables they systematically varied.
For researchers working with repeat-rich genomes, the decision to mask repeats before assembly is a required quality control step that directly determines whether the final assembly represents the true genome structure or a distorted version of it.
What Counts as a Repetitive Element
Repetitive DNA falls into two broad organizational categories. Satellite DNAs exist as long arrays of similar motifs at a small number of chromosomal sites, often located in centromeres or subtelomeric regions. Transposable elements, including DNA transposons and retrotransposons, are dispersed across the genome instead of clustered at specific loci. Both categories create assembly problems, but they require different masking considerations.
Transposable elements are the dominant repeat class in most eukaryotic genomes. In the Prinsepia uniflora genome, transposable elements occupied 777.28 Mb of the 1272.71 Mb assembly, constituting 61.07% of the entire genome. The total repetitive sequence content reached 875.99 Mb. In contrast, the gyrfalcon genome contained only 5,716 transposable elements that masked 7.61% of the assembly, with Chicken-repeat 1 (CR1) long interspersed nuclear elements being the most abundant family. These two examples bracket the range of repeat content that researchers encounter. A masking pipeline that works for a 7.61% repeat genome may need substantial adjustment for a 61% repeat genome.
The distinction between tandem repeats and dispersed repeats matters for assembly strategy. Tandem repeats, such as satellite DNA, produce reads that are nearly identical to each other and to the genomic region they came from. Dispersed repeats, such as transposable elements, produce reads that match multiple distant genomic locations. Both types confuse assemblers, but the failure modes differ. Tandem repeats often cause assembly to stall or produce collapsed arrays. Dispersed repeats often cause chimeric joins between unrelated regions.
How Repeat Masking Works
Repeat masking identifies repetitive sequences and replaces them with either hard masks (N characters) or soft masks (lowercase characters). Hard masking replaces every base in a repeat with N, which tells the assembler that the base is unknown. Soft masking converts the base to lowercase while preserving the original nucleotide identity. Most modern assembly pipelines use soft masking because it preserves sequence information that may be useful for downstream analysis while still signaling to the assembler that the region is repetitive.
The masking process has two main approaches. Homology-based masking compares the genome sequence against known repeat databases using tools such as RepeatMasker. This approach finds repeats that have been previously characterized in other species. De novo repeat identification uses tools such as RepeatModeler to build a species-specific repeat library from the genome itself. This approach finds repeats that are unique to the species being assembled or that have diverged too far from known families to be detected by homology alone.
The gyrfalcon genome project used a combined approach. They performed a de novo search for transposable elements that identified 5,716 TEs, then combined these with publicly available TE collections to mask 7.61% of the assembly. The combination of de novo and homology-based approaches is standard practice because each method catches repeats that the other misses.
At a Glance: Repeat Masking Decision Table
| Genome Repeat Content | Recommended Approach | Expected Impact on Assembly | Key Risk |
|---|---|---|---|
| Low (under 10%) | Homology-based masking with RepeatMasker against standard databases | Modest improvement in contiguity, prevents most chimeric joins | Overmasking if repeat databases contain sequences similar to functional genes |
| Moderate (10 to 40%) | Combined de novo RepeatModeler and homology-based RepeatMasker | Substantial improvement in N50 and reduction in collapsed regions | De novo library may miss young or highly diverged repeat families |
| High (over 40%) | Iterative de novo repeat modeling with manual curation of the repeat library | Critical for producing chromosome-level assemblies | Under-masking leaves enough repeats to cause misassembly, over-masking removes functional sequence |
| Satellite-rich (centromeric or telomeric) | Specialized tandem repeat finders in addition to standard masking | Prevents collapse of long tandem arrays | Tandem repeats may be impossible to assemble even with masking, may require long-read data |
Building a Species-Specific Repeat Library with RepeatModeler
RepeatModeler constructs a repeat library from the genome sequence itself. The tool identifies repetitive elements by looking for sequences that occur multiple times in the genome and then builds consensus models for each repeat family. The output is a FASTA file containing consensus sequences for the repeats found in your genome.
The input to RepeatModeler is the genome assembly itself. This creates a circular dependency problem. You need a genome assembly to build a repeat library, but you need a repeat library to produce a good genome assembly. The standard solution is to produce a preliminary assembly without masking, build a repeat library from that assembly, then re-assemble with masking. The preliminary assembly does not need to be chromosome-level. A draft assembly with reasonable contiguity is sufficient for repeat identification.
For the Prinsepia uniflora project, the researchers used the repeat-masked assembly for gene prediction and identified 49,261 protein-coding genes. Of these, 45,256 (91.87%) received functional annotations, 5,127 (10.41%) showed tandem duplication, and 2,373 (4.82%) were classified as transcription factor genes. The repeat masking step was essential for accurate gene prediction because unmasked repeats would have been misidentified as protein-coding genes or would have disrupted gene models.
When building a repeat library, consider the following parameters. The genome size determines how much compute time RepeatModeler requires. Larger genomes need more time and memory. The diversity of repeat families in your species determines whether the default settings are adequate. Species with complex repeat landscapes may need additional rounds of repeat modeling to capture all families. The divergence of your species from well-characterized relatives determines how much weight to give de novo versus homology-based approaches.
Running RepeatMasker with the Custom Library
RepeatMasker uses a repeat library to identify and mask repetitive sequences in the genome. The tool can run in several modes. The default mode produces a hard mask. The soft mask mode produces lowercase output. The cross-species mode uses a library from a related species. The custom library mode uses the library you built with RepeatModeler.
The standard workflow is to run RepeatMasker with the custom RepeatModeler library combined with known repeat databases. The gyrfalcon project combined de novo identified TEs with publicly available TE collections. This combined approach ensures that both species-specific and conserved repeats are masked.
The output of RepeatMasker includes a masked genome sequence, a table of repeat annotations, and a summary of repeat content by class and family. The summary table is your first quality check. It tells you the total percentage of the genome masked and the breakdown by repeat type. Compare this to published repeat content for related species. If your repeat content is dramatically different from expectations, investigate whether the repeat library is incomplete or the masking parameters are incorrect.
For the Ornithodoros turicata genome, 33% of the genome was identified as repeats and masked. This value is typical for arthropod genomes. If your arthropod genome shows 5% repeat content, your repeat library is likely missing major families. If it shows 80%, you may be overmasking and removing functional sequence.
Soft Masking Versus Hard Masking
Soft masking preserves the nucleotide sequence while marking it as repetitive. Hard masking replaces the sequence with N characters. The choice between them affects downstream analysis.
Soft masking is preferred for assembly because it preserves information. If a repeat contains a functional element, such as a gene fragment or a regulatory sequence, soft masking allows that information to remain available. Hard masking destroys the sequence information, which can be problematic if you later need to extract the sequence of a repeat family or examine variation within repeats.
Soft masking also allows assemblers to use the sequence information when it is unambiguous. Some assembly algorithms can use soft-masked sequence to resolve local structure while avoiding the use of repeats for long-range joins. Hard masking forces the assembler to treat the region as unknown, which can create gaps in the assembly.
The tradeoff is that soft masking requires the assembler to respect the mask. Not all assemblers handle soft masks correctly. Check the documentation for your assembler to confirm that it recognizes lowercase sequence as masked. If it does not, you may need to use hard masking or configure the assembler to treat lowercase as masked.
Integrating Masking into the Assembly Workflow
The placement of masking in the assembly workflow depends on the assembly strategy. For long-read assemblies using PacBio or Oxford Nanopore data, masking is typically applied before the assembly step. The reads are mapped to the repeat library, and repetitive reads or read segments are marked. The assembler then avoids using these segments for contig extension.
For hybrid assembly strategies that combine long reads with short reads, masking may be applied at multiple points. The long-read assembly may be produced with masking, then the short reads are used for polishing. The polishing step should also respect the mask to avoid introducing errors in repetitive regions.
The Carménère grape genome project tested about a thousand combinations of assembly parameters, sequencing coverage, error correction, and repeat masking methods. This systematic approach highlights an important point. The optimal masking strategy depends on the specific genome and the assembly algorithm. What works for one species may not work for another. Plan to test multiple masking approaches and compare the resulting assemblies.
Quality Checks After Masking
After masking and assembly, evaluate the assembly quality using multiple metrics. BUSCO completeness measures the fraction of conserved single-copy orthologs present in the assembly. The gyrfalcon assembly achieved 96.7% complete BUSCO genes out of 8,338 BUSCO groups in the aves_odb10 library. The Ornithodoros turicata assembly achieved 97.8% BUSCO completeness. These values indicate that masking did not remove functional sequence.
Contiguity metrics such as N50 measure the length of the assembled contigs and scaffolds. The Prinsepia uniflora assembly achieved contig and super-scaffold N50 values of 2.77 and 79.32 Mb, respectively. These values indicate that the assembly is highly contiguous, which is the expected outcome of effective repeat masking.
Check for collapsed repeats by comparing the number of repeat copies in the assembly to the expected copy number. If a repeat family is known to have 100 copies in the genome but the assembly contains only 20, the assembly has collapsed. This is a common failure mode in repeat-rich genomes. The Prinsepia uniflora genome contained 875.99 Mb of repetitive sequences, and the assembly successfully represented this content. If your assembly shows fewer repeat copies than expected, consider whether the masking was too aggressive or the assembly algorithm collapsed the repeats.
Common Failure Patterns in Repeat Masking
Under-masking occurs when the repeat library misses major repeat families. This leads to chimeric assemblies where unrelated genomic regions are joined through shared repeats. The assembly may have inflated contiguity metrics that do not reflect true genome structure. The fix is to improve the repeat library by running RepeatModeler with different parameters or adding known repeats from related species.
Over-masking occurs when the repeat library contains sequences similar to functional genes. This leads to the removal of genuine coding sequence from the assembly. The assembly may have inflated repeat content and reduced BUSCO completeness. The fix is to curate the repeat library to remove sequences that match known genes.
Library contamination occurs when the repeat library contains sequences that are not repetitive. This can happen if the input genome contains contamination from other species or if RepeatModeler misidentifies low-complexity sequence as a repeat family. The fix is to screen the repeat library against known gene databases and remove matches.
Parameter mismatch occurs when the masking parameters do not match the assembly algorithm requirements. Some assemblers expect hard masks, others expect soft masks. Some expect the mask to be applied to the reads, others to the assembly. Check the documentation for your specific tools.
Interpreting the Repeat Landscape
The repeat landscape provides information about the evolutionary history of the genome. The age distribution of transposable elements, estimated from their divergence from consensus sequences, indicates when repeat families were active. Young repeats have low divergence from their consensus. Old repeats have high divergence.
This information is useful for assembly decisions. If a genome has many young, active repeat families, the repeats are likely to be highly similar to each other and difficult to distinguish. This increases the risk of misassembly. If the repeats are old and diverged, they may be easier to distinguish because each copy has accumulated unique mutations.
The Prinsepia uniflora genome showed evidence of recent whole-genome duplication following its separation from Prunus salicina. This evolutionary event affects the repeat landscape because duplicated regions may contain paralogous sequences that resemble repeats. The repeat masking approach must distinguish between true repeats and duplicated genes.
Records and Documentation for Reproducibility
Document every step of the repeat masking process. Record the version of RepeatModeler and RepeatMasker used. Record the parameters for each run. Record the contents of the repeat library, including the number of families and the source of each family. Record the masking percentage and the breakdown by repeat class.
The nf-core documentation emphasizes the importance of reproducible workflow standards. Following community pipeline standards ensures that your masking process can be reproduced by other researchers. The Galaxy Training Network provides accessible workflow training that can help you structure your masking pipeline for reproducibility.
Store the repeat library as a versioned file. The library is a key output of the masking process and should be preserved alongside the assembly. If you update the library, create a new version instead of overwriting the old one. This allows you to trace the effect of library changes on assembly quality.
Tools and Training Resources
The NCBI provides access to sequence databases and analysis services that support repeat identification. The RepeatMasker output can be compared against NCBI databases to identify repeat families and check for contamination.
The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education. These resources can help you understand the underlying algorithms and interpret the output of masking tools.
Bioconductor provides official package documentation for reproducible genomic analysis. Several Bioconductor packages support repeat analysis and genome assembly quality assessment. The package documentation includes installation instructions and workflow examples.
The Galaxy Training Network offers accessible workflow training for genome assembly and repeat analysis. These tutorials provide step-by-step instructions that you can adapt to your specific genome.
The Carpentries Lessons provide foundational computing training that is useful for managing the computational requirements of repeat masking. Shell scripting, data management, and version control skills are essential for running masking pipelines efficiently.
Limitations of Repeat Masking
Repeat masking cannot solve all assembly problems. Some repeats are so long or so similar that no masking strategy can prevent misassembly. Long tandem arrays of satellite DNA, such as those found in centromeres, may be impossible to assemble even with masking. These regions may remain as gaps in the final assembly.
Masking removes sequence from consideration, which means that any functional elements within repeats are also removed. This is a particular concern for genes that have transposable element insertions in their introns or regulatory regions. The masking may remove these elements, affecting gene annotation.
The repeat library is a model of the repeats in the genome, not the repeats themselves. The library may miss rare or highly diverged repeats. It may also contain artifacts that do not represent true repeats. The quality of the library directly determines the quality of the masking.
Repeat masking is computationally intensive. RepeatModeler and RepeatMasker require substantial memory and processing time for large genomes. The Prinsepia uniflora genome is 1,272.71 Mb, and the repeat analysis required significant computational resources. Plan your compute budget accordingly.
Professional Escalation Criteria
Seek expert assistance when the repeat content of your genome is unexpectedly high or low compared to related species. This may indicate a problem with the repeat library, the assembly, or the sequencing data.
Consult a bioinformatics specialist when the assembly shows signs of collapse or chimerism that persist after multiple masking attempts. These problems may require specialized assembly strategies or additional sequencing data.
Request help from the tool developers or community when RepeatModeler or RepeatMasker produces errors or unexpected output. The documentation for these tools includes troubleshooting guidance, and community forums can provide practical advice.
Escalate to a genome assembly expert when the genome contains complex repeat structures such as large satellite arrays or recently amplified transposable element families. These genomes may require specialized assembly approaches beyond standard masking.
Safety and Data Management Context
Repeat masking involves working with large genomic datasets that require careful data management. Store raw sequencing data, intermediate files, and final assemblies in organized directory structures. Use version control for scripts and configuration files. Document file locations and parameters in a lab notebook or electronic record.
The computational requirements of repeat masking may exceed the capacity of a standard desktop computer. Use a computing cluster or cloud resource for large genomes. The nf-core documentation provides guidance on configuring workflows for different computing environments.
Back up the repeat library and the masked assembly. These files represent significant computational investment and may be difficult to reproduce. Store them in multiple locations and verify the integrity of the backups.
A Practical Decision Framework for Masking Stringency and Assembly Validation
Choosing the right masking stringency is not a one-time decision. It is a series of decisions that you revisit as you evaluate assembly quality. The framework below gives you a concrete sequence of checks and adjustments that connect masking choices to measurable assembly outcomes. It is designed for researchers who have already built a repeat library and run an initial masking pass, and who now need to decide whether their masking was too aggressive, too weak, or correctly calibrated.
Step 1: Establish Your Baseline Repeat Content Expectation
Before you adjust any parameters, you need a target range for repeat content. This target comes from three sources. First, published repeat content for the same species or the closest sequenced relative. Second, the repeat content reported in the RepeatModeler output, which estimates the repeat landscape from your own sequence data. Third, the repeat content you would predict from genome size, because genome size correlates with transposable element load in most eukaryotic lineages.
The Prinsepia uniflora genome provides a useful upper reference point. Its assembly reached 1272.71 Mb across 16 pseudochromosomes, and transposable elements occupied 777.28 Mb, or 61.07% of the genome. The total repetitive sequence content reached 875.99 Mb. If you are assembling a plant genome of similar size and your masking pipeline reports 10% repeat content, your repeat library is missing major families. At the other extreme, the gyrfalcon genome contained only 7.61% transposable elements, with 5,716 TEs identified by de novo search and combined with public collections. If you are assembling a bird genome and your masking reports 40% repeat content, you are likely overmasking and removing functional sequence.
Write your expected range into your laboratory notebook or project record before you run the masking pipeline. This prevents you from rationalizing an unexpected result after the fact. The range should be broad enough to accommodate genuine biological variation but narrow enough to flag real problems. A reasonable starting range is the published value for the closest relative plus or minus 10 percentage points.
Step 2: Run a Masking Matrix Instead of a Single Masking Pass
Most researchers run RepeatMasker once with default parameters and move on. The Carménère grape genome project demonstrates why this is insufficient. The authors tested approximately one thousand combinations of assembly parameters, sequencing coverage, error correction, and repeat masking methods to optimize contiguity and completeness. You do not need to test one thousand combinations, but you should test a small matrix of masking conditions before committing to a final assembly.
A practical masking matrix includes three conditions. Condition A is homology-based masking only, using RepeatMasker against standard databases. Condition B is combined masking, using your RepeatModeler library plus standard databases. Condition C is aggressive masking, using your RepeatModeler library with a lower cutoff score or with additional rounds of repeat modeling. Run all three conditions on the same input assembly and assemble each masked version with the same assembler and parameters.
This matrix approach gives you three assemblies to compare. The comparison reveals how sensitive your assembly is to masking choices. If all three assemblies produce nearly identical contiguity and BUSCO scores, your genome is not repeat-limited and masking stringency is not the critical variable. If the three assemblies differ substantially, masking stringency is a major driver of assembly quality and you need to choose carefully.
Step 3: Score Each Assembly with a Decision Matrix
For each of the three assemblies from your masking matrix, record four metrics. The first metric is BUSCO completeness, which measures the fraction of conserved single-copy orthologs present. The gyrfalcon assembly achieved 96.7% complete BUSCO genes out of 8,338 BUSCO groups in the aves_odb10 library. The Ornithodoros turicata assembly achieved 97.8% BUSCO completeness. These values represent the quality range you should target.
The second metric is contiguity, measured by contig N50 and scaffold N50. The Prinsepia uniflora assembly achieved contig and super-scaffold N50 values of 2.77 and 79.32 Mb respectively. Higher N50 values generally indicate better assembly, but only when BUSCO completeness remains high. An assembly with inflated N50 and collapsed BUSCO scores is likely chimeric.
The third metric is the number of complete repeat copies in the assembly compared to the expected copy number. This is the most direct check for repeat collapse. If a transposable element family is known from your RepeatModeler analysis to have 500 copies in the genome, but the assembly contains only 150 copies, the assembly has collapsed that family. The Prinsepia uniflora genome contained 875.99 Mb of repetitive sequences, and the assembly successfully represented this content. If your assembly shows fewer repeat copies than expected, your masking was either too aggressive or the assembler collapsed the repeats despite masking.
The fourth metric is the assembly size compared to the expected genome size. A substantially smaller assembly suggests collapse. A substantially larger assembly suggests contamination or assembly of haplotypic variation as separate contigs.
Record these four metrics in a table for each masking condition. The table becomes your decision record. It allows you to justify your final masking choice with data instead of intuition.
Step 4: Apply the Decision Rules
The decision rules below connect your metrics to a masking adjustment. These rules are heuristics, not absolute thresholds, because the optimal values depend on your genome and your research question.
If BUSCO completeness is below 90% and repeat content is higher than expected, you are likely overmasking. The repeat library contains sequences similar to functional genes, and masking has removed genuine coding sequence. The fix is to curate the repeat library. Screen the library against a protein database and remove any repeat models that match known genes. Then rerun the masking matrix.
If BUSCO completeness is high but contiguity is poor, you are likely under-masking. The assembler is stalling at repetitive regions because too many repeats remain unmasked. The fix is to increase masking stringency. Lower the RepeatMasker cutoff score, add more repeat families to the library, or run additional rounds of RepeatModeler to capture young or diverged families.
If the assembly size is substantially smaller than the expected genome size, you have repeat collapse. The assembler has merged multiple repeat copies into a single copy. This can happen even with masking if the repeats are young and highly similar. The fix is to increase masking stringency and consider whether your assembler handles soft masks correctly. Some assemblers ignore lowercase masking and use the sequence anyway.
If the assembly size is substantially larger than expected, check for contamination. The repeat library may contain sequences from other species, or the input assembly may contain contamination. Screen the assembly against NCBI databases to identify contaminating sequences. The NCBI provides search systems that can identify the taxonomic origin of contigs.
Step 5: Iterate with a Fixed Number of Rounds
Repeat masking improvement is iterative, but it needs a stopping rule. Without a stopping rule, you can spend weeks adjusting parameters with diminishing returns. A practical approach is to allow three rounds of masking adjustment. Round one is the initial masking matrix. Round two applies the decision rules from round one. Round three applies the decision rules from round two. After round three, accept the best assembly or escalate to a specialist.
Each round should change only one variable at a time. If you change the repeat library and the masking cutoff simultaneously, you cannot determine which change caused the improvement. The Carménère project tested about a thousand combinations, but they did so systematically, varying one parameter at a time to understand each variable's contribution.
Document each round in your project record. Record the masking condition, the four metrics, and the decision you made based on those metrics. This documentation serves two purposes. It allows you to reproduce your own work, and it allows reviewers to understand why you chose the final masking parameters.
Step 6: Validate the Final Assembly Beyond Standard Metrics
Standard metrics such as BUSCO and N50 are necessary but not sufficient. They do not detect all assembly errors. Additional validation steps are needed to confirm that the assembly represents the true genome structure.
Check the repeat landscape of the final assembly. The age distribution of transposable elements, estimated from their divergence from consensus sequences, should match the expected evolutionary history of the species. A genome with many young, active repeat families should show a peak of low-divergence repeats. A genome with old, inactive families should show a peak of high-divergence repeats. If the repeat landscape looks unusual, investigate whether masking removed or collapsed specific families.
Check for chimeric joins by examining the flanking sequences of contigs. If a contig contains a repeat at one end and a unique sequence at the other, the repeat may have caused an artificial join. The Galaxy Training Network provides accessible workflow training that includes assembly validation steps you can adapt to your genome.
Check the assembly against independent data. If you have genetic maps, optical maps, or Hi-C data, use them to validate the assembly structure. The Ornithodoros turicata genome used PacBio sequencing with an Illumina Hi-C library to produce a chromosome-level assembly. The Hi-C data provided independent validation of the chromosome-scale scaffolding. If you have similar data, use it to confirm that the assembly structure is correct.
Step 7: Record the Final Masking Decision
The final step is to record the masking decision in a way that another researcher can reproduce. Record the exact version of RepeatModeler and RepeatMasker used. Record all parameters, including cutoff scores and library contents. Record the masking percentage and the breakdown by repeat class. Record the four metrics from the final assembly.
The nf-core documentation emphasizes the importance of reproducible workflow standards. Following community pipeline standards ensures that your masking process can be reproduced by other researchers. The Bioconductor project provides official package documentation for reproducible genomic analysis, and several packages support repeat analysis and assembly quality assessment.
Store the repeat library as a versioned file. The library is a key output of the masking process and should be preserved alongside the assembly. If you update the library, create a new version instead of overwriting the old one. This allows you to trace the effect of library changes on assembly quality.
Common Failure Patterns in the Decision Framework
The decision framework fails in predictable ways. The most common failure is skipping the masking matrix and running a single masking pass. This gives you no information about how sensitive your assembly is to masking choices. The second most common failure is changing multiple variables between rounds, which makes it impossible to attribute improvements to specific changes. The third most common failure is ignoring the repeat copy number check and relying only on BUSCO and N50. These metrics can look good even when specific repeat families have collapsed.
The fourth failure pattern is over-iterating. Researchers who do not set a stopping rule can spend months adjusting masking parameters. The three-round rule prevents this. If three rounds of systematic adjustment do not produce an acceptable assembly, the problem is likely not masking stringency. It may be the assembly algorithm, the sequencing data, or the repeat structure of the genome itself.
The fifth failure pattern is treating the repeat library as fixed. The library should be updated as you learn more about the genome. If you identify a new repeat family during assembly validation, add it to the library and rerun the masking matrix. The gyrfalcon project combined de novo identified TEs with publicly available TE collections, recognizing that neither source alone was sufficient.
When to Escalate Beyond the Decision Framework
The decision framework has limits. Some genomes contain repeat structures that no masking strategy can fully resolve. Long tandem arrays of satellite DNA, such as those found in centromeres, may be impossible to assemble even with aggressive masking. These regions may remain as gaps in the final assembly. The NCBI provides access to sequence databases that can help you identify whether your unresolved regions correspond to known satellite arrays.
Escalate to a genome assembly specialist when the assembly shows signs of collapse or chimerism that persist after three rounds of masking adjustment. These problems may require specialized assembly strategies, additional sequencing data, or a different assembler. The Carménère project tested about a thousand combinations of parameters, which is beyond the scope of most projects. A specialist can help you identify which parameters matter most for your specific genome.
Escalate when the repeat content of your genome is dramatically different from related species and you cannot explain the difference. This may indicate a problem with the sequencing data, the assembly, or the repeat library. The Prinsepia uniflora genome at 61.07% transposable elements and the gyrfalcon genome at 7.61% transposable elements bracket the range of repeat content in eukaryotic genomes. If your genome falls outside this range for its taxonomic group, investigate before proceeding.
Integrating the Decision Framework with Existing Workflows
The decision framework is compatible with existing assembly workflows. It adds a systematic comparison step between the repeat library construction and the final assembly. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources that can help you understand the underlying algorithms and interpret the output of masking tools. The Carpentries Lessons provide foundational computing training that is useful for managing the computational requirements of running a masking matrix.
The framework also integrates with the records and documentation practices described earlier in this article. The masking matrix produces a natural set of records: the three masking conditions, the four metrics for each condition, and the decision rules applied. These records become part of the assembly documentation and support reproducibility.
The computational cost of the masking matrix is higher than a single masking pass. Running three masking conditions and three assemblies requires approximately three times the compute time of a single pass. For large genomes such as the 1272.71 Mb Prinsepia uniflora genome, this cost is substantial. However, the cost of producing a chimeric assembly that cannot be used for downstream analysis is far higher. The masking matrix is an investment in assembly quality.
The decision framework is not a replacement for understanding your genome. It is a structure for applying that understanding systematically. The repeat content, the repeat landscape, and the expected genome size all inform the framework. The framework then guides your parameter choices and tells you when to stop adjusting. The result is a masking decision that is justified by data instead of by habit or convenience.
Frequently Asked Questions
What is the difference between hard masking and soft masking?
Hard masking replaces every base in a repetitive region with the character N, which indicates an unknown base. Soft masking converts the bases to lowercase while preserving the original nucleotide sequence. Soft masking is generally preferred for assembly because it preserves sequence information that may be useful for downstream analysis. Hard masking may be necessary if your assembler does not recognize lowercase sequence as masked.
How do I know if my genome needs repeat masking?
Compare the repeat content of your genome to published values for related species. If your genome is a plant or animal with a large genome, it likely contains substantial repeat content. The Prinsepia uniflora genome contained 61.07% transposable elements. The gyrfalcon genome contained 7.61% transposable elements. If your genome falls in the higher range, masking is essential. Even low-repeat genomes benefit from masking to prevent the few repeats present from causing misassembly.
Can I use a repeat library from a related species instead of building my own?
You can use a library from a related species, but it will miss species-specific repeats and repeats that have diverged significantly from the reference. The gyrfalcon project combined de novo identified TEs with publicly available TE collections. This combined approach is recommended because it captures both conserved and species-specific repeats. Building your own library with RepeatModeler is the most thorough approach.
What is the minimum quality of the preliminary assembly needed for repeat modeling?
The preliminary assembly needs sufficient contiguity to contain full-length copies of the major repeat families. A draft assembly with an N50 of at least a few kilobases is usually adequate. The preliminary assembly does not need to be chromosome-level. The goal is to have enough sequence to identify the repeat families present in the genome.
How do I know if my repeat library is complete?
Compare the repeat content of your masked assembly to published values for related species. If your repeat content is much lower than expected, the library is likely incomplete. You can also examine the RepeatModeler output for the number of families identified. A small number of families in a repeat-rich genome suggests that the modeling was incomplete. Consider running RepeatModeler with different parameters or adding known repeats from related species.
What should I do if the assembly still has problems after masking?
Check whether the masking was applied correctly. Verify that the assembler recognized the mask and did not use masked sequence for contig extension. Check the repeat library for completeness and accuracy. Consider testing different masking parameters or a different assembler. If problems persist, consult a genome assembly specialist.
Does repeat masking affect gene annotation?
Repeat masking can affect gene annotation in both positive and negative ways. Masking prevents transposable elements from being misidentified as protein-coding genes, which improves annotation accuracy. However, masking can also remove functional elements that are located within repeats. The Prinsepia uniflora project used the repeat-masked assembly for gene prediction and identified 49,261 protein-coding genes. The masking was essential for accurate gene prediction.
How much computational resources do I need for repeat masking?
The computational requirements depend on the genome size and the repeat content. RepeatModeler and RepeatMasker require substantial memory and processing time for large genomes. The Prinsepia uniflora genome is 1,272.71 Mb and required significant computational resources for repeat analysis. Plan to use a computing cluster or cloud resource for genomes larger than a few hundred megabases.
Related Bioinformatics Guides
- Evaluating Genome Assembly Quality: Metrics and Tools
- Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- De Novo Genome Assembly with Long Reads: A Practical Workflow
- Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- A High-Quality Reference Genome Assembly of Prinsepia uniflora (Rosaceae).. Genes, 2023.
- The gyrfalcon (Falco rusticolus) genome.. G3 (Bethesda, Md.), 2023.
- Repetitive DNA in eukaryotic genomes.. Chromosome research : an international journal on the molecular, supramolecular and evolutionary aspects of chromosome biology, 2015.
- Diploid Genome Assembly of the Wine Grape Carménère.. G3 (Bethesda, Md.), 2019.
- Genome report: whole-genome assembly of the relapsing fever tick Ornithodoros turicata Dugès (Acari: Argasidae).. G3 (Bethesda, Md.), 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.