# Repeat Annotation in Genome Projects: How to Identify and Mask Transposable Elements with RepeatModeler and RepeatMasker


## Key Takeaways

- Repeat annotation is critical for accurate gene prediction and comparative genomics by identifying and masking transposable elements (TEs) and repetitive sequences, which can constitute a substantial fraction of eukaryotic genomes. Tools like RepeatModeler perform de novo discovery to build species-specific TE libraries, while RepeatMasker uses these libraries for genome-wide masking.
- RepeatModeler's `-LTRStruct` parameter is essential for improving the discovery of Long Terminal Repeat (LTR) retroelements, a major class of TEs often missed by simpler k-mer based approaches due to their size and complexity. The output of RepeatModeler includes consensus sequences classified by RepeatClassifier into known (e.g., DNA transposons, LINEs, SINEs) and unknown categories.
- RepeatMasker offers both soft masking (replacing repeats with lowercase letters) and hard masking (replacing with 'N' characters); soft masking is generally preferred for gene prediction pipelines like BRAKER3 as it preserves sequence information while signaling repetitive regions. The `.out` file generated by RepeatMasker provides detailed annotations of repeat matches, crucial for calculating genome-wide repeat content.
- Quality control is paramount, involving BUSCO analysis before and after masking to detect overmasking (where conserved genes are mistakenly masked) or undermasking (where TEs are missed). A significant drop in BUSCO completeness post-masking indicates issues with the repeat library, necessitating refinement.
- The composition of the repeat library for RepeatMasker should be tailored to taxonomic distance; for non-model organisms, the RepeatModeler-generated library is primary, while for well-studied taxa, combining it with curated databases (e.g., Dfam) can improve classification accuracy.
- Comprehensive record-keeping, including input assembly versions, exact tool parameters, output files (RepeatModeler library, RepeatMasker `.out` table), and BUSCO results, is vital for reproducibility and allows for future updates or re-analysis with improved methodologies.

---

A newly assembled genome contains far more than protein-coding genes. Transposable elements and other repetitive sequences often occupy a substantial fraction of eukaryotic genomes, and their presence directly interferes with gene prediction, sequence alignment, and comparative analyses. Repeat annotation is the process of identifying these repetitive families, building a species-specific library of consensus models, and masking the genome so that downstream gene annotation tools focus on unique, coding sequence. This article provides a practical workflow using RepeatModeler for de novo repeat discovery and RepeatMasker for classification and masking, with concrete guidance on input preparation, parameter choices, output interpretation, and quality control. The target reader is a researcher, student, or laboratory professional who has a genome assembly and needs to produce a repeat-masked version suitable for gene prediction and downstream analysis.

## At a Glance

The table below summarizes the core decisions in a repeat annotation project. Use it as a quick reference before starting the workflow.

| Workflow Stage | Primary Tool | Key Input | Main Output | Critical Decision |
| --- | --- | --- | --- | --- |
| Repeat discovery | RepeatModeler | Genome assembly FASTA | Species-specific repeat library | Whether to enable LTR structural discovery |
| Repeat classification | RepeatClassifier (within RepeatModeler) | Raw repeat families from discovery | Classified consensus sequences | Assignment of known versus unknown categories |
| Genome masking | RepeatMasker | Genome assembly plus repeat library | Masked FASTA and .out table | Choice of soft masking versus hard masking |
| Quality assessment | BUSCO, QUAST | Masked and unmasked assemblies | Completeness and assembly statistics | Comparison of gene-space completeness before and after masking |

The workflow proceeds from raw assembly to a repeat library, then to a masked genome. Each stage produces records that should be retained for reproducibility and reporting.

## Context and Scope of Repeat Annotation

### Why Repetitive Elements Matter in Genome Projects

Repetitive elements are biologically meaningful components that shape genome structure, influence gene regulation, and contribute to evolutionary dynamics. In the rock goby genome, repeat landscape characterization was an explicit part of the assembly report, alongside protein-coding gene prediction and functional annotation [<a href="#ref-1">1</a>]. In the Rhinogobio ventralis genome, repetitive elements accounted for 63.44% of the assembly, with DNA transposons alone representing 42.31% [<a href="#ref-2">2</a>]. These figures demonstrate that repeats can dominate a genome and cannot be ignored in any annotation project.

The practical problem is that gene prediction algorithms assume most of the genome is unique sequence. When repeats are left unmasked, gene predictors may produce spurious gene models from repetitive regions, inflate gene counts, and misassign functional annotations. Masking repeats before gene prediction reduces false positives and improves the accuracy of downstream analyses.

### The Difference Between Annotation and Masking

Repeat annotation and repeat masking are related but distinct tasks. Annotation produces a detailed description of each repeat family, including its classification, consensus sequence, and genomic distribution. Masking uses that annotation to replace repeat sequences in the genome with N characters (hard masking) or lowercase characters (soft masking). Soft masking preserves the original sequence information while signaling to downstream tools that the region is repetitive. Most gene prediction pipelines accept soft-masked genomes, because the lowercase bases remain available for alignment while the masking status is explicit.

### When to Perform Repeat Annotation

Repeat annotation should occur after assembly quality assessment and before gene prediction. The assembly must be complete enough to support repeat discovery. Highly fragmented assemblies with short contigs may produce incomplete repeat models, because a repeat family dispersed across many small contigs may not be reconstructed accurately. Assembly quality metrics such as contig N50 and BUSCO completeness provide a baseline for deciding whether the assembly is ready for repeat annotation [<a href="#ref-3">3</a>]. If the assembly fails basic quality checks, repeat annotation should be postponed until the assembly is improved.

## Core Principles of Transposable Element Biology

### Major Classes of Transposable Elements

Transposable elements fall into two broad classes based on their transposition mechanism. Class I elements, or retrotransposons, move through an RNA intermediate and require reverse transcription. Class II elements, or DNA transposons, move directly as DNA. Within these classes, elements are further divided into orders and superfamilies based on structural features, sequence motifs, and replication strategies.

Long terminal repeat (LTR) retroelements are a major order of Class I elements. They are widespread in eukaryotic genomes but difficult to identify automatically because of their size and sequence complexity [<a href="#ref-4">4</a>]. LTR elements contain long terminal repeats flanking internal coding regions, and their full-length copies can be several kilobases long. Automated discovery tools historically missed many LTR families because the elements are too large and complex for simple k-mer based approaches [<a href="#ref-4">4</a>].

Non-LTR retrotransposons include long interspersed nuclear elements (LINEs) and short interspersed nuclear elements (SINEs). DNA transposons include families such as Tc1-Mariner, hAT, and Helitrons. Each superfamily has distinct structural hallmarks that classification tools use to assign unknown sequences to known categories.

### Why Species-Specific Libraries Are Necessary

Transposable element sequences are highly variable across species [<a href="#ref-4">4</a>]. An element family that is active in one species may be absent or diverged beyond recognition in another. Public repeat databases such as RepBase and Dfam contain curated consensus sequences from many species, but they cannot cover every family in every genome. A newly assembled genome will contain species-specific families that are not represented in any database.

The solution is de novo repeat discovery. RepeatModeler2 was developed specifically to address this problem by automating the discovery of repeat families directly from the genome sequence [<a href="#ref-4">4</a>]. The pipeline identifies repetitive sequences, builds consensus models, and classifies them into known and unknown categories. Benchmarking on fruit fly, zebrafish, and rice showed that RepeatModeler2 identified approximately three times more consensus sequences matching curated libraries at high sequence identity and coverage than the original RepeatModeler [<a href="#ref-4">4</a>]. The greatest improvement was for LTR retroelements [<a href="#ref-4">4</a>].

### The Role of Curated Databases

Curated databases provide a reference for classification and a source of known families that may be missed by de novo discovery. RepeatMasker can use a combined library that includes both the species-specific models from RepeatModeler and a curated database appropriate for the taxonomic group. The curated database helps classify elements that are too diverged to be discovered de novo but still recognizable by homology.

The choice of curated database depends on the organism. RepeatMasker ships with databases for common model organisms, and additional databases are available from the Dfam consortium. For non-model organisms, the closest available taxonomic database is a reasonable starting point, but the results should be inspected for misclassification.

## Preparing the Input Genome Assembly

### Assembly Quality Requirements

Repeat discovery is computationally intensive and sensitive to assembly quality. The input should be a FASTA file containing the assembled contigs or chromosomes. The assembly should be free of vector contamination, adapter sequences, and other artifacts that could be mistaken for biological repeats.

Assembly quality metrics provide a useful preflight check. QUAST can assess contiguity, misassembly rates, and completeness [<a href="#ref-3">3</a>]. BUSCO can evaluate gene-space completeness by searching for conserved single-copy orthologs [<a href="#ref-3">3</a>]. An assembly with high BUSCO completeness and reasonable contiguity is suitable for repeat annotation. An assembly with very low contiguity or evidence of contamination should be cleaned or reassembled before proceeding.

### Formatting and Sanitizing the FASTA File

The FASTA file should use standard formatting with unambiguous sequence characters. Some repeat discovery tools are sensitive to non-ACGT characters, and large stretches of N characters can confuse k-mer counting. If the assembly contains many scaffolds with long N gaps, consider whether to mask or remove those gaps before repeat discovery.

Sequence headers should be short and unique. Some downstream tools truncate headers at whitespace, so headers with spaces or special characters can cause errors. Rename sequences to simple identifiers such as chr1, chr2, scaffold1, and contig1 before starting the workflow.

### Computational Resource Planning

RepeatModeler and RepeatMasker are resource-intensive. RepeatModeler uses multiple search engines and can run for days on large genomes. The memory and time requirements scale with genome size and repeat content. A chromosome-level assembly of a vertebrate genome may require significant compute resources, and the workflow should be run on a server or high-performance computing cluster instead of a laptop.

Containerized versions of RepeatModeler2 are available, which simplifies installation and ensures reproducibility across systems [<a href="#ref-4">4</a>]. The container includes all dependencies and can be run with standard container runtimes.

## Building a Species-Specific Repeat Library with RepeatModeler

### Overview of the RepeatModeler Pipeline

RepeatModeler2 automates the discovery of transposable element families from a genome assembly [<a href="#ref-4">4</a>]. The pipeline combines multiple algorithms to identify repetitive sequences, build multiple sequence alignments, and generate consensus models. A key innovation in RepeatModeler2 is the inclusion of a structural discovery module for LTR retroelements, which addresses the historical difficulty of identifying these large and complex elements [<a href="#ref-4">4</a>].

The pipeline produces a library of consensus sequences representing the repeat families in the genome. Each consensus sequence is a representative model for a family, and the library is used as input to RepeatMasker for genome-wide masking.

### Running RepeatModeler

The basic command structure for RepeatModeler is:

```
BuildDatabase -name mydb genome.fasta
RepeatModeler -database mydb -pa 8 -LTRStruct
```

The `BuildDatabase` step creates a BLAST database from the genome FASTA. The `-pa` flag specifies the number of parallel search jobs, and `-LTRStruct` enables the LTR structural discovery module. The output is a directory containing the repeat library and intermediate files.

The `-LTRStruct` flag is important for genomes with LTR retroelements. Because LTR elements are recalcitrant to automated identification due to their size and sequence complexity [<a href="#ref-4">4</a>], enabling structural discovery substantially improves recovery of these families.

### Interpreting RepeatModeler Output

The primary output is a FASTA file containing consensus sequences for each repeat family. The sequence headers include classification information assigned by RepeatClassifier. Families classified as known are assigned to categories such as DNA transposons, LINEs, SINEs, or LTR elements. Families that cannot be classified are labeled as unknown.

The number of families in the library varies by genome. A genome with high repeat content will produce more families than a genome with low repeat content. The library should be inspected for quality before use. Families with very short consensus sequences or very low copy numbers may be artifacts and can be removed.

### Common Failure Patterns in Repeat Discovery

RepeatModeler can fail or produce poor results for several reasons. Low complexity sequence and tandem repeats can consume computational resources and produce spurious models. Some pipelines recommend masking low complexity sequence before repeat discovery, but this must be balanced against the risk of removing genuine repeat families.

Highly fragmented assemblies can break repeat families across multiple contigs, preventing the pipeline from building complete consensus models. If the assembly has very low contiguity, consider whether repeat annotation should be deferred until the assembly is improved.

Genomes with extreme repeat content, such as the 63.44% repetitive Rhinogobio ventralis genome [<a href="#ref-2">2</a>], may require additional computational resources and longer run times. The pipeline should be monitored for progress and resource usage.

## Classifying Repeat Families

### How RepeatClassifier Works

RepeatClassifier is integrated into RepeatModeler2 and assigns each consensus sequence to a repeat class based on homology to known databases and structural features. The classifier searches each consensus against reference databases and uses the best hit to assign a classification. Sequences without significant hits are labeled unknown.

The classification step is critical because downstream analyses often filter by repeat class. For example, a study of DNA transposon evolution would need accurate classification of DNA transposon families. Misclassified families can lead to incorrect biological conclusions.

### Known and Unknown Families

The proportion of known versus unknown families varies by genome. Genomes from well-studied taxa will have more families classified as known, because reference databases contain related sequences. Genomes from understudied taxa will have more unknown families, reflecting the lack of related reference sequences.

Unknown families are not necessarily artifacts. They may represent genuine species-specific repeats that have no homology to known elements. These families should be retained in the library for masking, even though they cannot be assigned to a biological category.

### Refining the Library

The raw RepeatModeler library can be refined before masking. Short consensus sequences, typically under 100 base pairs, may represent low complexity repeats or artifacts. Some pipelines filter these out to reduce noise. However, some genuine small RNA-derived repeats are short, so filtering should be conservative.

The library can also be compared to curated databases to identify and remove sequences that are likely contamination. If the genome assembly was contaminated with sequences from another organism, the repeat library may contain families from that contaminant. Removing these families improves the specificity of masking.

## Masking the Genome with RepeatMasker

### Soft Masking Versus Hard Masking

RepeatMasker can produce two types of masked output. Hard masking replaces repetitive bases with N characters, destroying the original sequence information. Soft masking replaces repetitive bases with lowercase characters, preserving the original sequence while marking it as repetitive.

Soft masking is generally preferred for gene prediction. Gene predictors such as BRAKER3 can use soft-masked genomes and will ignore lowercase regions during model training while retaining the sequence for alignment [<a href="#ref-3">3</a>]. Hard masking is appropriate when the masked sequence will be used for applications that cannot tolerate lowercase characters, such as some alignment tools.

### Running RepeatMasker

The basic command structure for RepeatMasker is:

```
RepeatMasker -lib mylib.fa -xsmall -pa 8 genome.fasta
```

The `-lib` flag specifies the repeat library, `-xsmall` produces soft masking, and `-pa` sets the number of parallel jobs. The output includes a masked FASTA file, a .out table with detailed annotations, and summary statistics.

The repeat library used for masking should include both the species-specific models from RepeatModeler and a curated database. The combined library can be created by concatenating the RepeatModeler library with the appropriate RepeatMasker database.

### Interpreting the .out File

The .out file is a tab-delimited table with one row per repeat match. Each row includes the query sequence name, the start and end positions of the match, the repeat name, the repeat class, and a score. The file can be parsed to calculate genome-wide repeat content, the contribution of each repeat class, and the distribution of repeats across chromosomes.

The .out file is the primary record of the repeat annotation. It should be retained for downstream analyses and for reporting in publications. Many comparative genomics studies report repeat content as a percentage of the genome, and this value is calculated from the .out file.

### Summary Statistics

RepeatMasker produces a summary file that lists the total number of repeats, the total bases masked, and the percentage of the genome masked. This summary provides a quick check on the overall repeat content. The summary can be compared to published values for related species to assess whether the repeat content is plausible.

For example, the Muscovy duck genome had 22.14% repetitive sequence [<a href="#ref-5">5</a>], while the Rhinogobio ventralis genome had 63.44% [<a href="#ref-2">2</a>]. These values reflect real biological differences between species. A repeat content that is unexpectedly high or low relative to related species may indicate a problem with the assembly or the repeat library.

## Quality Control and Validation

### BUSCO Analysis Before and After Masking

BUSCO completeness should be assessed on both the unmasked and masked genomes. The unmasked BUSCO score provides a baseline for assembly quality. The masked BUSCO score should be similar, because masking should not remove conserved coding sequence. A large drop in BUSCO completeness after masking indicates that the repeat library contains sequences that match conserved genes, and the library should be refined.

The AquaaG pipeline integrates BUSCO evaluation into the annotation workflow, providing gene-space completeness reports that can be used to validate the masking step [<a href="#ref-3">3</a>]. This integration highlights the importance of quality assessment at each stage of the annotation process.

### Checking for Overmasking

Overmasking occurs when the repeat library contains sequences that match unique, coding regions. This can happen when a repeat family has inserted into a gene and the consensus model includes flanking coding sequence. Overmasking reduces the effective genome size for gene prediction and can cause genuine genes to be missed.

Overmasking can be detected by comparing gene predictions on the masked and unmasked genomes. If the masked genome produces substantially fewer gene models than expected, or if BUSCO completeness drops, the repeat library should be inspected for problematic families.

### Checking for Undermasking

Undermasking occurs when genuine repeats are not recognized by the repeat library. This can happen when a repeat family is too diverged from the consensus models to be detected, or when the family was not discovered by RepeatModeler. Undermasking leads to spurious gene predictions from repetitive regions.

Undermasking can be detected by examining the unmasked portions of the genome for high coverage of short reads or by comparing repeat content to published values for related species. If the repeat content is substantially lower than expected, the repeat library may be incomplete.

### Records and Measurements to Retain

A complete repeat annotation project should retain the following records:

- The input genome assembly FASTA file with version information
- The RepeatModeler output library and intermediate files
- The combined repeat library used for masking
- The RepeatMasker .out file and summary statistics
- BUSCO results for the unmasked and masked genomes
- Assembly quality metrics from QUAST or similar tools
- The exact commands and parameters used at each step

These records support reproducibility and allow the annotation to be updated when improved tools or databases become available.

## Integration with Gene Prediction Pipelines

### Using the Masked Genome with BRAKER3

BRAKER3 is a eukaryotic gene prediction pipeline that integrates RNA-seq evidence and protein homology to predict gene structures [<a href="#ref-3">3</a>]. The pipeline can use a soft-masked genome as input, and the masking status guides the gene predictor to focus on unique sequence.

The masked genome should be provided to BRAKER3 along with the appropriate evidence files. The pipeline will produce gene models that can be functionally annotated with tools such as EggNOG-mapper [<a href="#ref-3">3</a>]. The repeat annotation is a prerequisite for this step, because unmasked repeats would produce spurious gene models.

### Functional Annotation of Predicted Genes

After gene prediction, functional annotation assigns biological functions to predicted proteins. The AquaaG pipeline uses EggNOG-mapper for functional annotation [<a href="#ref-3">3</a>]. The quality of functional annotation depends on the quality of the gene models, which in turn depends on the quality of the repeat masking.

In the Muscovy duck genome project, 16,901 coding genes were predicted and 93.4% were functionally annotated [<a href="#ref-5">5</a>]. In the rock goby genome project, 23,493 protein-coding genes were identified and over 96% showed homology to known proteins [<a href="#ref-1">1</a>]. These high annotation rates reflect careful repeat masking and gene prediction.

### Comparative Genomics Applications

Repeat-masked genomes are used for comparative genomics analyses such as whole-genome alignment, synteny analysis, and phylogenetic reconstruction. Masking ensures that alignments focus on orthologous sequence instead of repetitive regions that may be shared by descent or by chance.

The repeat annotation itself is also a subject of comparative study. Repeat content and composition differ between species, and these differences can be linked to biological traits. The Rhinogobio ventralis genome, with its high DNA transposon content, provides a contrast to the Muscovy duck genome with lower overall repeat content [<a href="#ref-5">5</a>][<a href="#ref-2">2</a>]. These differences may reflect different evolutionary histories and selective pressures.

## Common Failure Patterns and Troubleshooting

### RepeatModeler Runs Out of Memory

RepeatModeler can consume large amounts of memory, especially on genomes with high repeat content. If the pipeline fails with out-of-memory errors, reduce the number of parallel jobs, increase available memory, or run the pipeline on a machine with more resources. Containerized versions may have memory limits that need to be adjusted [<a href="#ref-4">4</a>].

### RepeatMasker Produces No Matches

If RepeatMasker produces no matches, the repeat library may be empty or malformed. Check that the library FASTA file is valid and contains sequences. Also check that the library and genome use compatible sequence alphabets. If the library was built from a different assembly, it may not match the current genome.

### BUSCO Completeness Drops After Masking

A drop in BUSCO completeness after masking indicates overmasking. The repeat library likely contains sequences that match conserved genes. Inspect the library for families with high similarity to known proteins and remove or refine those families. Re-run RepeatMasker with the refined library.

### Repeat Content Is Unexpectedly Low

If the repeat content is much lower than expected for the taxonomic group, the repeat library may be incomplete. Re-run RepeatModeler with different parameters, or add a curated database to the masking library. Also check that the assembly contains the expected sequences and was not filtered during assembly.

### Classification Produces Many Unknown Families

A high proportion of unknown families is common for non-model organisms. The unknown families should be retained for masking, but they cannot be assigned to biological categories. If classification is important for the analysis, consider using additional classification tools or manually inspecting the unknown families.

## Limitations and Interpretation Boundaries

### Repeat Annotation Is Not Error-Free

Repeat annotation is a computational prediction, not a definitive biological characterization. The consensus sequences in the repeat library are models that represent the diversity of each family, but individual copies may diverge substantially from the consensus. Some families may be missed entirely, and some classified families may be assigned to the wrong category.

The repeat content reported for a genome is an estimate that depends on the tools, parameters, and databases used. Different pipelines may produce different repeat content values for the same genome. Comparisons between species should use consistent methods.

### De Novo Discovery Has Known Biases

De novo repeat discovery is biased toward high-copy, recently active families. Low-copy families and ancient, diverged families are more difficult to discover because they have fewer copies and less sequence similarity. The RepeatModeler2 improvement in LTR discovery addresses one major bias [<a href="#ref-4">4</a>], but other biases remain.

The benchmarking of RepeatModeler2 on fruit fly, zebrafish, and rice demonstrated improved recovery of curated families [<a href="#ref-4">4</a>], but these are model species with well-characterized repeat landscapes. Non-model species may present additional challenges.

### Curated Databases Are Incomplete

Curated repeat databases are valuable but incomplete. They are biased toward model organisms and economically important species. A repeat family that is common in a non-model species may be absent from all curated databases. The species-specific library from RepeatModeler is essential for capturing these families.

### Assembly Quality Limits Annotation Quality

Repeat annotation cannot compensate for poor assembly quality. If the assembly is fragmented or contains misassemblies, the repeat models will be incomplete or incorrect. The assembly quality should be assessed before repeat annotation, and the annotation should be revisited if the assembly is improved.

## Professional Escalation Criteria

### When to Seek Expert Assistance

Repeat annotation can be challenging, and some situations warrant consultation with a bioinformatics specialist or the tool developers. Seek assistance if:

- RepeatModeler fails repeatedly with errors that cannot be resolved by adjusting parameters
- The repeat library produces severe overmasking that cannot be fixed by filtering
- The repeat content is wildly different from published values for related species
- The genome assembly is highly fragmented and repeat annotation is needed before assembly improvement
- The analysis requires custom repeat classification beyond the capabilities of standard tools

### When to Revisit the Annotation

Repeat annotation should be revisited when the assembly is updated, when improved repeat databases become available, or when downstream analyses reveal problems. A new assembly version requires a new repeat annotation, because the repeat models are derived from the assembly sequence. Improved databases may allow better classification of unknown families.

### Documentation for Publication

Publications reporting genome assemblies should include a description of the repeat annotation methods. The description should specify the versions of RepeatModeler and RepeatMasker, the parameters used, the databases included in the masking library, and the resulting repeat content. This documentation supports reproducibility and allows readers to assess the quality of the annotation.

The genome reports for the Muscovy duck, rock goby, and Rhinogobio ventralis all include repeat content as part of the assembly characterization [<a href="#ref-5">5</a>][<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. This practice should be followed in any genome paper.

## A Practical Decision Framework for Repeat Masking Strategy

### Selecting the Masking Approach for Your Analysis Goal

The choice between soft masking and hard masking is not a single global decision. It depends on the downstream analysis that will consume the masked genome. A researcher running multiple downstream analyses may need to produce both masked versions from the same repeat annotation. The decision framework below maps analysis types to masking strategies.

For gene prediction with pipelines such as BRAKER3, soft masking is the standard choice because the lowercase bases remain available for alignment while the masking status is explicit to the predictor [<a href="#ref-3">3</a>]. For whole-genome alignment and synteny analysis, hard masking is often preferred because alignment tools treat lowercase and uppercase characters identically, which defeats the purpose of soft masking. For variant calling and short-read mapping, hard masking prevents reads from repetitive regions from aligning to multiple locations and producing spurious variant calls.

The decision should be recorded in the project metadata along with the rationale. A project that produces both masked versions from a single RepeatMasker run with the `-xsmall` flag can generate the hard-masked version by converting lowercase characters to N characters with a simple script. This approach avoids running RepeatMasker twice and ensures that both versions derive from the same repeat annotation.

### Matching Repeat Library Composition to Taxonomic Distance

The composition of the repeat library used for masking should reflect the taxonomic distance between the target species and the closest curated database. For a species with a well-curated database in RepeatMasker, such as human, mouse, or zebrafish, the masking library can include both the species-specific RepeatModeler library and the curated database. For a non-model species with no close curated database, the RepeatModeler library should be the primary masking library, and a curated database from a distantly related taxon may introduce more noise than signal.

The rock goby genome project characterized the repeat landscape as part of the assembly report, and the Rhinogobio ventralis project reported repetitive elements accounting for 63.44% of the genome with DNA transposons alone representing 42.31% [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. These projects demonstrate that repeat annotation is a standard component of genome reporting, but the specific library composition decisions are rarely documented in the final paper. The decision framework should be recorded in the project notebook or methods supplement.

### Parameter Selection for RepeatMasker

RepeatMasker parameters should be selected based on the repeat library composition and the desired sensitivity. The `-cutoff` parameter controls the minimum score for reporting a match. Lower cutoffs increase sensitivity but also increase false positives. The default cutoff is appropriate for most projects, but a project with a high-quality curated library may benefit from a higher cutoff to reduce noise.

The `-gccalc` parameter enables GC content calculation, which is useful for interpreting repeat content in GC-biased genomes. The `-nolow` parameter disables low complexity masking, which is sometimes desirable when the repeat library already contains low complexity models. The `-a` parameter produces an additional alignment file that can be used for detailed inspection of individual matches.

The choice of search engine is also a parameter decision. RepeatMasker can use RMBlast, cross_match, or HMMER for the search step. RMBlast is the default and is generally faster. HMMER is more sensitive for protein-coding repeat families such as SINEs and some LINEs. A project with many protein-coding repeat families may benefit from running RepeatMasker with HMMER in addition to RMBlast and comparing the results.

### Record Keeping for Masking Decisions

A repeat masking project should maintain a decision log that records the rationale for each parameter choice. The log should include the analysis goals, the masking strategy selected for each goal, the repeat library composition, the RepeatMasker parameters, and the date and version of each tool. This log supports reproducibility and allows the masking to be revisited when analysis goals change.

The decision log should also record the computational resources used, including the number of parallel jobs, the memory allocation, and the wall time. This information is useful for planning future repeat annotation projects on similar genomes. The AquaaG pipeline demonstrates the value of automated and reproducible annotation workflows that integrate assembly quality assessment, gene prediction, and functional annotation [<a href="#ref-3">3</a>]. A decision log provides the same reproducibility for the repeat masking step.

### Common Failure Patterns in Masking Strategy

A common failure pattern is using a single masked genome for all downstream analyses without considering the different requirements of each analysis. A researcher who soft-masks the genome for gene prediction and then uses the same soft-masked genome for whole-genome alignment may produce alignments that include repetitive regions, because the alignment tool does not distinguish lowercase from uppercase. The solution is to produce both masked versions and select the appropriate version for each analysis.

Another failure pattern is using a curated database from a distantly related species without inspecting the results. The curated database may contain families that match unique sequence in the target genome, causing overmasking. The solution is to compare the repeat content with and without the curated database and inspect the families that are unique to the curated database.

A third failure pattern is ignoring the GC content of the genome when interpreting repeat content. RepeatMasker reports repeat content as a percentage of the genome, but this percentage can be misleading in GC-biased genomes where some repeat families are enriched in GC-rich or GC-poor regions. The `-gccalc` parameter provides GC content information that should be used to interpret the repeat content in a genomic context.

### Validation of Masking Strategy

The masking strategy should be validated by comparing downstream analysis results on the masked and unmasked genomes. For gene prediction, the number of predicted genes and the BUSCO completeness should be compared. A large drop in BUSCO completeness after masking indicates overmasking, and the repeat library should be refined [<a href="#ref-3">3</a>]. For whole-genome alignment, the alignment coverage and the number of aligned blocks should be compared. A large reduction in alignment coverage after masking may indicate that the masking is too aggressive.

The validation results should be recorded in the decision log along with the masking decisions. This record allows the masking strategy to be adjusted when downstream analyses reveal problems. The validation step is often skipped in practice, but it is essential for ensuring that the masking strategy supports the analysis goals.

### Professional Escalation Criteria for Masking Strategy

Seek expert assistance if the masking strategy produces results that cannot be explained by the repeat biology of the species. For example, if the repeat content is dramatically different from published values for related species, or if the masked genome produces gene predictions that are clearly wrong, a bioinformatics specialist should review the repeat library and the masking parameters.

Seek assistance if the repeat library contains many families that match conserved proteins, because this indicates a problem with the RepeatModeler output or the library filtering step. A specialist can help identify the source of the problem and refine the library.

Seek assistance if the project requires a masking strategy that is not supported by standard RepeatMasker parameters, such as masking only specific repeat classes or masking with a custom scoring scheme. A specialist can help implement a custom masking workflow.

### Integration with Automated Pipelines

Repeat masking is increasingly integrated into automated genome annotation pipelines. The AquaaG pipeline integrates genome assembly retrieval, quality assessment, gene prediction, and functional annotation into a reproducible workflow [<a href="#ref-3">3</a>]. The nf-core community provides standardized pipeline documentation that supports reproducible workflow configuration [<a href="#ref-6">6</a>]. These pipelines typically include a repeat masking step, but the masking parameters are often fixed defaults.

A researcher who uses an automated pipeline should verify that the repeat masking step uses a species-specific library and appropriate parameters. Some pipelines use a generic repeat library that may not be appropriate for the target species. The decision framework described here can be used to evaluate whether the pipeline masking step meets the project requirements.

The Galaxy Training Network provides accessible workflow training that includes repeat annotation and masking tutorials [<a href="#ref-7">7</a>]. These tutorials demonstrate the standard parameters and common pitfalls. A researcher who is new to repeat masking can use these tutorials to build familiarity before running the workflow on their own genome.

### Records and Measurements for Masking Strategy

The following records should be retained for the masking strategy:

- The analysis goals and the masking strategy selected for each goal
- The repeat library composition, including the RepeatModeler library and any curated databases
- The RepeatMasker parameters, including the cutoff, search engine, and low complexity settings
- The masked genome files for each masking strategy
- The RepeatMasker .out file and summary statistics for each masking run
- The validation results comparing downstream analyses on masked and unmasked genomes
- The decision log with the rationale for each parameter choice

These records support reproducibility and allow the masking strategy to be revisited when analysis goals change or when improved tools become available. The records also support reporting in publications, where the masking strategy should be described with sufficient detail for readers to assess the quality of the annotation.

### Comparison of Masking Strategies Across Recent Genome Projects

Recent genome projects demonstrate different masking strategies in practice. The Muscovy duck genome project reported 22.14% repetitive sequence and used the repeat annotation as part of the genome characterization [<a href="#ref-5">5</a>]. The rock goby genome project characterized the repeat landscape and identified 23,493 protein-coding genes [<a href="#ref-1">1</a>]. The Rhinogobio ventralis genome project reported 63.44% repetitive elements with DNA transposons representing 42.31% [<a href="#ref-2">2</a>].

These projects used different masking strategies appropriate to their analysis goals. The Muscovy duck project focused on gene prediction and functional annotation, so soft masking was likely used for the gene prediction step. The Rhinogobio ventralis project reported detailed repeat class composition, which requires careful classification and a comprehensive repeat library. The rock goby project reported repeat landscape characterization alongside gene prediction, indicating that both masking and annotation were performed.

A researcher planning a new genome project should review the methods sections of recent genome papers for related species to understand the masking strategies that are standard in their field. The decision framework described here provides a structured approach to selecting a masking strategy, but the specific choices should be informed by the practices of the relevant research community.

## Frequently Asked Questions

### What is the difference between RepeatModeler and RepeatMasker?

RepeatModeler discovers repeat families de novo from a genome assembly and builds a library of consensus sequences. RepeatMasker uses a repeat library to search the genome and produce a masked version with annotations of each repeat match. RepeatModeler is typically run first to build the library, and RepeatMasker is run second to apply the library to the genome.

### Should I use soft masking or hard masking for gene prediction?

Soft masking is generally preferred for gene prediction. Soft masking preserves the original sequence in lowercase characters, allowing alignment tools to use the sequence while signaling that it is repetitive. Hard masking replaces repeats with N characters and destroys the sequence information. Gene predictors such as BRAKER3 can use soft-masked genomes effectively [<a href="#ref-3">3</a>].

### How long does RepeatModeler take to run?

Run time depends on genome size, repeat content, and available computational resources. Small genomes may complete in hours, while large vertebrate genomes can take days. The pipeline uses multiple search engines and parallel jobs, and the `-pa` parameter controls the number of parallel jobs. Monitor the pipeline progress and allocate sufficient time and resources.

### What should I do if my repeat library has many unknown families?

Unknown families are common for non-model organisms and do not necessarily indicate a problem. Retain the unknown families for masking, because they represent genuine repetitive sequence even if they cannot be classified. If classification is important for the analysis, consider additional classification tools or manual inspection.

### How do I calculate the repeat content of my genome?

The repeat content is calculated from the RepeatMasker .out file or summary statistics. The summary file reports the total bases masked and the percentage of the genome masked. The .out file can be parsed to calculate repeat content by class and family. The repeat content should be reported with the methods used to generate it.

### Can I use a repeat library from a related species?

A repeat library from a related species can be used as a supplement, but it should not replace a species-specific library. Transposable element sequences are highly variable across species [<a href="#ref-4">4</a>], and a related species library will miss species-specific families. Use the RepeatModeler library as the primary library and add a curated database for classification.

### What is the role of BUSCO in repeat annotation?

BUSCO assesses gene-space completeness by searching for conserved single-copy orthologs [<a href="#ref-3">3</a>]. It should be run before and after masking to ensure that masking does not remove conserved coding sequence. A large drop in BUSCO completeness after masking indicates overmasking and a need to refine the repeat library.

### How should I report repeat annotation in my publication?

Report the versions of RepeatModeler and RepeatMasker, the parameters used, the databases included in the masking library, and the resulting repeat content. Describe the repeat content by class if possible. This documentation supports reproducibility and allows readers to assess the annotation quality. Genome reports for recent assemblies include this information [<a href="#ref-5">5</a>][<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

## Related Bioinformatics Guides

- [Metagenome Assembled Genome Analysis: From Bins to Biological Insights](/knowledge/bioinformatics/metagenome-assembled-genome-analysis-from-bins-to-biological-insights)
- [Functional Metagenomics: From Gene Prediction to Pathway Reconstruction](/knowledge/bioinformatics/functional-metagenomics-from-gene-prediction-to-pathway-reconstruction)
- [Functional Annotation of Metagenomes: A Guide to Databases and Pipelines](/knowledge/bioinformatics/functional-annotation-of-metagenomes-a-guide-to-databases-and-pipelines)
- [Single-Cell Annotation: A Workflow for Cell Type Identification](/knowledge/bioinformatics/single-cell-annotation-a-workflow-for-cell-type-identification)
- [Machine Learning Bioinformatics Projects: From Idea to Publication](/knowledge/bioinformatics/machine-learning-bioinformatics-projects-from-idea-to-publication)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Highly contiguous chromosome-level assembly of the rock goby (Gobius paganellus) genome.](https://doi.org/10.1038/s41597-026-06659-9). 2026.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Chromosome-level genome assembly and annotation of the Rhinogobio ventralis, an endangered endemic fish from the Yangtze River.](https://doi.org/10.1038/s41597-026-06949-2). 2026.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [AquaaG: A comprehensive pipeline for quality assessment and annotation of genomes.](https://doi.org/10.1016/j.mex.2026.103955). 2026.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [RepeatModeler2 for automated genomic discovery of transposable element families.](https://pubmed.ncbi.nlm.nih.gov/32300014). Proceedings of the National Academy of Sciences of the United States of America, 2020.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [De novo chromosome-level genome assembly and annotation of the Muscovy duck (Cairina moschata).](https://doi.org/10.1016/j.psj.2026.106973). 2026.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.