# Assembly-Based vs. Read-Based Resistome Analysis: Which Approach Should You Use for Your Metagenomic Data?


## Key Takeaways

- Read-based resistome analysis directly maps short sequencing reads to ARG databases, offering rapid detection of known genes and moderate computational cost, but struggles with distinguishing closely related ARG variants and lacks genomic context.
- Assembly-based analysis reconstructs longer DNA contigs from reads, enabling higher confidence annotation, identification of genomic context (e.g., co-localization with mobile genetic elements like plasmids or transposons), and strain-level variant detection, albeit at a significantly higher computational and resource demand.
- The choice between approaches hinges on research objectives: assembly-based is crucial for inferring horizontal gene transfer potential and detailed variant analysis, while read-based is more suitable for broad screening of known ARGs, large cohorts, or low-depth data.
- Sensitivity and specificity are inversely related; read-based analysis with low identity thresholds increases false positives, while assembly-based analysis can miss low-abundance ARGs due to assembly fragmentation or failure, particularly in complex microbial communities.
- Both methods are critically dependent on the quality and comprehensiveness of the reference ARG database, with newer discoveries and specific resistance mechanisms (e.g., mutation-driven chromosomal resistance) requiring curated databases and specialized annotation tools.
- Sequencing depth is paramount for both approaches: higher depth improves read-based alignment confidence and assembly contiguity, with low-depth data strongly favoring read-based analysis due to the increased risk of assembly failure.

---

Metagenomic resistome analysis asks a deceptively simple question: which antibiotic resistance genes (ARGs) are present in your sample, and how abundant are they? The answer depends on a methodological fork in the road. You can map sequencing reads directly against reference databases, an approach called read-based analysis, or you can first assemble the reads into longer contiguous sequences (contigs) and then annotate those contigs for ARGs, an approach called assembly-based analysis. Each path produces different answers to the same biological question, and the choice between them shapes sensitivity, specificity, computational cost, and the types of downstream inferences you can make. This article compares the two approaches across those dimensions and provides a decision framework based on your research goals, data quality, and available computing resources. The target reader is a researcher, laboratory professional, or graduate student who has shotgun metagenomic sequencing data and needs to choose an analysis strategy that will produce defensible, reproducible results.

## The Core Distinction Between Read-Based and Assembly-Based Resistome Analysis

Read-based resistome analysis operates directly on the raw sequencing reads produced by your platform. Each read, typically 150 base pairs in length for Illumina data, is compared against a database of known ARG sequences using alignment tools. The comparison can be done with nucleotide aligners such as BWA or Bowtie2, or with translated protein aligners such as DIAMOND that compare the six-frame translation of each read against protein databases. The output is a table of ARG identities and abundances, usually normalized to sequencing depth or to a housekeeping gene.

Assembly-based resistome analysis takes a different route. The reads are first assembled into contigs using a de Bruijn graph assembler such as MEGAHIT, metaSPAdes, or metaFlye. The resulting contigs, which can range from a few hundred base pairs to hundreds of kilobases, are then searched for ARGs using annotation tools. Because contigs are longer than individual reads, they provide genomic context. You can determine whether an ARG is located near a mobile genetic element (MGE), whether multiple ARGs are clustered together, and whether the ARG is likely chromosomal or plasmid-borne.

The distinction matters because the two approaches answer different questions with different confidence levels. Read-based analysis answers the question of presence and approximate abundance quickly. Assembly-based analysis answers the question of presence, abundance, and genomic context, but at a higher computational cost and with the risk that assembly errors or fragmentation will obscure the true signal.

## What Each Approach Can and Cannot Tell You

Read-based analysis excels at detecting ARGs that have close matches in reference databases. When a read aligns with high identity to a known ARG sequence, you can be confident that the gene is present in your sample. The approach is computationally efficient, well documented, and supported by a mature ecosystem of tools and training materials. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflows for read-based taxonomic and functional profiling, and the [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal offers structured learning pathways for sequence analysis that include read-based approaches.

The limitation of read-based analysis is that short reads often cannot distinguish between closely related ARG variants. A 150-base-pair read may map equally well to several variants of a beta-lactamase gene, and the alignment tool will report the best match or a consensus. This ambiguity affects both sensitivity and specificity. If your sample contains a novel ARG variant that diverges from the database sequence, the read may fail to align at the identity threshold you have set, producing a false negative. Conversely, if a read aligns to a conserved region shared by multiple gene families, the annotation may assign it to the wrong family, producing a false positive.

Assembly-based analysis addresses some of these problems by providing longer context. A contig that spans the full length of an ARG, or a substantial portion of it, allows for more confident annotation. The longer sequence provides more information for the annotation tool to work with, and it allows you to inspect the flanking regions for MGEs, integrons, or other contextual features. The [Bioconductor](https://bioconductor.org/) project hosts packages for genomic annotation and analysis that can be integrated into assembly-based workflows, and the [nf-core](https://nf-co.re/docs) documentation describes community standards for reproducible assembly pipelines.

The cost of assembly is that it can fail. Metagenomic assemblies are complicated by the presence of multiple genomes at varying abundances, strain-level variation, and sequencing errors. Low-abundance organisms may not assemble into contigs long enough to contain a complete ARG, and repetitive regions can cause assembly breaks. The result is that assembly-based analysis can miss ARGs that read-based analysis would detect, particularly in complex communities or low-biomass samples.

## Sensitivity and Specificity Tradeoffs in Practice

The sensitivity of an approach is its ability to detect ARGs that are truly present. The specificity is its ability to avoid reporting ARGs that are not present. These two properties pull in opposite directions, and the choice of approach, along with the parameters you set, determines where you land on the tradeoff curve.

Read-based analysis with a low identity threshold, such as 70 percent, will detect more ARGs but will also report more false positives. A read from an unrelated gene that happens to share a conserved domain with an ARG may pass the threshold and be annotated as a resistance gene. Raising the threshold to 90 percent or higher reduces false positives but risks missing divergent variants. The [NCBI](https://www.ncbi.nlm.nih.gov/) hosts the reference databases and search tools that underpin many read-based analyses, and the choice of database version and search parameters is a documented part of the workflow.

Assembly-based analysis has a different sensitivity profile. The assembly step acts as a filter. Reads from low-abundance organisms may not assemble into contigs, and the ARGs they carry will be missed. However, the ARGs that are detected on contigs are typically supported by multiple reads and can be annotated with higher confidence. The longer context reduces the chance of misannotation, and the ability to inspect flanking regions provides additional evidence for or against a gene call.

A practical example comes from a study of raw cow and buffalo milk from the Brazilian Amazon. The researchers used shotgun metagenomic sequencing and generated over 3.1 million contigs. They found that buffalo milk had a higher abundance and diversity of ARG-associated contigs than cow milk, with 301 contigs in buffalo compared to 85 in cows. The study identified clinically relevant genes including AbaQ, ArnT, and KpnF, and found complex multi-AMR cassettes co-occurring with plasmids and viral sequences. This level of detail, including the co-occurrence of ARGs with plasmids, is only possible with assembly-based analysis. A read-based approach would have reported the presence of the genes but could not have established their genomic context or their association with mobile elements. The full study is available at [Resistome and Mobilome Profiling of Raw Cow and Buffalo Milk from the Brazilian Amazon via Shotgun Metagenomics](https://doi.org/10.3390/antibiotics15050454).

## Computational Cost and Resource Requirements

The computational cost of the two approaches differs by an order of magnitude or more. Read-based analysis is relatively inexpensive. The alignment of reads against a database is a parallelizable operation that can be completed on a standard workstation or a modest cloud instance. The memory requirements are moderate, and the runtime is measured in hours for a typical metagenomic sample.

Assembly-based analysis is substantially more demanding. The assembly step itself requires significant memory, particularly for complex communities with high microbial diversity. De Bruijn graph assemblers can require tens of gigabytes of RAM for a single sample, and the runtime can extend to days for large datasets. The subsequent annotation of contigs adds additional computational load. The [nf-core](https://nf-co.re/docs) documentation provides guidance on resource estimation for assembly pipelines, and the [Galaxy Training Network](https://training.galaxyproject.org/) offers tutorials that include practical advice on running assemblies within resource constraints.

The choice between the two approaches is therefore not purely scientific. It is also a question of what your computing environment can support. A laboratory with access to a high-performance computing cluster can reasonably run assembly-based analysis on large cohorts. A laboratory with only a desktop computer may need to use read-based analysis or to subsample their data before assembly.

The computational cost also affects reproducibility. Assembly algorithms are deterministic given the same input and parameters, but the choice of assembler, k-mer size, and other parameters can affect the output. The [Bioconductor](https://bioconductor.org/) project emphasizes reproducible research practices, and the [nf-core](https://nf-co.re/docs) framework provides versioned pipelines that lock down the software environment. If you are running assembly-based analysis, you should document the assembler version, parameters, and database versions to ensure that your results can be reproduced.

## Database Selection and Its Effect on Results

Both read-based and assembly-based approaches depend on the quality and completeness of the reference database. The [NCBI](https://www.ncbi.nlm.nih.gov/) hosts the National Database of Antibiotic Resistant Organisms (NDARO) and other sequence resources that are commonly used for ARG annotation. The Comprehensive Antibiotic Resistance Database (CARD) is another widely used resource, and it was used in a study of multidrug-resistant Escherichia coli from intensive swine production in Hungary, where researchers annotated resistance determinants using CARD and detected 82 distinct resistance determinants across 5,433 total occurrences in a cohort of 116 sequenced isolates. The study is available at [Inferred Mobility-Resolved Resistome Architecture Suggests Recurrent Co-Resistance Modules on a Conserved Chromosomal Backbone in Multidrug-Resistant Escherichia coli from Intensive Swine Production in Hungary](https://doi.org/10.3390/antibiotics15040367).

The choice of database affects both sensitivity and specificity. A database that includes a broad diversity of ARG variants will detect more genes but may also report more false positives. A curated database with high-quality sequences will be more specific but may miss novel variants. The database version matters as well. New ARGs are discovered continuously, and an outdated database will miss them.

The database also determines the types of resistance mechanisms you can detect. Most ARG databases focus on horizontally acquired resistance genes, such as those encoding beta-lactamases, aminoglycoside-modifying enzymes, and tetracycline efflux pumps. They are less comprehensive for mutation-driven resistance, where single nucleotide polymorphisms in chromosomal genes confer resistance. A study of Pseudomonas aeruginosa resistomes addressed this gap by developing PaREx, an open-source pipeline that includes 221 chromosomal genes associated with mutation-driven resistance, along with a PDC analyzer web tool for designating Pseudomonas-derived cephalosporinases. The pipeline was validated on 260 P. aeruginosa isolates and the PDC analyzer on more than 30,000 global genomes. The study is available at [PaREx: an open-source pipeline for the automated analysis of Pseudomonas aeruginosa resistomes from whole-genome sequences](https://doi.org/10.1128/aac.01326-25).

If your research question involves mutation-driven resistance, you need to ensure that your database and annotation tools can detect it. Read-based analysis against a database of acquired ARGs will miss chromosomal mutations entirely. Assembly-based analysis can detect them if you use a tool that compares contigs against a reference genome and identifies polymorphisms, but this is a different analysis than standard ARG annotation.

## Normalization Strategies and Their Implications

Quantification of ARG abundance requires normalization. The raw count of reads or contigs that map to an ARG is not directly comparable across samples because sequencing depth varies. The two common normalization strategies are to divide by total sequencing depth or to divide by the abundance of a single-copy housekeeping gene.

Normalization to total sequencing depth, often expressed as reads per kilobase per million mapped reads (RPKM) or transcripts per million (TPM), is straightforward and widely used. It allows comparison of ARG abundance across samples within a study. The limitation is that it does not account for differences in microbial load or community composition. A sample with high microbial biomass will have more total reads, and the ARG abundance will appear lower after normalization even if the proportion of resistant bacteria is the same.

Normalization to a single-copy housekeeping gene, such as rpoB, provides a measure of ARG abundance per bacterial cell. This approach was used in a study of agricultural soils from West Siberia, where the resistome load was normalized to rpoB and reported as a ratio. The study found normalized resistome loads ranging from 2.30 to 5.37, indicating moderate anthropogenic pressure, and detected class 1 integron integrase (intI1) in all 12 samples, with the intI1/rpoB ratio exceeding unity in 9 of 12 samples. The study is available at [West Siberian Soil Resistome: Mobile Antibiotic Resistance in Agricultural Microbiomes](https://doi.org/10.3390/antibiotics15050502).

The choice of normalization strategy interacts with the choice of analysis approach. Read-based analysis can normalize to total reads or to a housekeeping gene if the housekeeping gene is included in the reference database. Assembly-based analysis can normalize to rpoB by counting the number of contigs that contain rpoB and dividing the ARG count by that number. However, the assembly step can bias this normalization. If rpoB is present in multiple copies or if the assembly fails to reconstruct the full gene, the normalization will be inaccurate.

The practical implication is that you should decide on your normalization strategy before you run the analysis, and you should report it clearly in your methods. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials that include normalization steps, and the [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal offers courses on quantitative analysis of metagenomic data.

## Genomic Context and Mobility Inference

The most significant advantage of assembly-based analysis is the ability to infer genomic context. When an ARG is located on a contig, you can examine the flanking sequences for features that indicate mobility. These features include insertion sequences, transposases, integrases, and plasmid replication origins. The presence of these features suggests that the ARG can be transferred horizontally between bacteria, which has implications for the spread of resistance.

A study of agricultural soils from West Siberia used assembly-based analysis to characterize ARG-MGE associations. The researchers defined an association as co-localization within 10 kilobases on the same contig and found such associations in 11 of 12 samples. They also detected class 1 integron integrase in all samples, with the intI1/rpoB ratio exceeding unity in 9 of 12 samples. These findings indicate that the soil resistome is not static but is actively mobile, with ARGs positioned near MGEs where they can be transferred. The study is available at [West Siberian Soil Resistome: Mobile Antibiotic Resistance in Agricultural Microbiomes](https://doi.org/10.3390/antibiotics15050502).

The raw milk study from the Brazilian Amazon similarly used assembly-based analysis to identify ARG-MGE co-occurrence. The researchers found complex multi-AMR cassettes co-occurring with plasmids and widespread viral sequences dominated by Caudoviricetes. Integrons were ubiquitous in cattle and highly prevalent in buffalo samples. This level of detail is not available from read-based analysis. A read that maps to an ARG provides no information about what is adjacent to that gene in the genome.

The ability to infer mobility has practical implications for risk assessment. An ARG that is located on a plasmid or near a transposon poses a higher risk of horizontal transfer than an ARG that is located in a stable chromosomal region. If your research question involves the potential for resistance spread, assembly-based analysis is the appropriate choice. If your question is limited to the presence and abundance of ARGs, read-based analysis may suffice.

## Strain-Level Resolution and Variant Detection

Assembly-based analysis can provide strain-level resolution that is impossible with read-based analysis. When a contig spans a complete ARG, you can compare its sequence to reference alleles and determine which variant is present. This information is important for clinical and epidemiological applications, where the distinction between a narrow-spectrum and an extended-spectrum beta-lactamase can change the interpretation of the results.

The PaREx pipeline for Pseudomonas aeruginosa demonstrates the value of variant-level analysis. The pipeline was designed to detect mutation-driven resistance mechanisms, which are missed by standard ARG databases. It includes 221 chromosomal genes and can designate Pseudomonas-derived cephalosporinases (PDCs) using a web tool. The pipeline was validated on 260 isolates and the PDC analyzer on more than 30,000 global genomes. The study is available at [PaREx: an open-source pipeline for the automated analysis of Pseudomonas aeruginosa resistomes from whole-genome sequences](https://doi.org/10.1128/aac.01326-25).

Read-based analysis can also detect variants, but with lower confidence. A short read that spans a single nucleotide polymorphism can identify the variant if the polymorphism is within the read. However, the limited length of the read means that you cannot determine the phase of multiple polymorphisms. If two polymorphisms are separated by more than the read length, you cannot determine whether they are on the same allele.

Assembly-based analysis resolves this problem by reconstructing longer sequences. A contig that spans the full gene allows you to determine the complete allele sequence and to compare it against reference alleles. The tradeoff is that assembly can introduce errors, particularly in regions of high sequence similarity or in the presence of strain mixtures. A chimeric contig that joins sequences from two different strains can produce a false allele that does not exist in any single organism.

## The Role of Sequencing Depth and Coverage

The quality of both read-based and assembly-based analysis depends on sequencing depth. Higher depth provides more reads per genomic region, which improves the confidence of read-based alignments and the completeness of assemblies. Lower depth increases the risk of false negatives in both approaches.

For read-based analysis, the minimum depth depends on the abundance of the ARG you want to detect. A rare ARG present in a small fraction of the community will produce few reads, and those reads may be missed if the depth is too low. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides guidance on sequencing depth recommendations for various applications, and the [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal offers courses on experimental design for sequencing studies.

For assembly-based analysis, depth is even more critical. The assembler needs sufficient coverage to reconstruct contigs. Low-depth regions may not assemble at all, or they may assemble into short contigs that are too short to contain a complete ARG. The relationship between depth and assembly quality is nonlinear. Below a threshold, assembly quality degrades rapidly. Above the threshold, additional depth provides diminishing returns.

The practical implication is that you should assess the depth of your data before choosing an analysis approach. If your samples have low depth, read-based analysis is the safer choice. If your samples have high depth, assembly-based analysis can provide additional information. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tools for assessing sequencing depth and quality, and the [nf-core](https://nf-co.re/docs) documentation includes quality control modules that can be run before assembly.

## Quality Control and Preprocessing Steps

Both approaches require quality control of the raw reads before analysis. The standard steps include adapter trimming, quality filtering, and removal of host or contaminant sequences. These steps are documented in the [Galaxy Training Network](https://training.galaxyproject.org/) tutorials and in the [nf-core](https://nf-co.re/docs) pipeline documentation.

Adapter trimming removes the sequencing adapters that are ligated to the fragments during library preparation. If adapters are not removed, they can cause spurious alignments and assembly errors. Quality filtering removes reads with low average quality scores, which are more likely to contain sequencing errors. Host removal is important for clinical or agricultural samples where host DNA is abundant. If host reads are not removed, they consume computational resources and can interfere with assembly.

The choice of quality control parameters affects the downstream analysis. Aggressive quality filtering removes more reads but improves the average quality of the remaining reads. Lenient filtering retains more reads but includes more errors. The optimal parameters depend on the sequencing platform and the library preparation method.

For assembly-based analysis, quality control is particularly important. Sequencing errors create false k-mers in the de Bruijn graph, which can cause assembly breaks or chimeric contigs. The assembler can tolerate some errors, but excessive errors degrade the assembly quality. The [Bioconductor](https://bioconductor.org/) project provides packages for quality assessment and filtering, and the [nf-core](https://nf-co.re/docs) framework includes quality control modules that can be configured for different data types.

## Common Failure Patterns and How to Avoid Them

Several failure patterns recur in metagenomic resistome analysis. Recognizing these patterns can help you avoid them or diagnose them when they occur.

The first failure pattern is the use of an outdated or incomplete database. If your database does not include recently discovered ARGs, you will miss them regardless of whether you use read-based or assembly-based analysis. The solution is to use a current database version and to document the version in your methods. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides versioned releases of its databases, and the [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal offers guidance on database selection.

The second failure pattern is the use of an inappropriate identity threshold. A threshold that is too high will miss divergent ARGs. A threshold that is too low will report false positives. The optimal threshold depends on your research question. If you are interested in clinically relevant resistance, a higher threshold is appropriate. If you are interested in the full resistome, including distant homologs, a lower threshold may be necessary.

The third failure pattern is the failure to normalize properly. If you compare raw counts across samples without normalization, you will confound ARG abundance with sequencing depth. The solution is to choose a normalization strategy before analysis and to apply it consistently. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on normalization, and the [Bioconductor](https://bioconductor.org/) project hosts packages for differential abundance analysis.

The fourth failure pattern is the failure to validate assembly quality. Assembly-based analysis can produce chimeric contigs or fragmented assemblies that lead to incorrect conclusions. The solution is to assess assembly quality using metrics such as N50, the number of contigs, and the fraction of reads that map back to the assembly. The [nf-core](https://nf-co.re/docs) documentation describes quality assessment tools that can be integrated into assembly pipelines.

The fifth failure pattern is the failure to account for the limitations of the approach in the interpretation. Read-based analysis cannot infer genomic context. Assembly-based analysis can miss low-abundance ARGs. If you do not acknowledge these limitations, you may overinterpret your results. The [Carpentries](https://carpentries.org/lessons) lessons on data analysis and interpretation provide a foundation for critical evaluation of results.

## Decision Framework for Choosing Between Approaches

The choice between read-based and assembly-based analysis depends on four factors: your research question, your data quality, your computational resources, and your tolerance for false positives and false negatives.

If your research question is limited to the presence and abundance of known ARGs, read-based analysis is the appropriate choice. It is faster, cheaper, and less prone to assembly-related errors. It is also the better choice for low-depth data, where assembly is unlikely to produce useful contigs.

If your research question involves the genomic context of ARGs, the potential for horizontal gene transfer, or the identification of specific variants, assembly-based analysis is necessary. The additional information provided by contigs justifies the higher computational cost.

If your data quality is high, with deep sequencing and good read quality, assembly-based analysis is feasible. If your data quality is low, read-based analysis is the safer choice.

If your computational resources are limited, read-based analysis is the practical choice. If you have access to a high-performance computing cluster, assembly-based analysis is feasible.

The following decision table summarizes the tradeoffs:

| Factor | Read-Based Analysis | Assembly-Based Analysis |
|--------|---------------------|------------------------|
| Sensitivity for low-abundance ARGs | Higher, because individual reads are counted | Lower, because low-abundance organisms may not assemble |
| Specificity for ARG identity | Lower, because short reads cannot distinguish close variants | Higher, because longer contigs provide more information |
| Genomic context and mobility inference | Not possible | Possible, including ARG-MGE co-localization |
| Computational cost | Low to moderate | High, particularly for complex communities |
| Suitability for low-depth data | Good | Poor, assembly requires sufficient coverage |
| Variant-level resolution | Limited, depends on read length | Higher, full gene sequences can be compared |
| Reproducibility | High, with documented parameters | Moderate, assembly algorithms can vary |

## At a Glance: Key Differences and Decision Criteria

| Criterion | Read-Based Analysis | Assembly-Based Analysis |
|-----------|---------------------|------------------------|
| Primary output | ARG presence and abundance per read | ARG presence, abundance, and genomic context per contig |
| Computational requirement | Moderate CPU and memory | High CPU and memory, particularly for complex communities |
| Best suited for | Large cohorts, low-depth data, presence/absence screening | Small cohorts, high-depth data, mobility and variant analysis |
| Main limitation | Cannot infer genomic context or mobility | May miss low-abundance ARGs due to assembly failure |
| Typical runtime | Hours per sample | Hours to days per sample |
| Database dependence | High, identity thresholds determine sensitivity | High, but contig length provides additional annotation evidence |
| Recommended when | Research question is presence and abundance | Research question involves HGT, MGEs, or variants |

## Practical Workflow for Read-Based Resistome Analysis

A read-based workflow begins with quality control of the raw reads. You should trim adapters, filter low-quality reads, and remove host sequences. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials for each of these steps, and the [nf-core](https://nf-co.re/docs) documentation describes quality control modules that can be run as part of a pipeline.

After quality control, you align the reads against your ARG database. The alignment tool and parameters depend on the database and the analysis goal. You should document the tool version, the database version, and the alignment parameters in your methods. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides access to reference databases and alignment tools, and the [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal offers courses on sequence alignment.

After alignment, you process the alignment results to produce a table of ARG abundances. This step involves filtering alignments by identity and coverage thresholds, counting the reads that map to each ARG, and normalizing the counts. The normalization strategy should be chosen before analysis and applied consistently across samples.

The final step is interpretation. You should assess whether the detected ARGs are consistent with the expected biology of your sample type. If you detect an ARG that is unexpected, you should verify the alignment manually before reporting it. The [Carpentries](https://carpentries.org/lessons) lessons on data analysis provide guidance on critical evaluation of results.

## Practical Workflow for Assembly-Based Resistome Analysis

An assembly-based workflow begins with the same quality control steps as a read-based workflow. After quality control, you assemble the reads into contigs. The choice of assembler depends on your data type. MEGAHIT is fast and memory-efficient, making it suitable for large datasets. metaSPAdes produces higher-quality assemblies but requires more memory. metaFlye is designed for long-read data but can also handle short reads.

After assembly, you assess the quality of the assembly. The N50 statistic, which is the length at which half of the assembled bases are in contigs of that length or longer, provides a summary measure of assembly contiguity. You should also check the number of contigs and the fraction of reads that map back to the assembly. The [nf-core](https://nf-co.re/docs) documentation describes quality assessment tools for assemblies.

After quality assessment, you annotate the contigs for ARGs. The annotation tool compares the contigs against an ARG database and reports matches. You should set identity and coverage thresholds that are appropriate for your research question. The [Bioconductor](https://bioconductor.org/) project provides packages for functional annotation of genomic sequences.

After annotation, you analyze the genomic context of the detected ARGs. You should examine the flanking sequences for MGEs, integrases, and other mobility features. This analysis can be done manually or with automated tools. The study of agricultural soils from West Siberia used a 10-kilobase window around each ARG to define ARG-MGE associations, and you can adopt a similar approach. The study is available at [West Siberian Soil Resistome: Mobile Antibiotic Resistance in Agricultural Microbiomes](https://doi.org/10.3390/antibiotics15050502).

## Records and Measurements to Maintain

Regardless of the approach you choose, you should maintain detailed records of your analysis. The records should include the version of every tool and database used, the parameters for every step, and the quality metrics for every sample. These records are essential for reproducibility and for troubleshooting when results are unexpected.

The [nf-core](https://nf-co.re/docs) framework provides a structured approach to recording analysis parameters. Each pipeline run produces a log file that records the tool versions, parameters, and inputs. The [Galaxy Training Network](https://training.galaxyproject.org/) provides similar functionality through its history system, which records every step of an analysis.

You should also record the quality metrics for each sample. These metrics include the number of raw reads, the number of reads after quality control, the fraction of reads that map to the ARG database, and the assembly statistics if you are using assembly-based analysis. These metrics allow you to identify samples that are outliers and to assess whether the analysis was successful.

The [Carpentries](https://carpentries.org/lessons) lessons on reproducible research provide guidance on record keeping and documentation. The lessons emphasize the importance of documenting also what you did but also why you made the choices you did.

## Common Failure Patterns in Resistome Analysis

The first common failure pattern is the use of a database that does not match the research question. If you are studying Pseudomonas aeruginosa and use a database that does not include mutation-driven resistance mechanisms, you will miss the most relevant resistance determinants. The PaREx study addressed this gap by developing a pipeline that includes 221 chromosomal genes associated with mutation-driven resistance. The study is available at [PaREx: an open-source pipeline for the automated analysis of Pseudomonas aeruginosa resistomes from whole-genome sequences](https://doi.org/10.1128/aac.01326-25).

The second common failure pattern is the use of an identity threshold that is too high or too low. A threshold that is too high will miss divergent ARGs. A threshold that is too low will report false positives. The optimal threshold depends on the database and the research question. You should test multiple thresholds on a subset of your data to determine the sensitivity and specificity of each.

The third common failure pattern is the failure to normalize properly. If you compare raw counts across samples, you will confound ARG abundance with sequencing depth. The solution is to choose a normalization strategy before analysis and to apply it consistently. The study of agricultural soils from West Siberia normalized to rpoB, a single-copy housekeeping gene, and reported the resistome load as a ratio. The study is available at [West Siberian Soil Resistome: Mobile Antibiotic Resistance in Agricultural Microbiomes](https://doi.org/10.3390/antibiotics15050502).

The fourth common failure pattern is the failure to validate assembly quality. Assembly-based analysis can produce chimeric contigs or fragmented assemblies that lead to incorrect conclusions. The solution is to assess assembly quality using metrics such as N50 and the fraction of reads that map back to the assembly. The [nf-core](https://nf-co.re/docs) documentation describes quality assessment tools for assemblies.

The fifth common failure pattern is the failure to account for the limitations of the approach in the interpretation. Read-based analysis cannot infer genomic context. Assembly-based analysis can miss low-abundance ARGs. If you do not acknowledge these limitations, you may overinterpret your results. The [Carpentries](https://carpentries.org/lessons) lessons on data analysis and interpretation provide a foundation for critical evaluation of results.

## Limitations of Both Approaches

Both read-based and assembly-based approaches have inherent limitations that you should acknowledge in your interpretation.

Read-based analysis cannot distinguish between ARGs that are located on the chromosome and ARGs that are located on plasmids. This distinction matters for risk assessment, because plasmid-borne ARGs are more likely to be transferred horizontally. Read-based analysis also cannot determine whether multiple ARGs are located in the same genomic region, which is relevant for understanding co-resistance.

Assembly-based analysis can provide this context, but it is limited by the completeness of the assembly. Low-abundance organisms may not assemble into contigs, and the ARGs they carry will be missed. The assembly may also produce chimeric contigs that join sequences from different organisms, leading to false associations between ARGs and MGEs.

Both approaches are limited by the reference database. ARGs that are not in the database will not be detected, regardless of the approach. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides access to reference databases, and the [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal offers guidance on database selection and versioning.

Both approaches are also limited by the sequencing platform. Short-read platforms such as Illumina produce reads that are typically 150 base pairs in length. These reads are sufficient for read-based analysis and for assembly of simple communities, but they may not be sufficient for assembly of complex communities or for resolving repetitive regions. Long-read platforms such as Oxford Nanopore and PacBio produce longer reads that improve assembly quality but have higher error rates.

## Welfare and Safety Context for Agricultural Samples

When the resistome analysis involves agricultural samples, the results have implications for animal welfare and public health. The detection of ARGs in livestock or agricultural environments indicates the presence of selective pressure from antibiotic use. The study of multidrug-resistant Escherichia coli from intensive swine production in Hungary found that 78 of 203 isolates (38.4 percent) met the definition of multidrug resistance, with marked between-farm variation. The study is available at [Inferred Mobility-Resolved Resistome Architecture Suggests Recurrent Co-Resistance Modules on a Conserved Chromosomal Backbone in Multidrug-Resistant Escherichia coli from Intensive Swine Production in Hungary](https://doi.org/10.3390/antibiotics15040367).

The detection of ARGs in raw milk from cows and buffalo in the Brazilian Amazon highlights the food safety risk associated with unpasteurized dairy consumption. The study found that buffalo milk had a higher abundance and diversity of ARG-associated contigs than cow milk, and identified clinically relevant genes including AbaQ, ArnT, and KpnF. The study is available at [Resistome and Mobilome Profiling of Raw Cow and Buffalo Milk from the Brazilian Amazon via Shotgun Metagenomics](https://doi.org/10.3390/antibiotics15050454).

The detection of ARGs in agricultural soils from West Siberia indicates that soil microbiomes are reservoirs of resistance genes and mobile genetic elements. The study found ARG-MGE co-localizations in 11 of 12 samples and detected class 1 integron integrase in all samples. The study is available at [West Siberian Soil Resistome: Mobile Antibiotic Resistance in Agricultural Microbiomes](https://doi.org/10.3390/antibiotics15050502).

If your analysis detects ARGs that are associated with mobile genetic elements or that are clinically relevant, you should consider the implications for farm management and public health. The presence of mobile ARGs suggests that resistance can spread between bacteria, potentially including pathogens. The presence of clinically relevant ARGs in agricultural products suggests a food safety risk.

## Professional Escalation Criteria

You should escalate your findings to a qualified professional when the results have implications beyond your immediate research question. The following situations warrant escalation:

If you detect ARGs that confer resistance to antibiotics that are critically important for human medicine, you should consult with a clinical microbiologist or an infectious disease specialist. The detection of such ARGs in agricultural samples may indicate a public health risk.

If you detect ARGs that are associated with mobile genetic elements in a context where horizontal gene transfer is likely, you should consult with a molecular epidemiologist. The potential for resistance spread has implications for infection control and farm management.

If your results are inconsistent with the expected biology of your sample type, you should consult with a bioinformatics specialist before reporting the results. The inconsistency may indicate a technical problem with the analysis, such as a database error or a misconfiguration of the pipeline.

If you are analyzing clinical samples and the results will inform patient care, you should follow the reporting protocols of your institution. The [NCBI](https://www.ncbi.nlm.nih.gov/) provides resources for clinical sequence analysis, and the [EMBL-EBI Training](https://www.ebi.ac.uk/training) portal offers courses on clinical bioinformatics.

## Frequently Asked Questions

### What is the main difference between read-based and assembly-based resistome analysis?

Read-based analysis aligns individual sequencing reads directly against a database of known ARG sequences. It is fast and computationally efficient, but it cannot provide information about the genomic context of the detected genes. Assembly-based analysis first assembles the reads into longer contigs and then annotates those contigs for ARGs. It is slower and more computationally demanding, but it allows you to determine whether an ARG is located near a mobile genetic element, whether multiple ARGs are clustered together, and whether the ARG is likely chromosomal or plasmid-borne.

### Which approach is more sensitive for detecting ARGs?

Read-based analysis is generally more sensitive for detecting ARGs from low-abundance organisms. Individual reads from a rare organism can align to the database even if the organism does not assemble into a contig. Assembly-based analysis can miss these ARGs because low-abundance organisms may not produce enough reads to assemble into contigs. However, the ARGs that are detected by assembly-based analysis are typically supported by more evidence and can be annotated with higher confidence.

### Which approach is more specific for identifying ARG variants?

Assembly-based analysis is generally more specific for identifying ARG variants. A contig that spans the full length of a gene provides more sequence information than a short read, allowing the annotation tool to distinguish between closely related variants. Read-based analysis can identify variants if the polymorphism is within the read, but it cannot determine the phase of multiple polymorphisms that are separated by more than the read length.

### Can read-based analysis detect mobile genetic elements?

Read-based analysis can detect the presence of MGE sequences if they are included in the reference database, but it cannot determine whether an ARG is located near an MGE. The association between an ARG and an MGE requires genomic context, which is only available from assembly-based analysis. If your research question involves the potential for horizontal gene transfer, you should use assembly-based analysis.

### What computational resources do I need for each approach?

Read-based analysis can be run on a standard workstation or a modest cloud instance. The memory requirements are moderate, and the runtime is measured in hours for a typical metagenomic sample. Assembly-based analysis requires substantially more resources. The assembly step can require tens of gigabytes of RAM, and the runtime can extend to days for large datasets. If you have access to a high-performance computing cluster, assembly-based analysis is feasible. If not, read-based analysis is the practical choice.

### How does sequencing depth affect the choice of approach?

Sequencing depth is a critical factor. Read-based analysis can work with lower depth because individual reads are counted. Assembly-based analysis requires sufficient coverage to reconstruct contigs. Below a threshold, assembly quality degrades rapidly, and low-abundance organisms may not assemble at all. If your samples have low depth, read-based analysis is the safer choice. If your samples have high depth, assembly-based analysis can provide additional information.

### What normalization strategy should I use?

The two common strategies are normalization to total sequencing depth and normalization to a single-copy housekeeping gene such as rpoB. Normalization to total depth allows comparison of ARG abundance across samples within a study. Normalization to rpoB provides a measure of ARG abundance per bacterial cell. The choice depends on your research question. If you are interested in the proportion of resistant bacteria, normalization to rpoB is appropriate. If you are interested in the total ARG load, normalization to total depth is appropriate.

### How should I report my methods to ensure reproducibility?

You should report the version of every tool and database used, the parameters for every step, and the quality metrics for every sample. The [nf-core](https://nf-co.re/docs) framework provides a structured approach to recording analysis parameters, and the [Galaxy Training Network](https://training.galaxyproject.org/) provides similar functionality through its history system. You should also report the normalization strategy and the identity and coverage thresholds used for ARG annotation.

## Related Bioinformatics Guides

- [Evaluating Metagenomic Assembly Tools: A Benchmarking Framework for Short-Read and Long-Read Data](/knowledge/bioinformatics/evaluating-metagenomic-assembly-tools-a-benchmarking-framework-for-short-read-and-long-read-data)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Longitudinal Microbiome Data Analysis: Methods and Best Practices](/knowledge/bioinformatics/longitudinal-microbiome-data-analysis-methods-and-best-practices)
- [Mass Spectrometry-Based Proteomics: Data Analysis Pipelines and Tools](/knowledge/bioinformatics/mass-spectrometry-based-proteomics-data-analysis-pipelines-and-tools)
- [Metagenomics Data Analysis: From Raw Reads to Biological Insights](/knowledge/bioinformatics/metagenomics-data-analysis-from-raw-reads-to-biological-insights)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [West Siberian Soil Resistome: Mobile Antibiotic Resistance in Agricultural Microbiomes.](https://doi.org/10.3390/antibiotics15050502). 2026.
- [Resistome and Mobilome Profiling of Raw Cow and Buffalo Milk from the Brazilian Amazon via Shotgun Metagenomics.](https://doi.org/10.3390/antibiotics15050454). 2026.
- [Inferred Mobility-Resolved Resistome Architecture Suggests Recurrent Co-Resistance Modules on a Conserved Chromosomal Backbone in Multidrug-Resistant &lt,i&gt,Escherichia coli&lt,/i&gt, from Intensive Swine Production in Hungary.](https://doi.org/10.3390/antibiotics15040367). 2026.
- [Environmental Altitude and Host Genetics Shape Divergent Microbiota and a Conserved Resistome in Porcine Intestinal Niches.](https://doi.org/10.3390/microorganisms14040832). 2026.
- [&lt,i&gt,Pa&lt,/i&gt,REx: an open-source pipeline for the automated analysis of &lt,i&gt,Pseudomonas aeruginosa&lt,/i&gt, resistomes from whole-genome sequences.](https://doi.org/10.1128/aac.01326-25). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.