# MultiQC for RNA-seq: How to Aggregate and Visualize QC Metrics Across Multiple Samples


## Key Takeaways

- MultiQC aggregates quality control (QC) metrics from multiple RNA-seq analysis steps (e.g., FastQC, STAR alignment, featureCounts quantification) into a single, interactive HTML report, enabling efficient assessment of large cohorts by presenting per-sample data side-by-side.
- Examination of per-base quality scores, GC content, and adapter contamination via FastQC outputs within MultiQC is critical for identifying samples with uniformly low quality or adapter enrichment, guiding decisions on trimming or exclusion prior to downstream analysis.
- STAR alignment metrics, particularly the uniquely mapped percentage, are vital for detecting issues such as sample contamination, index hopping, or poor reference quality, with rates below 70% warranting investigation.
- featureCounts quantification completeness, assessed by the assigned fraction of reads, reveals potential problems with annotation, alignment, or library preparation, where fractions below 60% necessitate further scrutiny.
- MultiQC facilitates the detection of batch effects by visualizing sample clustering based on sequencing run, lane, or preparation date across various QC metrics, allowing researchers to distinguish technical artifacts from biological variation.
- The general statistics table within MultiQC provides a consolidated, sortable overview of key metrics (e.g., raw reads, alignment rate, assigned reads) for all samples, enabling rapid outlier detection and serving as a documented record of QC decisions.

---

RNA sequencing produces per-sample quality control files that quickly become unmanageable as cohort sizes grow. A typical experiment generates separate reports from read quality assessment, alignment, and quantification steps, and inspecting each file individually across dozens or hundreds of samples is impractical. MultiQC solves this problem by scanning output directories, parsing the relevant log files, and producing a single interactive HTML report that places all samples side by side. This article explains how to run MultiQC on FastQC, STAR, and featureCounts outputs, how to interpret the aggregated report for large cohorts, and how to use the results to make defensible decisions about sample inclusion, trimming, and downstream analysis.

The practical outcome for researchers is a repeatable quality control workflow that identifies outlier samples, reveals batch effects, and documents the evidence used to retain or exclude data before differential expression analysis. The workflow described here follows the pattern used in published RNA-seq studies, where FastQC and MultiQC provide the initial quality gate, STAR performs alignment, and featureCounts generates count matrices for downstream statistical testing.

## At a Glance

| Decision Point | What MultiQC Shows | Action to Take |
| --- | --- | --- |
| Read quality before trimming | Per-base quality scores, GC content, adapter contamination across all samples | Identify samples with uniformly low quality or adapter enrichment and trim or exclude them |
| Alignment performance | STAR mapping rates, uniquely mapped reads, multi-mapping reads per sample | Flag samples with low unique mapping rates and investigate causes before proceeding |
| Quantification completeness | featureCounts assigned, unassigned, and multi-mapping counts per sample | Confirm that most reads are assigned to features and check for systematic assignment failures |
| Batch and lane effects | Clustering of samples by sequencing run, lane, or preparation date in summary plots | Track sample metadata alongside QC metrics to distinguish biological variation from technical artifacts |
| Outlier detection | Samples that fall outside the distribution of the cohort on any metric | Review raw data for those samples and decide whether to re-sequence, trim more aggressively, or exclude |

## The Problem of Scale in RNA-seq Quality Control

RNA-seq experiments generate multiple quality control outputs for every sample. The raw sequencing data arrive as FASTQ files, and the first assessment step typically runs FastQC to evaluate per-base sequence quality, GC content, sequence duplication levels, overrepresented sequences, and adapter content. After trimming and alignment, additional metrics emerge from the aligner, including total reads, uniquely mapped reads, multi-mapping reads, and reads unmapped for various reasons. Quantification tools then report how many reads were assigned to genes or transcripts and how many were left unassigned.

For a small experiment with six samples, inspecting each FastQC HTML report and each alignment log manually is tedious but feasible. For a study with dozens or hundreds of samples, the approach collapses. The published RNA-seq datasets that use MultiQC reflect this reality. A transcriptome dataset from beef heifers at weaning used FastQC and MultiQC for quality control before STAR alignment and DESeq2 analysis, with samples sequenced on the Illumina NovaSeq platform [<a href="#ref-1">1</a>]. An Arabidopsis dataset examining the response of imbibed seeds to nitrogen-containing compounds used FastQC for read quality checks, Salmon for quasi-mapping alignment, and MultiQC to merge mapping rate results into a single report across 45 samples [<a href="#ref-2">2</a>]. A study of human cell lines infected with influenza A virus applied a pipeline that included quality control with FastQC and MultiQC, adapter and quality trimming with Cutadapt, filtering to the influenza genome with STAR, and transcript quantification with Salmon [<a href="#ref-3">3</a>].

These examples share a common structure. MultiQC sits at the aggregation point where per-sample metrics become cohort-level information. The tool does not replace FastQC, STAR, or featureCounts. It collects their outputs and presents them in a unified view so that a researcher can assess the entire experiment in one pass.

## What MultiQC Aggregates and How It Works

MultiQC searches a specified directory for log files and report files produced by supported bioinformatics tools. It recognizes the output formats of FastQC, STAR, featureCounts, Salmon, Cutadapt, and many other tools, parses the relevant metrics, and generates a single HTML report with interactive plots and tables. The report groups metrics by tool, so all FastQC modules appear together, all STAR alignment statistics appear together, and all featureCounts results appear together.

The aggregation is particularly valuable for spotting patterns that are invisible when examining one sample at a time. A single FastQC report might show a modest drop in quality at the 3-prime end of reads, which is common and often acceptable. When the same pattern appears across all samples from one sequencing lane but not from another lane, the cause is likely technical instead of biological. MultiQC makes this comparison straightforward because it plots all samples on the same axes.

The tool also generates a general statistics table that combines key metrics from multiple tools into one view. This table typically includes the number of reads per sample, the percentage of reads that passed filtering, the alignment rate, and the number of reads assigned to features. Sorting this table by any column reveals outliers immediately.

## Running MultiQC on FastQC Outputs

The first step in the workflow is to run FastQC on all raw FASTQ files. FastQC produces one HTML report and one ZIP archive per sample. The HTML report is for human inspection, and the ZIP archive contains the underlying data in a machine-readable format. MultiQC reads the data from the ZIP archives, so both outputs should be preserved.

After FastQC has processed every sample, run MultiQC from the directory that contains the FastQC output files. The command is simple:

```
multiqc .
```

MultiQC scans the current directory recursively, finds all FastQC ZIP archives, and produces a report named `multiqc_report.html` along with a data directory containing the parsed metrics in text format. The data directory is useful for downstream analysis because it allows the metrics to be imported into R or Python for statistical testing.

The FastQC section of the MultiQC report displays the per-base sequence quality plot for all samples overlaid on one graph. This view shows whether quality degrades at read ends across the cohort and whether any sample deviates from the general pattern. The per-sequence GC content plot reveals whether the GC distribution matches the expected distribution for the organism under study. A sample with an abnormal GC profile may indicate contamination, library preparation problems, or a biological anomaly.

The sequence duplication level plot identifies samples with excessive duplication, which can arise from low input RNA, over-amplification during library preparation, or sequencing depth that saturates the transcriptome. The overrepresented sequences table flags adapter contamination and other artifacts. When adapters appear at high frequency, trimming is required before alignment.

## Interpreting FastQC Metrics Across a Cohort

The interpretation of FastQC metrics changes when viewed across a cohort instead of for a single sample. A per-base quality score that drops below Q30 at the final bases of reads is common with Illumina sequencing and does not automatically disqualify a sample. The relevant question is whether the drop is consistent across samples and whether downstream tools can tolerate it. Trimming tools such as Cutadapt remove low-quality bases and adapter sequences, and the published influenza dataset demonstrates this workflow, where FastQC and MultiQC quality control preceded adapter and quality trimming with Cutadapt [<a href="#ref-3">3</a>].

GC content requires organism-specific interpretation. The expected GC distribution depends on the species, the transcriptome composition, and whether the library is stranded. A sample that deviates sharply from the cohort distribution warrants investigation. The deviation could indicate contamination with another species, a problem with the library preparation, or a genuine biological difference such as the expression of a highly GC-biased set of genes.

Sequence duplication levels deserve particular attention in RNA-seq because duplication arises from both biological and technical sources. Highly expressed genes produce many identical reads, which is expected. Excessive duplication across the entire library, however, suggests that the input RNA was degraded or that the PCR amplification steps introduced bias. The MultiQC report shows the duplication curve for each sample, and samples with curves that rise steeply and plateau at high levels may need to be flagged.

## Running MultiQC on STAR Alignment Outputs

STAR is a splice-aware aligner commonly used in RNA-seq workflows. It produces a Log.final.out file for each sample that contains the key alignment statistics, including the total number of reads, the number of reads mapped to multiple loci, the number mapped to too many loci, the number unmapped due to short reads, and the percentage of reads uniquely mapped. STAR also produces a splice junctions file and a BAM file of aligned reads.

MultiQC parses the STAR Log.final.out files and presents the alignment statistics in a table and a series of bar plots. The most important metric is the uniquely mapped percentage, which indicates what fraction of reads aligned to exactly one location in the reference genome. Low unique mapping rates can result from repetitive genomes, contaminated samples, poor reference quality, or reads that are too short to be placed uniquely.

The published beef heifer dataset used STAR for read alignment after FastQC and MultiQC quality control [<a href="#ref-1">1</a>]. The Arabidopsis dataset used Salmon for quasi-mapping alignment and MultiQC to merge the mapping rate results into a single report [<a href="#ref-2">2</a>]. Both approaches are valid, and the choice between STAR and Salmon depends on whether the analysis requires splice-aware alignment to a genome or transcript-level quantification against a reference transcriptome.

When running MultiQC on STAR outputs, the command remains the same:

```
multiqc .
```

MultiQC finds the STAR logs in the current directory and adds the alignment metrics to the report. The general statistics table now includes the STAR alignment rate alongside the FastQC metrics, giving a single view of read quality and alignment performance.

## Interpreting STAR Alignment Metrics Across a Cohort

The uniquely mapped percentage is the first alignment metric to examine in the MultiQC report. For a typical RNA-seq experiment against a well-annotated reference genome, the uniquely mapped percentage should be high, often above 80 percent. The published breast cancer workflow using Salmon reported mapping rates ranging from 92.80 percent to 94.13 percent across six libraries [<a href="#ref-4">4</a>]. These figures provide a reference point for what is achievable with good-quality data and a suitable reference.

Samples with uniquely mapped percentages well below the cohort median require investigation. Possible causes include sample contamination, index hopping during multiplexed sequencing, a mismatch between the sample species and the reference genome, or RNA degradation that produced fragments too short to align uniquely. The MultiQC report also shows the percentage of reads mapped to multiple loci, which is elevated in genomes with high repeat content or in samples with contamination from closely related species.

The number of reads unmapped due to short length is another useful metric. If a large fraction of reads are too short to map, the library may have suffered degradation, or the trimming step may have been too aggressive. The published Zika virus study trimmed reads with a Phred quality threshold of 30 and a minimum length of 70 base pairs before alignment with STAR [<a href="#ref-5">5</a>]. These parameters represent a reasonable starting point, but the optimal values depend on read length, insert size, and the reference genome.

## Running MultiQC on featureCounts Outputs

featureCounts assigns aligned reads to genomic features such as genes or exons. It produces a summary file for each sample that reports the number of reads successfully assigned to features, the number unassigned due to no features, the number unassigned due to ambiguity, and the number unassigned due to other reasons. These summary files are the input that MultiQC parses for the quantification section of the report.

The featureCounts summary appears in the MultiQC report as a stacked bar plot showing the proportion of reads in each assignment category for every sample. The assigned fraction is the key metric. A high assigned fraction indicates that the annotation is appropriate for the data and that the alignment placed reads within annotated features. A low assigned fraction suggests problems with the annotation, the alignment, or the library preparation.

The published RNA-seq workflows that use featureCounts typically run it after STAR alignment. The Zika virus study used STAR for alignment and featureCounts for counting, with Ensembl release 109 as the annotation source [<a href="#ref-5">5</a>]. The beef heifer dataset used STAR for alignment and DESeq2 for differential expression, following the same general structure [<a href="#ref-1">1</a>]. MultiQC aggregates the featureCounts summaries so that the assigned fractions for all samples appear in one table.

## Interpreting featureCounts Metrics Across a Cohort

The assigned fraction should be consistent across samples from the same experiment. If most samples show 70 to 80 percent assigned reads and one sample shows 40 percent, that sample likely has a problem. The cause could be contamination, a library preparation failure, or an annotation mismatch. The MultiQC report makes this inconsistency visible immediately.

The unassigned due to ambiguity category reflects reads that overlap multiple features and cannot be confidently assigned to a single gene. This fraction depends on the annotation density and the read length. A high ambiguity fraction may indicate that the annotation includes many overlapping transcripts or that the reads are longer than the average feature spacing.

The unassigned due to no features category includes reads that align to intergenic regions or to features not present in the annotation. A high fraction in this category can indicate that the reference annotation is incomplete, that the sample contains sequences from another organism, or that the alignment placed reads in regions not covered by the annotation.

## The General Statistics Table as a Cohort Overview

The general statistics table in the MultiQC report combines the most important metrics from all tools into a single sortable table. Each row represents one sample, and each column represents one metric. The table typically includes the number of raw reads, the percentage of reads that passed quality filtering, the alignment rate, and the number of reads assigned to features.

Sorting this table by any column reveals the distribution of that metric across the cohort. Sorting by the number of reads shows whether all samples achieved the target sequencing depth. Sorting by the alignment rate shows whether any sample has an unusually low mapping percentage. Sorting by the assigned reads shows whether quantification succeeded uniformly.

The general statistics table also serves as a record of the quality control decisions made during the analysis. A researcher can export the table, annotate it with decisions about trimming and sample inclusion, and include it as supplementary material in a publication. This documentation supports the reproducibility requirements emphasized in bioinformatics training resources [<a href="#ref-6">6</a>][<a href="#ref-7">7</a>].

## Using MultiQC to Detect Batch Effects

Batch effects are systematic technical differences between groups of samples processed at different times, on different instruments, or by different operators. They are a major concern in RNA-seq because they can confound biological comparisons. MultiQC does not perform statistical tests for batch effects, but it provides the visualization needed to detect them.

The per-base quality plots, GC content plots, and alignment metrics in the MultiQC report can all reveal batch structure. If samples from sequencing run A cluster together on the GC content plot and samples from run B cluster separately, the difference is likely technical. The same logic applies to alignment rates and assigned fractions.

To use MultiQC for batch effect detection, the researcher must know which samples belong to which batch. This information comes from the sample metadata, not from MultiQC itself. The researcher can color the MultiQC plots by metadata categories using the configuration options, or can export the parsed metrics and plot them in R with batch labels.

The published workflows that use MultiQC typically combine it with principal component analysis for batch effect detection. The Zika virus study used PCA to show clear stage separation between experimental groups, with one sham outlier retained on PC2 [<a href="#ref-5">5</a>]. The breast cancer workflow used DESeq2 and visualization tools for exploratory analysis after quantification [<a href="#ref-4">4</a>]. MultiQC provides the quality gate before these downstream analyses, and the batch structure visible in the QC metrics should be considered when interpreting the PCA results.

## Practical Workflow for Running MultiQC

The following steps describe a complete MultiQC workflow for a typical RNA-seq experiment. The workflow assumes that raw FASTQ files have been generated and that the researcher has access to a reference genome and annotation.

**Step 1: Run FastQC on all raw FASTQ files.**

Run FastQC on every sample and store the HTML and ZIP outputs in a dedicated directory. The command for a single file is `fastqc sample.fastq.gz`, and the command for multiple files is `fastqc sample1.fastq.gz sample2.fastq.gz`. FastQC can also process a directory of files.

**Step 2: Run MultiQC on the FastQC outputs.**

Navigate to the directory containing the FastQC outputs and run `multiqc .`. The resulting report shows the read quality metrics for all samples. Inspect the per-base quality, GC content, duplication, and adapter content plots. Identify samples that deviate from the cohort and decide whether trimming is needed.

**Step 3: Trim adapters and low-quality bases.**

If the FastQC report shows adapter contamination or poor quality at read ends, trim the reads with Cutadapt or a similar tool. The published influenza dataset used Cutadapt for adapter and quality trimming after the FastQC and MultiQC quality control step [<a href="#ref-3">3</a>]. The Zika virus study used Trim Galore with a Phred threshold of 30 and a minimum length of 70 base pairs [<a href="#ref-5">5</a>]. The trimming parameters should be recorded for reproducibility.

**Step 4: Run FastQC on the trimmed reads and run MultiQC again.**

Repeat the FastQC and MultiQC steps on the trimmed reads to confirm that the trimming resolved the problems identified in the raw data. The second MultiQC report provides evidence that the trimming step was effective.

**Step 5: Align the trimmed reads with STAR.**

Run STAR on each sample to produce aligned reads and the Log.final.out alignment statistics. The STAR command requires the reference genome index, the read files, and output file names. The published workflows use STAR for splice-aware alignment to a reference genome [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>][<a href="#ref-5">5</a>].

**Step 6: Run MultiQC on the STAR outputs.**

Run `multiqc .` again to add the STAR alignment metrics to the report. Examine the uniquely mapped percentage, the multi-mapping percentage, and the unmapped read categories. Flag samples with alignment rates well below the cohort median.

**Step 7: Quantify reads with featureCounts.**

Run featureCounts on the STAR BAM files to assign reads to genomic features. The featureCounts command requires the annotation file and the BAM files. The output includes the count matrix and the summary files.

**Step 8: Run MultiQC on the featureCounts outputs.**

Run `multiqc .` a final time to include the featureCounts assignment statistics in the report. The general statistics table now contains the complete quality control picture for every sample.

**Step 9: Document the decisions.**

Export the general statistics table and record the decisions made at each step. Which samples were trimmed? Which samples were excluded? What thresholds were applied? This documentation supports the reproducibility of the analysis and provides the evidence needed for publication.

## Records and Measurements to Maintain

The MultiQC report is a record of the quality control process, but it does not capture every decision made during the analysis. The researcher should maintain additional records that document the reasoning behind each decision.

The sample metadata should include the sample identifier, the biological condition, the sequencing run, the lane, the library preparation date, and any other relevant batch information. This metadata is essential for interpreting the MultiQC plots and for detecting batch effects.

The trimming parameters should be recorded for each sample. If different samples require different trimming parameters, the rationale should be documented. The published workflows typically apply the same trimming parameters to all samples to avoid introducing sample-specific biases [<a href="#ref-3">3</a>][<a href="#ref-5">5</a>].

The alignment parameters should be recorded, including the reference genome version, the annotation version, and any STAR options that differ from the defaults. The published workflows use specific reference versions, such as GRCh38 for human data [<a href="#ref-3">3</a>][<a href="#ref-4">4</a>] and GRCm39 for mouse data [<a href="#ref-5">5</a>].

The featureCounts parameters should be recorded, including the annotation file, the feature type, and the attribute type. These parameters determine how reads are assigned to features and can affect the downstream differential expression results.

The MultiQC report itself should be archived with the analysis. The HTML report is self-contained and can be opened in any web browser. The data directory contains the parsed metrics in text format and can be imported into R or Python for further analysis.

## Common Failure Patterns in MultiQC Reports

Several failure patterns recur in MultiQC reports from RNA-seq experiments. Recognizing these patterns allows the researcher to diagnose problems quickly and take corrective action.

**Uniformly low per-base quality across all samples.** This pattern indicates a problem with the sequencing run itself, such as a failed reagent cartridge or an instrument malfunction. The affected samples may need to be re-sequenced. If only one sample shows low quality, the problem is more likely sample-specific, such as degradation during RNA extraction or library preparation.

**Adapter contamination in a subset of samples.** Adapter contamination appears as overrepresented sequences in the FastQC report. The affected samples require trimming before alignment. The published influenza dataset used Cutadapt for this purpose [<a href="#ref-3">3</a>]. If adapter contamination persists after trimming, the trimming parameters may need to be adjusted.

**Low unique mapping rate in one sample.** A single sample with a low unique mapping rate may be contaminated with another species, may have experienced index hopping during multiplexed sequencing, or may have been mislabeled. The researcher should check the sample metadata and, if necessary, examine the aligned reads for evidence of contamination.

**Low unique mapping rate across all samples.** This pattern suggests a systematic problem with the reference genome, the annotation, or the alignment parameters. The reference may be inappropriate for the species, or the reads may be too short to map uniquely. The researcher should verify the reference genome version and consider whether a different aligner or different parameters would improve the mapping rate.

**Inconsistent assigned fractions in featureCounts.** If most samples show a similar assigned fraction and one sample deviates sharply, the outlier sample likely has a problem. The cause could be contamination, a library preparation failure, or an annotation mismatch. The researcher should investigate the sample before including it in the downstream analysis.

**Batch structure visible in multiple metrics.** If samples cluster by sequencing run, lane, or preparation date across multiple QC metrics, a batch effect is present. The researcher should record the batch structure and consider including batch as a covariate in the differential expression model.

## Limitations of MultiQC

MultiQC is an aggregation and visualization tool, not a decision-making tool. It does not tell the researcher which samples to exclude or which thresholds to apply. The thresholds for quality metrics depend on the organism, the sequencing platform, the library preparation method, and the downstream analysis requirements.

MultiQC also does not perform statistical tests. It cannot determine whether the differences between samples are statistically significant or whether the batch structure is confounded with the biological conditions. These questions require downstream analysis with tools such as DESeq2, which the published workflows use for differential expression analysis [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>][<a href="#ref-2">2</a>][<a href="#ref-4">4</a>][<a href="#ref-5">5</a>].

The MultiQC report is only as good as the input data. If the FastQC, STAR, or featureCounts outputs are incomplete or incorrectly formatted, MultiQC may miss samples or produce misleading plots. The researcher should verify that every sample produced the expected output files before running MultiQC.

MultiQC does not replace the need for human inspection of individual reports. The aggregated view is designed to highlight outliers and patterns, but a researcher may need to open the individual FastQC report for a flagged sample to understand the specific problem. The aggregated view guides the inspection, and the individual reports provide the detail.

## Reproducibility and Workflow Integration

MultiQC fits naturally into reproducible RNA-seq workflows. The tool is command-line based, so it can be incorporated into shell scripts, Snakemake pipelines, or Nextflow pipelines. The nf-core community provides standardized pipelines for RNA-seq analysis that include MultiQC as a standard component, and the documentation describes how to configure and run these pipelines on local, cluster, or cloud resources [<a href="#ref-8">8</a>][<a href="#ref-9">9</a>].

The published workflows demonstrate the integration of MultiQC into larger pipelines. The beef heifer dataset used a bioinformatic workflow based on FastQC and MultiQC for quality control, STAR for read alignment, and DESeq2 for differential expression analysis [<a href="#ref-1">1</a>]. The Arabidopsis dataset used FastQC for read quality checks, Salmon for quasi-mapping alignment, and MultiQC to merge mapping rate results into a single report [<a href="#ref-2">2</a>]. The influenza dataset used FastQC and MultiQC for quality control, Cutadapt for trimming, STAR for filtering to the influenza genome, and Salmon for transcript quantification [<a href="#ref-3">3</a>].

The reproducibility of these workflows depends on version control for the analysis scripts, documentation of the parameters, and archiving of the input and output files. The Carpentries lessons provide foundational training in the shell, Git, and programming skills needed to manage reproducible analyses [<a href="#ref-7">7</a>]. The EMBL-EBI training resources offer learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-6">6</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-10">10</a>]. The Bioconductor project offers official package and workflow documentation for reproducible genomic analysis [<a href="#ref-11">11</a>].

## Quality Control Thresholds and Decision Criteria

The decision to trim, retain, or exclude a sample should be based on the specific context of the experiment. The following criteria provide a starting point for making these decisions, but the researcher should adjust them based on the organism, the sequencing platform, and the downstream analysis requirements.

**Per-base quality.** Reads with a Phred quality score below 20 at the 3-prime end are commonly trimmed. The published Zika virus study used a Phred threshold of 30 for trimming [<a href="#ref-5">5</a>]. The choice depends on the read length and the downstream analysis. Longer reads can tolerate more aggressive trimming because sufficient length remains for alignment.

**Adapter contamination.** If adapter sequences appear in more than a small percentage of reads, trimming is required. The published influenza dataset used Cutadapt for adapter and quality trimming [<a href="#ref-3">3</a>]. The trimming parameters should be recorded and applied consistently across samples.

**Unique mapping rate.** A uniquely mapped percentage below 70 percent warrants investigation. The published breast cancer workflow reported mapping rates above 92 percent [<a href="#ref-4">4</a>]. The acceptable threshold depends on the genome complexity and the read length. Genomes with high repeat content may have lower unique mapping rates.

**Assigned fraction.** An assigned fraction below 60 percent warrants investigation. The acceptable threshold depends on the annotation quality and the library preparation method. The researcher should compare the assigned fractions across samples and flag any sample that deviates sharply from the cohort.

**Sequencing depth.** The number of reads per sample should be sufficient for the intended analysis. Differential expression analysis requires more reads than simple gene presence or absence detection. The researcher should check the read counts in the general statistics table and ensure that all samples meet the minimum depth.

## Professional Escalation Criteria

Some quality control problems require consultation with a bioinformatics specialist, a sequencing facility, or a statistician. The following situations warrant escalation.

**Systematic failure across all samples.** If all samples show poor quality, low mapping rates, or low assigned fractions, the problem is likely technical and affects the entire experiment. The sequencing facility should be consulted to determine whether the sequencing run failed and whether re-sequencing is possible.

**Unexplained outlier samples.** If one or a few samples deviate sharply from the cohort and the cause cannot be identified, a bioinformatics specialist should review the data. The specialist may identify issues with the library preparation, the alignment, or the annotation that are not apparent from the MultiQC report.

**Batch effects confounded with biological conditions.** If the batch structure aligns perfectly with the biological conditions, the experiment cannot distinguish technical from biological effects. A statistician should be consulted to determine whether the analysis can be adjusted or whether additional samples are needed.

**Contamination suspected.** If the QC metrics suggest contamination with another species, the researcher should verify the sample identity and consult with the sequencing facility. Contamination can arise during sample collection, RNA extraction, library preparation, or sequencing.

## Safety and Regulatory Context

RNA-seq data from human samples are subject to privacy and ethical regulations. The published datasets that use MultiQC include human cell lines infected with influenza virus [<a href="#ref-3">3</a>] and human breast cancer cell lines [<a href="#ref-4">4</a>]. Researchers working with human data must ensure that the samples were collected with appropriate consent and that the data are handled in compliance with applicable regulations.

The NCBI provides official descriptions of the databases and search systems used to deposit and access RNA-seq data [<a href="#ref-12">12</a>]. The published datasets are deposited in public repositories such as the European Nucleotide Archive, the Gene Expression Omnibus, and the Sequence Read Archive [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Researchers should follow the repository guidelines for data deposition and metadata submission.

The NCBI Sequence Read Archive and Gene Expression Omnibus provide the infrastructure for sharing RNA-seq data. The Arabidopsis dataset is available at NCBI with the data identification number GSE221567 [<a href="#ref-2">2</a>]. The beef heifer dataset is available at GEO with the accession GSE221903 [<a href="#ref-1">1</a>]. The influenza dataset is available at the European Nucleotide Archive via the ArrayExpress partner repository with the accession E-MTAB-9511 [<a href="#ref-3">3</a>]. Researchers should deposit their own data in these repositories to support reproducibility and data sharing.

## Frequently Asked Questions

### What is the difference between FastQC and MultiQC?

FastQC evaluates the quality of a single FASTQ file and produces a detailed report for that sample. MultiQC aggregates the outputs of FastQC and other tools across many samples into a single report. FastQC answers the question of whether one sample is good, and MultiQC answers the question of whether the entire cohort is consistent. The published workflows use both tools together, with FastQC providing the per-sample detail and MultiQC providing the cohort overview [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

### Do I need to run MultiQC before or after trimming?

MultiQC should be run both before and after trimming. The first run on the raw reads identifies adapter contamination and poor-quality bases that require trimming. The second run on the trimmed reads confirms that the trimming resolved the problems. The published influenza dataset used FastQC and MultiQC for quality control before trimming with Cutadapt [<a href="#ref-3">3</a>]. The comparison of the two reports documents the effectiveness of the trimming step.

### Can MultiQC replace manual inspection of FastQC reports?

MultiQC cannot fully replace manual inspection of individual FastQC reports. The aggregated view highlights outliers and patterns across the cohort, but the individual reports provide the detail needed to diagnose specific problems. A researcher should use the MultiQC report to identify samples that warrant closer inspection and then open the individual FastQC reports for those samples.

### What alignment metrics should I examine in the MultiQC report?

The uniquely mapped percentage is the most important alignment metric. It indicates what fraction of reads aligned to exactly one location in the reference genome. The multi-mapping percentage and the unmapped read categories provide additional context. The published breast cancer workflow reported mapping rates above 92 percent [<a href="#ref-4">4</a>], and the beef heifer dataset used STAR for alignment after FastQC and MultiQC quality control [<a href="#ref-1">1</a>].

### How do I detect batch effects with MultiQC?

MultiQC does not perform statistical tests for batch effects, but the plots can reveal batch structure. If samples from one sequencing run cluster together on the GC content plot or show different alignment rates than samples from another run, a batch effect is likely. The researcher should track sample metadata alongside the QC metrics and use the metadata to interpret the plots.

### What should I do if one sample has a much lower mapping rate than the rest?

Investigate the sample before deciding whether to exclude it. Check the sample metadata for mislabeling, verify that the sample is the correct species, and examine the FastQC report for contamination or degradation. If the cause cannot be identified, consult a bioinformatics specialist. The decision to exclude a sample should be documented and reported.

### Can MultiQC be used with Salmon instead of STAR?

Yes. MultiQC supports the output formats of both STAR and Salmon. The published Arabidopsis dataset used Salmon for quasi-mapping alignment and MultiQC to merge the mapping rate results into a single report [<a href="#ref-2">2</a>]. The influenza dataset used STAR for filtering to the influenza genome and Salmon for transcript quantification [<a href="#ref-3">3</a>]. The choice between STAR and Salmon depends on whether the analysis requires genome alignment or transcript-level quantification.

### How do I include MultiQC results in a publication?

The MultiQC report can be included as supplementary material, and the general statistics table can be exported and included as a table in the manuscript. The researcher should describe the quality control workflow in the methods section, including the versions of FastQC, MultiQC, STAR, and featureCounts, and the parameters used for trimming and alignment. The published datasets describe their quality control workflows in the methods sections of the data articles [<a href="#ref-3">3</a>][<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

## Related Bioinformatics Guides

- [RNA-Seq Batch Effect Detection and Correction](/knowledge/bioinformatics/rna-seq-batch-effect-detection-and-correction)
- [RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform](/knowledge/bioinformatics/rna-seq-data-analysis-in-galaxy-a-user-friendly-platform)
- [RNA-Seq Data Analysis Workflow: From Raw Reads to Insights](/knowledge/bioinformatics/rna-seq-data-analysis-workflow-from-raw-reads-to-insights)
- [RNA-Seq Quality Control: Essential Checks and Tools](/knowledge/bioinformatics/rna-seq-quality-control-essential-checks-and-tools)
- [Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics](/knowledge/bioinformatics/single-cell-rna-sequencing-quality-control-a-practical-guide-to-filtering-and-metrics)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Transcriptomic dataset from peripheral white blood cells of beef heifers at weaning.](https://pubmed.ncbi.nlm.nih.gov/36969977). Data in brief, 2023.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Arabidopsis transcriptome dataset of the response of imbibed wild-type and glucosinolate-deficient seeds to nitrogen-containing compounds.](https://pubmed.ncbi.nlm.nih.gov/37006386). Data in brief, 2023.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [RNA-Seq transcriptome data of human cells infected with influenza A/Puerto Rico/8/1934 (H1N1) virus.](https://pubmed.ncbi.nlm.nih.gov/33318985). Data in brief, 2020.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [An End-to-End Reproducible RNA-Seq Workflow from Raw Sequencing Reads to Differential Expression, Pathway Enrichment, and Biological Interpretation](https://doi.org/10.21203/rs.3.rs-10682168/v1). 2026.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [TRANSCRIPTOMIC AND FUNCTIONAL PATHWAY ANALYSIS OF ZIKA VIRUS-INDUCED PARALYSIS AND RECOVERY IN A MOUSE MODEL OF GUILLAIN-BARRÉ SYNDROME](https://doi.org/10.4238/9ssb1b14). Genetics and Molecular Research, 2026.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [RNA-seq: From FASTQ to Counts](https://doi.org/10.61700/s0isdk5tplhw82438). 2025.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-11"></a>[<a href="#ref-11">11</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-12"></a>[<a href="#ref-12">12</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.