Quality Control Metrics for Small RNA-seq: What to Check and How to Interpret Them

By Dr. Zubair Khalid, DVM, MS, PhD ·

Quality Control Metrics for Small RNA-seq: What to Check and How to Interpret Them

Key Takeaways

  • The read length distribution is the primary diagnostic for small RNA-seq, with distinct peaks expected around 21-23 nt for miRNAs and broader distributions for piRNAs, necessitating comparison to organism-specific literature rather than generic mRNA-seq standards.
  • High adapter content is expected and normal in small RNA-seq due to short inserts; the critical QC check is accurate adapter trimming, confirmed by re-examining the read length distribution post-trimming.
  • Mapping rates for small RNA-seq are inherently lower (60-90% for well-annotated genomes) than mRNA-seq due to short reads and potential mapping to repetitive elements, requiring adjusted alignment parameters and careful interpretation of unmapped reads.
  • Duplication rates are expected to be high for abundant small RNAs like miRNAs, but uniform high duplication across all size classes may indicate PCR over-amplification, requiring input RNA amount and PCR cycle number assessment.
  • rRNA and tRNA contamination should be low after size selection and depletion steps; a high fraction of reads mapping to these structural RNAs reduces effective sequencing depth for target small RNAs and necessitates review of library preparation protocols.
  • Pre-analytical factors such as specimen collection, RNA integrity (RIN is less informative for small RNAs), and extraction method significantly impact small RNA-seq success, with minimally invasive samples posing particular challenges due to limited RNA yield.

Small RNA sequencing requires a distinct quality control framework because the molecules under study are short, often carry specific terminal modifications, and originate from diverse biogenesis pathways. Standard mRNA-seq QC metrics, such as fragment length distributions centered on hundreds of bases or exon-intron junction rates, do not transfer directly to small RNA libraries. This article defines the QC metrics that matter for small RNA-seq, explains how to interpret them in the context of library preparation and biological questions, and provides concrete thresholds and troubleshooting steps for researchers managing their own datasets.

Scope and Reader Context

This guidance applies to researchers who have generated or plan to generate small RNA-seq data from total RNA, including microRNA (miRNA), small interfering RNA (siRNA), Piwi-interacting RNA (piRNA), and other short non-coding RNA species. The QC decisions described here apply at multiple stages: after sequencing, after adapter trimming, after alignment, and before differential expression analysis. The intended reader is a biology student, laboratory researcher, or life-science practitioner who needs to determine whether a small RNA-seq dataset is reliable enough for downstream interpretation.

The central problem addressed is that generic RNA-seq QC pipelines often flag small RNA libraries as problematic or fail to flag genuine problems. Read lengths are short by design, mapping rates may be low because small RNAs map to repetitive or unannotated regions, and adapter content is expected to be high because the insert is shorter than the read length. Each of these features requires a different interpretive frame than standard mRNA-seq.

Why Small RNA-seq QC Differs from mRNA-seq

Small RNA-seq libraries are prepared by ligating adapters to RNA molecules that are typically 18 to 30 nucleotides long. The sequencing read often extends past the insert into the 3' adapter, producing high adapter content that would indicate a problem in a standard mRNA-seq library but is a normal feature of small RNA libraries. The read length distribution is therefore a primary diagnostic, not a secondary check.

The biological composition of small RNA populations also differs by tissue, developmental stage, and organism. A library from a tissue with high miRNA content will show a dominant peak around 21 to 23 nucleotides. A library from germline tissue may show a broader distribution with piRNA-sized reads around 24 to 30 nucleotides. A library from a plant sample will show different size classes than a mammalian sample because plant small RNA biogenesis pathways differ. These expectations must be established before QC thresholds are applied.

RNA integrity also matters differently for small RNA-seq. Small RNAs are more stable than long mRNAs under some degradation conditions, but the ligation-based library preparation is sensitive to the presence of degraded fragments, ribosomal RNA remnants, and other contaminants. Pre-analytical handling, including specimen collection, storage, and RNA extraction, influences sequencing success rates and data quality. Minimally invasive specimens such as small biopsies and cytologic samples present particular challenges because of limited RNA yield and distinct workflows. These pre-analytical factors affect whether sequencing succeeds and the accuracy of the results.

At a Glance: Core QC Metrics for Small RNA-seq

The table below summarizes the primary QC metrics, what they indicate, typical values for acceptable small RNA libraries, and the first troubleshooting step when a metric falls outside the expected range.

MetricWhat It IndicatesTypical Acceptable RangeFirst Troubleshooting Step
Read length distributionSize class composition of small RNA populationDominant peak at 21 to 23 nt for miRNA libraries, broader distribution for piRNA or total small RNACheck RNA size selection and adapter ligation efficiency
Adapter contentFraction of reads containing 3' adapter sequenceHigh adapter content expected, nearly all reads may contain adapter sequenceConfirm adapter trimming parameters match the ligation chemistry
Mapping rateFraction of reads aligning to the reference genomeVariable by organism and annotation quality, often 60% to 90% for well-annotated genomesVerify reference genome version and alignment parameters for short reads
Duplication rateFraction of reads with identical sequenceHigh duplication expected for abundant miRNAs, inspect distribution across size classesAssess whether duplication reflects biological abundance or PCR over-amplification
Nucleotide composition biasPosition-specific base frequencies at read endsStable composition within size classes, no strong positional biasCheck for ligation bias or contamination from other RNA species
rRNA and tRNA contaminationFraction of reads mapping to structural RNAsShould be low after size selection, varies by library preparation methodReview size selection and depletion steps

Read Length Distribution as the Primary Diagnostic

The read length distribution is the first metric to examine after sequencing and adapter trimming. For a typical mammalian miRNA library, the distribution should show a sharp peak at 21 to 23 nucleotides, corresponding to mature miRNA lengths. A library with a broad distribution across 18 to 30 nucleotides may indicate that the size selection step captured a wider range of small RNAs, which may be appropriate for studies of piRNAs or other small RNA classes but should be documented and interpreted accordingly.

A library with a dominant peak at unexpected lengths, such as 30 to 40 nucleotides, may indicate contamination with tRNA fragments or other RNA degradation products. A library with a peak at very short lengths, below 18 nucleotides, may indicate adapter dimer formation or excessive RNA degradation. The expected distribution should be established from the literature for the organism and tissue under study before thresholds are applied.

The read length distribution also interacts with adapter trimming. If the insert is shorter than the read length, the 3' adapter will be sequenced. The position of the adapter sequence within the read depends on the insert length. After trimming, the distribution of insert lengths should match the expected small RNA size classes. A distribution that shifts after trimming in unexpected ways may indicate that the trimming parameters do not match the adapter sequence or ligation chemistry used.

Adapter Content and Trimming Decisions

Adapter content is a critical QC metric for small RNA-seq because the insert is often shorter than the read length. The fraction of reads containing adapter sequence will be high, and this is expected. The relevant question is whether the adapter sequence is correctly identified and trimmed.

The adapter sequence used in the library preparation must be known exactly. Common adapter sequences are provided by kit manufacturers, and the trimming tool must be configured with the correct sequence and allowed error rate. A mismatch between the actual adapter and the sequence provided to the trimming tool will leave adapter sequence in the reads, which will interfere with alignment and quantification.

After trimming, the read length distribution should be re-examined. If the distribution still shows a peak at the read length, adapter trimming may have failed for a substantial fraction of reads. If the distribution shows a peak at very short lengths, the trimming may have been too aggressive, removing part of the insert.

Adapter dimer contamination is a specific failure mode. Adapter dimers form when the 3' and 5' adapters ligate to each other without an insert. These molecules are typically 40 to 50 base pairs in length and produce reads that consist entirely of adapter sequence. After trimming, these reads have zero length and should be removed. A high fraction of adapter dimer reads indicates a problem in the ligation step, and the library should be considered compromised.

Mapping Rate and Alignment Parameters

The mapping rate for small RNA-seq is typically lower than for mRNA-seq because small RNAs are short and may map to multiple locations in the genome. Repetitive elements, tRNA genes, and rRNA genes produce multi-mapping reads that are difficult to assign unambiguously. The mapping rate also depends on the completeness of the reference genome and the quality of the annotation.

For well-annotated genomes such as human and mouse, a mapping rate of 60% to 90% is common for small RNA libraries. Lower mapping rates may indicate contamination with other RNA species, adapter trimming failure, or a reference genome that lacks sequences corresponding to the small RNA population under study. For less well-annotated genomes, lower mapping rates may be acceptable, but the reasons should be documented.

Alignment parameters must be adjusted for short reads. The default parameters in many aligners are optimized for reads of 50 to 150 nucleotides and may be too stringent for reads of 18 to 30 nucleotides. The allowed number of mismatches and the seed length should be adjusted to accommodate short reads. Multi-mapping reads should be handled consistently, either by assigning them to one location at random, by distributing them across all locations, or by excluding them from quantification. The choice affects downstream analysis and should be documented.

The mapping rate should also be examined by size class. If reads of a particular length map at a much lower rate than other lengths, this may indicate a specific contamination or a biological feature such as a class of small RNAs that is not represented in the reference genome.

Duplication Rate and PCR Artifacts

Duplication rate measures the fraction of reads with identical sequence. In small RNA-seq, high duplication rates are expected because abundant miRNAs produce many identical reads. The duplication rate should be interpreted in the context of the read length distribution and the biological sample.

A high duplication rate across all size classes may indicate PCR over-amplification during library preparation. This occurs when too many PCR cycles are used or when the input RNA amount is too low. PCR duplicates introduce bias because they inflate the count of abundant sequences and may create false sequence variants.

A high duplication rate concentrated in a few size classes is more likely to reflect biological abundance. For example, a few highly expressed miRNAs may account for a large fraction of the reads in a library. The duplication rate should be examined separately for each size class to distinguish biological from technical duplication.

The input RNA amount and the number of PCR cycles should be recorded for each library. If the duplication rate is high and the input amount was low, the library should be considered for re-preparation. If the duplication rate is high but the input amount was adequate, the duplication is more likely to reflect biological abundance.

Nucleotide Composition Bias

Nucleotide composition bias refers to position-specific deviations in base frequencies across reads. In small RNA-seq, composition bias can arise from ligation bias, where the adapter ligation step favors certain sequences at the ligation junction. This bias can distort the measured abundance of small RNAs that have unfavorable sequences at their ends.

Composition bias can also arise from contamination with other RNA species. For example, tRNA fragments have characteristic sequence features that differ from miRNAs. A library with a high fraction of tRNA fragments will show a composition profile that differs from a library dominated by miRNAs.

The nucleotide composition should be examined separately for each size class. Position-specific base frequencies should be relatively stable within a size class. Large deviations at the first and last positions may indicate ligation bias. The biological relevance of composition bias depends on the study question. For studies of miRNA abundance, ligation bias is a known limitation that should be acknowledged. For studies of small RNA sequence variants, composition bias can create false variants and should be investigated.

rRNA and tRNA Contamination

Ribosomal RNA and transfer RNA contamination is a common problem in small RNA-seq libraries. rRNA and tRNA are abundant in total RNA and can survive size selection if their fragments fall within the selected size range. tRNA fragments, in particular, are in the same size range as miRNAs and can be difficult to remove.

The fraction of reads mapping to rRNA and tRNA genes should be quantified. A high fraction indicates that the size selection or depletion steps were not effective. The acceptable fraction depends on the library preparation method and the biological question. For studies focused on miRNAs, a high tRNA fraction reduces the effective sequencing depth for miRNAs and may require deeper sequencing.

Some library preparation methods include steps to deplete rRNA and tRNA. If these steps were not used, the expected contamination should be estimated from the literature for the organism and tissue under study. The contamination fraction should be reported in the methods section of any publication.

Pre-Analytical Factors Affecting Small RNA-seq Quality

The quality of small RNA-seq data depends on the quality of the input RNA. Pre-analytical factors, including specimen collection, processing, and storage, influence sequencing success rates and the accuracy of results. This is particularly important for minimally invasive specimens such as small biopsies and cytologic samples, which have limited RNA yield and distinct workflows.

RNA degradation is a primary concern. Small RNAs are more stable than long mRNAs under some conditions, but degradation can still affect the small RNA population. The RNA integrity number (RIN) is commonly used to assess RNA quality, but it is based on the ratio of ribosomal RNA peaks and may not reflect the quality of small RNAs. A sample with a low RIN may still yield usable small RNA-seq data, but the results should be interpreted with caution.

The RNA extraction method also affects small RNA recovery. Some extraction methods are optimized for total RNA and may not efficiently recover small RNAs. The extraction method should be validated for the small RNA population of interest. The yield and size distribution of the extracted RNA should be assessed before library preparation.

Specimen storage conditions affect RNA quality. Frozen specimens should be stored at appropriate temperatures, and freeze-thaw cycles should be minimized. RNAlater preservation is an alternative to freezing and may be preferable for some specimen types. The storage conditions should be recorded for each specimen.

Library Preparation and Size Selection

The library preparation method determines the size range of small RNAs that are captured. Most methods include a size selection step that enriches for RNAs in the 18 to 30 nucleotide range. The efficiency of size selection affects the composition of the library and the fraction of reads that map to small RNA loci.

Size selection can be performed using gel extraction or bead-based methods. Gel extraction provides precise size selection but is labor-intensive and may have variable recovery. Bead-based methods are faster but may have broader size ranges. The size selection method should be recorded, and the resulting size distribution should be verified using a bioanalyzer or similar instrument before sequencing.

The number of PCR cycles used during library preparation affects the duplication rate. Fewer cycles reduce PCR duplicates but may not produce enough library for sequencing. More cycles increase the library yield but introduce more duplicates and bias. The optimal number of cycles depends on the input amount and the library preparation method. The cycle number should be recorded for each library.

Sequencing Depth and Coverage

The required sequencing depth for small RNA-seq depends on the biological question and the complexity of the small RNA population. For studies of abundant miRNAs, a few million reads per sample may be sufficient. For studies of rare small RNAs or for detecting novel small RNAs, deeper sequencing is required.

The relationship between sequencing depth and the number of detected small RNAs should be examined. A saturation curve, where the number of detected small RNAs is plotted against the number of reads, indicates whether additional sequencing would identify new small RNAs. If the curve has plateaued, additional sequencing is unlikely to add new information.

The sequencing depth should also be considered in the context of the duplication rate. If a large fraction of reads are duplicates, the effective sequencing depth is lower than the raw read count. The effective depth, after removing duplicates, should be used to assess whether the sequencing was sufficient.

Quality Control Workflow for Small RNA-seq

A structured QC workflow ensures that all relevant metrics are examined and documented. The workflow should be applied consistently to all samples in a study.

Step 1: Examine Raw Read Quality

The first step is to examine the raw read quality using a tool such as FastQC. The per-base quality scores should be high across the read length. Low quality at the 3' end of reads is common and may be acceptable if the quality is high in the region that contains the small RNA insert. The per-base quality should be examined separately for each size class after trimming.

Step 2: Trim Adapters and Low-Quality Bases

Adapter trimming should be performed using a tool that is configured with the correct adapter sequence and allowed error rate. Low-quality bases at the 3' end of reads should be trimmed. The trimming parameters should be recorded. After trimming, the read length distribution should be re-examined to confirm that the expected size classes are present.

Step 3: Assess Read Length Distribution

The read length distribution after trimming should show the expected size classes for the organism and tissue under study. A dominant peak at 21 to 23 nucleotides is expected for miRNA libraries. The distribution should be compared to the expected distribution from the literature. Deviations should be investigated.

Step 4: Align to Reference Genome

The trimmed reads should be aligned to the reference genome using parameters optimized for short reads. The mapping rate should be calculated and compared to the expected range for the organism. Multi-mapping reads should be handled consistently. The alignment parameters should be recorded.

Step 5: Quantify Small RNA Expression

The aligned reads should be quantified against an annotation of small RNA loci. The quantification should be performed separately for each small RNA class. The number of reads assigned to each class should be examined. A high fraction of unassigned reads may indicate an incomplete annotation or a biological feature such as novel small RNAs.

Step 6: Examine Duplication and Contamination

The duplication rate should be calculated and examined by size class. The fraction of reads mapping to rRNA and tRNA should be quantified. The nucleotide composition should be examined for positional bias. These metrics should be interpreted in the context of the library preparation method and the biological sample.

Step 7: Document and Report QC Metrics

All QC metrics should be documented in a consistent format. The metrics should be reported in the methods section of any publication. The raw data and the QC metrics should be deposited in a public database such as the NCBI Sequence Read Archive, which provides official infrastructure for sequence data storage and retrieval.

Records and Measurements for QC Documentation

Consistent record keeping is essential for QC interpretation and for reproducing the analysis. The following records should be maintained for each library:

Record CategorySpecific Data to CapturePurpose in QC Interpretation
Specimen informationIdentifier, tissue type, collection date, storage conditionsLinks QC outcomes to pre-analytical handling variables
RNA metricsExtraction method, yield, RIN if measuredIdentifies whether input quality explains library failures
Library preparationKit version, size selection method, PCR cycle number, adapter sequencesEnables troubleshooting of ligation and amplification steps
Sequencing parametersPlatform, read length, raw read countEstablishes expected adapter content and depth calculations
Post-trimming metricsRead count after trimming, length distributionConfirms adapter removal and size class preservation
Alignment and quantificationMapping rate, duplication rate, rRNA/tRNA fraction, assigned read fractionProvides the core technical quality indicators for downstream analysis

These records enable the researcher to identify the source of QC failures and to compare QC metrics across samples and studies.

Common Failure Patterns and Troubleshooting

Several failure patterns recur in small RNA-seq datasets. Recognizing these patterns and knowing the first troubleshooting step reduces the time spent on problematic libraries.

Failure Pattern 1: No Peak in the Read Length Distribution

A library with no clear peak in the read length distribution may indicate that the size selection failed or that the RNA was degraded. The first step is to examine the raw read quality and the adapter content. If the reads contain high adapter content and the trimmed lengths are highly variable, the size selection likely failed. The library should be re-prepared with a more precise size selection method.

Failure Pattern 2: Dominant Peak at Unexpected Length

A library with a dominant peak at a length that does not match the expected small RNA class may indicate contamination. For example, a peak at 30 to 40 nucleotides may indicate tRNA fragments. The first step is to quantify the fraction of reads mapping to tRNA and rRNA. If contamination is confirmed, the library preparation should be repeated with additional depletion steps.

Failure Pattern 3: Very Low Mapping Rate

A library with a very low mapping rate may indicate adapter trimming failure, contamination with non-small RNA species, or a reference genome that does not contain the small RNA sequences. The first step is to examine the reads that do not map. If these reads contain adapter sequence, the trimming parameters should be adjusted. If the reads map to other organisms, the sample may be contaminated.

Failure Pattern 4: Very High Duplication Rate

A library with a very high duplication rate across all size classes may indicate PCR over-amplification. The first step is to check the input RNA amount and the number of PCR cycles. If the input was low or the cycle number was high, the library should be re-prepared. If the duplication is concentrated in a few size classes, it may reflect biological abundance.

Failure Pattern 5: High rRNA and tRNA Contamination

A library with a high fraction of reads mapping to rRNA and tRNA may indicate that the size selection or depletion steps were not effective. The first step is to review the library preparation protocol and confirm that the depletion steps were performed correctly. If the contamination is high, the library should be re-prepared.

Interpretation Limits and Biological Context

QC metrics provide information about the technical quality of a library, but they do not directly indicate biological validity. A library that passes all QC metrics may still produce results that do not reflect the biological state of the sample. Conversely, a library that fails some QC metrics may still yield biologically meaningful results if the failures are understood and documented.

The interpretation of QC metrics depends on the biological question. For a study of miRNA abundance, the read length distribution and the fraction of reads mapping to miRNA loci are the most important metrics. For a study of novel small RNAs, the fraction of unassigned reads and the mapping rate to unannotated regions are more relevant.

The choice of thresholds for QC metrics should be based on the expected distribution for the organism and tissue under study. Fixed thresholds, such as a universal mapping rate cutoff, are not appropriate because the expected values vary widely across biological contexts. The thresholds should be established from the literature and from the distribution of QC metrics across all samples in the study.

Reproducibility and Workflow Standards

Reproducibility in small RNA-seq analysis requires that the analysis workflow is documented and version-controlled. The use of established workflow frameworks improves reproducibility and reduces the risk of errors. Community-developed pipelines provide standardized implementations of QC, alignment, and quantification steps.

The nf-core documentation describes community pipeline standards for reproducible analysis workflows. These pipelines are implemented in a common workflow language and include version tracking and containerization. Using a community pipeline ensures that the analysis steps are transparent and reproducible.

The Galaxy Training Network provides accessible tutorials for small RNA-seq analysis, including QC steps. These tutorials are useful for researchers who are new to small RNA-seq analysis and for those who want to verify that their analysis steps follow established practices.

The Bioconductor project provides packages for small RNA-seq analysis within the R programming environment. These packages include tools for QC, alignment, and quantification, and they are documented with reproducible examples.

Tools and Resources for Small RNA-seq QC

Several tools and resources are available for small RNA-seq QC. The choice of tools depends on the researcher's familiarity with command-line interfaces and programming languages.

FastQC is a widely used tool for examining raw read quality. It provides per-base quality scores, adapter content, and other metrics. FastQC is suitable for an initial assessment of raw reads, but the metrics must be interpreted in the context of small RNA libraries.

MultiQC aggregates QC metrics from multiple tools into a single report. This is useful for comparing QC metrics across many samples. MultiQC can be used to identify samples that deviate from the expected distribution.

The plantDARIO web tool provides small RNA-seq analysis for plant species, including QC, quantification, and prediction of novel small RNAs. This tool is useful for researchers working with Arabidopsis thaliana, Beta vulgaris, and Solanum lycopersicum.

The NCBI provides databases for sequence data and analysis services. The Sequence Read Archive stores raw sequencing data, and the Gene Expression Omnibus stores processed data. Depositing data in these databases ensures that the data are available for re-analysis and verification.

Welfare and Safety Context for Laboratory Practice

Small RNA-seq involves the use of chemical reagents and laboratory equipment that require appropriate safety precautions. The following considerations apply to laboratory work associated with small RNA-seq.

RNA extraction and library preparation involve the use of organic solvents, including phenol and chloroform. These reagents are hazardous and should be handled in a fume hood with appropriate personal protective equipment. The safety data sheets for all reagents should be reviewed before use.

The use of RNase-free reagents and consumables is essential for RNA work. RNase contamination degrades RNA and compromises the quality of the library. Gloves should be worn at all times, and surfaces should be cleaned with RNase decontamination solutions.

The sequencing instrument uses lasers and other equipment that require appropriate training. The instrument should be operated according to the manufacturer's instructions, and the operator should be trained before use.

Professional Escalation Criteria

Some QC failures require consultation with a bioinformatics specialist or a core facility. The following situations warrant escalation:

  • The read length distribution does not match the expected pattern, and the cause cannot be identified from the library preparation records.
  • The mapping rate is very low, and the unaligned reads do not match any known contamination source.
  • The duplication rate is very high, and the input RNA amount and PCR cycle number do not explain the duplication.
  • The QC metrics vary substantially across samples from the same biological condition, and the variation cannot be attributed to biological factors.
  • The analysis workflow produces different results when run with different parameter settings, and the source of the discrepancy is unclear.

A bioinformatics specialist can help identify the source of QC failures and recommend appropriate analysis parameters. A core facility can provide guidance on library preparation and sequencing best practices.

A Practical Decision Framework for Small RNA-seq Library Acceptance

QC metrics only become useful when they feed into a clear decision about whether to proceed with a library, re-sequence it, or discard it. Many researchers collect QC numbers without a structured path for acting on them. This section provides a decision framework that separates libraries into three categories: accept, conditional accept, and reject. The framework uses the metrics described in the previous sections but organizes them into a sequence of checks that reflect their relative importance for downstream analysis.

The Three-Tier Decision Framework

The framework operates at three tiers. Tier 1 contains the metrics that determine whether a library is fundamentally usable. Tier 2 contains metrics that affect interpretation but do not necessarily invalidate the library. Tier 3 contains metrics that should be documented but rarely trigger rejection on their own.

Tier 1: Gatekeeping Metrics

These metrics determine whether the library enters downstream analysis at all. A library that fails a Tier 1 metric should be rejected or re-prepared unless the failure has a documented biological explanation.

The first gatekeeping metric is the read length distribution after adapter trimming. A library with no clear peak in the expected small RNA size range indicates that the library preparation failed to capture the intended molecules. For a miRNA-focused study, the trimmed read length distribution should show a dominant peak at 21 to 23 nucleotides. The absence of this peak means the library does not represent the biological population of interest.

The second gatekeeping metric is the fraction of reads that survive adapter trimming. If more than half of the raw reads are lost during trimming because they consist entirely of adapter sequence, the library likely contains a high proportion of adapter dimers. Adapter dimers form when adapters ligate to each other without an insert, and they indicate a problem in the ligation step. A library with a high adapter dimer fraction should be re-prepared.

The third gatekeeping metric is the mapping rate to the reference genome. For well-annotated genomes, a mapping rate below 40% warrants investigation before proceeding. This threshold is lower than the typical acceptable range of 60% to 90% because the decision to reject should account for organisms with incomplete annotations. The key question is whether the unmapped reads can be explained by known factors such as adapter trimming failure, contamination, or missing genomic sequences.

Tier 2: Interpretation Metrics

These metrics do not trigger rejection on their own but affect how the data should be interpreted. A library that passes Tier 1 but shows concerning Tier 2 metrics may be used with appropriate caveats.

The duplication rate is a Tier 2 metric. High duplication concentrated in a few size classes is expected for abundant miRNAs and reflects biological abundance. High duplication spread uniformly across all size classes may indicate PCR over-amplification. The distinction matters because PCR duplicates inflate counts of abundant sequences and can create false sequence variants. If the duplication rate is high and the input RNA amount was low, the library should be flagged for cautious interpretation.

The rRNA and tRNA contamination fraction is also a Tier 2 metric. A high contamination fraction reduces the effective sequencing depth for the small RNAs of interest but does not necessarily invalidate the library. The acceptable contamination fraction depends on the biological question. A study of miRNA abundance can tolerate higher contamination than a study of rare small RNA species.

Tier 3: Documentation Metrics

These metrics should be recorded and reported but rarely trigger rejection. Nucleotide composition bias is a Tier 3 metric because ligation bias is a known limitation of small RNA-seq that affects all libraries to some degree. The bias should be documented and acknowledged in the interpretation of results, but it does not warrant discarding a library that passes Tier 1 and Tier 2 checks.

Applying the Framework in Practice

The framework should be applied consistently across all samples in a study. The first step is to establish the expected read length distribution and mapping rate for the organism and tissue under study. These expectations come from the literature and from the distribution of QC metrics across all samples in the study.

The second step is to apply the Tier 1 checks to each library. A library that fails a Tier 1 check should be examined for the specific failure pattern described in the troubleshooting section. If the failure cannot be explained and corrected, the library should be re-prepared.

The third step is to apply the Tier 2 checks to libraries that pass Tier 1. Libraries with concerning Tier 2 metrics should be flagged in the analysis records. The flag should note which metric is concerning and how it affects interpretation.

The fourth step is to document all Tier 3 metrics in the study records. These metrics should be reported in the methods section of any publication.

A Record System for QC Decisions

A structured record system ensures that QC decisions are transparent and reproducible. The following table provides a template for recording the decision for each library:

Library IDTier 1 Pass or FailTier 1 NotesTier 2 FlagsTier 3 NotesDecisionDecision Rationale
Sample 1PassPeak at 22 nt, 78% mappingHigh duplication in 22 nt classLigation bias at position 1AcceptDuplication reflects miRNA abundance
Sample 2FailNo peak, 35% adapter dimersNot assessedNot assessedRejectAdapter dimer contamination
Sample 3PassPeak at 22 nt, 65% mapping25% tRNA contaminationNoneConditional acceptContamination reduces depth for miRNAs

The record should include the specific values for each metric, beyond the pass or fail designation. This allows the decision to be reviewed and challenged if necessary.

Common Failure Patterns in the Decision Framework

The decision framework identifies several recurring patterns that researchers should recognize.

Pattern A: Tier 1 Failure with Adapter Dimers

A library fails Tier 1 because more than half of the reads are lost during adapter trimming. The read length distribution after trimming shows no clear peak. This pattern indicates adapter dimer contamination from the ligation step. The library should be re-prepared with attention to the adapter concentration and the ligation conditions.

Pattern B: Tier 1 Failure with No Size Peak

A library passes the adapter trimming check but shows no clear peak in the read length distribution after trimming. The reads are spread across a wide range of lengths. This pattern may indicate that the size selection step failed or that the RNA was degraded. The library should be re-prepared with a more precise size selection method.

Pattern C: Tier 2 Flag for PCR Duplication

A library passes Tier 1 but shows high duplication across all size classes. The input RNA amount was low, and the PCR cycle number was high. This pattern indicates PCR over-amplification. The library may be used with caution, but the duplication should be noted in the analysis records. If the study requires precise quantification of abundant small RNAs, the library should be re-prepared with more input RNA or fewer PCR cycles.

Pattern D: Tier 2 Flag for tRNA Contamination

A library passes Tier 1 but shows a high fraction of reads mapping to tRNA genes. This pattern indicates that the size selection or depletion steps did not remove tRNA fragments. The library may be used if the biological question focuses on miRNAs, but the reduced effective depth for miRNAs should be considered. If the study requires detection of rare small RNAs, the library should be re-prepared with additional depletion steps.

When to Escalate to Professional Support

The decision framework includes clear escalation criteria. A researcher should consult a bioinformatics specialist or core facility when:

  • A library fails Tier 1 and the failure pattern does not match any of the known patterns described above.
  • The QC metrics vary substantially across replicate libraries from the same biological condition, and the variation cannot be attributed to biological factors.
  • The decision framework produces different outcomes when applied by different members of the research team, indicating that the criteria are not being applied consistently.
  • The library passes all QC checks but produces results that contradict established biological knowledge for the system under study.

The escalation should include the complete QC record for the library, the decision framework output, and a description of the specific concern. This information allows the specialist to identify the source of the problem and recommend appropriate action.

Integration with Reproducible Workflow Standards

The decision framework should be implemented within a reproducible analysis workflow. The nf-core documentation describes community pipeline standards that include version tracking and containerization. Implementing the QC decision framework within such a workflow ensures that the same criteria are applied consistently across samples and studies.

The Galaxy Training Network provides accessible tutorials for small RNA-seq analysis that include QC steps. These tutorials can help researchers implement the decision framework in a reproducible manner. The Bioconductor project provides packages for small RNA-seq analysis within the R programming environment, including tools for QC visualization and documentation.

The decision framework should be documented in the methods section of any publication. The documentation should include the specific thresholds used for each metric, the rationale for those thresholds, and the decision rules for each tier. This transparency allows other researchers to evaluate the QC decisions and to compare results across studies.

Frequently Asked Questions

What read length distribution should I expect for a small RNA-seq library?

The expected read length distribution depends on the organism and tissue. For mammalian miRNA libraries, a dominant peak at 21 to 23 nucleotides is typical. Libraries from germline tissue may show a broader distribution with piRNA-sized reads around 24 to 30 nucleotides. Plant libraries may show different size classes because plant small RNA biogenesis pathways differ. The expected distribution should be established from the literature before thresholds are applied.

Why is the adapter content so high in my small RNA-seq data?

High adapter content is expected in small RNA-seq because the insert is often shorter than the read length. The sequencing read extends past the insert into the 3' adapter. The relevant question is whether the adapter sequence is correctly identified and trimmed. If the adapter content remains high after trimming, the trimming parameters may not match the adapter sequence or ligation chemistry.

What mapping rate should I expect for small RNA-seq?

The mapping rate depends on the organism and the quality of the reference genome and annotation. For well-annotated genomes such as human and mouse, a mapping rate of 60% to 90% is common. Lower mapping rates may indicate contamination, adapter trimming failure, or a reference genome that lacks sequences corresponding to the small RNA population. The mapping rate should be examined by size class to identify specific problems.

How do I distinguish biological duplication from PCR duplication?

Biological duplication occurs when abundant small RNAs produce many identical reads. PCR duplication occurs when the same molecule is amplified multiple times during library preparation. The duplication rate should be examined by size class. If the duplication is concentrated in a few size classes, it is more likely to reflect biological abundance. If the duplication is high across all size classes, PCR over-amplification is more likely.

What should I do if my library has high rRNA and tRNA contamination?

High rRNA and tRNA contamination indicates that the size selection or depletion steps were not effective. The library preparation protocol should be reviewed to confirm that the depletion steps were performed correctly. If the contamination is high, the library should be re-prepared with additional depletion steps. The acceptable contamination fraction depends on the biological question and the library preparation method.

How much sequencing depth do I need for small RNA-seq?

The required sequencing depth depends on the biological question and the complexity of the small RNA population. For studies of abundant miRNAs, a few million reads per sample may be sufficient. For studies of rare small RNAs or for detecting novel small RNAs, deeper sequencing is required. A saturation curve, where the number of detected small RNAs is plotted against the number of reads, indicates whether additional sequencing would identify new small RNAs.

Can I use standard RNA-seq QC tools for small RNA-seq data?

Standard RNA-seq QC tools can be used for small RNA-seq data, but the metrics must be interpreted in the context of small RNA libraries. Read length distributions, adapter content, and mapping rates have different expected values for small RNA-seq than for mRNA-seq. The QC metrics should be examined separately for each size class, and thresholds should be established from the expected distribution for the organism and tissue under study.

What records should I keep for small RNA-seq QC documentation?

Records should include specimen identifier and tissue type, collection and storage conditions, RNA extraction method and yield, RNA quality metrics, library preparation method, size selection method, PCR cycle number, adapter sequences, sequencing platform and read length, and all QC metrics including read counts, read length distribution, mapping rate, duplication rate, and contamination fractions. These records enable the researcher to identify the source of QC failures and to compare QC metrics across samples and studies.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.