Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Section: Molecular Diagnostics

Understanding Sequencing Quality Control: Metrics and Best Practices

Sequencing quality control (QC) covers the measurements and decisions that determine whether a sequencing library and the data it produces are trustworthy for downstream analysis. For laboratory students, technicians, researchers, and diagnostic professionals, QC is the difference between a confident variant call and a costly false result. This article explains the core QC metrics for sequencing libraries and data, including concentration, fragment size, quality scores, and GC bias, and describes practical best practices for implementing a QC workflow in a diagnostic or research setting.

The Role of QC in the Sequencing Workflow

Quality control in sequencing is not a single step but a continuous process that spans the entire workflow from sample receipt to final variant interpretation. The World Health Organization Laboratory Quality Management System Handbook emphasizes that a quality management system requires documented procedures, records, and corrective actions across all phases of laboratory work. In sequencing, this translates to defined acceptance criteria for samples, libraries, and data at each stage.

The FDA-led Sequencing and Quality Control Phase 2 (SEQC2) project was established to develop standard analysis protocols and quality control metrics for DNA testing to enhance scientific research and precision medicine. This effort reflects a broader recognition that sequencing results are only as reliable as the quality systems surrounding them. For clinical applications, the stakes are higher because results directly inform patient management decisions.

Preanalytical variables deserve particular attention. A 2022 methods review on preanalytical variables and sample quality control for clinical variant analysis notes that broad molecular profiling by next-generation sequencing of solid tumors has become a critical tool for clinical decision-making. The review emphasizes that accurate and reliable assays are needed to assess the genetic makeup of tumor cells and guide clinicians in therapy decisions. To obtain high-quality sequencing data for variant detection, certain preanalytical steps and quality metrics should be followed, including assessment of sample types, choice of extraction method, library preparation technology, sequencing platform, and finally sequencing quality control. Each of these steps has challenges and pitfalls that need to be addressed.

Core Library QC Metrics

Before a library is loaded onto a sequencer, it must pass several quality checks. These library-level metrics predict whether the sequencing run will produce usable data.

DNA and RNA Concentration

Library concentration is the most basic QC metric. Accurate quantification enables accurate loading of libraries onto the sequencer and improves sequencing performance by reducing under and overloading errors. Digital PCR is a valuable tool to quantify next-generation sequencing libraries precisely and accurately. Accurate quantification also benefits users by enabling uniform loading of indexed or barcoded libraries, which improves sequencing uniformity across samples.

Concentration measurements can be obtained through several methods, including fluorometric assays, spectrophotometry, and quantitative PCR. Each method has different sensitivity and specificity for double-stranded DNA versus adapter-dimers or other contaminants. The choice of quantification method should match the intended use of the measurement.

Fragment Size Distribution

Library fragment size affects sequencing performance on most platforms. Fragment size assessment is typically performed using capillary electrophoresis or automated electrophoresis systems. The size distribution reveals whether the library contains the expected insert sizes and whether adapter-dimers or other artifacts are present.

For hybrid capture enrichment workflows, fragment size also affects capture efficiency. Libraries with excessive short fragments may produce low-complexity data, while libraries with very long fragments may cluster poorly on some platforms. The acceptable size range depends on the sequencing platform and the application.

Library Purity and Adapter Content

Adapter contamination is a common library quality issue. Adapter-dimers form when adapter molecules ligate to each other without an insert. These artifacts consume sequencing capacity and produce reads that contain only adapter sequence. Many bioinformatics tools can detect adapter contamination in raw sequencing data, but preventing it during library preparation is more efficient than removing it afterward.

Library Quantification by Digital PCR

Droplet digital PCR provides precise and accurate library quantification in addition to size quality assessment. This approach enables users to QC their sequencing libraries with confidence. The advantages of digital PCR-based quantification include absolute quantitation without a standard curve and the ability to distinguish amplifiable library molecules from total DNA.

Sequencing Data QC Metrics

Once sequencing is complete, the raw data must be evaluated before alignment and variant calling. Quality control at this stage identifies run-level problems that affect all samples in the flow cell.

Quality Scores

Quality scores, typically reported as Phred scores, represent the probability that a base call is incorrect. A Phred score of Q20 corresponds to an error rate of 1 in 100, while Q30 corresponds to 1 in 1,000. Quality scores are calculated by the sequencing instrument during base calling and are stored in the sequencing output files.

Per-base quality scores should be examined across the length of the reads. Many sequencing platforms show declining quality toward the 3-prime end of reads. The pattern of quality decline can indicate specific instrument or chemistry problems.

Read Depth and Coverage

Coverage depth refers to the number of times a particular nucleotide position is sequenced. Coverage uniformity describes how evenly reads are distributed across the target region. Both metrics are essential for confident variant detection.

A 2019 study presented EphaGen, a computational methodology for bioinformatics quality control that estimates the probability of missing any variant from a defined spectrum within a particular sequencing dataset. The study found that this performance measure was superior to conventional bioinformatics metrics such as coverage depth and coverage uniformity for assessing diagnostic sensitivity. This finding suggests that coverage metrics alone may not fully capture the clinical sensitivity of a sequencing run.

GC Bias

GC bias refers to the uneven representation of genomic regions based on their GC content. Some library preparation methods and sequencing platforms amplify or sequence regions with extreme GC content less efficiently than regions with moderate GC content. GC bias can be assessed by plotting coverage or read counts against GC content.

Severe GC bias can lead to poor coverage in GC-rich or GC-poor regions, potentially causing false-negative variant calls in those areas. The acceptable level of GC bias depends on the application and the genomic regions of interest.

Duplication Rate

PCR duplication occurs when multiple copies of the same original DNA fragment are sequenced. High duplication rates reduce the effective coverage and can introduce bias in variant allele frequency estimates. Duplication rates are typically calculated after alignment by identifying reads with identical start positions and insert sizes.

For targeted sequencing panels, some level of duplication is expected. Unique molecular identifiers (UMIs) can distinguish true biological duplicates from PCR duplicates, enabling more accurate variant calling.

Multi-Stage QC Strategies

Quality control should be conducted at multiple stages of sequencing data analysis, also on raw data. A 2014 review in Briefings in Bioinformatics discusses proper quality control procedures and parameters for Illumina technology-based human DNA re-sequencing at three different stages: raw data, alignment, and variant calling. Monitoring quality control metrics at each of the three stages provides unique and independent evaluations of data quality from differing perspectives. Properly conducting quality control protocols at all three stages and correctly interpreting the results are crucial to ensure a successful and meaningful study.

Similarly, a 2014 study in Genomics describes QC3, a quality control tool that monitors quality control metrics at each stage of DNA sequencing data analysis: raw data, alignment, and variant detection. QC3 offers unique features such as detection of batch effects and cross-contamination. The study emphasizes that proper quality control should be conducted at all stages of DNA sequencing data analysis to correctly interpret the quality of the data.

Raw Data QC

Raw data QC examines the sequencing output files before alignment. Key metrics include per-base quality scores, per-sequence quality scores, GC content distribution, adapter content, and duplication levels. Tools that generate these metrics are widely available and should be run on every sequencing dataset.

Alignment QC

After reads are aligned to a reference genome, additional metrics become available. These include mapping rate, coverage depth and uniformity, insert size distribution, and the distribution of reads across chromosomes. Unexpected patterns in these metrics can indicate sample contamination, misidentification, or alignment problems.

Variant Calling QC

At the variant calling stage, QC metrics include transition-to-transversion ratio, heterozygous-to-homozygous ratio, and the distribution of variant allele frequencies. These metrics can reveal systematic errors in variant calling or sample quality issues that were not apparent at earlier stages.

QC for Multi-Sample Experiments

Multi-sample experiments, such as sequencing of tumor-normal pairs, require additional QC metrics to ensure validity of results. A 2017 study in Bioinformatics suggests a workflow for QC of DNA sequencing of tumor-normal pairs that combines well-known single-sample QC metrics with additional metrics specific for tumor-normal pairs. These multi-sample QC metrics still lack standardization, and the study proposes a flexible workflow using tools that produce qcML, a generic XML format for QC of omics experiments.

For tumor-normal pairs, key additional metrics include:

  • Contamination estimates for both tumor and normal samples
  • Concordance of germline variants between tumor and normal samples
  • Sex chromosome copy number consistency
  • Sample identity verification using common SNPs

QC Without a Reference Genome

Not all sequencing projects have a high-quality reference genome available. A 2014 study in Frontiers in Genetics demonstrates that by generating a rapid, non-optimized draft assembly of raw reads, it is possible to obtain reliable and informative QC metrics, thus removing the need for a high-quality reference. The study used benchmark datasets generated from control samples across a range of genome sizes to illustrate that QC inferences made using draft assemblies are broadly equivalent to those made using a well-established reference.

This approach is valuable for non-model organisms and for projects where the reference genome is incomplete or of poor quality. The QC tools described in the study are routinely used in the authors' production facility to assess the quality of sequencing data from non-model organisms.

Internal Standards and Spike-In Controls

Internal standards and spike-in controls provide a way to monitor assay performance within each sample. The SEQC2 study designed a synthetic internal standard spike-in for each actionable mutation target, suitable for use in NGS following hybrid capture enrichment and unique molecular index (UMI) or non-UMI library preparation. When mixed with contrived ctDNA reference samples, internal standards enabled calculation of technical error rate, limit of blank, and limit of detection for each variant at each nucleotide position in each sample. True-positive mutations with variant allele fraction too low for detection by current practice were detected with this method, thereby increasing sensitivity.

The EuroClonality-NGS Working Group established two types of QC to accompany NGS-based immunoglobulin and T cell receptor assays. First, a central polytarget QC is used to monitor the primer performance of each of the EuroClonality multiplex NGS assays. Second, a standardized human cell line-based DNA control is spiked into each patient DNA sample to work as a central in-tube QC and calibrator for minimal residual disease quantification. Having integrated those two reference standards in the ARResT/Interrogate bioinformatic platform, EuroClonality-NGS provides a complete protocol for standardized immunoglobulin and T cell receptor gene rearrangement analysis by NGS with high reproducibility, accuracy, and precision for valid marker identification and quantification in diagnostics of lymphoid malignancies.

Sample Type Considerations

Sample type significantly affects sequencing quality. A 2025 study compared formalin-fixed paraffin-embedded (FFPE) and fresh-frozen (FF) samples using the Illumina TruSight Oncology 500 assay for comprehensive genomic profiling. The study performed 138 DNA and 138 RNA analyses on 69 paired samples. Results demonstrated the significant potential of FF tissue as a primary source of higher-quality genetic material to detect small variants, microsatellite instabilities, and tumor mutational burden compared with FFPE samples. The study found lower concordance in the detection of splice variants, fusions, and copy number variants in paired samples.

The degradation of nucleic acids that occurs during FFPE fixation can lead to unreliable results or hinder analysis. Laboratories that receive FFPE samples should expect lower library quality and should adjust their QC thresholds accordingly. In some cases, samples that fail standard QC metrics may still produce clinically useful data if the relevant regions are adequately covered.

At a Glance: Key QC Metrics and Actions

QC Metric What It Measures Typical Action When Out of Range
Library concentration Quantity of amplifiable library molecules Re-quantify with an alternative method or repeat library preparation
Fragment size distribution Insert size and presence of adapter-dimers Size-select the library or repeat library preparation
Per-base quality scores Base call accuracy across read length Trim low-quality bases or consider the run failed
Coverage depth and uniformity Completeness of target region sequencing Increase sequencing depth or investigate capture efficiency
GC bias Evenness of coverage across GC content Adjust library preparation or enrichment conditions
Duplication rate Proportion of PCR duplicates Use UMI-based methods or reduce PCR cycles
Mapping rate Proportion of reads that align to reference Check for contamination or reference mismatch
Contamination estimate Presence of DNA from other samples Investigate sample handling and re-extract if needed

Practical Implementation of a QC Workflow

Implementing a robust QC workflow requires defined procedures, trained personnel, and documented records. The following steps provide a framework for establishing QC in a sequencing laboratory.

Step 1: Define Acceptance Criteria

Before implementing QC, define what constitutes acceptable quality for each metric. Acceptance criteria should be based on the intended use of the data. A research project exploring novel variants may tolerate lower quality than a clinical diagnostic assay reporting results to physicians.

Document the rationale for each acceptance criterion. When criteria are changed, record the reason for the change and the date it took effect.

Step 2: Establish Sample Receipt QC

Sample receipt QC includes verifying sample identity, assessing sample quantity and quality, and documenting any issues. For tissue samples, assess the tumor content and necrosis percentage. For blood samples, verify the volume and check for hemolysis.

The choice of DNA extraction method affects downstream library quality. Document the extraction method used for each sample and record any deviations from the standard protocol.

Step 3: Implement Library Preparation QC

After library preparation, measure concentration and fragment size for every library. Record these values in the laboratory information management system or a spreadsheet. Compare each library to the acceptance criteria and document any libraries that fail.

For targeted panels, consider adding a quantitative PCR step to measure the amount of amplifiable library molecules. This step is particularly important for samples with degraded DNA, such as FFPE tissues.

Step 4: Perform Sequencing Run QC

After sequencing, review the run-level metrics provided by the instrument. These include cluster density, quality scores, and the percentage of bases above Q30. Compare these metrics to historical runs to identify trends or sudden changes.

If the run-level metrics indicate a problem, determine whether the issue affects all samples or only specific samples. This distinction helps identify whether the problem is instrument-related or sample-related.

Step 5: Conduct Data Analysis QC

After alignment and variant calling, review the data-level QC metrics. These include coverage statistics, mapping rates, and variant calling metrics. Compare these metrics to the acceptance criteria and document any samples that fail.

For clinical samples, consider using a metric that estimates the probability of missing clinically relevant variants, as described in the EphaGen study. This approach provides a more direct measure of diagnostic sensitivity than coverage metrics alone.

Step 6: Document and Review

Document all QC results, including both passing and failing samples. Record the actions taken for failing samples, such as re-extraction, re-library preparation, or re-sequencing. Review QC data regularly to identify trends that may indicate systematic problems.

The World Health Organization Laboratory Quality Management System Handbook describes the requirements for a quality management system, including document control, records, and internal audits. These principles apply to sequencing laboratories as much as to other diagnostic laboratories.

Records and Measurements

Maintaining accurate records is essential for sequencing QC. The following records should be maintained for each sequencing project:

  • Sample receipt forms with sample identity and condition
  • DNA or RNA extraction records with quantity and quality measurements
  • Library preparation records with concentration and fragment size
  • Sequencing run records with instrument metrics
  • Data analysis records with alignment and variant calling metrics
  • Corrective action records for any failures

Records should be legible, indelible, and stored securely. Electronic records should have backup systems and access controls. The retention period for records should be defined in the laboratory quality manual.

Common Failure Patterns

Recognizing common failure patterns helps laboratories troubleshoot problems quickly.

Low Library Yield

Low library yield can result from poor input DNA quality, inefficient ligation, or excessive purification losses. For FFPE samples, DNA degradation is a common cause. If the input DNA quantity was low, consider repeating the library preparation with more input material or using a library preparation method optimized for low-input samples.

Adapter-Dimer Contamination

Adapter-dimers appear as a distinct peak at approximately 120 to 130 base pairs on a fragment size trace. They consume sequencing capacity and reduce the amount of useful data. Adapter-dimers can result from excessive adapter concentration, inefficient AMPure bead purification, or degraded input DNA. Reducing adapter concentration or increasing the number of purification steps can help.

High Duplication Rate

High duplication rates reduce effective coverage and can bias variant allele frequencies. This problem often results from too many PCR cycles during library amplification or from low input DNA quantity. Using UMI-based library preparation methods can help distinguish true duplicates from PCR duplicates.

GC Bias

Severe GC bias can cause poor coverage in GC-rich or GC-poor regions. This problem can result from the library preparation method, the PCR polymerase used, or the enrichment method. If GC bias is consistently observed, consider switching to a different library preparation method or adjusting the PCR conditions.

Sample Contamination

Sample contamination can be detected through unexpected allele frequencies, sex chromosome discrepancies, or the presence of variants from multiple individuals. Contamination can occur during sample collection, DNA extraction, or library preparation. Implementing strict sample handling procedures and using separate work areas for pre- and post-PCR steps can reduce contamination risk.

Batch Effects

Batch effects are systematic differences between groups of samples processed at different times or with different reagent lots. A 2014 study describes QC3, which offers detection of batch effects and cross-contamination. Batch effects can be identified by comparing QC metrics across batches and by using statistical methods to detect clustering by batch.

Limitations of QC Metrics

QC metrics have limitations that users should understand.

Quality Scores Are Estimates

Quality scores are estimates of base call accuracy, not guarantees. The actual error rate can differ from the predicted error rate, particularly for systematic errors that affect specific sequence contexts. Quality score calibration should be monitored over time.

Coverage Metrics Do Not Capture All Problems

As demonstrated by the EphaGen study, conventional metrics such as coverage depth and coverage uniformity may not fully capture the diagnostic sensitivity of a sequencing dataset. A dataset can have adequate coverage by conventional metrics yet still miss clinically relevant variants due to context-specific sequencing errors.

QC Thresholds Are Application-Specific

QC thresholds that are appropriate for one application may not be appropriate for another. A research project using whole-genome sequencing may tolerate lower coverage than a clinical targeted panel. Each laboratory should define thresholds based on the intended use of the data and validate those thresholds with appropriate reference materials.

Reference-Free QC Has Limitations

While draft assembly-based QC can provide useful metrics without a reference genome, this approach has limitations. Draft assemblies may not capture all the information available from a well-annotated reference, and some QC metrics may be less reliable.

Safety and Regulatory Context

Sequencing laboratories must operate within a framework of safety and regulatory requirements.

Biosafety

The World Health Organization Laboratory Biosafety Manual provides guidance on safe handling of biological materials. Sequencing laboratories handle human samples that may contain infectious agents. Standard precautions should be followed when handling blood, tissue, and other biological specimens. The biosafety level of the laboratory should match the risk group of the organisms being handled.

Quality Management Systems

The World Health Organization Laboratory Quality Management System Handbook describes the components of a quality management system, including organization, personnel, equipment, purchasing and inventory, process control, information management, documents and records, occurrence management, assessment, process improvement, customer service, and facilities and safety. Sequencing laboratories should implement these components to ensure reliable results.

Method Validation

The U.S. Food and Drug Administration Bioanalytical Method Validation Guidance describes the expectations for validating bioanalytical methods used in regulatory submissions. While this guidance is primarily directed at pharmacokinetic and toxicokinetic studies, the principles of accuracy, precision, selectivity, sensitivity, reproducibility, and stability apply to sequencing-based assays as well.

The National Center for Advancing Translational Sciences Assay Guidance Manual provides guidance on developing and validating assays for drug discovery and translational research. This resource includes information on quality control, reference standards, and data analysis.

Professional Escalation Criteria

Laboratory personnel should know when to escalate QC failures to supervisors or medical directors. Escalation is appropriate when:

  • A sample fails QC and the result would affect patient management
  • QC metrics indicate a systematic problem that may affect multiple samples
  • A new reagent lot or instrument produces unexpected QC results
  • There is evidence of sample contamination or misidentification
  • QC metrics are borderline and the interpretation is uncertain

The decision to report results from a sample that fails QC should be made by the laboratory director or medical director, not by individual technologists.

Best Practices for Sequencing QC

The following best practices summarize the key recommendations for sequencing quality control.

Use Multiple QC Methods

No single QC metric captures all aspects of library and data quality. Use multiple complementary methods, including fluorometric quantification, electrophoretic size assessment, and quantitative PCR or digital PCR for library quantification.

Monitor Trends Over Time

Individual QC results are most informative when compared to historical data. Track QC metrics over time to identify trends that may indicate gradual changes in reagent quality, instrument performance, or technician technique.

Use Reference Materials

Reference materials with known variants provide a way to monitor assay performance over time. Include reference materials in each sequencing run or at regular intervals to verify that the assay continues to perform as expected.

Document Everything

Document all QC results, including passing and failing results. Record the actions taken for failing samples and the rationale for those actions. This documentation is essential for troubleshooting and for regulatory compliance.

Validate QC Thresholds

QC thresholds should be validated using samples with known quality characteristics. The validation should demonstrate that samples passing the thresholds produce reliable results and that samples failing the thresholds produce unreliable results.

Train Personnel

All personnel involved in sequencing should receive training on QC procedures. Training should include the rationale for each QC step, the interpretation of QC metrics, and the actions to take when QC fails. Competency should be assessed regularly.

Frequently Asked Questions

What is the most important QC metric for sequencing libraries?

No single metric is most important because library quality depends on multiple factors. Concentration and fragment size are the two most basic metrics, and both should be assessed for every library. Concentration determines how much library to load onto the sequencer, while fragment size affects clustering and sequencing performance. For targeted panels, the amount of amplifiable library molecules measured by quantitative PCR or digital PCR may be more informative than total DNA concentration.

How do I choose QC thresholds for my sequencing application?

QC thresholds should be based on the intended use of the data and validated with appropriate reference materials. For clinical applications, thresholds should be set to ensure that clinically relevant variants are detected with acceptable sensitivity. The EphaGen study demonstrates that metrics estimating the probability of missing variants from a defined spectrum can be more informative than conventional coverage metrics. Thresholds should be documented and reviewed regularly.

What should I do when a library fails QC?

When a library fails QC, first determine the cause of the failure. Common causes include poor input DNA quality, inefficient enzymatic reactions, and contamination. Depending on the cause, options include repeating the library preparation with more input material, adjusting reagent concentrations, or using a different library preparation method. For samples with limited material, such as FFPE biopsies, it may not be possible to repeat the library preparation, and the decision to proceed with a failing library should be made by the laboratory director.

How does FFPE sample quality affect sequencing QC?

FFPE samples have degraded nucleic acids due to the fixation process. A 2025 study found that fresh-frozen tissue produced higher-quality genetic material for detecting small variants, microsatellite instabilities, and tumor mutational burden compared with FFPE samples. FFPE samples may produce libraries with lower concentration, smaller fragment sizes, and higher duplication rates. Laboratories should expect lower QC metrics for FFPE samples and should validate their QC thresholds accordingly.

What is the difference between raw data QC and alignment QC?

Raw data QC examines the sequencing output files before alignment and includes metrics such as per-base quality scores, GC content distribution, and adapter content. Alignment QC examines the aligned reads and includes metrics such as mapping rate, coverage depth and uniformity, and insert size distribution. Both stages provide unique information, and a 2014 review recommends conducting QC at raw data, alignment, and variant calling stages for a complete assessment.

How can I detect sample contamination in sequencing data?

Sample contamination can be detected through several methods. Unexpected allele frequencies, such as heterozygous calls at positions expected to be homozygous, may indicate contamination. Sex chromosome discrepancies between the reported sample sex and the observed data can also indicate contamination. Some QC tools, such as QC3 described in a 2014 study, offer specific features for detecting cross-contamination. For tumor-normal pairs, concordance of germline variants between the two samples can help verify sample identity.

What are internal standards and why are they useful for QC?

Internal standards are known sequences that are added to samples before library preparation or sequencing. The SEQC2 study designed synthetic internal standard spike-ins for each actionable mutation target, which enabled calculation of technical error rate, limit of blank, and limit of detection for each variant. The EuroClonality-NGS Working Group uses a standardized human cell line-based DNA control spiked into each patient sample as a central in-tube QC and calibrator. Internal standards provide a way to monitor assay performance within each sample and to detect problems that might not be apparent from sample-level metrics alone.

When should I escalate a QC failure to a supervisor or director?

Escalate a QC failure when the result would affect patient management, when the failure suggests a systematic problem affecting multiple samples, when a new reagent lot or instrument produces unexpected results, or when there is evidence of contamination or sample misidentification. Borderline results that are difficult to interpret should also be escalated. The decision to report results from a sample that fails QC should be made by the laboratory director or medical director.

Related Diagnostic Guides

References and Further Reading

This article is educational and does not replace validated laboratory procedures, institutional biosafety review, manufacturer instructions, or professional interpretation.