Sequencing Library Preparation: Key Steps and Quality Control
Next generation sequencing (NGS) library preparation converts purified DNA or RNA into a format compatible with a sequencing instrument. The process involves fragmenting nucleic acids, attaching platform-specific adapters, and amplifying the resulting molecules. Quality control checkpoints at each stage determine whether a library will produce interpretable data or fail during sequencing. This article provides a practical walkthrough of the library preparation workflow for laboratory students, technicians, researchers, and diagnostic professionals, with emphasis on the decisions that affect downstream data quality.
At a Glance
Library preparation is the bridge between raw biological material and the sequencing instrument. Every step introduces potential bias, and the choices made during preparation directly influence the accuracy of variant detection, gene expression measurements, and other downstream analyses. The table below summarizes the main stages, their purpose, and the quality control considerations that apply at each point.
| Workflow Stage | Primary Purpose | Key Quality Control Considerations |
|---|---|---|
| Nucleic Acid Extraction and Assessment | Obtain pure, intact DNA or RNA from the sample | Yield, purity ratios, integrity assessment, and contamination checks |
| Fragmentation | Generate fragments of appropriate size for the sequencing platform | Fragment size distribution, uniformity of fragmentation, input amount |
| End Repair and A-Tailing | Create blunt ends and add a single adenine for adapter ligation | Enzyme efficiency, reaction completeness, carryover of reagents |
| Adapter Ligation | Attach platform-specific sequences to fragment ends | Ligation efficiency, adapter dimer formation, adapter concentration |
| Size Selection | Remove unwanted fragments and adapter dimers | Recovery of target fragments, reproducibility of selection |
| Amplification | Increase library concentration for sequencing | Number of PCR cycles, amplification bias, duplicate reads |
| Library Quantification and QC | Confirm concentration and fragment size before pooling | Quantification method accuracy, fragment size distribution, contamination |
Core Principles of Library Preparation
The Purpose of a Sequencing Library
A sequencing library is a collection of DNA fragments that have been modified so a sequencing instrument can read their sequences. The modifications include the addition of adapter sequences that provide binding sites for the sequencing primers and the flow cell surface, as well as index sequences that allow multiple samples to be pooled and sequenced together. The library must accurately represent the nucleic acid content of the original sample, because any bias introduced during preparation will be reflected in the final data.
NGS has become a standard tool in modern molecular biology and clinical diagnostics. The preparation of libraries in which DNA or RNA fragments are fused with adapters followed by PCR amplification and sequencing is a fundamental requirement for all sequencing applications. Robust library preparation methods that produce a representative, non-biased source of nucleic acid material are of crucial importance, yet it has become clear that libraries for all types of applications contain biases that compromise data quality and can lead to erroneous interpretation. A detailed understanding of these biases is essential for careful interpretation of sequencing data and for finding ways to improve library quality. Almost all steps of the various protocols have been reported to introduce bias, particularly in RNA sequencing, which is technically more challenging than DNA sequencing. For each type of bias, methods for improvement exist, and researchers should be aware of these options when converting raw nucleic acid into a sequencing library. 9
Input Material Considerations
The quality and quantity of input nucleic acid determine which library preparation strategy is appropriate. High molecular weight DNA from fresh or frozen tissue behaves differently from degraded DNA extracted from formalin-fixed paraffin embedded (FFPE) specimens. RNA inputs require additional considerations, including the need to remove ribosomal RNA or select for polyadenylated transcripts.
For FFPE samples, the crosslinking and degradation that occur during fixation create challenges for library preparation. Some commercial kits have been developed to address these challenges directly. One international performance evaluation study validated a library preparation kit designed for FFPE specimens that eliminates the separate pre-analytical steps of DNA extraction, purification, and isolation. In that study, eleven institutions tested eight FFPE samples previously assessed with standard protocols, and 92.8% of samples were successfully analyzed on both Thermo Fisher Scientific and Illumina platforms. Compared with the standard workflow, the kit detected 90.5% of the variants. The authors concluded that the kit combined with a targeted sequencing panel constituted a convenient, practical, and robust cost-saving solution for FFPE NGS analysis in routine practice. 6
For RNA applications, the choice between poly(A) selection and ribosomal RNA depletion affects both the cost and the information content of the resulting data. A standardized RNA extraction study in Entamoeba species compared six extraction protocols and found that a combined TRIzol and column-based approach provided the highest purity ratios. The study then evaluated two library preparation strategies and found that poly(A) selection was more efficient, yielding higher RNA concentrations and low residual ribosomal RNA, whereas rRNA depletion remained inefficient with approximately 87% rRNA remaining. 18
Bias and Representation
Every enzymatic step in library preparation has the potential to introduce bias. Fragmentation methods differ in their sequence preferences, ligation efficiency varies with fragment ends, and PCR amplification preferentially amplifies certain sequences over others. These biases can lead to uneven coverage across the genome or transcriptome, which in turn affects the sensitivity of variant detection and the accuracy of quantification.
The choice of library preparation kit has a measurable impact on efficiency. A systematic comparison of nine commercially available DNA library preparation kits used droplet digital PCR to quantify the amount of DNA remaining after each protocol step. The study found important variations between kits, with kits that combined several steps into a single reaction exhibiting final yields four to seven times higher than other kits. Adapter ligation yield itself varied by more than a factor of ten between kits, and certain ligation efficiencies were so low that they could impair the original library complexity and impoverish the sequencing results. When a PCR enrichment step was necessary, lower adapter-ligated DNA inputs led to greater amplification yields, hiding the latent disparity between kits. 21
Practical Workflow Steps
Nucleic Acid Extraction and Quality Assessment
The library preparation workflow begins with nucleic acid extraction. The extraction method must be appropriate for the sample type and the downstream application. For DNA sequencing, the extracted DNA should be free of proteins, RNA, and chemical contaminants that could inhibit enzymatic reactions. For RNA sequencing, the RNA must be intact enough to represent the transcriptome accurately.
Quality assessment of the input nucleic acid includes measurement of concentration, purity, and integrity. Spectrophotometric measurements provide absorbance ratios that indicate protein and chemical contamination. Fluorometric methods using intercalating dyes provide more accurate concentration measurements for double-stranded DNA. Automated electrophoresis systems provide fragment size distributions and RNA integrity scores.
The importance of input quality is particularly evident for challenging sample types. For small RNA sequencing from urinary exosomes, RNA quality was assessed by spectrophotometric quantification and Bioanalyzer software analysis before library preparation. The study demonstrated that good quality sequencing libraries could be prepared following an optimized small RNA library preparation protocol, but only when the input RNA was adequately characterized. 20
Fragmentation
Fragmentation reduces high molecular weight nucleic acids to sizes compatible with the sequencing platform. The optimal fragment size depends on the sequencing instrument and the application. Whole genome sequencing typically requires larger fragments than amplicon-based targeted sequencing.
Three main fragmentation approaches are used in library preparation. Enzymatic fragmentation uses endonucleases to cut DNA at specific sequences or randomly. Mechanical fragmentation uses sonication or nebulization to shear DNA physically. Transposase-based methods combine fragmentation and adapter insertion in a single step.
The choice of fragmentation method affects the uniformity of coverage and the representation of difficult genomic regions. Enzymatic methods are more convenient and require less input DNA, but they may introduce sequence bias. Mechanical methods are more random but require more input material and additional cleanup steps.
For mitochondrial genome sequencing, one optimization study used enzymatic fragmentation after long-range amplification of the mitochondrial genome. The study designed custom primers after initial attempts with pre-made commercial primers were unsuccessful. The optimized protocol produced clear, specific amplicons of 9.8 and 8.5 kilobases, and sequencing demonstrated high-quality reads with an average coverage depth of 742x and a GC content of 43 to 45%. 25
End Repair and A-Tailing
After fragmentation, the DNA fragments have heterogeneous ends that are not compatible with adapter ligation. End repair converts these ends into blunt ends by filling in 5-prime overhangs and removing 3-prime overhangs. A-tailing then adds a single adenine to the 3-prime ends, creating a compatible overhang for ligation to adapters that have a thymine at their 3-prime ends.
These enzymatic steps are efficient but not perfect. Incomplete end repair reduces ligation efficiency, and excessive A-tailing enzyme can remove nucleotides from the fragment ends. The reactions are typically performed in a single tube with a cleanup step afterward to remove enzymes and buffers.
Adapter Ligation
Adapter ligation attaches the platform-specific sequences to the fragment ends. The adapters contain sequences required for cluster generation, sequencing primer binding, and sample indexing. The ligation reaction uses DNA ligase to join the adapter to the A-tailed fragment.
Ligation efficiency is a critical determinant of library complexity. The droplet digital PCR study described earlier found that adapter ligation yield varied by more than a factor of ten between kits. Low ligation efficiency means that many fragments are lost before amplification, reducing the complexity of the final library and potentially introducing bias. 21
Adapter dimer formation is a common problem during ligation. Adapter dimers are molecules consisting of two adapters ligated together without an intervening insert. These molecules sequence efficiently and consume sequencing capacity without providing useful data. Size selection steps are designed to remove adapter dimers, but the effectiveness of removal depends on the size difference between the dimers and the target fragments.
For small RNA libraries, adapter dimer contamination is a particular concern because the inserts are short. A study optimizing small RNA library preparation from urinary exosomes found that when a size selection by gel purification step was included within the workflow, adapter dimers were completely removed from cDNA libraries. The inclusion of this modification step augmented the small RNA mapped reads, with a significant 37% increase in miRNA reads, and the gel purification step made no difference to the tagged miRNA population. 20
Size Selection
Size selection removes fragments that are too short or too long for the sequencing platform. Short fragments, including adapter dimers, waste sequencing capacity. Long fragments may not cluster efficiently or may produce reads that span beyond the read length.
Two main approaches are used for size selection. Gel-based methods provide precise size selection but are labor-intensive and have variable recovery. Bead-based methods using solid phase reversible immobilization (SPRI) beads are more convenient and scalable but provide broader size ranges.
The choice of size selection method affects the reproducibility of the library preparation. Bead-based methods are more commonly used in high-throughput workflows because they are amenable to automation and produce consistent results across samples.
Amplification
PCR amplification increases the concentration of the library to levels required for sequencing. The number of PCR cycles must be optimized for the input amount and the application. Too few cycles produce insufficient library, while too many cycles introduce amplification bias and increase the proportion of duplicate reads.
Amplification bias occurs because PCR preferentially amplifies certain sequences based on their GC content and secondary structure. This bias can distort the representation of the original sample, particularly for RNA sequencing where transcript abundances are being measured.
The number of PCR cycles should be minimized whenever possible. Lower input amounts require more cycles, which increases bias. Some library preparation methods have been developed to reduce the number of steps and therefore reduce the opportunities for bias to be introduced.
One approach to streamlining the workflow uses real-time quantitative PCR (qPCR) to combine amplification and quantification in a single step. This fluorescent amplification for NGS (FA-NGS) method replaced conventional PCR and quantification with qPCR using SYBR Green I. The qPCR enabled individual library quantification for pooling in a single tube without the need for additional reagents, and a melting curve analysis was implemented as an intermediate quality control test to confirm successful amplification. Sequencing analysis showed comparable percent reads for each indexed library, demonstrating that pooling calculations based on qPCR allow for an even representation of sequencing reads. The modified workflow had fewer overall steps and therefore less risk of user error. 7
Library Quantification and Quality Control
Before pooling and sequencing, each library must be quantified and assessed for quality. Quantification determines the concentration of amplifiable library molecules, which is used to calculate pooling volumes. Quality assessment confirms that the fragment size distribution is appropriate and that adapter dimers and other contaminants are minimal.
Quantification methods include fluorometric measurement of double-stranded DNA, qPCR, and droplet digital PCR. Fluorometric methods measure total DNA concentration, including non-amplifiable molecules. qPCR and droplet digital PCR measure only molecules that can be amplified, providing a more accurate estimate of the sequencing-ready library concentration.
The FA-NGS study demonstrated that qPCR-based quantification could be used directly for pooling calculations, eliminating the need for separate quantification steps. The melting curve analysis provided an intermediate quality control check that confirmed successful amplification before proceeding to pooling. 7
Options and Tradeoffs in Library Preparation
Manual versus Automated Workflows
Manual library preparation is flexible and requires minimal specialized equipment, but it is labor-intensive and prone to user error. Automated workflows reduce hands-on time and improve reproducibility but require capital investment and technical expertise.
The challenges of manual workflows have driven the development of automated systems. One study described an enclosed automated targeted NGS library preparation system that could produce qualified targeted amplicon libraries in three steps with only 15 minutes of hands-on time. Rigorous cross-contamination testing using simulated contaminant plasmids confirmed that the design of disposable cassettes guaranteed zero sample cross-contamination. The system showed 100% accuracy and precision in detecting germ-line and somatic mutations, and the panels showed 100% concordance with verified methods in a prospective cohort study enrolling 363 patients and a cohort of 45 pan-cancer samples. The authors concluded that the automated platform could overcome major challenges for implementing NGS assays clinically. 8
For smaller laboratories with low to medium throughput, large-scale liquid handlers may be prohibitively expensive. A proof-of-concept study demonstrated library preparation on a commercially available and open lab-on-a-chip platform that provided an alternative automation approach. The platform covered common library preparation steps optimized to a microfluidic environment, including customizable PCR for target enrichment, end repair, adapter ligation, nucleic acid purification via magnetic beads, and an integrated quantification step. Processing reference cell-free DNA samples in the cartridge revealed highly comparable results to manual processing with a Pearson correlation of 0.94 based on amplicon sequencing. 13
Ligation-Based versus Transposase-Based Methods
Ligation-based methods fragment the DNA first and then ligate adapters in a separate step. Transposase-based methods use a transposase enzyme that simultaneously fragments the DNA and inserts adapter sequences. Transposase-based methods are faster and require less input DNA, but they may introduce sequence bias at the insertion sites.
The choice between these methods depends on the application and the input material. For challenging samples with low DNA input, transposase-based methods may be the only practical option. For applications requiring uniform coverage, ligation-based methods may provide better results.
A comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of poliovirus amplicons evaluated ligation-based (native barcoding) and transposase-based (rapid barcoding) approaches. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility. Despite these differences, both methods produced identical consensus sequences across all samples, with comparable read quality. Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time from 200 to 55 minutes and per-sample cost from $16.54 to $12.82, while simplifying the workflow and reducing technical complexity. The authors concluded that sequencing yield may not be a determinant of downstream analytical outcomes for this application, and that rapid barcoding represents a cost-effective and efficient approach for routine surveillance. 15
Targeted versus Whole Genome Approaches
Targeted library preparation methods enrich for specific genomic regions before sequencing. This approach reduces the amount of sequencing required and allows higher coverage of the regions of interest. Whole genome methods sequence the entire genome, providing a more comprehensive view but requiring more sequencing capacity.
Target enrichment can be performed using amplicon-based methods, where PCR primers amplify specific regions, or hybridization-based methods, where probes capture specific regions from a fragmented library. Amplicon-based methods are simpler and faster but may miss variants in regions with primer binding site mutations. Hybridization-based methods are more comprehensive but require more input DNA and longer workflows.
A study adapting a target-enrichment library preparation workflow for use on FFPE gastric biopsies to investigate Helicobacter pylori demonstrated the utility of this approach for clinical samples. The study modified the Agilent SureSelect XT protocol for implementation on an automated system and used RNA probes targeting key genes associated with virulence, antibiotic resistance, and multilocus sequence typing. Mutations linked to macrolide resistance, levofloxacin resistance, and rifamycin resistance were accurately detected, and the multilocus sequence typing profiles were consistent with those obtained via Sanger sequencing. 22
Observations and Measurements
Quality Control Metrics
Several metrics are used to assess library quality at different stages of the workflow. These metrics provide objective criteria for deciding whether to proceed with sequencing or to repeat the preparation.
DNA concentration measurements are used throughout the workflow to track recovery and to calculate input amounts for subsequent steps. Fluorometric methods are preferred for double-stranded DNA because they are specific and sensitive. Spectrophotometric methods provide purity ratios that indicate contamination.
Fragment size distribution is assessed using automated electrophoresis systems. The distribution should show a peak at the expected fragment size with minimal adapter dimer contamination. The presence of a large adapter dimer peak indicates that size selection was ineffective.
Library concentration is measured after amplification and cleanup. The concentration should be sufficient for the sequencing platform and application. Low concentrations may indicate problems with ligation or amplification efficiency.
Records and Documentation
Accurate record keeping is essential for troubleshooting and for quality assurance in diagnostic applications. Records should include the sample identifier, extraction method, input amount, library preparation kit and lot number, fragmentation method, PCR cycle number, size selection method, and all quality control measurements.
The World Health Organization Laboratory Quality Management System Handbook provides guidance on the documentation practices that support reliable laboratory testing. 1 The handbook emphasizes the importance of standard operating procedures, records, and quality control in producing reliable results.
For diagnostic applications, the U.S. Food and Drug Administration Bioanalytical Method Validation Guidance describes the expectations for method validation and documentation. 4 While this guidance is focused on bioanalytical methods, the principles of validation, documentation, and quality control apply to NGS library preparation in regulated settings.
Common Failure Patterns
Library preparation failures can be categorized based on where in the workflow the problem occurs. Recognizing the pattern of failure helps identify the root cause and determine the appropriate corrective action.
Low library yield after amplification can result from insufficient input DNA, inefficient ligation, excessive cleanup losses, or too few PCR cycles. The droplet digital PCR study demonstrated that ligation efficiency varies substantially between kits, and low ligation efficiency can impair library complexity. 21
Adapter dimer contamination appears as a peak at approximately 120 to 130 base pairs in the fragment size distribution. This problem is more common in small RNA libraries where the insert size is short. The small RNA optimization study found that gel purification completely removed adapter dimers from cDNA libraries. 20
Uneven coverage across the target regions can result from amplification bias, fragmentation bias, or inefficient capture. The review of library preparation bias noted that almost all steps of the various protocols have been reported to introduce bias, especially in RNA sequencing. 9
High duplicate rates in the sequencing data indicate that the library complexity was low. This can result from insufficient input DNA, excessive PCR amplification, or loss of fragments during cleanup steps.
Quality and Welfare Controls
Biosafety Considerations
Library preparation involves handling biological samples that may contain infectious agents. The World Health Organization Laboratory Biosafety Manual provides guidance on the safe handling of biological materials. 2 Laboratory personnel should follow institutional biosafety policies and use appropriate personal protective equipment when handling samples.
The biosafety considerations for library preparation include the risk of exposure to bloodborne pathogens when processing blood samples, the risk of aerosol generation during mechanical fragmentation, and the risk of contamination when handling amplified products. Work areas should be organized to separate pre-amplification and post-amplification steps to prevent contamination.
Contamination Control
Contamination is a major concern in library preparation because it can lead to incorrect results. Sources of contamination include cross-contamination between samples, contamination from previously amplified products, and contamination from the laboratory environment.
The enclosed automated system described earlier was designed to address contamination concerns. The disposable cassette design guaranteed zero sample cross-contamination in rigorous testing using simulated contaminant plasmids. 8 For manual workflows, the use of dedicated pipettes, filtered tips, and separate work areas for pre- and post-amplification steps is essential.
The National Center for Advancing Translational Sciences Assay Guidance Manual provides guidance on assay development and quality control that is applicable to NGS library preparation. 3 The manual emphasizes the importance of appropriate controls and quality metrics in developing reliable assays.
Professional Escalation Criteria
Laboratory personnel should know when to escalate problems to a supervisor or seek technical support from the kit manufacturer. Escalation is appropriate when quality control metrics consistently fail despite troubleshooting, when the root cause of a failure cannot be identified, or when results are needed for clinical decisions.
The National Center for Biotechnology Information provides literature resources that can help troubleshoot library preparation problems. 5 Searching the literature for specific failure patterns can identify solutions that have been reported by other laboratories.
Limitations and Interpretation
Sources of Bias
All library preparation methods introduce some degree of bias. The review of library preparation bias emphasized that robust library preparation methods that produce a representative, non-biased source of nucleic acid material are of crucial importance, yet libraries for all types of applications contain biases that compromise the quality of NGS datasets and can lead to erroneous interpretation. 9
The practical implication is that sequencing data should be interpreted with an understanding of the limitations of the library preparation method. Variants detected at low allele frequencies may be artifacts of amplification bias. Differences in gene expression between samples may reflect differences in library preparation efficiency instead of true biological differences.
Input Amount Limitations
The amount of input nucleic acid required for library preparation varies by method and application. Most standard protocols require microgram amounts of DNA or nanogram amounts of RNA. Clinical samples often provide much less material.
Microfluidic approaches have been developed to address the gap between the tiny quantities of biomaterials provided by clinical samples and the large DNA input required by most assays. One study presented a microfluidic droplet-based system for NGS library preparation capable of reducing the number of pipetting steps significantly, reducing reagent consumption by tenfold, and automating much of the process, while supporting an extremely low DNA input requirement of 10 picograms per library. This semiautomated technology allowed for low-input preparations of eight libraries simultaneously while reducing batch-to-batch variation and operator hands-on time. 12
Platform-Specific Considerations
Different sequencing platforms require different library preparation approaches. The adapter sequences, fragment size ranges, and quantification methods differ between platforms. Laboratories should select library preparation methods that are compatible with their sequencing platform.
For Oxford Nanopore sequencing, the choice between rapid and native barcoding methods involves tradeoffs between yield, cost, and hands-on time. The poliovirus study found that native barcoding generated more sequencing output and demonstrated higher run-to-run reproducibility, but rapid barcoding reduced hands-on time and per-sample cost while producing identical consensus sequences. 15
Specialized Applications
Small RNA Library Preparation
Small RNA sequencing requires specialized library preparation methods because the inserts are short and the input amounts are often limited. The choice of library preparation kit significantly affects the results.
A comparative study of four commercial small RNA library preparation kits evaluated their performance in profiling miRNAs from cell-free saliva, plasma, and their extracellular vesicles. Using both synthetic reference and biological samples, the study assessed the kits efficiency in handling low RNA input, minimizing bias, and detecting diverse miRNAs. One kit outperformed the others, showing the highest miRNA mapping rates, minimal adapter dimers, and the broadest miRNA detection, particularly in saliva. The study underscored the critical impact of library preparation on miRNA sequencing outcomes and offered guidance for selecting optimal protocols for biomarker discovery from non-invasive sample matrices. 23
RNA Sequencing from Biofluids
Extracellular or cell-free RNAs derived from biofluids are utilized in biomarker studies to study health and disease. A protocol for total RNA sequencing analysis of blood plasma exRNA using the Switching Mechanism At 5-prime End of RNA Template (SMART) cDNA synthesis technology describes all steps from blood plasma preparation to sequencing data analysis. In addition to blood plasma, this protocol can also be applied to exRNA purified from other human, murine, and rat biofluids. 16
Ribosome Profiling
Ribosome profiling (RIBO-seq) is a technique for studying protein synthesis in vivo by precisely mapping the position and number of ribosomes on an mRNA transcript. The key step is rapid inhibition of translation and adequate disintegration of cells. A protocol for bacteria suggests filtration and flash-freezing in liquid nitrogen for cell harvesting with an optional pretreatment with chloramphenicol to arrest translation. For disintegration, the protocol proposes grinding frozen cells with mortar and pestle in the presence of aluminum oxide to mechanically disrupt the cell wall. For library preparation, the protocol recommends using a commercially available small RNA kit for Illumina sequencing, following manufacturer guidelines with some degree of optimization. 19
Pseudouridine Sequencing
Pseudouridine is an abundant RNA modification present in diverse non-coding RNA species and also exists in mammalian mRNA. Bisulfite-induced deletion sequencing (BID-seq) provides a quantitative method to map RNA pseudouridine distribution transcriptome-wide at single-base resolution. The optimized protocol includes fast steps in both library preparation and data analysis and generates highly reproducible results by inducing high deletion ratios at pseudouridine modifications within diverse sequence contexts while displaying almost zero background deletions at unmodified uridines. The workflow takes five days to complete and includes RNA preparation, library construction, next-generation sequencing, and data analysis. Library construction can be completed by researchers who have basic knowledge and skills in molecular biology and genetics. 11
piRNA Sequencing
The piRNA pathway in planarian flatworms is operated by three PIWI proteins, and the identification of piRNA sequences by next-generation sequencing is imperative for understanding their function. A bioinformatics analysis pipeline for processing and systematic characterization of planarian piRNAs includes steps for the removal of PCR duplicates based on unique molecular identifier sequences and accounts for piRNA multimapping to different loci in the genome. The protocol includes a fully automated pipeline that is freely available at GitHub. 10
T Cell Receptor Sequencing
T cell receptor sequencing enables high-resolution characterization of the adaptive immune repertoire. A practical workflow for bulk TCR sequencing using buffy coat as the starting material does not require T cell enrichment. The study assessed key pre-analytical factors affecting RNA yield and integrity, including red blood cell contamination, DMSO concentration during cryopreservation, and blood collection tube type, and identified conditions that minimize RNA degradation. The study also introduced a TRAC-targeting RT-qPCR assay as a cost-effective quality control step to quantify T cell-specific RNA prior to library preparation. A comparison of three commercial library preparation kits found that one kit reproducibly yielded the highest clonotype recovery and lowest noise, even without enrichment. 14
Frequently Asked Questions
What is the most important quality control step in library preparation?
The most important quality control step depends on the application, but library quantification before pooling is critical because it determines the proportion of reads assigned to each sample. Inaccurate quantification leads to uneven coverage across samples and wasted sequencing capacity. The FA-NGS study demonstrated that qPCR-based quantification could be used directly for pooling calculations, eliminating the need for separate quantification steps and reducing the risk of user error. 7
How much input DNA is required for library preparation?
The required input amount varies by method and application. Standard protocols typically require nanogram to microgram amounts of DNA, but specialized methods have been developed for low-input samples. One microfluidic system supported an extremely low DNA input requirement of 10 picograms per library. 12 Laboratories should consult the kit manufacturer instructions for the specific input range and adjust their extraction and quantification methods accordingly.
What causes adapter dimers and how can they be prevented?
Adapter dimers form when adapters ligate to each other without an intervening insert. This problem is more common in small RNA libraries where the inserts are short. Size selection is the primary method for removing adapter dimers. A study of small RNA library preparation from urinary exosomes found that gel purification completely removed adapter dimers from cDNA libraries and increased miRNA reads by 37%. 20
How many PCR cycles should be used for library amplification?
The optimal number of PCR cycles depends on the input amount and the application. The goal is to produce enough library for sequencing while minimizing amplification bias. Lower input amounts require more cycles, which increases bias. The FA-NGS study used qPCR to combine amplification and quantification, allowing the reaction to be monitored in real time and stopped at the appropriate point. 7
What is the difference between ligation-based and transposase-based library preparation?
Ligation-based methods fragment the DNA first and then ligate adapters in a separate step. Transposase-based methods use a transposase enzyme that simultaneously fragments the DNA and inserts adapter sequences. Transposase-based methods are faster and require less input DNA, but they may introduce sequence bias. A comparison of these approaches for poliovirus sequencing found that the ligation-based method generated more sequencing output but required more hands-on time and cost more per sample. 15
Can library preparation be performed on FFPE samples?
Yes, but FFPE samples present challenges due to crosslinking and degradation. Some commercial kits have been developed specifically for FFPE specimens. An international study validated a library preparation kit that eliminates the separate pre-analytical steps of DNA extraction, purification, and isolation from FFPE specimens. The study found that 92.8% of samples were successfully analyzed and the kit detected 90.5% of variants compared with standard protocols. 6
How does library preparation affect the detection of low-frequency variants?
Library preparation bias can affect the detection of low-frequency variants by distorting the representation of the original sample. Amplification bias preferentially amplifies certain sequences, which can either mask true variants or create artifacts. The review of library preparation bias noted that almost all steps of the various protocols have been reported to introduce bias. 9 Using fewer PCR cycles and optimizing the library preparation method for the specific application can reduce this bias.
What should be done when library preparation fails?
When library preparation fails, the first step is to review the quality control data to identify where in the workflow the problem occurred. Low yield after amplification may indicate problems with input DNA, ligation, or PCR. Adapter dimer contamination indicates problems with size selection. If the root cause cannot be identified, consult the kit manufacturer technical support or search the literature for similar problems. The National Center for Biotechnology Information provides literature resources that can help troubleshoot library preparation problems. 5
Related Diagnostic Guides
- DNA Shearing for NGS Library Preparation: Methods and Quality Control
- How to Interpret DNA Sequencing Chromatograms: Peaks, Quality, and Heterozygotes
- Procedure for Quality Control: Step-by-Step Implementation in a Molecular Lab
- DNA Ligation Troubleshooting: Common Problems and Solutions for Cloning Success
- Gel Electrophoresis Quality Control: Assessing DNA Integrity and Purity
References and Further Reading
- Laboratory Quality Management System Handbook. World Health Organization.
- Laboratory Biosafety Manual. World Health Organization.
- Assay Guidance Manual. National Center for Advancing Translational Sciences.
- Bioanalytical Method Validation Guidance. U.S. Food and Drug Administration.
- NCBI Literature Resources. National Center for Biotechnology Information.
- TargetPlex FFPE-Direct DNA Library Preparation Kit for SiRe NGS panel: an international performance evaluation study.. Journal of clinical pathology, 2022.
- Fluorescent amplification for next generation sequencing (FA-NGS) library preparation.. BMC genomics, 2020.
- Development and clinical applications of an enclosed automated targeted NGS library preparation system.. Clinica chimica acta, international journal of clinical chemistry, 2023.
- Library preparation methods for next-generation sequencing: tone down the bias.. Experimental cell research, 2014.
- Genome-Wide Analysis of Planarian piRNAs.. Methods in molecular biology (Clifton, N.J.), 2023.
- BID-seq for transcriptome-wide quantitative sequencing of mRNA pseudouridine at base resolution.. Nature protocols, 2024.
- Microfluidic Platform for Next-Generation Sequencing Library Preparation with Low-Input Samples.. Analytical chemistry, 2020.
- Automation of customizable library preparation for next-generation sequencing into an open microfluidic platform.. Scientific reports, 2024.
- A practical and accessible workflow for bulk TCR sequencing from buffy coat samples without T cell enrichment.. 2026.
- Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.. 2026.
- Protocol for total RNA sequencing analysis of extracellular RNA from biofluids.. 2026.
- Dataset on the transcriptomic responses of 'Fuji Hubrax' apples to deep seawater application.. 2026.
- Standardized RNA extraction protocol for <,i>,Entamoeba<,/i>, species: advancing molecular diagnostics and amebiasis control.. 2026.
- RIBO-seq in Bacteria: a Sample Collection and Library Preparation Protocol for NGS Sequencing.. Journal of Visualized Experiments, 2021.
- Optimization of small RNA library preparation protocol from human urinary exosomes. Journal of Translational Medicine, 2020.
- Quantitation of next generation sequencing library preparation protocol efficiencies using droplet digital PCR assays - a systematic comparison of DNA library preparation kits for Illumina sequencing. BMC Genomics, 2016.
- Assessment of target-enrichment library preparation and next-generation sequencing of paraffin-embedded gastric biopsies for H. pylori diagnosis and evaluation of virulome and resistome. Scientific Reports, 2025.
- Optimizing Small RNA Sequencing for Salivary Biomarker Identification: A Comparative Study of Library Preparation Protocols. International Journal of Molecular Sciences, 2025.
- Efficient In-house Coupling of Sample and Library Preparation for ChIP-Seq of Histone Modifications in Complex Plant Tissues.. Methods in molecular biology, 2025.
- OPTIMIZATION OF ILLUMINA® NEXTERA™ XT LIBRARY PREPARATION FOR THE MITOCHONDRIAL GENOME SEQUENCING AND CONFIRMATORY SANGER SEQUENCING.. Medicinski glasnik, 2025.
- Library Construction for NGS. Learning Materials in Biosciences, 2021.
This article is educational and does not replace validated laboratory procedures, institutional biosafety review, manufacturer instructions, or professional interpretation.