How to Prepare Samples for Sequencing: A Step-by-Step Guide

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Prepare Samples for Sequencing: A Step-by-Step Guide

Introduction to Sample Preparation for Sequencing

Sample preparation for sequencing is the multi-step process of converting a biological specimen—blood, tissue, cultured cells, soil, or even single cells—into a purified nucleic acid library that is compatible with a next-generation sequencing (NGS) platform. The process encompasses everything from cell lysis and nucleic acid extraction to enzymatic manipulation, amplification, and quality assessment. The goal is to produce a library of DNA fragments, each flanked by platform-specific adapter sequences, at a concentration and size distribution that the sequencer can read efficiently.

The quality of your sequencing data is determined long before the sequencer runs. Every step in sample preparation introduces potential biases, losses, or contaminants that propagate through the workflow. A poorly prepared sample cannot be rescued by deeper sequencing; it will simply produce more data of the same poor quality. Understanding the mechanistic basis of each step—why certain buffers are used, why specific enzymes are chosen, why temperature thresholds matter—allows you to troubleshoot effectively and make informed decisions when protocols deviate from the ideal.

Why Sample Prep Matters

Sequencing platforms, whether Illumina, Ion Torrent, or Oxford Nanopore, do not sequence native DNA directly. They require a library: a collection of DNA fragments of defined length, flanked by known adapter sequences that serve as priming sites for amplification and as anchors for surface attachment. For RNA sequencing, the RNA must first be converted to complementary DNA (cDNA) via reverse transcription. The library preparation process determines:

  • Read length and coverage uniformity: Fragment size distribution dictates how many bases are read from each molecule.
  • Quantitative accuracy: PCR amplification bias can distort the true representation of sequences in the original sample.
  • Error rates: Damaged or chemically modified nucleic acids can be misread by polymerases.
  • Multiplexing capacity: Index sequences added during library prep allow multiple samples to be sequenced in a single run.

A typical NGS workflow proceeds from sample collection through nucleic acid extraction, quality assessment, library construction, quantification, pooling, and finally loading onto the sequencer. Each stage has specific quality metrics that predict success at the next stage.

Overview of NGS Workflow

The canonical workflow for DNA sequencing is as follows:

  1. Sample collection and lysis: Cells are disrupted to release nucleic acids.
  2. Purification: DNA or RNA is separated from proteins, lipids, polysaccharides, and other cellular debris.
  3. Quality control: Concentration, purity, and integrity are assessed.
  4. Fragmentation: High-molecular-weight DNA is sheared into fragments of 200–600 bp.
  5. End repair and A-tailing: Blunt-ended fragments are generated and given a single 3′ adenine overhang.
  6. Adapter ligation: Y-shaped or forked adapters with complementary thymine overhangs are ligated to both ends.
  7. Size selection: Fragments of the desired length are selected, typically by bead-based methods.
  8. PCR amplification: Adapter-ligated fragments are amplified to add full-length adapter sequences and indices.
  9. Final QC and quantification: Libraries are quantified by qPCR or fluorometry and assessed for size distribution.
  10. Pooling and sequencing: Libraries are combined in equimolar ratios and loaded onto the flow cell.

For RNA, an additional reverse transcription step converts RNA to cDNA before adapter ligation. For specialized applications like Bisulfite Sequencing or CHIP Sequencing, additional chemical or immunoprecipitation steps are interposed.

Nucleic Acid Extraction and Purification

The first step in preparing a sample for sequencing is isolating nucleic acids in a form that is free of inhibitors and intact enough to yield full-length sequence information. The choice of extraction method depends on the sample type, the nucleic acid of interest, and the downstream application.

Choosing the Right Extraction Method

Three broad categories of extraction methods dominate: organic extraction, silica column-based kits, and magnetic bead-based methods.

Organic extraction (the classic phenol-chloroform method) remains the gold standard for yield and quality, particularly for difficult samples such as muscle, skin, or plant tissues rich in polysaccharides. The mechanism is phase partitioning: cells are lysed in a chaotropic buffer containing guanidinium thiocyanate or SDS, then mixed with phenol and chloroform. Proteins partition into the organic phase, lipids into the interphase, and nucleic acids remain in the aqueous phase. The aqueous phase is then precipitated with ethanol or isopropanol in the presence of salt (typically 0.3 M sodium acetate, pH 5.2) to neutralize the negatively charged phosphate backbone and drive DNA out of solution. The DNA pellet is washed with 70% ethanol to remove residual salt and chaotropes, then resuspended in a low-ionic-strength buffer such as 10 mM Tris-Cl, pH 8.0, or nuclease-free water.

Column-based kits (e.g., Qiagen DNeasy, Zymo Quick-DNA) use silica membranes that bind nucleic acids in the presence of high concentrations of chaotropic salts (typically 4–6 M guanidinium hydrochloride or guanidinium thiocyanate). The mechanism is electrostatic: at high ionic strength, the negatively charged phosphate backbone of DNA displaces water molecules and binds to the silanol groups on the silica surface. Contaminants—proteins, polysaccharides, and salts—pass through the column during centrifugation or vacuum steps. After washing with an ethanol-containing buffer, the nucleic acid is eluted in a low-salt buffer or water, which disrupts the electrostatic interaction. Column methods are faster and more reproducible than organic extraction, but can suffer from lower yields for very small or very large fragments, and some columns have a binding capacity limit (typically 10–100 µg depending on the product).

Magnetic bead-based methods (e.g., Agencourt AMPure XP, MagAttract) use carboxylated paramagnetic beads that reversibly bind nucleic acids in the presence of polyethylene glycol (PEG) and salt. The PEG acts as a crowding agent, excluding water and promoting the precipitation of DNA onto the bead surface. The beads are captured with a magnet, washed with 70% ethanol, and eluted in water or low-salt buffer. The critical advantage of bead-based methods is that the bead-to-sample ratio controls the size cutoff: higher PEG concentrations favor binding of smaller fragments, allowing size selection and purification in a single step. This is why bead-based cleanups are the standard in library preparation.

For RNA extraction, the same principles apply, but with additional considerations. RNA is single-stranded and highly susceptible to degradation by ubiquitous RNases. The chaotropic agent guanidinium thiocyanate is a potent RNase inhibitor, which is why the single-step acid-phenol method (Chomczynski and Sacchi) and its column-based derivatives (e.g., Qiagen RNeasy) are standard. For mRNA applications, poly(A) tail selection using oligo-dT magnetic beads enriches for messenger RNA and depletes ribosomal RNA, which constitutes up to 80% of total RNA.

Assessing Nucleic Acid Quality and Quantity

After extraction, you must verify that the nucleic acid is intact, pure, and at sufficient concentration. The key metrics are:

  • A260/A280 ratio: Absorbance at 260 nm measures nucleic acids; absorbance at 280 nm measures protein (aromatic amino acids). A pure DNA sample has a ratio of 1.8; pure RNA has a ratio of 2.0. Lower ratios indicate protein contamination.
  • A260/A230 ratio: Absorbance at 230 nm detects chaotropic salts, phenol, and carbohydrates. Values should be 2.0–2.2. Lower values indicate carryover of extraction reagents, which can inhibit downstream enzymes.
  • Integrity: Assessed by gel electrophoresis or microfluidic analysis (see below).

It is important to note that spectrophotometric quantification measures total nucleic acid, including degraded fragments and single-stranded DNA, which may not be amplifiable. For accurate library quantification, fluorometric methods are preferred.

Quality Control and Quantification

Accurate quantification and integrity assessment are critical because they determine how much input material to use for library preparation and whether the material is suitable at all. Under-quantification leads to overloading the library prep reaction, which can cause adapter dimers and PCR artifacts; over-quantification leads to low library yield.

Spectrophotometry vs. Fluorometry

Spectrophotometry (NanoDrop or similar) measures absorbance at 260 nm and calculates concentration using the Beer-Lambert law. It is fast, requires only 1–2 µL of sample, and provides purity ratios. However, it cannot distinguish between DNA, RNA, free nucleotides, and degraded fragments. A sample with significant degradation will show a normal A260 reading despite containing very little amplifiable high-molecular-weight DNA.

Fluorometry (Qubit, DeNovix) uses intercalating dyes that bind specifically to double-stranded DNA (e.g., PicoGreen), single-stranded DNA (e.g., RiboGreen for RNA), or total nucleic acids. The dye only fluoresces when bound to the target, so the measurement is specific to intact, double-stranded molecules. Fluorometry is the method of choice for quantifying input material for library preparation and for final library quantification, because it measures only the molecules that will actually be sequenced.

RNA Integrity Number (RIN) and DNA Integrity Number (DIN)

Integrity assessment is performed on microfluidic platforms such as the Agilent Bioanalyzer or TapeStation, which separate nucleic acids by size in a microchannel and detect them by laser-induced fluorescence. The output is an electropherogram showing the size distribution of fragments.

For RNA, the RNA Integrity Number (RIN) is calculated by an algorithm that considers the ratio of 28S to 18S ribosomal RNA peaks, the presence of degradation products, and the overall shape of the trace. RIN values range from 1 (fully degraded) to 10 (intact). For standard RNA-seq, a RIN of 7 or higher is generally acceptable, though some protocols can tolerate lower integrity. For single-cell RNA-seq, where the starting material is limiting, RIN is less informative because the RNA is already fragmented by the lysis process.

For DNA, the DNA Integrity Number (DIN) ranges from 1 to 10 and is calculated from the size distribution of genomic DNA. High-molecular-weight DNA (greater than 20 kb) yields a DIN of 9–10. For whole-genome sequencing, a DIN of 7 or higher is recommended, as heavily degraded DNA produces short reads that map poorly to repetitive regions and increase coverage non-uniformity. For amplicon-based approaches, lower integrity may be acceptable because the target regions are short.

Library Preparation Fundamentals

Library preparation is the conversion of purified nucleic acids into a sequencer-ready format. The core steps are fragmentation, end repair, A-tailing, adapter ligation, and PCR amplification. Each step is mechanistically distinct and has specific failure modes.

Fragmentation Methods

Fragmentation is required because sequencing platforms produce reads of limited length (typically 150–300 bp for Illumina, up to 4 kb for Oxford Nanopore). The input DNA must be fragmented to a size distribution that matches the platform's optimal read length, with a small overhang to account for adapter sequences.

Enzymatic fragmentation uses endonucleases such as Fragmentase (a mixture of a nicking enzyme and a dsDNA-specific endonuclease) or NEBNext dsDNA Fragmentase. These enzymes introduce double-strand breaks at random positions, producing fragments with 5′ phosphate and 3′ hydroxyl ends. The reaction is controlled by incubation time and temperature (typically 37°C for 10–30 minutes) and is stopped by adding EDTA, which chelates the magnesium cofactor required by the enzymes. Enzymatic fragmentation is gentle and works well on low-input samples, but can introduce sequence bias if the enzyme has sequence preferences.

Mechanical shearing using focused ultrasonication (Covaris) or nebulization uses acoustic energy to break DNA. The Covaris system uses focused acoustic waves to create cavitation bubbles that mechanically shear DNA into fragments of a defined size, typically 150–500 bp. The size distribution is controlled by the acoustic duty cycle, peak incident power, and treatment time. Mechanical shearing produces a more uniform size distribution than enzymatic methods and does not introduce sequence bias, but requires specialized equipment and higher input amounts (typically 100 ng–1 µg).

Tagmentation (used in Illumina's Nextera kits) combines fragmentation and adapter ligation in a single step. A hyperactive transposase (Tn5) is loaded with adapter sequences and simultaneously fragments the DNA and inserts the adapters at the cut sites. The reaction is fast (5–10 minutes at 55°C) and requires very low input (1–50 ng), making it ideal for clinical samples. The trade-off is that tagmentation can introduce a slight insertion bias at open chromatin regions and requires careful optimization of the transposase-to-DNA ratio.

Adapter Ligation and Indexing

After fragmentation, the ends of the DNA fragments are not compatible with adapter ligation. Fragments generated by mechanical shearing have 5′ overhangs or blunt ends with 5′ phosphates; fragments generated by enzymatic fragmentation have 3′ overhangs. The end repair step converts all ends to blunt, 5′-phosphorylated ends using a cocktail of enzymes: T4 DNA polymerase (fills in 5′ overhangs and removes 3′ overhangs), T4 polynucleotide kinase (adds a phosphate to 5′ hydroxyls), and Klenow fragment (fills in 5′ overhangs). The reaction is performed in the presence of dNTPs at 20°C for 30 minutes, then the enzymes are removed by bead purification.

A-tailing adds a single adenine to the 3′ end of each blunt fragment using Klenow fragment (3′→5′ exo-minus) or Taq polymerase in the presence of dATP only. This creates a 3′ overhang that is complementary to the thymine overhang on the adapter. The A-tailing reaction is performed at 37°C for 30 minutes, followed by heat inactivation at 65°C for 20 minutes. The purpose of the A/T overhang is to prevent adapter concatemerization: the adapters have a 3′ T overhang, so they can only ligate to A-tailed fragments, not to each other.

Adapter ligation uses T4 DNA ligase to covalently join the A-tailed fragment to the Y-shaped adapter. The Y-shape is critical: the top strand contains the P5 flow cell binding sequence and a sequencing primer binding site; the bottom strand contains the P7 sequence and a different sequencing primer binding site. The Y-shape ensures that the two strands are not complementary to each other in the middle region, which prevents the formation of fold-back structures during cluster amplification. The ligation reaction is performed at 20°C for 15–30 minutes with a high concentration of ligase (200 U/µL) and a 10- to 100-fold molar excess of adapters.

Indexing is achieved by incorporating a unique 6–10 base pair sequence into the adapter. During PCR amplification, the index sequence is copied into the final library molecule. This allows multiple libraries to be pooled and sequenced simultaneously, with the index read identifying which sample each read came from. The number of possible indices scales as 4^n, so a 6-base index yields 4096 combinations, while an 8-base index yields 65,536.

PCR Enrichment and Its Pitfalls

PCR amplification serves two purposes: it adds the full-length adapter sequences (including the P5 and P7 flow cell binding sites) and it amplifies the library to a sufficient concentration for sequencing. The PCR reaction uses a high-fidelity polymerase such as Phusion (a proofreading polymerase with 3′→5′ exonuclease activity) or Q5, with primers that anneal to the adapter sequences.

The number of PCR cycles must be carefully controlled. Each cycle introduces a risk of:

  • Amplification bias: GC-rich regions amplify less efficiently than AT-rich regions, leading to uneven coverage.
  • Duplicate reads: PCR duplicates arise when the same original molecule is amplified multiple times. These are computationally removed during analysis, but they reduce the effective sequencing depth.
  • Mutation accumulation: Even high-fidelity polymerases have error rates of approximately 5 × 10⁻⁷ per base per cycle. After 15 cycles, the error rate is approximately 7.5 × 10⁻⁶, which is acceptable for most applications but problematic for variant detection at low allele frequencies.

The optimal cycle number depends on the input amount: 100 ng of DNA requires 4–6 cycles, 10 ng requires 8–10 cycles, and 1 ng requires 12–15 cycles. The PCR reaction is typically run for 15–30 seconds of denaturation at 98°C, 30 seconds of annealing at 60–65°C, and 30 seconds of extension at 72°C, for a total of 4–15 cycles.

After PCR, the library is purified to remove primers, nucleotides, and polymerase. A double-sided bead purification (e.g., 0.6× then 0.8× AMPure XP) is used to select fragments in the desired size range: the first bead ratio removes large fragments (greater than 600 bp), and the second removes small fragments (less than 200 bp), including adapter dimers.

Target Enrichment and Amplification Strategies

Not all sequencing applications require whole-genome or whole-transcriptome coverage. Targeted approaches reduce cost, increase depth, and simplify analysis by focusing sequencing on specific genomic regions.

Amplicon Sequencing

Amplicon sequencing uses PCR to amplify specific regions of interest before library preparation. The primers are designed to flank the target region, and the PCR product is then processed through the standard library prep workflow. This approach is used for:

  • 16S rRNA sequencing for microbiome analysis, where the V3–V4 hypervariable region is amplified.
  • Targeted mutation panels for cancer diagnostics, where specific exons or hotspots are amplified.
  • HLA typing, where highly polymorphic regions are amplified and sequenced.

The advantages of amplicon sequencing are simplicity, low input requirements (as little as 1 ng), and high depth. The disadvantages are that PCR bias can distort the relative abundance of different amplicons, and the approach is limited to known target regions. For amplicon sequencing, the PCR primers themselves serve as the first step of library preparation, and the amplicons are then processed through end repair, A-tailing, and adapter ligation.

Hybrid Capture Approaches

Hybrid capture (also called solution-based capture or SureSelect) uses biotinylated RNA or DNA probes complementary to the target regions. The library is prepared first (fragmentation, end repair, A-tailing, adapter ligation, and PCR), then denatured and hybridized to the probes in solution. The probe-target hybrids are pulled down with streptavidin-coated magnetic beads, washed to remove non-target sequences, and the enriched library is amplified by PCR.

Hybrid capture is more expensive and time-consuming than amplicon sequencing, but it offers several advantages:

  • Larger target regions: Can capture exomes (approximately 30 Mb) or custom panels of hundreds of genes.
  • Lower bias: Hybridization is less biased than PCR, so coverage is more uniform.
  • Detection of structural variants: Because the entire target region is captured, breakpoints can be detected.

The hybridization reaction is performed at 65°C for 16–24 hours in the presence of a blocking agent (Cot-1 DNA to block repetitive sequences) and a high concentration of formamide to lower the melting temperature. The capture efficiency is typically 60–80%, meaning that 20–40% of reads will be off-target.

Pooling, Normalization, and Final QC

Once libraries are prepared, they must be quantified, normalized, and pooled before loading onto the sequencer. The goal is to load an equimolar mixture of libraries so that each sample receives approximately the same number of reads.

Library Quantification by qPCR

The most accurate method for quantifying libraries is quantitative PCR (qPCR) using primers that anneal to the P5 and P7 adapter sequences. The qPCR reaction measures the number of amplifiable molecules, which is the relevant metric for cluster generation. A standard curve is generated using a library of known concentration, and the unknown libraries are interpolated from the curve. The qPCR reaction is typically run with SYBR Green detection, with an initial denaturation at 95°C for 3 minutes, followed by 35 cycles of 95°C for 15 seconds and 60°C for 45 seconds.

qPCR quantification is essential because fluorometric quantification (Qubit) measures total DNA, including non-amplifiable molecules such as adapter dimers and single-stranded DNA. If the library contains a significant fraction of adapter dimers, the Qubit concentration will overestimate the effective concentration, leading to underloading of the flow cell.

Pooling Strategies

Libraries are pooled based on their qPCR concentrations. The pooling strategy depends on the sequencing platform and the desired read depth per sample:

  • For whole-genome sequencing at 30× coverage, each sample requires approximately 90 Gb of data (for a 3 Gb human genome). With a NovaSeq S4 flow cell producing 300 Gb, up to 3 samples can be pooled.
  • For RNA-seq, each sample typically requires 20–50 million reads, so 10–20 samples can be pooled on a single lane.
  • For targeted panels, each sample requires 500–1000× depth, so 50–100 samples can be pooled.

The pooling ratio is calculated by dividing the desired number of reads per sample by the total reads available, then adjusting for the library concentration. For example, if a flow cell produces 400 million reads and you want 20 million reads per sample, you can pool 20 samples. If each library is at 10 nM, you would pool 2 µL of each library in a total volume of 40 µL.

After pooling, the final library mix is denatured with 0.2 N NaOH for 5 minutes at room temperature, then neutralized with 200 mM Tris-HCl, pH 7.0, and diluted to the loading concentration (typically 100–300 pM for Illumina platforms). The denaturation step is critical: only single-stranded molecules can hybridize to the flow cell surface.

Common Pitfalls and Troubleshooting

Even experienced researchers encounter failures in sample preparation. The most common issues are contamination, degradation, adapter dimers, and index hopping.

Contamination and Degradation

Contamination can arise from the extraction reagents, the laboratory environment, or cross-contamination between samples. The most insidious form is microbial contamination, which introduces foreign DNA that can be mistaken for the sample. This is particularly problematic for low-biomass samples such as biopsies or ancient DNA. Prevention strategies include:

  • Using dedicated pipettes and filter tips for pre-PCR and post-PCR work.
  • Performing extractions in a laminar flow hood with UV sterilization.
  • Including negative controls (extraction blanks) in every batch.
  • Using the Sanger Sequencing Protocol to verify the identity of a single amplicon before proceeding to NGS.

Degradation is the loss of high-molecular-weight nucleic acids due to nuclease activity, repeated freeze-thaw cycles, or prolonged storage. DNA is relatively stable, but RNA is highly labile. Prevention strategies include:

  • Storing nucleic acids at −80°C in small aliquots to avoid repeated freeze-thaw.
  • Adding RNase inhibitors (e.g., RNasin, SUPERase•In) to RNA samples.
  • Processing samples as quickly as possible after collection.
  • Using RNA-stabilizing reagents (e.g., RNAlater) for tissue samples.

Adapter Dimers and Index Hopping

Adapter dimers are the most common library preparation artifact. They form when adapters ligate to each other instead of to the insert DNA. Adapter dimers are typically 120–130 bp in length and contain no insert sequence. They sequence efficiently, wasting reads and reducing the effective depth for the sample. The primary cause is an excess of adapters relative to insert DNA, which can occur when the input DNA is over-quantified or when the A-tailing reaction is incomplete. Prevention strategies include:

  • Using the correct adapter-to-insert ratio (typically 10:1 to 100:1 molar excess).
  • Performing a double-sided bead purification after ligation to remove small fragments.
  • Running a Bioanalyzer or TapeStation trace to check for a peak at 120–130 bp.

Index hopping (also called index switching) occurs when free index primers in the pooled library anneal to the adapter sequences of other libraries during cluster amplification, causing a read to be assigned to the wrong sample. This is a particular problem on patterned flow cells (NovaSeq, NextSeq) where the exclusion amplification chemistry allows free primers to diffuse between clusters. Prevention strategies include:

  • Using unique dual indices (UDIs), where both the i5 and i7 indices are unique to each sample.
  • Reducing the amount of free index primers by performing an extra bead purification after PCR.
  • Including a PhiX control spike-in to monitor index hopping rates.

Summary and Best Practices

Sample preparation is the most critical determinant of sequencing success. The following checklist summarizes the key steps and best practices.

Checklist for Sample Prep

  1. Extract nucleic acids using a method appropriate for the sample type, with negative controls included.
  2. Quantify by fluorometry (Qubit) and assess purity by spectrophotometry (NanoDrop).
  3. Assess integrity by Bioanalyzer or TapeStation (RIN for RNA, DIN for DNA).
  4. Fragment DNA to the appropriate size distribution (200–500 bp for Illumina).
  5. End repair and A-tail using the manufacturer's recommended enzyme concentrations and incubation times.
  6. Ligate adapters at the correct molar ratio, with a 10- to 100-fold excess of adapters.
  7. Purify with double-sided bead selection to remove adapter dimers and large fragments.
  8. PCR amplify for the minimum number of cycles required to reach the desired yield.
  9. Quantify the final library by qPCR, not just fluorometry.
  10. Pool libraries in equimolar ratios, and verify the pool by a final Bioanalyzer trace.
  11. Denature and dilute the pool to the loading concentration immediately before sequencing.

Importance of Controls

Every batch of sample preparations should include:

  • Extraction blank: A tube containing all reagents but no sample, processed through the entire workflow. This detects reagent contamination.
  • Positive control: A well-characterized sample (e.g., a reference DNA such as NA12878) processed in parallel. This detects systematic failures.
  • Library prep negative: A tube containing water instead of DNA, processed through library preparation. This detects adapter contamination.

Documentation is essential. Record the sample ID, extraction date, extraction method, quantification values, integrity values, library prep kit lot number, PCR cycle number, and final library concentration for every sample. This information is invaluable for troubleshooting when a sequencing run fails.

Frequently Asked Questions

How do I prepare a DNA sample for sequencing?

To prepare a DNA sample for sequencing, first extract and purify the DNA from your sample type using an appropriate method (column-based kits for most applications, organic extraction for difficult samples). Quantify the DNA by fluorometry and assess integrity by gel electrophoresis or Bioanalyzer. Then construct a sequencing library by fragmenting the DNA (enzymatically or mechanically), repairing the ends, adding A-tails, ligating platform-specific adapters, and amplifying by PCR. Finally, quantify the library by qPCR, pool with other libraries if multiplexing, and load onto the sequencer.

What is the first step in preparing a sample for sequencing?

The first step is nucleic acid extraction and purification. This involves lysing the cells to release DNA or RNA, removing proteins and other contaminants, and concentrating the nucleic acids into a clean, nuclease-free buffer. The quality of this step determines the success of all downstream steps.

How much DNA is needed for whole genome sequencing?

For whole-genome sequencing on Illumina platforms, the recommended input is 100 ng to 1 µg of high-molecular-weight DNA. Low-input kits (e.g., Nextera DNA Flex) can work with as little as 1 ng, but the lower the input, the higher the risk of amplification bias and duplicate reads. For Oxford Nanopore sequencing, 1–5 µg of high-molecular-weight DNA is recommended for optimal read lengths.

How do I prepare RNA for sequencing?

RNA sequencing requires converting RNA to cDNA before library preparation. The workflow is: extract total RNA, deplete ribosomal RNA or select for poly(A)-tailed mRNA, fragment the RNA (typically 200–300 bp), reverse transcribe to cDNA using random hexamers or oligo-dT primers, synthesize the second strand, then proceed with the standard library preparation steps of end repair, A-tailing, adapter ligation, and PCR amplification. The RNA integrity number (RIN) should be assessed before starting.

What is library preparation in sequencing?

Library preparation is the process of converting purified nucleic acids into a format compatible with the sequencing platform. It involves fragmenting the DNA or cDNA to the appropriate size, adding platform-specific adapter sequences that enable surface attachment and sequencing priming, and amplifying the library to sufficient concentration. The library is the actual material that is loaded onto the sequencer.

How do I check the quality of my DNA before sequencing?

Check DNA quality by three complementary methods: spectrophotometry (NanoDrop) for purity ratios (A260/A280 and A260/A230), fluorometry (Qubit) for accurate double-stranded DNA concentration, and microfluidic analysis (Bioanalyzer or TapeStation) for integrity and size distribution. The DNA integrity number (DIN) should be above 7 for whole-genome sequencing.

What are common mistakes in sample preparation for sequencing?

Common mistakes include: using too much or too little input DNA, incomplete removal of extraction reagents (which inhibit downstream enzymes), over-amplification during PCR (causing bias and duplicates), adapter dimers from excessive adapter concentration, index hopping from free index primers, and contamination from the laboratory environment. Each of these can be prevented by following the manufacturer's protocols precisely and including appropriate controls.

Key Takeaways

  • Sample preparation is the most critical determinant of sequencing data quality; errors at this stage cannot be corrected by deeper sequencing.
  • Nucleic acid extraction must yield pure, intact DNA or RNA; assess purity by spectrophotometry and integrity by microfluidic analysis.
  • Library preparation involves fragmentation, end repair, A-tailing, adapter ligation, and PCR amplification, each with specific mechanistic requirements and failure modes.
  • Accurate quantification by fluorometry (for input) and qPCR (for final libraries) is essential for successful pooling and loading.
  • Adapter dimers and index hopping are the most common library preparation artifacts; they can be minimized by correct adapter ratios and unique dual indices.
  • Always include negative controls and positive controls in every batch to distinguish systematic failures from sample-specific issues.
  • Document every step, including lot numbers, cycle numbers, and quantification values, to enable effective troubleshooting.

Further Reading

  • Teufel M, Sobetzko P. Reducing costs for DNA and RNA sequencing by sample pooling using a metagenomic approach. BMC genomics. 2022. PubMed 35999507
  • Gonye ALK et al. Protocol for bulk RNA sequencing of enriched human neutrophils from whole blood and estimation of sample purity. STAR protocols. 2023. PubMed 36853705
  • Paegel BM, Blazej RG, Mathies RA. Microfluidic devices for DNA sequencing: sample preparation and electrophoretic analysis. Current opinion in biotechnology. 2003. PubMed 1256600100004-6)
  • Valle-Silva GD et al. Applicability of the SNPforID 52-plex panel for human identification and ancestry evaluation in a Brazilian population sample by next-generation sequencing. Forensic science international. Genetics. 2019. PubMed 30889526

Related Clinical & Scientific Guides