Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

Dna Fingerprinting

DNA fingerprinting is a laboratory technique that identifies individuals by comparing patterns in their DNA sequences. It relies on analyzing specific regions of the genome that vary greatly between people, known as polymorphic markers. This guide is for forensic scientists, molecular biologists, genetic counselors, law enforcement professionals, and students who need a practical, source bounded understanding of how DNA fingerprinting works, when to use it, and how to interpret results correctly. The foundational principles are explained in authoritative references such as the NCBI Bookshelf which offers free biomedical texts covering molecular genetics.

Modern DNA fingerprinting uses Short Tandem Repeats (STRs) or Single Nucleotide Polymorphisms (SNPs) to generate a unique profile. The method has expanded from crime scene analysis to paternity testing, wildlife trafficking investigations, and even quality control in cell therapy manufacturing. For instance, the EMBL EBI Training resources outline the bioinformatics approaches used to analyze such genetic data. This guide will walk you through core concepts, decision points, a practical workflow, common pitfalls, and the inherent limitations of DNA fingerprinting.

At a Glance

Aspect Key Details
What it is A method to identify individuals based on unique DNA patterns at selected loci.
Primary markers Short Tandem Repeats (STRs), Single Nucleotide Polymorphisms (SNPs), mitochondrial DNA, Y chromosome markers.
Core applications Forensic identification, paternity/maternity testing, missing person cases, wildlife conservation, cell line authentication.
Typical workflow DNA extraction, quantification, multiplex PCR amplification, capillary electrophoresis, data analysis, statistical interpretation.
Common platforms Applied Biosystems Genetic Analyzers, Illumina sequencing (for SNPs), Oxford Nanopore (for rapid field use).
Quality checks Positive/negative controls, allele ladders, reproducibility tests, stutter band filters.
Major limitation Cannot distinguish identical twins by standard STR methods, mixed samples require statistical deconvolution.

Core Concepts and Decision Criteria

What Makes a DNA Fingerprint Unique

Every individual (except identical twins) has a unique combination of alleles at the tested loci. The discriminatory power comes from analyzing multiple independent markers. The smaller the combined probability of a random match, the stronger the identification. The Bioconductor project provides software packages for calculating these probabilities and performing quality control in high throughput genotyping.

Selecting the Right Marker System

Your choice of markers depends on the sample type, required resolution, and available resources.

  • STRs: The gold standard for forensic casework. They require small amounts of DNA (around 0.5 1 ng) and are highly discriminatory when 13 or more loci are used. Commercial kits such as GlobalFiler or PowerPlex cover core loci. Use STRs when you need a profile that is admissible in court and compatible with law enforcement databases like CODIS.
  • SNPs: Useful when DNA is degraded (e.g., ancient bones, telogen hairs). SNPs can be typed from shorter amplicons (50 100 bp) and are suitable for massively parallel sequencing. They offer lower discrimination per locus but can be combined in large panels (500+ SNPs) to achieve high power. The Galaxy Training Network offers workflows for SNP calling from raw sequencing data.
  • Mitochondrial DNA (mtDNA): Applied when nuclear DNA is too degraded or scarce, such as from hair shafts or old skeletal remains. mtDNA is maternally inherited and does not recombine, so it cannot distinguish siblings or maternal relatives. It is best used for lineage tracking, not individual identification.
  • Y chromosome markers: For male specific identification, especially in sexual assault cases where a female background is present. Y STRs trace paternal lineage.

Decision Criteria for Method Choice

  • Sample condition: Intact biological stains (blood, saliva) work well with STRs. Degraded samples favor mini STRs or SNPs.
  • Required throughput: Single samples can go through capillary electrophoresis. Large population studies are better served by sequencing on platforms described in the NCBI Sequence Read Archive, which hosts raw data for reproducibility.
  • Legal standards: Many jurisdictions require a minimum number of STR loci and adherence to validated kits. Check local guidelines before adopting an alternative method.
  • Mixed samples: STR analysis of mixtures requires probabilistic genotyping software. For highly complex mixtures (three or more contributors), SNP sequencing may offer better resolution.

When DNA Fingerprinting Is Not Appropriate

Do not use standard DNA fingerprinting for:

  • Determining physical traits or disease risk (use dedicated genotyping arrays).
  • Identifying identical twins (requires whole genome sequencing or epigenetic markers).
  • Rapid field decisions without proper laboratory controls (use preliminary screening tools like bar coded panels, but confirm with validated methods).

Practical Workflow or Implementation Steps

A reliable DNA fingerprinting experiment follows a structured sequence. Each step must be documented and include controls. The following workflow is adapted from guidelines found in the NCBI Bookshelf forensic science chapters.

Step 1: Sample Collection and Preservation

Collect biological material using sterile swabs, punches, or cutting tools. Avoid cross contamination by changing gloves between samples. Store blood stains dry at room temperature. Liquid blood or tissue should be frozen at 20°C or lower. Saliva samples on FTA cards can be kept indefinitely at ambient conditions.

Step 2: DNA Extraction

Choose a method that yields sufficient quantity and purity. Common techniques:

  • Organic extraction (phenol chloroform): High yield, but uses hazardous chemicals. Good for difficult tissues.
  • Silica membrane columns: Quick and safe, ideal for routine samples. Use commercial kits validated for forensic use.
  • Chelex resin: Simple and low cost, suitable for buccal swabs. May leave inhibitors in the extract.

Measure DNA concentration using fluorometry (e.g., Qubit). Use agarose gel electrophoresis or a microfluidic chip to assess degradation.

Step 3: Quantification and Quality Assessment

Quantify human DNA specifically using real time PCR kits that target human specific sequences (e.g., Quantifiler Trio). This step also indicates degradation (by comparing short and long amplicons) and presence of PCR inhibitors. For non human applications (wildlife forensics), use species specific qPCR assays.

Step 4: Multiplex PCR Amplification

Perform PCR using a validated commercial STR kit. Follow the manufacturer’s thermal cycling parameters exactly. Include a positive control (known DNA), a negative control (water instead of DNA), and an extraction blank. The negative control must show no peaks.

A recent application of DNA guided Argonaute enzymes for RNA cleavage, as described in A DNA guided prokaryotic Argonaute enables programmable RNA cleavage for sequencing and quality control of in vitro transcribed RNAs, demonstrates how novel nucleases could be adapted to cut DNA at specific sites. While not yet routine in forensic labs, such technologies may eventually improve the precision of marker selection.

Step 5: Capillary Electrophoresis and Data Collection

Separate the PCR products by size using a capillary electrophoresis instrument. Include an internal size standard and an allelic ladder in each run. Software automatically assigns alleles based on fragment length.

Step 6: Data Analysis and Profile Interpretation

Review electropherograms manually. Check that peaks are above the analytical threshold (typically 50 150 relative fluorescent units). Examine each locus for:

  • Stutter peaks (indicating PCR slippage, should be less than 15% of the main peak height).
  • Allele dropout (especially in degraded or low template samples).
  • Mixture indications (more than two peaks at a locus).

Use statistical software to calculate the random match probability (RMP) or likelihood ratio (LR). Population databases are available through the EMBL EBI Training and the Galaxy Training Network. The Bioconductor package forensic offers functions for LR calculations.

Step 7: Reporting

A complete report includes the laboratory case identifier, sample description, the profile (list of allele calls), the RMP or LR, and a statement of conclusions (e.g., “the suspect cannot be excluded as the source”). Always note any artefacts or sub optimal features.

Common Mistakes and How to Avoid Them

  • Inadequate negative controls: A contaminated negative control invalidates the entire batch. Always run an extraction blank and a PCR negative control. If the negative shows amplification, repeat from extraction.
  • Ignoring stutter peaks: A stutter peak at the parent position can be misinterpreted as a true allele in a mixture. Use software filters and manually verify. The threshold should be set based on validation data.
  • Assuming single source from a clean electropherogram: Forensic samples often contain multiple contributors even if only two peaks appear at each locus. Look for peak height imbalances and tri allele patterns. Use mixture detection software.
  • Using the wrong population database: RMP calculations depend on allele frequencies. Databases should match the ethnic or geographic origin of the sample. The NCBI Sequence Read Archive hosts population level sequencing data that can be used to update local allele frequencies.
  • Overinterpreting partial profiles: A low template or degraded sample may yield only 4 8 loci. The RMP becomes much higher, reducing discrimination. Do not exclude a suspect solely on a partial profile unless the loci present are inconsistent.
  • Failing to validate new kits or instruments: Laboratories must perform internal validation studies (sensitivity, precision, mixture studies) before using any new reagent or instrument. Guidelines are available from the National Institute of Standards and Technology (NIST) and referenced in NCBI Bookshelf.

Limits of Interpretation and Uncertainty

No DNA fingerprinting method provides absolute certainty. Key limitations include:

  • Identical twins: Standard STR and SNP methods cannot distinguish monozygotic twins. For forensic cases where a twin is a suspect, additional techniques such as whole genome sequencing looking for de novo mutations or epigenetic markers may be attempted, but they are not routine.
  • DNA mixtures: Interpreting mixtures with more than two contributors is challenging. Probabilistic genotyping programs (e.g., STRmix, TrueAllele) produce likelihood ratios, but the results are only as good as the model and the quality of the data. Mixed samples can yield false inclusions if not carefully deconvoluted.
  • Low template DNA: Samples with less than 100 pg of DNA are prone to allele dropout, drop in (artifact alleles), and increased stochastic effects. Some courts restrict the use of low template profiles as sole evidence. The Galaxy Training Network has tutorials on setting filtering thresholds for low coverage data.
  • Population substructure: Even with large databases, allele frequencies may not perfectly represent the population of a given individual. Bayesian methods that incorporate theta (substructure correction) are recommended.
  • Contamination: Trace amounts of DNA from the collector, the environment, or previous samples can result in a mixed profile that confounds interpretation. Strict adherence to sterile technique and negative controls is imperative.
  • Novel markers and platforms: The use of new markers (e.g., microhaplotypes, insertion/deletion polymorphisms) or sequencing based typing is promising but may lack validation for courtroom presentation. For example, the FINDER system described in FINDER converts zero background kinetic fingerprinting into area scalable attomolar biomarker detection is aimed at biomarker detection, not forensic identification, but it illustrates how innovative fingerprinting methods are emerging. Always verify that any new approach meets the Daubert standard or equivalent admissibility criteria.

Furthermore, DNA fingerprinting cannot reveal when or how a sample was deposited. A person’s DNA can persist on an object for years. A match only indicates that the individual is a possible source of the DNA, not that they were present at the crime scene at the relevant time. For animal crime investigations, the same caution applies. The paper The contribution of bloodstain pattern analysis (BPA) in the investigation of crimes against domestic and wild animals emphasizes integrating DNA analysis with other physical evidence to build a solid case.

Frequently Asked Questions

1. How accurate is DNA fingerprinting?
When using a validated 13 20 STR kit and proper controls, the random match probability can be less than 1 in a trillion for a single source profile. Accuracy depends on sample quality, laboratory protocols, and statistical interpretation. The NCBI Bookshelf provides detailed statistical models for calculating error rates.

2. Can DNA fingerprinting be fooled by mixing someone else’s DNA?
Yes, intentional contamination is a known concern. Laboratories should test for evidence of multiple contributors and use probabilistic genotyping to account for mixtures. In casework, chain of custody and independent verification reduce the risk of fabrication.

3. Why are 13 or 20 STR loci used instead of just one?
One locus does not provide enough discrimination. The combined power of multiple independently assorting loci dramatically lowers the chance of a random match. For example, two individuals might share the same allele at one locus by chance, but the probability of sharing alleles at all 13 is minuscule.

4. Do I need a specific software to analyze DNA fingerprints?
Yes. Capillary electrophoresis data require both instrument software (e.g., GeneMapper ID X) and statistical packages (e.g., STRmix, FamLink). For SNP based fingerprinting, you can use open source tools from Bioconductor or workflows from Galaxy Training Network. The NCBI Sequence Read Archive allows you to download reference data for benchmarking.

References and Further Reading

Related Articles