DNA Profiling: Steps, Techniques, and Applications Explained
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to DNA Profiling
What is DNA Profiling?
DNA profiling, also known as DNA fingerprinting or DNA typing, is a forensic and diagnostic technique that identifies individuals by analyzing unique patterns in their DNA. The method exploits the fact that, with the exception of identical twins, no two humans share the same genomic sequence. By examining specific polymorphic regions of the genome, DNA profiling generates a genetic barcode that can distinguish one individual from another with near-certainty.
The core purpose of DNA profiling is comparative: a biological sample of unknown origin is analyzed and its profile is compared against a reference profile from a known individual or a database. This comparison answers questions of identity—whether a suspect's blood was found at a crime scene, whether a man is the biological father of a child, or whether human remains belong to a missing person. The technique's power lies in its combination of high discrimination power (the ability to distinguish unrelated individuals) and high sensitivity (the ability to work with minute quantities of biological material).
DNA profiling should not be confused with DNA sequencing. Profiling examines a limited set of highly variable genetic markers rather than reading the entire genome. This targeted approach makes the technique rapid, cost-effective, and amenable to automation, which is why it remains the standard in forensic laboratories worldwide.
History and Development
The foundations of DNA profiling were laid in 1985 by Sir Alec Jeffreys at the University of Leicester. Jeffreys discovered that certain regions of human DNA contain repetitive sequences that vary dramatically between individuals. His original method, called restriction fragment length polymorphism (RFLP) analysis, used restriction enzymes to cut DNA at specific recognition sites, followed by Southern blotting to visualize the resulting fragment patterns. The first practical application came in 1985 when Jeffreys used the technique to resolve an immigration dispute, proving that a British boy was the biological son of Ghanaian parents.
RFLP analysis had significant limitations: it required relatively large amounts of high-quality DNA (microgram quantities), took weeks to complete, and was poorly suited to degraded samples. The advent of the polymerase chain reaction (PCR) in the late 1980s revolutionized the field. PCR allowed forensic scientists to amplify specific DNA regions from nanogram quantities of template, making it possible to profile samples that were previously too small or too degraded.
In the 1990s, the forensic community standardized on short tandem repeat (STR) analysis as the method of choice. STRs are short (2–6 base pair) repeated sequences that are highly polymorphic. The FBI launched the Combined DNA Index System (CODIS) in 1998, establishing a set of 13 core STR loci (later expanded to 20) that all US forensic laboratories use, enabling nationwide database searching. Today, commercial multiplex kits allow simultaneous amplification of 15–24 STR loci in a single reaction, generating profiles with discrimination powers exceeding one in a quadrillion.
Biological Basis of DNA Profiling
Short Tandem Repeats (STRs)
Short tandem repeats, also called microsatellites, are regions of the genome where a short nucleotide sequence (typically 2–6 base pairs) is repeated in tandem. For example, the locus D3S1358 contains the repeat unit AGAT, and individuals differ in how many times this unit is repeated—some may have 14 copies on one chromosome and 16 on the other, while others have 15 and 17.
STRs are ideal markers for DNA profiling for several reasons. First, they are abundant: the human genome contains hundreds of thousands of STR loci distributed across all chromosomes. Second, they are highly polymorphic: the number of repeats at a given locus varies widely across the population, with heterozygosity values (the probability that an individual carries two different alleles at a locus) typically exceeding 70%. Third, STR alleles are small—usually 100–400 base pairs in length—which means they can be amplified by PCR even from degraded DNA samples where larger fragments would fail.
The mutation rate of STRs is relatively high, approximately 10⁻³ to 10⁻⁴ per locus per generation, due to replication slippage. This property is useful for tracing recent ancestry and for paternity testing, but it also means that occasional mutations must be accounted for when interpreting kinship results.
Single Nucleotide Polymorphisms (SNPs)
Single nucleotide polymorphisms are positions in the genome where a single base pair differs between individuals. With over 10 million SNPs in the human genome, they represent the most abundant source of genetic variation. However, most SNPs are biallelic—they have only two possible variants (e.g., A or G)—which limits their discrimination power individually. A single SNP can only distinguish between two alleles, whereas a single STR with 15 alleles can distinguish between 15 variants.
SNPs are nevertheless valuable in specific applications. They are more stable than STRs (lower mutation rates, approximately 10⁻⁸ per generation), making them useful for analyzing very old or highly degraded DNA. SNPs also provide information about ancestry and phenotypic traits (such as eye color) that STRs cannot. In practice, forensic laboratories use SNP panels of 50–100 markers to achieve discrimination power comparable to a standard STR panel, but SNP analysis is generally reserved for challenging samples where STR amplification fails.
Sample Collection and Preparation
Sources of DNA
DNA for profiling can be obtained from virtually any biological material containing nucleated cells. Common sources in forensic casework include:
- Blood and bloodstains: A rich source of DNA from white blood cells; dried stains on fabric or surfaces are stable for years.
- Saliva and buccal swabs: Epithelial cells from the inner cheek; the standard reference sample for suspects and victims.
- Semen: Sperm cells are the primary source in sexual assault cases; differential extraction separates sperm DNA from epithelial cell DNA.
- Hair: Only hair roots (with follicular tissue) contain nuclear DNA; hair shafts contain mitochondrial DNA but no nuclear DNA.
- Tissue and bone: Used for postmortem identification; bone and teeth are often the only sources in decomposed remains.
- Urine and feces: Contain shed epithelial cells, though DNA yields are low and contamination risk is high.
The choice of sample source depends on the nature of the case and the condition of the evidence. Collection protocols emphasize preventing contamination: gloves must be worn, sterile swabs used, and samples dried thoroughly before packaging in breathable paper bags (plastic bags trap moisture and promote microbial growth that degrades DNA).
DNA Extraction Methods
DNA extraction separates genomic DNA from cellular proteins, lipids, and other debris while preserving DNA integrity. Three main approaches are used:
Organic extraction (phenol-chloroform) : The classic method. Cells are lysed with a buffer containing SDS (sodium dodecyl sulfate, a detergent that disrupts cell membranes) and proteinase K (a serine protease that digests proteins). The lysate is mixed with phenol:chloroform:isoamyl alcohol (25:24:1), which causes proteins to partition into the organic phase while DNA remains in the aqueous phase. After centrifugation, the aqueous layer is collected and DNA is precipitated with cold ethanol and salt. This method yields high-quality DNA but uses hazardous chemicals and is time-consuming.
Solid-phase extraction (silica-based): The most common method in modern forensic laboratories. DNA binds to silica membranes or magnetic beads in the presence of high concentrations of chaotropic salts (e.g., guanidinium thiocyanate, typically 4–6 M). Contaminants are washed away with ethanol-based buffers, and pure DNA is eluted in a low-salt buffer or water. Commercial kits (e.g., QIAGEN QIAamp, Promega Maxwell) automate this process, processing 96 samples simultaneously in under an hour.
Chelex extraction: A simple method where samples are boiled in a 5% Chelex resin slurry. Chelex is a chelating ion-exchange resin that binds metal ions (which can inhibit PCR) and, combined with boiling, causes cell lysis and protein denaturation. The supernatant contains single-stranded DNA suitable for PCR. This method is rapid and inexpensive but yields lower-quality DNA than organic or silica-based extraction.
Following extraction, DNA quantity is measured using real-time PCR or fluorometric methods (e.g., PicoGreen assay), and quality is assessed by agarose gel electrophoresis to check for degradation.
PCR Amplification in DNA Profiling
Primer Design
Polymerase chain reaction (PCR) amplifies specific DNA regions exponentially, generating millions of copies from a single template molecule. The reaction requires two oligonucleotide primers—short single-stranded DNA sequences (18–25 nucleotides) that flank the target region—along with a thermostable DNA polymerase (typically Taq polymerase from Thermus aquaticus), deoxynucleotide triphosphates (dNTPs), and a buffer containing magnesium chloride (typically 1.5–2.5 mM MgCl₂).
Primer design is critical for successful amplification. Primers must be:
- Specific: They should anneal only to the intended target sequence, not to other genomic regions. This is ensured by selecting unique flanking sequences and checking them against genome databases.
- Balanced in melting temperature (Tm): Forward and reverse primers should have similar Tms (within 2–5°C) to ensure they anneal to their targets at the same temperature. Typical annealing temperatures are 55–65°C.
- Free of secondary structure: Primers should not form hairpins or primer-dimers (primer pairs annealing to each other), which would compete with the target amplification.
In STR analysis, primers are designed to anneal to conserved flanking regions on either side of the repeat region. Because the repeat length varies between individuals, the PCR product size varies correspondingly, and this size difference is what the profiling system detects.
Multiplex PCR
Multiplex PCR amplifies multiple loci simultaneously in a single reaction tube. This is essential for DNA profiling because it conserves limited sample material, reduces cost, and increases throughput. A typical forensic multiplex amplifies 15–24 STR loci plus a sex-determining marker (amelogenin) in one reaction.
Achieving successful multiplex PCR requires careful optimization. All primer pairs must work under the same buffer conditions and thermal cycling parameters. Primers are designed to generate non-overlapping product size ranges so that alleles from different loci can be distinguished by size. Additionally, primers for different loci are labeled with different fluorescent dyes (e.g., 6-FAM, VIC, NED, PET), allowing loci with overlapping size ranges to be distinguished by color.
A typical thermal cycling protocol for STR amplification involves:
- Initial denaturation: 95°C for 11 minutes (to activate the hot-start Taq polymerase)
- 28–30 cycles of:
- Denaturation: 94°C for 30 seconds
- Annealing: 59°C for 90 seconds
- Extension: 72°C for 60 seconds
- Final extension: 60°C for 60 minutes (to promote complete adenylation of products)
The final extension step is crucial: Taq polymerase adds a single adenine to the 3' end of amplified products, and incomplete adenylation causes "split peaks" that complicate analysis. The extended final incubation ensures all products are fully adenylated.
Capillary Electrophoresis and Fragment Analysis
Electrophoresis Basics
Capillary electrophoresis (CE) separates DNA fragments by size as they migrate through a polymer-filled capillary under an electric field. DNA is negatively charged due to its phosphate backbone, so it migrates toward the anode (positive electrode). The polymer matrix (typically a linear polyacrylamide solution) acts as a molecular sieve: smaller fragments migrate faster through the matrix than larger fragments.
In forensic CE, the amplified PCR products are mixed with:
- Formamide: A denaturing agent that keeps DNA single-stranded during electrophoresis.
- Internal size standard: A mixture of fluorescently labeled DNA fragments of known sizes (e.g., 50–500 base pairs) that is run alongside the sample in every capillary. This standard allows precise sizing of sample fragments by comparing their migration times to the standard's.
The sample is electrokinetically injected into the capillary by applying a voltage (typically 1–3 kV for 5–30 seconds). Electrophoresis then proceeds at high voltage (typically 15 kV), with fragments separated over 20–40 minutes. As fragments pass a laser detection window, the fluorescent labels are excited and emitted light is recorded by a charge-coupled device (CCD) camera. The instrument records both the color (identifying which dye/locus) and the migration time (converted to fragment size).
Interpreting Electropherograms
The raw data from CE is displayed as an electropherogram: a plot of fluorescence intensity (y-axis) versus fragment size in base pairs (x-axis), with different colors representing different dye labels. Each peak in the electropherogram represents a PCR product from a specific allele.
Interpretation involves several steps:
- Peak detection: Software identifies peaks above a fluorescence threshold (typically 50–150 relative fluorescence units, RFU).
- Size calling: The internal size standard is used to convert migration time to fragment size in base pairs.
- Allele calling: Fragment sizes are compared to an allelic ladder—a mixture of all known alleles for each locus—run in a separate capillary. Each sample allele is assigned a number corresponding to its repeat count (e.g., allele 14 at D3S1358 means 14 AGAT repeats).
- Quality assessment: Peaks are evaluated for height, shape, and the presence of artifacts (see Common Pitfalls).
A heterozygous individual shows two peaks at a locus (one from each chromosome), while a homozygous individual shows a single peak (two copies of the same allele). The relative peak heights should be approximately equal for heterozygotes (within 70% of each other), and the peak height for homozygotes should be roughly double that of a single heterozygote allele.
Short Tandem Repeat (STR) Analysis
CODIS Loci
The Combined DNA Index System (CODIS), maintained by the FBI, defines the standard set of STR loci used in the United States for forensic DNA profiling. The original 13 core loci were expanded to 20 in 2017 to increase discrimination power and improve compatibility with international databases. The current CODIS 20 loci are:
| Locus | Chromosome | Repeat Motif | Number of Alleles |
|---|---|---|---|
| CSF1PO | 5 | TAGA | 10–15 |
| D3S1358 | 3 | AGAT | 8–21 |
| D5S818 | 5 | AGAT | 7–15 |
| D7S820 | 7 | GATA | 6–15 |
| D8S1179 | 8 | TCTA | 8–19 |
| D13S317 | 13 | TATC | 5–16 |
| D16S539 | 16 | GATA | 5–15 |
| D18S51 | 18 | AGAA | 7–27 |
| D21S11 | 21 | TCTA/TCTG | 24–38 |
| D1S1656 | 1 | TAGA | 11–21 |
| D2S441 | 2 | TCTA | 8–17 |
| D2S1338 | 2 | TGCC/TTCC | 15–28 |
| D10S1248 | 10 | GGAA | 8–19 |
| D12S391 | 12 | AGAT | 15–27 |
| D19S433 | 19 | AAGG | 9–17 |
| D22S1045 | 22 | ATT | 8–19 |
| FGA | 4 | CTTT/TTCC | 17–51 |
| TH01 | 11 | TCAT | 3–14 |
| TPOX | 2 | AATG | 4–16 |
| vWA | 12 | TCTA/TCTG | 10–25 |
The combined power of discrimination across all 20 loci exceeds one in a quintillion for unrelated individuals. This means that the probability of two unrelated individuals sharing the same profile at all 20 loci is astronomically small—far less than the number of humans who have ever lived.
Commercial multiplex kits (e.g., Applied Biosystems GlobalFiler, Promega PowerPlex Fusion) amplify all 20 CODIS loci plus amelogenin (the sex marker) in a single reaction. These kits are rigorously validated and quality-controlled, ensuring consistent performance across laboratories.
Allele Calling and Genotyping
Allele calling is the process of assigning a numeric designation to each detected peak. This is accomplished by comparing the sample's fragment sizes to an allelic ladder—a mixture of all known alleles for each locus, generated by the kit manufacturer. The ladder is run in a separate capillary on the same instrument, under identical conditions, and the software calibrates the sample's fragment sizes against the ladder.
For example, if a sample produces a peak at 174 base pairs at the D3S1358 locus, and the allelic ladder shows that allele 16 runs at 174 base pairs, the sample is assigned allele 16. If the sample shows two peaks at this locus (e.g., alleles 14 and 16), the individual is heterozygous. If only one peak appears, the individual is homozygous for that allele.
The resulting genotype—the complete set of alleles across all loci—constitutes the DNA profile. This profile is typically reported as a string of numbers, such as:
D3S1358: 14, 16 | vWA: 17, 18 | FGA: 22, 24 | ... and so on for all loci.
This numeric format allows profiles to be stored in databases (such as CODIS) and compared electronically. A match between two profiles requires identical alleles at every locus tested.
Statistical Interpretation of DNA Profiles
Population Genetics
The statistical weight of a DNA match depends on how rare the profile is in the general population. This requires knowledge of allele frequencies at each locus in relevant populations. Forensic laboratories maintain allele frequency databases for major population groups (e.g., African American, Caucasian, Hispanic, Asian), compiled from thousands of reference samples.
Several population genetics principles underpin the calculations:
- Hardy-Weinberg equilibrium: In a large, randomly mating population, genotype frequencies remain constant across generations. This allows calculation of expected genotype frequencies from allele frequencies. For a heterozygous genotype (A, B), the expected frequency is 2pq, where p and q are the frequencies of alleles A and B. For a homozygous genotype (A, A), the expected frequency is p².
- Linkage equilibrium: Alleles at different loci are inherited independently (unless the loci are physically close on the same chromosome). This allows multiplication of genotype frequencies across loci to obtain the combined profile frequency.
- Population substructure: Human populations are not perfectly random-mating; there is genetic differentiation between subgroups. To account for this, forensic calculations incorporate a correction factor (θ, theta) that adjusts allele frequencies upward, making the statistical estimate conservative (favoring the defendant).
Random Match Probability
The random match probability (RMP) is the probability that a randomly selected unrelated individual from the population would have the same DNA profile as the evidence sample. It is calculated by multiplying the genotype frequencies across all loci.
For example, consider a profile with the following genotypes at three loci:
- Locus 1: Heterozygous (alleles 14, 16); allele 14 frequency = 0.10, allele 16 frequency = 0.08
- Genotype frequency = 2 × 0.10 × 0.08 = 0.016
- Locus 2: Homozygous (allele 17); allele 17 frequency = 0.15
- Genotype frequency = 0.15² = 0.0225
- Locus 3: Heterozygous (alleles 22, 24); allele 22 frequency = 0.05, allele 24 frequency = 0.03
- Genotype frequency = 2 × 0.05 × 0.03 = 0.003
The combined RMP = 0.016 × 0.0225 × 0.003 = 1.08 × 10⁻⁶, or approximately 1 in 925,000.
With 20 loci, each typically having genotype frequencies of 0.01–0.05, the combined RMP routinely reaches values of 10⁻¹⁸ to 10⁻²⁰—equivalent to one in a quintillion or more.
In practice, the RMP is often expressed as a likelihood ratio (LR) , which compares two hypotheses: that the evidence DNA came from the suspect (Hp, prosecution hypothesis) versus that it came from an unknown unrelated individual (Hd, defense hypothesis). The LR equals 1/RMP when the suspect's profile matches the evidence. An LR of 10¹⁵ means the evidence is 10¹⁵ times more likely if the suspect is the source than if an unknown individual is the source.
Applications of DNA Profiling
Forensic Science
Forensic DNA profiling serves multiple roles in the criminal justice system:
- Suspect identification: DNA from crime scene evidence (blood, semen, saliva, touch DNA) is compared to suspects' reference samples. A match provides powerful evidence of involvement.
- Exoneration: DNA testing has exonerated hundreds of wrongfully convicted individuals through organizations like the Innocence Project. Post-conviction DNA testing can prove that a convicted person was not the source of biological evidence.
- Cold case resolution: DNA databases allow unsolved cases to be revisited when new technology or database entries produce a match.
- Familial searching: When a crime scene profile has no direct match in the database, partial matches may identify close relatives of the perpetrator, providing investigative leads.
- Rapid DNA analysis: Portable instruments now allow DNA profiling at booking stations, enabling quick comparison of arrestees' profiles against crime scene databases.
The chain of custody is critical in forensic applications: every transfer of evidence must be documented to ensure the sample's integrity and admissibility in court.
Paternity and Kinship Testing
DNA profiling is the gold standard for establishing biological relationships. In paternity testing, the child's profile is compared to the alleged father's. Because a child inherits one allele at each locus from each parent, the child must share one allele with the biological father at every locus. If the alleged father does not share an allele with the child at two or more loci, paternity is excluded. If he shares alleles at all loci, the probability of paternity is calculated using the RMP framework, typically yielding probabilities exceeding 99.99%.
Kinship testing extends this principle to other relationships: siblingship, grandparentage, and identification of missing persons through family reference samples. In mass disaster victim identification, DNA profiles from remains are compared to reference samples from family members or from personal items (e.g., toothbrush, razor) known to contain the missing person's DNA.
Ancestry and Genealogy
Direct-to-consumer ancestry testing uses SNP panels (typically 500,000–1 million markers) to estimate biogeographical ancestry. These tests compare an individual's SNP genotypes to reference populations from around the world, producing percentage estimates of ancestry from major continental groups. While not used for forensic identification, ancestry testing has been applied in forensic contexts to predict a suspect's likely ethnic background, and in investigative genetic genealogy—where crime scene DNA is uploaded to public genealogy databases to identify relatives of the perpetrator.
Quality Control and Common Pitfalls in DNA Profiling
Contamination Prevention
Contamination is the most significant threat to DNA profiling accuracy. Even a single skin cell from an investigator can overwhelm a low-level evidence sample. Prevention strategies include:
- Physical separation: Pre-PCR areas (sample processing, extraction, PCR setup) must be physically separated from post-PCR areas (amplification, electrophoresis). Airflow systems maintain positive pressure in pre-PCR rooms and negative pressure in post-PCR rooms.
- Protective equipment: Analysts wear gloves, lab coats, face masks, and hair covers. Gloves are changed frequently, especially between handling different samples.
- Negative controls: Every batch of extractions includes a reagent blank (all reagents but no sample) to detect contamination in the extraction process. Every PCR run includes a negative amplification control (water instead of DNA) to detect contamination in the PCR setup.
- UV irradiation and cleaning: Work surfaces and equipment are regularly cleaned with bleach (sodium hypochlorite) and UV-irradiated to destroy residual DNA.
- Sample handling: Evidence samples are processed one at a time where possible, and dedicated equipment is used for reference samples versus evidence samples.
Degraded DNA and Low-Template DNA
Degraded DNA presents a major challenge in forensic casework. Environmental factors—heat, humidity, UV light, microbial activity—fragment DNA into smaller pieces over time. Since STR primers require intact flanking regions, degradation preferentially affects larger loci. A sample degraded to fragments of 200 base pairs may still amplify at small loci (e.g., TH01, product size ~180 bp) but fail at large loci (e.g., FGA, product size ~350 bp). This results in partial profiles with missing alleles, reducing discrimination power.
Low-template DNA (also called touch DNA) refers to samples containing fewer than 100 pg of DNA (approximately 15 cells). Such samples are prone to:
- Allele drop-out: Failure to amplify one allele in a heterozygote, making the individual appear homozygous.
- Allele drop-in: Spurious peaks from contamination or polymerase errors that appear as real alleles.
- Stutter artifacts: During PCR, Taq polymerase can slip on repeat sequences, generating products that are one repeat unit shorter (or occasionally longer) than the true allele. Stutter peaks typically appear at 5–10% of the true allele's height and can be mistaken for genuine alleles, especially in mixtures.
Laboratories address these challenges through:
- Enhanced amplification: Increasing PCR cycle number (from 28 to 34 cycles) to amplify low-template samples.
- Replicate analysis: Amplifying the sample multiple times and only reporting alleles that appear consistently across replicates.
- Consensus profiling: Combining results from multiple amplifications to generate a consensus profile, requiring an allele to appear in at least two replicates to be scored.
- Interpretation thresholds: Setting analytical thresholds (minimum peak height) and stochastic thresholds (below which heterozygote balance cannot be assumed) to guide allele calling.
Mixture interpretation—where two or more individuals contribute to a sample—is among the most complex tasks in forensic DNA analysis. Modern probabilistic genotyping software (e.g., STRmix, TrueAllele) uses Bayesian models to evaluate the likelihood of different genotype combinations, providing quantitative support for inclusion or exclusion of contributors.
Frequently Asked Questions
What are the steps of DNA profiling?
The standard workflow involves: (1) sample collection from biological evidence; (2) DNA extraction and purification; (3) quantification of extracted DNA; (4) PCR amplification of STR loci using multiplex kits; (5) capillary electrophoresis to separate amplified fragments by size; (6) allele calling by comparison to allelic ladders; (7) statistical interpretation to calculate the random match probability; and (8) reporting and comparison against reference profiles or databases.
How does DNA profiling work?
DNA profiling exploits naturally occurring variations in the number of short tandem repeats at specific genomic locations. PCR amplifies these regions, and capillary electrophoresis separates the products by size. Because the number of repeats varies between individuals, the sizes of the amplified fragments differ, generating a unique pattern of peaks (the profile) for each person. A match between two profiles indicates that they originated from the same individual (or an identical twin).
What is the difference between DNA profiling and DNA sequencing?
DNA profiling examines a limited set of highly variable markers (typically 20 STR loci) to generate a pattern for identity comparison. It does not read the actual DNA sequence—it only measures fragment sizes. DNA sequencing determines the exact order of nucleotides in a DNA molecule. Sequencing provides far more information (the entire genome or targeted regions) but is more expensive and time-consuming. Profiling is optimized for speed, cost, and statistical power in identity testing; sequencing is used when actual sequence information is needed, such as in medical genetics or ancestry analysis.
Why are STRs used in DNA profiling?
STRs are used because they combine several desirable properties: they are highly polymorphic (many alleles exist in the population), their small size (100–400 bp) allows amplification from degraded DNA, they are widely distributed across the genome, and they can be amplified in multiplex reactions. Their high mutation rate (10⁻³–10⁻⁴ per generation) is manageable for identity testing but useful for kinship analysis. No other marker type offers this combination of discrimination power, robustness, and technical simplicity.
Can DNA profiling be wrong?
DNA profiling is highly reliable when performed correctly, but errors can occur. Sources of error include: sample contamination (introducing extraneous DNA), sample mix-up (mislabeling), interpretation errors (especially with mixtures or low-template samples), and laboratory errors. The probability of a false match due to coincidence is astronomically small with full 20-locus profiles (less than one in a quintillion). However, partial profiles from degraded samples have lower discrimination power, and the risk of error increases with sample complexity. Rigorous quality control, accreditation standards, and independent review mitigate these risks.
How is DNA profiling used in paternity testing?
In paternity testing, the child's DNA profile is compared to the alleged father's at 15–20 STR loci. A child inherits one allele at each locus from the biological father. If the alleged father does not share an allele with the child at two or more loci, paternity is excluded. If he shares alleles at all tested loci, the probability of paternity is calculated using population allele frequencies. This calculation incorporates the prior probability of paternity (often assumed to be 0.5) and the likelihood ratio from the DNA data, typically yielding probabilities of 99.99% or higher when paternity is confirmed.
Key Takeaways
- DNA profiling identifies individuals by analyzing short tandem repeat (STR) polymorphisms at 20 standardized CODIS loci, achieving discrimination powers exceeding one in a quintillion.
- The workflow comprises sample collection, DNA extraction, PCR amplification, capillary electrophoresis, allele calling, and statistical interpretation.
- PCR multiplexing allows simultaneous amplification of all 20 loci plus the sex marker from nanogram quantities of DNA, even from degraded samples.
- Random match probability and likelihood ratios quantify the strength of a DNA match, incorporating population genetics principles and conservative corrections for population substructure.
- DNA profiling is applied in forensic casework, paternity and kinship testing, missing person identification, and ancestry analysis.
- Contamination, degradation, low-template artifacts (stutter, drop-out, drop-in), and mixture interpretation are the principal challenges, addressed through strict quality controls, replicate analysis, and probabilistic genotyping software.
- DNA profiling differs fundamentally from DNA sequencing: it measures fragment sizes at polymorphic markers rather than reading nucleotide sequences, optimizing for speed, cost, and statistical power in identity testing.
Further Reading
- Behrouzi R et al. Cell-free and extrachromosomal DNA profiling of small cell lung cancer. Trends in molecular medicine. 2025. PubMed 39232927
- Monckton DG, Jeffreys AJ. DNA profiling. Current opinion in biotechnology. 1993. PubMed 776533390046-y)
- Chakravarty D, Solit DB. Clinical cancer genomic profiling. Nature reviews. Genetics. 2021. PubMed 33762738
- Stevens J. DNA profiling. Medicine, science, and the law. 1991. PubMed 2062203
- DNA profiling in India. Nature methods. 2015. PubMed 26824100
- Goodwin W. DNA profiling: The first 30years. Science & justice : journal of the Forensic Science Society. 2015. PubMed 26654069