# The Role of Segregation Data in Variant Interpretation: How to Apply PP1 and BS4 in ACMG/AMP Classification


## Key Takeaways

- Family segregation data are crucial for establishing a causal link between a genetic variant and a Mendelian disease, directly informing ACMG/AMP classification through PP1 (co-segregation supporting pathogenicity) and BS4 (lack of segregation supporting benign classification).
- The strength of segregation evidence is quantified using LOD scores, which measure the likelihood of variant-disease co-segregation versus chance, with higher LOD scores (e.g., ≥3.0) indicating stronger evidence for pathogenicity.
- PP1 evidence strength can be upgraded from supporting to moderate or strong based on the number of informative meioses and calculated LOD scores, with approximate thresholds of 1.5-2.5 for moderate and ≥3.0 for strong, requiring careful adjustment for pedigree structure, incomplete penetrance, and disease prevalence.
- BS4 is applied when affected individuals do not carry the variant, arguing against pathogenicity, and its strength can be considered strong with multiple affected non-carriers across different family branches, distinct from BS2 which relies on population observations in unaffected individuals.
- Accurate segregation analysis necessitates meticulous pedigree confirmation, orthogonal genotype verification (e.g., Sanger sequencing), precise definition of the disease model (inheritance, penetrance, phenocopy rate), and robust documentation of all parameters and calculations.
- Incomplete penetrance and phenocopies significantly complicate segregation analysis by reducing the informativeness of unaffected carriers or affected non-carriers, respectively, requiring their explicit incorporation into LOD score calculations and evidence strength assignments.

---

Family segregation data provide direct evidence for determining whether a genetic variant causes a Mendelian disease. When a variant tracks with affected status across multiple meioses in a pedigree, the probability that it is pathogenic increases. When affected individuals in a family do not carry the variant, the evidence argues against pathogenicity. The American College of Medical Genetics and Genomics (ACMG) and the Association for Molecular Pathology (AMP) formalized these observations as two evidence categories: PP1 for co-segregation supporting pathogenicity and BS4 for lack of segregation supporting a benign classification. This article explains how to quantify segregation evidence using LOD scores, how to apply PP1 and BS4 at appropriate strength levels, and how to adjust for pedigree structure, incomplete penetrance, and disease prevalence. The practical outcome for clinical geneticists and laboratory professionals is a defensible, reproducible method for converting family observations into ACMG/AMP evidence assignments.

## The Rationale for Segregation Evidence in Variant Classification

Segregation analysis answers a direct question: does the variant behave as if it causes the disease? In a dominant disorder, every affected individual should carry the variant and every unaffected individual should not, assuming complete penetrance and accurate phenotyping. In a recessive disorder, affected individuals should carry two pathogenic alleles and unaffected carriers should carry one or none. When these patterns hold across multiple family members, the evidence accumulates.

The ACMG/AMP framework assigns PP1 as a supporting level of evidence for pathogenicity when segregation data are present. The 2015 standards state that PP1 applies when the variant co-segregates with disease in multiple affected family members. BS4 applies when the variant fails to segregate with disease, meaning an affected individual does not carry the variant in a situation where the variant would be expected to cause the phenotype. Both criteria require careful pedigree interpretation because phenocopies, reduced penetrance, and de novo variants can create misleading patterns.

The strength of segregation evidence depends on the number of informative meioses. A single affected parent and affected child provide one informative meiosis. Three affected individuals in a row provide two informative meioses. The more meioses observed, the higher the LOD score and the stronger the evidence. The ACMG/AMP framework allows PP1 to be upgraded from supporting to moderate or strong when segregation data are extensive, although the 2015 publication does not specify exact thresholds. Subsequent recommendations from the ClinGen Sequence Variant Interpretation Working Group provide guidance on converting LOD scores to evidence strength levels.

## Calculating LOD Scores from Pedigree Data

The logarithm of odds (LOD) score measures the likelihood that a variant and a disease phenotype co-segregate due to linkage versus the likelihood that they co-segregate by chance. A LOD score of 3 means the odds are 1000 to 1 in favor of linkage. For variant interpretation, the LOD score quantifies how much the pedigree data support the variant causing the disease.

### The Basic Formula and Assumptions

The LOD score is calculated as log10 of the ratio of two likelihoods. The numerator is the probability of observing the pedigree genotypes and phenotypes given that the variant is linked to the disease. The denominator is the probability of observing the same data given that the variant is unlinked and segregates independently. The calculation requires specifying the disease model, including mode of inheritance, allele frequency, penetrance, and phenocopy rate.

For a simple dominant disease with complete penetrance and no phenocopies, each affected individual who carries the variant contributes to the numerator. Each affected individual who does not carry the variant reduces the likelihood. The calculation becomes more complex with reduced penetrance because an unaffected carrier does not necessarily contradict the linkage hypothesis.

### Simple Pedigree LOD Score Calculation

For small pedigrees, LOD scores can be calculated manually using the formula for a dominant trait. Consider a pedigree with an affected parent and two affected children, all carrying the variant. Under complete penetrance and a rare disease allele, the probability of this observation given linkage is 1. The probability given no linkage depends on the chance that each affected child inherited the variant from the affected parent, which is 0.5 for each child. The LOD score is log10(1 / 0.25) = 0.602.

Each additional affected child who carries the variant adds log10(2) = 0.301 to the LOD score. Three affected children give a LOD of 0.903, four give 1.204, and five give 1.505. To reach a LOD of 3 using only affected individuals in a dominant pedigree, approximately 10 informative meioses are needed. This explains why PP1 rarely reaches strong evidence strength in small families.

### Using Software for Complex Pedigrees

Manual calculation becomes impractical for pedigrees with multiple branches, unaffected carriers, consanguinity, or reduced penetrance. Several software packages perform parametric linkage analysis and LOD score calculation. Merlin, LINKAGE, and SUPERLINK are commonly used in research settings. The choice of software depends on pedigree size, marker density, and the need for integration with variant calling pipelines.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials for genomic analysis workflows that can be adapted for family-based variant analysis. These tutorials emphasize reproducible analysis steps and documentation, which are essential for clinical reporting. The [nf-core documentation](https://nf-co.re/docs) describes community standards for pipeline configuration and reproducibility that apply to family-based sequencing projects.

## Applying PP1 at Different Evidence Strengths

The ACMG/AMP framework designates PP1 as supporting evidence, but the ClinGen Sequence Variant Interpretation Working Group has published recommendations for upgrading PP1 based on the number of informative meioses and the calculated LOD score. These recommendations are not part of the original 2015 standards but represent the current consensus approach.

### Supporting Strength

PP1 at supporting strength applies when segregation data exist but are limited. A single affected parent and affected child pair provides one informative meiosis. Two affected siblings and an affected parent provide two informative meioses. The LOD score in these situations is typically below 1.0. Supporting strength is appropriate when the segregation pattern is consistent with pathogenicity but the number of observations is too small to provide strong statistical evidence.

### Moderate Strength

PP1 at moderate strength requires more extensive segregation data. The ClinGen recommendations suggest that a LOD score of approximately 1.5 to 2.0 supports moderate strength. This typically requires three to four informative meioses in a dominant pedigree. The exact threshold depends on the disease model and the pedigree structure.

### Strong Strength

PP1 at strong strength requires a LOD score of approximately 3.0 or higher. This corresponds to roughly 10 informative meioses in a simple dominant pedigree. Such extensive pedigrees are uncommon in clinical practice but occur in research cohorts and in populations with founder effects or large extended families.

### The PP1_Moderate Example from Lynch Syndrome Literature

A 2020 report in Genes described a novel MLH1 in-frame deletion in a Slovenian Lynch syndrome family. The authors applied PP1 at moderate strength as part of the evidence supporting a likely pathogenic classification. The variant, LRG_216t1:c.2236_2247delCTGCCTGATCTA p.(Leu746_Leu749del), was associated with an uncommon isolated loss of PMS2 immunohistochemistry staining in tumor tissue. The segregation analysis contributed to the classification alongside PM1, PM2, and PM4. This example illustrates how PP1 at moderate strength can be combined with other evidence categories to reach a clinically actionable classification. The full report is available at [PubMed](https://pubmed.ncbi.nlm.nih.gov/32197529).

## Applying BS4 for Lack of Segregation

BS4 applies when the variant does not segregate with disease in a way that contradicts pathogenicity. The most straightforward application is when an affected individual in a family does not carry the variant despite the variant being present in other affected relatives. This observation argues that the variant is not the sole cause of the disease in that family.

### Criteria for Applying BS4

BS4 requires that the affected individual who lacks the variant has a phenotype consistent with the disease being investigated. If the affected individual has a different condition, the observation does not count as evidence against segregation. The phenotyping must be reliable, and the possibility of a phenocopy must be considered. In diseases with high phenocopy rates, such as some cancers and neuropsychiatric conditions, an affected non-carrier is less informative.

### Strength Levels for BS4

The ACMG/AMP framework lists BS4 as supporting evidence for a benign classification. Unlike PP1, the original standards do not describe upgrading BS4 to higher strength levels. However, extensive segregation data showing multiple affected non-carriers can provide strong evidence against pathogenicity. Some laboratories apply BS4 at strong strength when multiple affected individuals in different branches of a family lack the variant.

### Distinguishing BS4 from Other Benign Criteria

BS4 is distinct from BS2, which applies when the variant is observed in a healthy adult individual in a recessive disease or in an unaffected adult in a dominant disease. BS2 uses population observations of unaffected individuals, while BS4 uses family segregation data. Both criteria can contribute to a benign classification, but they address different types of evidence.

## Adjusting for Incomplete Penetrance

Incomplete penetrance complicates segregation analysis because an unaffected carrier does not contradict the pathogenicity hypothesis. A variant with 80% penetrance means that 20% of carriers will never develop the disease. In a small pedigree, an unaffected carrier might simply be a non-penetrant individual instead of evidence against pathogenicity.

### Incorporating Penetrance into LOD Score Calculations

Parametric linkage analysis requires specifying the penetrance value in the disease model. A lower penetrance reduces the LOD score for a given pedigree because the observation of an unaffected carrier is more compatible with linkage. Conversely, a higher penetrance increases the LOD score because unaffected carriers are less expected.

For example, consider a pedigree with an affected parent, an affected child, and an unaffected child, all carrying the variant. With complete penetrance, the unaffected child is unexpected and reduces the LOD score. With 70% penetrance, the unaffected child is more compatible with the linkage hypothesis, and the LOD score is higher.

### Practical Guidance for Penetrance Assumptions

The penetrance value used in the analysis should be based on published data for the specific gene and disease. If no published penetrance estimates exist, a conservative approach is to use a range of values and report the LOD score as a range. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) resources provide guidance on selecting appropriate disease models and penetrance values for genetic analysis.

### The MYORG Compound Heterozygote Example

A 2025 report in the Chinese Journal of Medical Genetics described a primary familial brain calcification family with compound heterozygous MYORG variants. The proband carried c.337_348dup (p.Leu113_Arg116dup), a known pathogenic variant, and c.1268T>G (p.Val423Gly). Segregation analysis showed that the father carried the first variant and the mother carried the second. The authors applied PP1 at supporting strength as part of the evidence for classifying the novel variant as likely pathogenic. The classification also used PM2_Supporting, PM3_Supporting, PP3_Moderate, and PP4_Supporting. This example demonstrates how PP1 contributes to recessive disease classification when the segregation pattern confirms compound heterozygosity. The full report is available at [PubMed](https://pubmed.ncbi.nlm.nih.gov/40555662).

## Handling Phenocopies and Reduced Penetrance in Pedigree Interpretation

Phenocopies are individuals who have the disease phenotype but do not carry the genetic variant under investigation. In diseases with a high phenocopy rate, an affected non-carrier in a family does not necessarily contradict the pathogenicity of the variant. The phenocopy rate must be incorporated into the LOD score calculation and into the decision to apply BS4.

### Estimating Phenocopy Rates

Phenocopy rates vary by disease. For highly penetrant monogenic disorders with distinctive phenotypes, the phenocopy rate is low. For common diseases with genetic heterogeneity, the phenocopy rate can be substantial. Published literature and population databases provide estimates for many conditions. When no estimate exists, a sensitivity analysis using different phenocopy rates can show how the LOD score changes.

### The Factor VII Deficiency Example

A 2025 report in the Chinese Journal of Medical Genetics described a family with hereditary factor VII deficiency due to compound heterozygous F7 gene variants. The proband presented with menorrhagia and had a significantly prolonged prothrombin time of 33.1 seconds. The family included 12 members across three generations. Segregation analysis confirmed that the proband inherited one variant from each parent. The authors used the ACMG guidelines to classify the variants. This example illustrates how segregation data in recessive disorders require confirming the carrier status of both parents and the affected status of the proband. The full report is available at [PubMed](https://pubmed.ncbi.nlm.nih.gov/41451501).

## At a Glance: Segregation Evidence Strength and LOD Score Thresholds

The following table summarizes the relationship between informative meioses, LOD scores, and recommended PP1 evidence strength. These values assume a dominant disease with complete penetrance and no phenocopies. Adjustments are needed for recessive diseases, reduced penetrance, and phenocopy rates.

| Evidence Strength | Approximate Informative Meioses | Approximate LOD Score | Typical Pedigree Structure |
|---|---|---|---|
| PP1 Supporting | 1 to 2 | 0.3 to 0.9 | Affected parent and one or two affected children |
| PP1 Moderate | 3 to 5 | 1.5 to 2.5 | Three generations with multiple affected individuals |
| PP1 Strong | 8 to 12 | 3.0 or higher | Extended family with multiple affected branches |
| BS4 Supporting | 1 affected non-carrier | Not applicable | Single affected individual lacking the variant |

The thresholds in this table are approximate and should be adjusted based on the specific disease model. Laboratories should document the assumptions used in their LOD score calculations and apply consistent thresholds across cases.

## Practical Workflow for Incorporating Segregation Data

The following workflow provides a structured approach to applying PP1 and BS4 in variant classification. This workflow assumes that the laboratory has already identified a candidate variant through sequencing and filtering.

### Step 1: Confirm Pedigree Structure and Relationships

Before performing segregation analysis, confirm the pedigree structure and the relationships between family members. Misattributed parentage or adoption can invalidate segregation results. If relationship testing has not been performed, consider whether the pedigree structure is reliable enough for segregation analysis.

### Step 2: Verify Variant Genotypes in All Available Family Members

Confirm the variant genotype in all available family members using an orthogonal method such as Sanger sequencing. Sequencing artifacts can create false genotype calls that distort segregation patterns. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference sequences and variation databases that support variant confirmation.

### Step 3: Define the Disease Model

Specify the mode of inheritance, penetrance, phenocopy rate, and allele frequency for the disease. Use published estimates when available. Document the sources for these parameters. If no published estimates exist, use conservative assumptions and perform sensitivity analysis.

### Step 4: Calculate the LOD Score

Calculate the LOD score using appropriate software or manual methods for small pedigrees. Record the software version, the input parameters, and the output. The [Bioconductor project](https://bioconductor.org/) provides R packages for genetic analysis that can be used for LOD score calculation and pedigree analysis.

### Step 5: Assign PP1 or BS4 Evidence Strength

Based on the LOD score and the segregation pattern, assign PP1 at supporting, moderate, or strong strength, or assign BS4 if the variant fails to segregate. Document the rationale for the assignment, including the number of informative meioses and the LOD score.

### Step 6: Integrate with Other Evidence Categories

Combine the segregation evidence with other ACMG/AMP evidence categories to reach a final classification. Segregation evidence rarely stands alone. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on variant annotation and interpretation that support the integration of multiple evidence types.

### Step 7: Document and Report

Document the segregation analysis in the clinical report, including the pedigree, the genotypes, the LOD score, and the evidence strength assignment. This documentation supports transparency and allows other laboratories to reproduce the analysis.

## Records and Measurements for Segregation Analysis

Accurate records are essential for reproducible segregation analysis. The following data should be recorded for each family included in the analysis.

### Pedigree Records

The pedigree should include all available family members, their affected status, their ages, and their relationship to the proband. The pedigree should distinguish between confirmed affected status, reported affected status, and unknown status. The source of the phenotype information should be documented.

### Genotype Records

The genotype of each family member at the variant locus should be recorded, including the method used for genotyping and the quality metrics. Sanger sequencing traces or equivalent data should be archived. The reference sequence and the variant nomenclature should follow standard conventions.

### LOD Score Calculation Records

The LOD score calculation should record the software used, the version, the disease model parameters, and the input pedigree file. The output should include the LOD score at the variant locus and any sensitivity analyses performed.

### Evidence Assignment Records

The assignment of PP1 or BS4 should record the evidence strength, the number of informative meioses, the LOD score, and the rationale for any adjustments for penetrance or phenocopy rate. This record supports the final variant classification.

## Common Failure Patterns in Segregation Analysis

Several recurring problems can compromise segregation analysis. Recognizing these patterns helps laboratories avoid incorrect evidence assignments.

### Failure to Confirm Relationships

When pedigree relationships are assumed instead of confirmed, misattributed parentage can create false segregation patterns. An affected child who does not carry the variant might be the biological child of a different father. Relationship testing using polymorphic markers can identify these situations.

### Incomplete Phenotyping

When affected status is based on self-report instead of clinical evaluation, misclassification can occur. An individual reported as unaffected might have mild or subclinical disease. Conversely, an individual reported as affected might have a different condition. Clinical evaluation by a qualified professional reduces this risk.

### Ignoring Reduced Penetrance

Applying PP1 without adjusting for reduced penetrance can overstate the evidence. An unaffected carrier in a family with a highly penetrant variant is more informative than an unaffected carrier in a family with a variant that has 50% penetrance. The LOD score calculation must incorporate the penetrance estimate.

### Confusing Segregation with Linkage

Segregation analysis at a variant locus is not the same as linkage analysis at a marker locus. The variant itself must be genotyped in all family members. Using linked markers as a proxy for the variant introduces uncertainty and can lead to incorrect conclusions.

### Applying BS4 Without Considering Phenocopies

An affected non-carrier in a disease with a high phenocopy rate provides weaker evidence against pathogenicity than an affected non-carrier in a disease with a low phenocopy rate. The phenocopy rate must be considered before applying BS4.

### Overinterpreting Small Pedigrees

A single affected parent and affected child pair provides limited evidence. The LOD score from one informative meiosis is 0.301, which is far below the threshold for strong evidence. Laboratories should resist the temptation to assign PP1 at moderate or strong strength based on small pedigrees.

## The Darier Disease Example: PP1 in a Small Pedigree

A 2025 report in the Chinese Journal of Medical Genetics described a Chinese Han pedigree with Darier disease. The proband was a 67-year-old female with keratotic papules in sebaceous areas. Whole exome sequencing identified a missense variant, c.68G>A (p.Gly23Glu), in exon 1 of the ATP2A2 gene. Sanger sequencing confirmed that the proband and her eldest daughter carried the variant, while other family members did not. The authors classified the variant as pathogenic using PS1, PM1, PM2_Supporting, PP1, PP3, and PP4.

This example illustrates how PP1 at supporting strength contributes to a pathogenic classification when combined with other evidence. The segregation data showed co-segregation of the variant with the disease phenotype in the pedigree, but the number of informative meioses was small. The PP1 evidence was supporting instead of moderate or strong. The full report is available at [PubMed](https://pubmed.ncbi.nlm.nih.gov/40350400).

## Limitations of Segregation Evidence

Segregation evidence has inherent limitations that laboratories must acknowledge when applying PP1 and BS4.

### Small Family Size

Most clinical families are too small to provide strong segregation evidence. A typical nuclear family with an affected parent and two affected children provides a LOD score below 1.0. Reaching a LOD of 3.0 requires an extended pedigree that is uncommon in clinical practice.

### Genetic Heterogeneity

In diseases caused by multiple genes, a variant in one gene might segregate with disease in some families but not in others. The segregation pattern in a single family reflects the specific genetic cause in that family, not the general relationship between the gene and the disease.

### Phenotypic Variability

Variable expressivity can complicate segregation analysis. An individual who carries the variant and has mild symptoms might be classified as unaffected if the symptoms are not recognized. This misclassification reduces the apparent segregation and can lead to an incorrect BS4 assignment.

### Population Stratification

In populations with founder effects or high rates of consanguinity, the background allele frequency of the variant affects the interpretation. A variant that is common in the population is more likely to co-segregate with disease by chance. The allele frequency should be considered when interpreting segregation data.

### The VUS Challenge

The [article titled "Not 'just a VUS'"](https://doi.org/10.1016/j.gimo.2026.104401) addresses the challenge of variants of uncertain significance in clinical practice. Segregation data can help reclassify VUSs, but the limitations described above mean that many VUSs remain uncertain even after segregation analysis. The article emphasizes the importance of systematic approaches to VUS resolution.

## Quality Controls for Segregation Analysis

Quality controls ensure that segregation analysis is accurate and reproducible. The following controls should be implemented in laboratories performing variant classification.

### Genotype Quality Control

All variant genotypes used in segregation analysis should be confirmed by an orthogonal method. Sanger sequencing is the standard confirmation method. The quality of the sequencing traces should be reviewed by a qualified professional.

### Pedigree Quality Control

The pedigree structure should be reviewed for consistency with the reported family relationships. Discrepancies between reported and genetic relationships should be investigated before proceeding with segregation analysis.

### Software Validation

The software used for LOD score calculation should be validated against known pedigrees with known LOD scores. The validation results should be documented. Software updates should be tested before use in clinical cases.

### Double Review

Segregation analysis should be reviewed by at least two qualified professionals. The review should confirm the pedigree structure, the genotype calls, the disease model parameters, and the evidence strength assignment.

## Professional Escalation Criteria

Laboratories should have clear criteria for escalating segregation analysis to more experienced professionals or to external experts. The following situations warrant escalation.

### Complex Pedigree Structures

Pedigrees with consanguinity, multiple branches, or uncertain relationships require specialized expertise. The [The Carpentries Lessons](https://carpentries.org/lessons) provide foundational training in data handling and reproducible analysis that supports rigorous pedigree analysis.

### Discrepant Results

When segregation analysis produces results that conflict with other evidence categories, the case should be reviewed by a molecular genetics expert. For example, a variant with strong functional evidence but apparent lack of segregation requires careful evaluation.

### High-Stakes Classifications

Classifications that will be used for predictive testing, prenatal diagnosis, or surgical decision-making warrant additional review. The potential consequences of an incorrect classification justify the additional scrutiny.

### Research-Only Findings

Segregation data generated in a research setting should be reviewed before use in clinical classification. The research protocols, consent processes, and data quality should be evaluated.

## Regulatory and Reporting Context

Variant classification laboratories operate under regulatory frameworks that require documentation and transparency. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide access to reference sequences, variation databases, and clinical resources that support variant interpretation. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) resources provide guidance on data management and analysis best practices.

Clinical reports should include the segregation analysis methods, the LOD score, and the evidence strength assignment. The report should distinguish between confirmed and reported phenotypes and between confirmed and inferred genotypes. The limitations of the segregation analysis should be acknowledged.

## Integrating Segregation Data with Variant Calling Workflows

Segregation analysis depends on accurate variant calling. The variant calling workflow must produce reliable genotypes for all family members before segregation analysis can proceed.

### Germline Variant Calling

Germline variant calling identifies variants present in the constitutional genome. For family-based analysis, the variant caller should be run consistently across all family members. The [nf-core documentation](https://nf-co.re/docs) describes community standards for germline variant calling pipelines that support reproducible family analysis.

### Variant Filtering

Variant filtering removes artifacts and low-quality calls before segregation analysis. The filtering criteria should be applied consistently across all family members. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on variant filtering that support rigorous analysis.

### Somatic Variant Calling

Somatic variant calling is relevant when the disease involves somatic mutations, such as in cancer. Segregation analysis is generally not applicable to somatic variants because they are not inherited. The distinction between germline and somatic variants is critical for segregation analysis.

### Reproducibility

Reproducible variant calling requires version control, containerization, and documentation. The [nf-core documentation](https://nf-co.re/docs) describes best practices for reproducible pipelines. The [Bioconductor project](https://bioconductor.org/) provides R packages that support reproducible genomic analysis.

## The Role of Training and Education

Segregation analysis requires specialized skills in pedigree analysis, statistical genetics, and variant interpretation. Training programs should cover these topics.

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) provides learning pathways for bioinformatics and genetic analysis. The [Galaxy Training Network](https://training.galaxyproject.org/) offers hands-on tutorials for variant analysis workflows. The [The Carpentries Lessons](https://carpentries.org/lessons) provide foundational training in computing and data skills. The [Bioconductor project](https://bioconductor.org/) offers documentation and workflows for statistical analysis in genomics.

Laboratories should ensure that personnel performing segregation analysis have appropriate training and demonstrated competence. Continuing education should cover updates to ACMG/AMP guidelines and new tools for segregation analysis.

## A Decision Framework for Segregation Evidence When Pedigree Data Are Ambiguous

Segregation analysis in clinical practice rarely produces clean patterns that map directly onto PP1 or BS4 assignments. Affected individuals may carry the variant but have atypical presentations, unaffected carriers may appear in the pedigree, or the family may include individuals with unknown phenotype status. These ambiguous situations require a structured decision framework that prevents both overinterpretation of weak evidence and dismissal of informative data. The framework below provides a stepwise method for evaluating segregation evidence when the pedigree does not fit the idealized patterns described in the ACMG/AMP standards.

### Step 1: Classify Each Family Member into One of Six Observation Categories

Before any LOD score calculation or evidence assignment, assign every genotyped family member to a discrete observation category. This classification forces explicit documentation of what each person contributes to the analysis and prevents vague descriptions such as "seems to segregate" from entering the clinical record.

The six categories are:

1. **Affected carrier**: The individual has a confirmed clinical phenotype consistent with the disease and carries the variant. This observation supports PP1.
2. **Unaffected non-carrier**: The individual has no evidence of disease and does not carry the variant. This observation supports PP1.
3. **Affected non-carrier**: The individual has a confirmed clinical phenotype but does not carry the variant. This observation supports BS4, unless a phenocopy explanation applies.
4. **Unaffected carrier**: The individual carries the variant but shows no evidence of disease. This observation is neutral or weakly supports PP1 depending on the penetrance estimate.
5. **Unknown phenotype carrier**: The individual carries the variant but their clinical status cannot be determined due to age, incomplete evaluation, or unavailable medical records. This observation is uninformative for segregation.
6. **Unknown phenotype non-carrier**: The individual does not carry the variant but their clinical status is unknown. This observation is uninformative for segregation.

The distinction between categories 4 and 5 is critical. An unaffected carrier with a documented clinical evaluation at an age beyond the typical onset range provides different information than a carrier whose phenotype has never been assessed. The [EMBL-EBI Training](https://www.ebi.ac.uk/training) resources emphasize the importance of structured phenotype data collection in genetic analysis, and this classification system applies that principle to segregation evidence.

### Step 2: Determine Whether the Pedigree Is Informative for Segregation

A pedigree is informative for segregation analysis only when it contains at least one observation from categories 1, 2, or 3. A pedigree with only unaffected carriers and unknown phenotype individuals provides no segregation evidence because no observation distinguishes between the linkage and no-linkage hypotheses.

The minimum informative pedigree for PP1 is one affected carrier and one unaffected non-carrier in a dominant disease model. This provides one informative meiosis. For BS4, the minimum is one affected non-carrier in a family where the variant is present in other affected relatives.

Pedigrees that contain only affected carriers without any unaffected individuals are partially informative. The LOD score calculation can proceed, but the absence of unaffected individuals means the analysis cannot distinguish between complete segregation and a situation where the variant is present in all family members regardless of disease status.

### Step 3: Evaluate the Phenotype Confidence for Each Affected Individual

The reliability of the affected status determination directly affects the weight of each segregation observation. An affected individual whose phenotype was confirmed by clinical examination, biochemical testing, or imaging provides stronger evidence than an individual whose affected status is based on family report.

Assign a phenotype confidence level to each affected individual:

- **High confidence**: The phenotype was confirmed by a qualified clinician using objective diagnostic criteria. Examples include a prolonged prothrombin time of 33.1 seconds in the factor VII deficiency family described in [PubMed](https://pubmed.ncbi.nlm.nih.gov/41451501) or the keratotic papules in sebaceous areas documented in the Darier disease proband from [PubMed](https://pubmed.ncbi.nlm.nih.gov/40350400).
- **Moderate confidence**: The phenotype was confirmed by clinical examination but without objective diagnostic testing, or the phenotype was reported by a clinician but not directly examined by the laboratory.
- **Low confidence**: The phenotype is based on family report, medical records of uncertain quality, or the individual has not been examined.

Observations with low phenotype confidence should be excluded from LOD score calculations or included only in sensitivity analyses. Including low-confidence affected individuals can inflate the LOD score if they are carriers or deflate it if they are non-carriers, leading to incorrect PP1 or BS4 assignments.

### Step 4: Apply the Decision Rules for PP1 Assignment

Once each family member is classified and phenotype confidence is established, apply the following decision rules to determine whether PP1 applies and at what strength.

**Rule A: PP1 is not applicable when the pedigree contains an affected non-carrier with high or moderate phenotype confidence.** This observation contradicts co-segregation and should trigger evaluation for BS4 instead. The only exception is when the affected non-carrier has a phenotype that is clearly a phenocopy, such as a different disease subtype or an environmental cause, and this determination is documented by a clinician.

**Rule B: PP1 at supporting strength applies when the pedigree contains at least one affected carrier and the segregation pattern is consistent with the inheritance model.** This includes pedigrees with one affected parent and one affected child, two affected siblings with an affected parent, or an affected individual with an unaffected non-carrier relative. The LOD score will typically be below 1.0.

**Rule C: PP1 at moderate strength requires a calculated LOD score of approximately 1.5 or higher and no conflicting observations.** Conflicting observations include affected non-carriers, unaffected carriers in a disease with high penetrance, or unknown phenotype individuals who could alter the LOD score if their status were determined.

**Rule D: PP1 at strong strength requires a calculated LOD score of approximately 3.0 or higher and no conflicting observations.** The pedigree must include multiple affected carriers across at least three generations or multiple branches. The [ClinGen Sequence Variant Interpretation Working Group](https://www.ncbi.nlm.nih.gov/) recommendations provide the framework for converting LOD scores to evidence strength levels.

**Rule E: When the pedigree contains unaffected carriers, calculate the LOD score under at least two penetrance assumptions.** Use the published penetrance estimate for the gene and disease if available, and a conservative lower estimate. If the LOD score remains above the threshold for a given evidence strength under both assumptions, assign that strength. If the LOD score crosses a threshold between assumptions, assign the lower strength and document the sensitivity analysis.

### Step 5: Apply the Decision Rules for BS4 Assignment

BS4 requires a different decision pathway because it addresses the absence of segregation instead of its presence.

**Rule F: BS4 at supporting strength applies when at least one affected non-carrier with high or moderate phenotype confidence is identified in a family where the variant is present in other affected relatives.** The affected non-carrier must have a phenotype consistent with the disease being investigated. If the affected non-carrier has a different condition, the observation does not support BS4.

**Rule G: BS4 should be upgraded to moderate or strong strength only when multiple affected non-carriers are identified across different branches of the family.** A single affected non-carrier in a small family could represent a phenocopy or a misdiagnosis. Multiple affected non-carriers in different branches provide more robust evidence against pathogenicity.

**Rule H: BS4 is not applicable when the affected non-carrier has low phenotype confidence or when the disease has a high phenocopy rate.** In diseases such as hereditary cancer syndromes where phenocopies are common, an affected non-carrier provides weaker evidence. The phenocopy rate should be estimated from published literature and documented in the analysis.

**Rule I: BS4 should not be applied when the affected non-carrier is the only affected individual in the family.** In this situation, the variant may be a de novo event in another family member, or the affected non-carrier may have a different genetic cause. The segregation pattern is uninformative without additional affected relatives.

### Step 6: Document the Decision Path in the Clinical Record

The decision framework produces a documented trail that supports the final PP1 or BS4 assignment. The clinical record should include:

- The classification of each family member into one of the six observation categories
- The phenotype confidence level for each affected individual
- The penetrance assumptions used in LOD score calculations
- The calculated LOD score and the software or method used
- The specific decision rule that led to the evidence strength assignment
- Any sensitivity analyses performed

This documentation serves two purposes. First, it allows another laboratory to reproduce the analysis and reach the same conclusion. Second, it provides a clear rationale if the classification is challenged during clinical review or by an external auditor. The [Galaxy Training Network](https://training.galaxyproject.org/) emphasizes reproducible analysis documentation, and this framework applies that principle to the specific context of segregation evidence.

### Common Decision Traps in Ambiguous Pedigrees

Several recurring decision errors occur when laboratories encounter ambiguous segregation data. Recognizing these traps improves the consistency of PP1 and BS4 assignments.

**Trap 1: Treating an unaffected carrier as evidence against pathogenicity.** An unaffected carrier in a disease with incomplete penetrance does not contradict the pathogenicity hypothesis. The carrier may simply be non-penetrant. The LOD score calculation must incorporate the penetrance estimate instead of treating the observation as a segregation failure.

**Trap 2: Ignoring young unaffected individuals.** An unaffected child who has not reached the typical age of onset provides limited information. The child may develop the disease later. The analysis should either exclude young individuals from the unaffected category or apply an age-dependent penetrance model.

**Trap 3: Combining segregation evidence across unrelated families.** PP1 and BS4 apply to segregation within a single family. Combining observations from multiple unrelated families into a single LOD score calculation is not appropriate because the families may have different genetic causes or different phenocopy rates. Each family should be analyzed separately, and the evidence from each family should be documented independently.

**Trap 4: Applying BS4 when the affected non-carrier has a milder or atypical phenotype.** An affected individual with an atypical presentation may have a different condition that mimics the disease under investigation. The phenotype must be carefully evaluated before the observation is used as BS4 evidence. The factor VII deficiency family from [PubMed](https://pubmed.ncbi.nlm.nih.gov/41451501) illustrates the importance of objective biochemical testing, in this case coagulation tests showing a prolonged prothrombin time, to confirm the phenotype.

**Trap 5: Using segregation data from a single family to override strong functional evidence.** Segregation evidence is one component of the ACMG/AMP framework. A variant with strong functional evidence and a LOD score below 1.0 should not be downgraded based on limited segregation data. Conversely, a variant with apparent segregation failure should not be classified as pathogenic based on functional evidence alone if the segregation data are robust. The integration of evidence categories requires judgment and should be documented.

### A Worked Example Using the Decision Framework

Consider a hypothetical dominant disease pedigree with the following observations: the proband is affected and carries the variant, the proband's affected sibling carries the variant, the proband's unaffected mother carries the variant, and the proband's unaffected father does not carry the variant. The disease has a published penetrance of 70%.

Applying the framework: the affected carrier sibling supports PP1, the unaffected non-carrier father supports PP1, and the unaffected carrier mother is neutral or weakly supports PP1 depending on the penetrance assumption. The pedigree contains two informative meioses from the affected carrier and unaffected non-carrier observations. The LOD score calculation under 70% penetrance produces a value below 1.0. Under Rule B, PP1 at supporting strength applies. The unaffected carrier mother does not trigger BS4 because she is unaffected, not affected. The documentation should note that the LOD score was calculated under the published 70% penetrance estimate and that a sensitivity analysis using 50% penetrance produced a similar result.

This framework provides a reproducible method for handling the ambiguous pedigrees that dominate clinical practice. By classifying each observation, evaluating phenotype confidence, applying explicit decision rules, and documenting the analysis, laboratories can apply PP1 and BS4 consistently across cases. The framework does not replace the need for clinical judgment, but it ensures that judgment is applied systematically and transparently.

## Frequently Asked Questions

### How many informative meioses are needed to apply PP1 at moderate strength?

Approximately three to five informative meioses are typically needed for PP1 at moderate strength, corresponding to a LOD score of about 1.5 to 2.0. The exact number depends on the disease model, including penetrance and phenocopy rate. A pedigree with an affected grandparent, affected parent, and two affected children provides three informative meioses. Laboratories should calculate the LOD score instead of relying on the number of meioses alone.

### Can PP1 be applied in recessive disease families?

Yes, PP1 applies to recessive diseases when the segregation pattern confirms that affected individuals carry two pathogenic alleles and unaffected individuals carry one or none. The compound heterozygote example from the MYORG family illustrates this application. The LOD score calculation for recessive diseases differs from dominant diseases because the expected segregation ratios are different.

### What is the difference between PP1 and BS4?

PP1 applies when the variant co-segregates with disease, meaning affected individuals carry the variant and unaffected individuals do not. BS4 applies when the variant fails to segregate, meaning an affected individual does not carry the variant. PP1 supports pathogenicity, while BS4 supports a benign classification. Both criteria require careful pedigree interpretation and adjustment for penetrance and phenocopy rates.

### How does incomplete penetrance affect the LOD score?

Incomplete penetrance reduces the LOD score for a given pedigree because an unaffected carrier is more compatible with the linkage hypothesis. The penetrance value must be specified in the disease model used for LOD score calculation. A lower penetrance value makes the observation of unaffected carriers less informative.

### When should BS4 be applied?

BS4 should be applied when an affected individual in a family does not carry the variant despite the variant being present in other affected relatives. The affected individual must have a phenotype consistent with the disease being investigated. The phenocopy rate should be considered before applying BS4.

### Can segregation evidence alone classify a variant?

Segregation evidence alone rarely provides enough evidence for a definitive classification. PP1 at strong strength requires a LOD score of approximately 3.0, which requires an extended pedigree. Most clinical families provide only supporting or moderate evidence. Segregation evidence should be combined with other ACMG/AMP evidence categories.

### How should laboratories document segregation analysis?

Laboratories should document the pedigree structure, the genotype of each family member, the disease model parameters, the LOD score, and the evidence strength assignment. The documentation should include the software used and the version. This documentation supports transparency and reproducibility.

### What are the common errors in segregation analysis?

Common errors include failing to confirm pedigree relationships, incomplete phenotyping, ignoring reduced penetrance, confusing segregation with linkage, applying BS4 without considering phenocopies, and overinterpreting small pedigrees. Laboratories should implement quality controls to prevent these errors.

## Related Bioinformatics Guides

- [Lipidomic Analysis: A Beginner's Guide to Workflows and Data Interpretation](/knowledge/bioinformatics/lipidomic-analysis-a-beginner-s-guide-to-workflows-and-data-interpretation)
- [Spatial Transcriptomics Data Analysis: A Guide to Preprocessing, Integration, and Interpretation](/knowledge/bioinformatics/spatial-transcriptomics-data-analysis-a-guide-to-preprocessing-integration-and-interpretation)
- [Spatial Omics Data Analysis: From Image Processing to Biological Interpretation](/knowledge/bioinformatics/spatial-omics-data-analysis-from-image-processing-to-biological-interpretation)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Microbiome Data Analysis in R: A Practical Guide for Compositional Data](/knowledge/bioinformatics/microbiome-data-analysis-in-r-a-practical-guide-for-compositional-data)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [[Analysis of clinical feature and genetic variant in a Chinese Han pedigree affected with Darier's disease].](https://pubmed.ncbi.nlm.nih.gov/40350400). Zhonghua yi xue yi chuan xue za zhi = Zhonghua yixue yichuanxue zazhi = Chinese journal of medical genetics, 2025.
- [A Novel Germline MLH1 In-Frame Deletion in a Slovenian Lynch Syndrome Family Associated with Uncommon Isolated PMS2 Loss in Tumor Tissue.](https://pubmed.ncbi.nlm.nih.gov/32197529). Genes, 2020.
- [[Analysis of a Chinese pedigree affected with hereditary factor Ⅶ deficiency due to compound heterozygous variants of F7 gene].](https://pubmed.ncbi.nlm.nih.gov/41451501). Zhonghua yi xue yi chuan xue za zhi = Zhonghua yixue yichuanxue zazhi = Chinese journal of medical genetics, 2025.
- [[A case report of a family with Primary familial brain calcification caused by a novel MYORG gene variants].](https://pubmed.ncbi.nlm.nih.gov/40555662). Zhonghua yi xue yi chuan xue za zhi = Zhonghua yixue yichuanxue zazhi = Chinese journal of medical genetics, 2025.
- [Not "just a VUS".](https://doi.org/10.1016/j.gimo.2026.104401). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.