How to Determine the Number of Principal Components to Retain
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Parallel analysis is the statistically grounded primary method for determining principal component retention in biological data, comparing observed eigenvalues to those from random data to identify components representing real structure beyond chance, unlike the Kaiser criterion or subjective scree plots.
- Eigenvalues, representing the variance explained by each principal component, are the core metric, with the Kaiser criterion retaining components with eigenvalues > 1, though this can overestimate in noisy biological datasets with many variables.
- The scree plot offers a visual check by plotting eigenvalues in descending order to identify an "elbow", but its subjectivity makes it ambiguous in biological data where eigenvalues often decline gradually rather than showing a sharp break.
- Variance explained is foundational; retain components that capture a meaningful proportion of total variance while remaining interpretable for the biological question, acknowledging that arbitrary thresholds like 70-80% may not align with underlying biological structure.
- Data preparation, including handling missing values and scaling variables (especially for gene expression data where genes have different baseline expression levels), is critical before PCA, as scaling ensures variables contribute equitably and prevents high-variance variables from dominating early components.
- Stability testing via resampling (e.g., sample-based subsampling) is crucial for biological data to ensure retention decisions are robust across batches or subsets, as technical artifacts can influence component retention and downstream analyses like clustering.
Quick Answer
- Retain principal components that together explain a meaningful proportion of total variance while each retained component remains interpretable for your biological question.
- Use parallel analysis as your primary decision method, then confirm with a scree plot and the Kaiser criterion for components with eigenvalues above one.
- No single rule fits every dataset, so report the method you used and justify your cutoff with the variance explained and downstream analysis requirements.
Understanding Principal Component Retention in Biological Data
Principal component analysis reduces high dimensional biological data into a smaller set of uncorrelated variables that capture the dominant patterns of variation. The central problem for biologists is deciding how many of these components to keep for downstream analysis. Retaining too few components discards biological signal, while retaining too many reintroduces noise and complicates interpretation.
The decision matters because principal components are ordered by the amount of variance they explain. The first component captures the largest share of total variance, the second captures the next largest share, and so on. After a certain point, additional components explain progressively smaller amounts of variance that may represent technical noise instead of meaningful biological structure.
For gene expression data, the number of samples is often small relative to the number of genes measured. This creates a situation where the total variance is spread across many dimensions, and the distinction between signal and noise becomes less obvious. The retention decision therefore requires a structured approach instead of a default rule.
The choice of how many components to retain affects every downstream step. Clustering results, visualization, and any subsequent statistical modeling all depend on which components you keep. A poorly chosen number can create false biological conclusions or obscure real patterns that exist in the data.
At a Glance
| Method | What It Does | Strengths | Limitations |
|---|---|---|---|
| Kaiser criterion | Retains components with eigenvalues greater than 1 | Simple, automatic, widely used | Can overestimate or underestimate in noisy biological data |
| Scree plot | Visualizes eigenvalues in descending order to find the elbow | Quick visual check, no computation required | Subjective, ambiguous when no clear elbow exists |
| Parallel analysis | Compares observed eigenvalues to those from random data | Data driven, statistically grounded, recommended for biological data | Requires computational resampling, sensitive to sample size |
Core Principles of Principal Component Retention
Variance Explained as the Foundation
The total variance in a dataset is the sum of the variances of all original variables. Each principal component accounts for a portion of this total variance, and the proportion explained by each component decreases as you move from the first to the last. The cumulative proportion of variance explained tells you how much of the total variation your retained components capture.
For biological data, the goal is to capture enough variance to represent the underlying structure without including noise. A common practice is to retain enough components to explain a target proportion of variance, often 70 to 80 percent. However, this threshold is arbitrary and may not align with the actual biological structure in your data.
The variance explained by each component depends on the scaling of your variables. If you standardize your data so each variable has unit variance, the eigenvalues sum to the number of variables. If you do not standardize, variables with larger variances dominate the first components. For gene expression data, standardization is often appropriate because genes have different baseline expression levels and variances.
Eigenvalues as the Core Metric
Each principal component has an associated eigenvalue that equals the variance explained by that component. The eigenvalue is the sum of squared loadings for that component across all variables. Larger eigenvalues indicate components that capture more variance.
The Kaiser criterion retains components with eigenvalues greater than 1. The logic is that a component with an eigenvalue less than 1 explains less variance than a single standardized variable. This rule is simple and widely used, but it has limitations. In datasets with many variables, the criterion can retain too many components because the average eigenvalue is 1 when variables are standardized.
The Kaiser criterion assumes that each variable contributes equally to the total variance. This assumption holds when variables are standardized, but it does not account for the correlation structure among variables. In gene expression data, where genes are often highly correlated, the criterion can be misleading.
The Scree Plot as a Visual Check
The scree plot displays eigenvalues in descending order against the component number. The plot typically shows a steep decline in eigenvalues for the first few components, followed by a gradual leveling off. The point where the curve levels off is called the elbow, and components before the elbow are considered meaningful.
The scree plot is a visual method that requires judgment. Different observers may identify different elbows, especially when the curve declines gradually. The plot is most useful when there is a clear break between the steep and flat portions of the curve.
For biological data, the scree plot can be ambiguous. Gene expression data often produces a gradual decline in eigenvalues instead of a sharp break. This makes the elbow difficult to identify and reduces the reliability of the visual method.
Parallel Analysis as the Statistical Standard
Parallel analysis compares the eigenvalues from your observed data to eigenvalues from random data with the same dimensions. The random data is generated with the same number of observations and variables but with no underlying structure. Components with eigenvalues greater than the corresponding random eigenvalues are retained.
The logic is that any component with an eigenvalue greater than what would be expected by chance is likely to represent real structure. Parallel analysis is considered more robust than the Kaiser criterion or the scree plot because it uses a statistical comparison instead of a fixed threshold or visual judgment.
Parallel analysis requires you to generate random datasets and compute their eigenvalues. The number of random datasets affects the stability of the comparison. More random datasets produce a more stable estimate of the random eigenvalue distribution. The method is computationally intensive but feasible for most biological datasets.
Practical Workflow for Determining Component Retention
Step 1: Prepare Your Data
Before you can determine the number of components to retain, you must prepare your data. This includes handling missing values, deciding whether to scale variables, and checking for outliers.
Missing values must be addressed because principal component analysis requires a complete data matrix. Common approaches include removing observations with missing values, imputing missing values, or using algorithms that handle missing data. The choice depends on the proportion of missing values and the nature of your data.
Scaling is important because principal component analysis is sensitive to the scale of the variables. If you do not scale, variables with larger variances dominate the first components. For gene expression data, scaling is often recommended because genes have different baseline expression levels and variances.
Outliers can distort the variance structure and affect the eigenvalues. You should check for outliers before running the analysis and decide whether to remove them or use robust methods.
Step 2: Run Principal Component Analysis
Once your data is prepared, you run the principal component analysis to obtain the eigenvalues and eigenvectors. The output includes the eigenvalues for each component, the proportion of variance explained, and the cumulative proportion.
The eigenvalues are the basis for all retention methods. You should record the eigenvalues for all components, beyond the first few. This allows you to apply the Kaiser criterion, create a scree plot, and run parallel analysis.
The output also includes the loadings, which tell you how much each original variable contributes to each component. Loadings are important for interpreting the biological meaning of the retained components.
Step 3: Apply the Kaiser Criterion
The Kaiser criterion retains components with eigenvalues greater than 1. This is the simplest method and can be applied directly to the eigenvalues from your analysis.
To apply the criterion, you count the number of components with eigenvalues greater than 1. This number is the number of components to retain.
The criterion is easy to implement but has limitations. In datasets with many variables, the criterion may retain too many components. In datasets with few variables, it may retain too few. The criterion does not account for the correlation structure among variables.
Step 4: Examine the Scree Plot
Create a scree plot by plotting the eigenvalues against the component number. Look for the elbow where the curve levels off. The number of components before the elbow is the number to retain.
The scree plot is a visual method that requires judgment. You should look for a clear break in the curve. If the curve declines gradually, the elbow may be difficult to identify.
The scree plot is most useful when the data has a clear structure. For biological data with gradual declines, the plot may be ambiguous.
Step 5: Run Parallel Analysis
Parallel analysis compares your eigenvalues to those from random data. You generate random datasets with the same number of observations and variables, compute their eigenvalues, and compare them to your observed eigenvalues.
The number of random datasets affects the stability of the comparison. A larger number of random datasets provides a more stable estimate of the random eigenvalue distribution. A common choice is to generate 100 or more random datasets.
Retain components whose observed eigenvalues exceed the corresponding random eigenvalues. The number of components that meet this criterion is the number to retain.
Parallel analysis is more robust than the Kaiser criterion or the scree plot. It uses a statistical comparison instead of a fixed threshold or visual judgment.
Step 6: Compare the Methods
The three methods may not agree on the number of components to retain. The Kaiser criterion may suggest a different number than the scree plot or parallel analysis.
When the methods disagree, you should examine the data to understand why. The Kaiser criterion may be influenced by the number of variables. The scree plot may be ambiguous. Parallel analysis may be affected by the sample size.
In general, parallel analysis is considered the most reliable method. The Kaiser criterion is a simple approximation that may be useful when the data is well structured. The scree plot is a visual check that can help you understand the variance structure.
Step 7: Consider the Downstream Analysis
The number of components you retain affects the downstream analysis. If you are using the components for clustering, you need enough components to capture the biological structure. If you are using the components for regression, you need to avoid overfitting.
The choice of the number of components should be guided by the downstream analysis. You may need to try different numbers of components and evaluate the results.
For example, if you are using the components for clustering, you can compare the clustering results for different numbers of components. The number of components that produces the most biologically meaningful clusters is the best choice.
Options and Tradeoffs in Retention Methods
Kaiser Criterion
The Kaiser criterion is the simplest method for determining the number of components to retain. It is based on the idea that a component should explain at least as much variance as a single variable. The criterion is easy to apply and requires no additional computation.
The criterion is most appropriate when the variables are standardized and the number of variables is moderate. In this case, the average eigenvalue is 1, and the criterion retains components that explain more than the average variance.
The criterion has limitations. In datasets with many variables, the criterion may retain too many components. In datasets with few variables, it may retain too few. The criterion does not account for the correlation structure among variables.
Scree Plot
The scree plot is a visual method that displays the eigenvalues against the component number. The plot shows the variance explained by each component and helps you identify the point where the variance levels off.
The scree plot is useful for understanding the variance structure of your data. It can help you see whether there is a clear break between the components that capture biological structure and those that capture noise.
The scree plot is subjective. Different people may identify different elbows in the same plot. The plot is most useful when there is a clear break in the curve.
Parallel Analysis
Parallel analysis is a statistical method that compares the eigenvalues from your data to those from random data. The method is more robust than the Kaiser criterion or the scree plot because it uses a statistical comparison.
The tradeoff is that parallel analysis requires additional computation. You need to generate random data and compute the eigenvalues. The number of random datasets affects the stability of the comparison.
Parallel analysis is the recommended method for biological data. It is more reliable than the other methods and provides a statistical basis for the decision.
Cumulative Variance Explained
The cumulative variance explained is the proportion of total variance captured by the retained components. A common rule is to retain enough components to explain 70 to 80 percent of the total variance.
This rule is simple and easy to apply. However, the threshold is arbitrary and may not align with the biological structure in your data. The rule does not account for the correlation structure among variables.
The cumulative variance rule is most useful when you have a clear idea of how much variance you need to capture. For example, if you are using the components for a regression, you may need to capture a certain proportion of the variance to avoid overfitting.
Observations and Measurements
Eigenvalues
The eigenvalues are the primary measurement for determining the number of components to retain. You should record the eigenvalues for all components, beyond the first few.
The eigenvalues tell you how much variance each component explains. The first component has the largest eigenvalue, and the eigenvalues decrease as you move to the next components.
The eigenvalues are used in all retention methods. The Kaiser criterion uses the eigenvalue threshold of 1. The scree plot uses the eigenvalues to create the plot. Parallel analysis compares the eigenvalues to random data.
Proportion of Variance Explained
The proportion of variance explained is the eigenvalue divided by the total variance. This tells you the percentage of total variance captured by each component.
The proportion of variance explained is useful for understanding the variance structure. It helps you see how much variance is captured by the first few components and how much is left for the remaining components.
The cumulative proportion of variance explained is the sum of the proportions for the retained components. This tells you the total variance captured by the retained components.
Loadings
The loadings tell you how much each original variable contributes to each component. The loadings are the coefficients of the linear combination that defines the component.
The loadings are useful for interpreting the biological meaning of the components. A component with high loadings for a group of genes may represent a biological pathway or process.
The loadings are also useful for checking the stability of the components. If the loadings change dramatically when you change the number of components, the components may not be stable.
Records and Documentation
Record the Eigenvalues
You should record the eigenvalues for all components. This allows you to apply the retention methods and to reproduce the analysis.
The eigenvalues should be recorded in a table with the component number and the eigenvalue. You should also record the proportion of variance explained and the cumulative proportion.
Document the Retention Method
You should document the method you used to determine the number of components to retain. This includes the method, the number of components retained, and the justification.
The documentation should be included in your analysis report. This allows others to understand your decision and to reproduce the analysis.
Record the Downstream Analysis
You should record the downstream analysis and the results. This includes the clustering results, the regression results, or any other analysis you performed.
The downstream results help you evaluate the retention decision. If the downstream analysis produces meaningful results, the retention decision is likely appropriate.
Quality Controls and Checks
Check the Variance Structure
You should check the variance structure of your data before you run the principal component analysis. This includes checking the distribution of the variables and the correlation structure.
The variance structure affects the eigenvalues and the retention decision. If the variables have very different variances, you may need to scale the data.
Check the Eigenvalues
You should check the eigenvalues for any anomalies. The eigenvalues should be positive and decreasing. If the eigenvalues are negative, there may be a problem with the data.
Check the Loadings
You should check the loadings for the retained components. The loadings should be interpretable and biologically meaningful.
If the loadings are not interpretable, the components may not be capturing biological structure. You may need to consider a different number of components.
Check the Downstream Results
You should check the downstream results to see if they are biologically meaningful. The clustering results should separate the samples into meaningful groups. The regression results should be interpretable.
If the downstream results are not meaningful, you may need to reconsider the number of components.
Common Failure Patterns
Retaining Too Many Components
Retaining too many components is a common failure pattern. This happens when you use the Kaiser criterion or the cumulative variance rule with a high threshold.
Retaining too many components adds noise to the downstream analysis. The noise can obscure the biological signal and lead to incorrect conclusions.
Retaining Too Few Components
Retaining too few components is another common failure pattern. This happens when you use the scree plot and identify the elbow too early.
Retaining too few components discards biological signal. The downstream analysis may miss important biological patterns.
Ignoring the Correlation Structure
Ignoring the correlation structure is a common failure pattern. The correlation structure affects the eigenvalues and the retention decision.
If the variables are highly correlated, the eigenvalues may be larger than expected. This can lead to retaining too many components.
Using a Single Method
Using a single method is a common failure pattern. The Kaiser criterion, the scree plot, and parallel analysis may produce different results.
Using a single method may lead to a biased decision. You should use multiple methods and compare the results.
Limitations and Interpretation
The Retention Decision Is Not Absolute
The number of components to retain is not a fixed number. It depends on the data, the downstream analysis, and the biological question.
The retention decision should be based on the data and the analysis. You should not use a single rule without considering the context.
The Methods Have Limitations
Each retention method has limitations. The Kaiser criterion is simple but may be biased. The scree plot is subjective. Parallel analysis is robust but requires computation.
You should be aware of the limitations and use the methods appropriately.
The Biological Interpretation Is Important
The retention decision should be guided by the biological interpretation. The retained components should be biologically meaningful.
If the retained components are not biologically meaningful, the retention decision may be wrong. You should consider the biological context when making the decision.
Professional Escalation Criteria
When to Seek Help
You should seek help when you are unsure about the retention decision. This includes when the methods disagree, when the scree plot is ambiguous, or when the downstream results are not meaningful.
You should also seek help when you are not familiar with the methods. The retention decision is important and should be made carefully.
When to Consult a Statistician
You should consult a statistician when the data is complex or when the retention decision is critical. A statistician can help you choose the appropriate method and interpret the results.
You should also consult a statistician when you are not sure about the assumptions of the methods. The methods have assumptions that may not be met in your data.
A Decision Framework for Retention Stability Across Resampling and Subset Validation
The methods described above answer the question of how many components to retain for one complete dataset. A separate practical problem arises when you need to know whether that retention decision is stable. Biological datasets are often collected in batches, from different laboratories, or across time points. A retention decision that holds for one subset of samples may not hold for another. This section provides a decision framework for testing retention stability, a record system for tracking those checks, and a troubleshooting method for diagnosing instability when it appears.
Why Stability Testing Matters for Biological Data
Gene expression datasets are rarely collected in a single uninterrupted run. Samples arrive in batches, sequencing depth varies, and technical artifacts can differ between processing dates. These factors create variability that is not biological. Principal component analysis will capture this technical variability in some components, and the retention decision will shift depending on how much technical noise is present.
A retention decision based on a single run of the full dataset can be misleading. If you retain components based on parallel analysis of the full dataset, you have no information about whether the same number of components would be retained if you removed a batch of samples or if you analyzed a subset. This matters because downstream analyses such as clustering or differential expression testing are sensitive to the number of components used. A decision that is not stable across subsets can produce results that do not generalize to new data.
The framework below treats retention as a decision that must be validated, not a single calculation. It uses resampling to test whether the number of retained components changes when the input data changes. This approach is consistent with the principle that a robust analytical decision should not depend on a small number of samples or a single technical artifact.
Step 1: Define the Resampling Plan
Before you run any stability checks, you must define how you will resample the data. The resampling plan determines what kind of instability you can detect. There are two main approaches: sample-based resampling and variable-based resampling.
Sample-based resampling involves drawing subsets of the observations. This tests whether the retention decision is stable when you remove or duplicate samples. For gene expression data, this is the most relevant approach because sample composition often varies between batches. You can draw random subsets of samples, or you can draw subsets that correspond to known batches.
Variable-based resampling involves drawing subsets of the genes or features. This tests whether the retention decision depends on a small number of highly influential genes. This is useful when you suspect that a few genes with very high variance are driving the first few components.
For most biological datasets, sample-based resampling is the more informative choice. It directly addresses the question of whether your retention decision will hold when you collect new samples. Variable-based resampling is a secondary check that can be useful when you have a small number of genes with extreme variance.
Step 2: Run Parallel Analysis on Each Resampled Dataset
For each resampled dataset, you must run the full retention analysis. This means applying parallel analysis to each subset, beyond computing the eigenvalues. The reason is that the number of components retained by parallel analysis can change when the sample size changes.
Parallel analysis compares observed eigenvalues to eigenvalues from random data with the same dimensions. When you resample the data, the dimensions change. A subset with fewer samples will have a different random eigenvalue distribution than the full dataset. This means the retention decision can change even if the underlying biological structure is the same.
You should run parallel analysis on each resampled dataset using the same number of random datasets that you used for the full dataset. This ensures that the comparison is consistent. Record the number of components retained for each resampled dataset.
The number of resampled datasets depends on the size of your data and the computational resources available. A minimum of 100 resampled datasets is a reasonable starting point. For datasets with many samples, you may need fewer resamples because the subsets are more similar to the full dataset. For datasets with few samples, you may need more resamples to get a stable estimate of the retention distribution.
Step 3: Record the Retention Distribution
The output of the resampling procedure is a distribution of retention counts. For each resampled dataset, you have a number of components retained. This distribution tells you how stable the retention decision is.
The most useful summary is the mode, which is the most common number of components retained across the resampled datasets. If the mode matches the number of components you retained for the full dataset, the decision is stable. If the mode is different, the full dataset decision may be an artifact of the specific samples you have.
You should also record the range of retention counts. A narrow range, such as 4 to 6 components, indicates moderate stability. A wide range, such as 2 to 10 components, indicates that the retention decision is highly sensitive to the samples included.
The distribution should be recorded in a table with the number of components retained and the frequency of that count across the resampled datasets. This table is the primary record of the stability check.
Step 4: Compare the Full Dataset Decision to the Resampling Distribution
The full dataset retention decision is the number of components you retained from the original analysis. This number should be compared to the distribution from the resampling procedure.
If the full dataset decision is the mode of the resampling distribution, you can proceed with confidence. The decision is representative of the data structure and is not dependent on a specific set of samples.
If the full dataset decision is not the mode, you need to investigate. The full dataset may include samples that are not representative of the broader population. You should examine the samples that are driving the difference and consider whether they should be included in the analysis.
If the resampling distribution is wide, you should consider whether the retention decision is meaningful at all. A wide distribution means that the number of components is not a stable property of the data. In this case, you may need to use a different approach, such as retaining components based on the downstream analysis instead of on a statistical criterion.
Step 6: Use the Stability Check to Adjust the Retention Decision
The stability check is beyond a diagnostic. It can be used to adjust the retention decision when the full dataset decision is not stable.
If the resampling distribution has a clear mode that is different from the full dataset decision, you should consider using the mode as the retention number. This is because the mode is the most common result across different subsets of the data, and it is more likely to generalize to new samples.
If the resampling distribution is wide and does not have a clear mode, you should not rely on a single retention number. Instead, you should consider a range of components and evaluate the downstream analysis across that range. This is a more conservative approach that acknowledges the uncertainty in the retention decision.
The stability check should be reported alongside the retention decision. This allows other researchers to see how stable the decision is and to interpret the downstream results accordingly.
Records and Measurements for Stability Checks
The stability check requires a specific set of records. These records are separate from the eigenvalues and loadings recorded in the main analysis.
The first record is the resampling protocol. This includes the type of resampling used, the number of resampled datasets, and the size of each resampled dataset. This information is necessary for reproducing the stability check.
The second record is the retention distribution. This is the table of retention counts across the resampled datasets. This table should include the number of components retained and the frequency of that count.
The third record is the comparison between the full dataset decision and the resampling distribution. This includes the mode of the distribution and the range of retention counts.
These records should be kept with the main analysis records. They provide the evidence for the stability of the retention decision and allow others to evaluate the robustness of the analysis.
Common Failure Patterns in Stability Checks
The Full Dataset Decision Is Not the Mode
This is the most common failure pattern. The full dataset retention decision is different from the most common retention count across the resampled datasets. This indicates that the full dataset includes samples that are not representative of the broader population.
The cause is often a batch effect or a small number of outlier samples. These samples can influence the eigenvalues and change the retention decision. The solution is to identify these samples and determine whether they should be included in the analysis.
The Resampling Distribution Is Wide
A wide distribution means that the retention count varies substantially across the resampled datasets. This indicates that the retention decision is not a stable property of the data.
The cause is often a dataset with a weak biological signal. When the signal is weak, the eigenvalues are close to the random eigenvalues, and small changes in the data can change the retention decision. The solution is to consider whether the dataset has enough signal for principal component analysis to be meaningful.
The Resampling Distribution Has Multiple Modes
A distribution with multiple modes means that the data supports more than one retention decision. This can happen when the data has multiple distinct structures, such as a strong batch effect and a weaker biological effect.
The cause is often a dataset with a strong technical artifact. The first components capture the artifact, and the biological signal is in later components. The retention decision depends on whether you want to include the artifact components. The solution is to consider the downstream analysis and whether the artifact components should be retained.
The Stability Check Is Computationally Expensive
Running parallel analysis on 100 or more resampled datasets can be computationally intensive. This is especially true for large datasets with many genes.
The cause is the number of resampled datasets and the size of the dataset. The solution is to reduce the number of resampled datasets or to use a smaller subset of the data for the stability check. The stability check does not need to use the full dataset. A subset of the samples can provide a reasonable estimate of the retention distribution.
When to Escalate to a Statistician
The stability check is a diagnostic tool, and it can reveal problems that require statistical expertise to resolve. You should escalate to a statistician when the stability check reveals a wide distribution or multiple clusters.
A wide distribution indicates that the retention decision is not stable, and a statistician can help you determine whether the data is suitable for principal component analysis. A multiple cluster distribution indicates that the data has multiple structures, and a statistician can help you determine how to handle the technical artifacts.
You should also escalate when the stability check is computationally infeasible. A statistician can help you design a more efficient resampling protocol or recommend an alternative approach.
Integration with the Main Retention Decision
The stability check is not a replacement for the main retention methods. It is a validation step that should be used after you have applied parallel analysis, the scree plot, and the Kaiser criterion to the full dataset.
The stability check provides additional evidence for the retention decision. If the full dataset decision is stable across resampled datasets, you can report the decision with confidence. If the decision is not stable, you need to investigate the cause and adjust the decision accordingly.
The stability check should be reported in the methods section of your analysis. You should report the resampling protocol, the retention distribution, and the comparison to the full dataset decision. This allows others to evaluate the robustness of your retention decision.
The stability check is particularly important for gene expression data because of the batch effects and technical artifacts that are common in this type of data. A retention decision that is not stable across batches may not be reproducible in a new experiment. The stability check provides a way to test the reproducibility of the decision before you commit to a downstream analysis.
Frequently Asked Questions
What is the Kaiser criterion for principal component retention?
The Kaiser criterion retains components with eigenvalues greater than 1. The logic is that a component with an eigenvalue less than 1 explains less variance than a single original variable. The criterion is simple and easy to apply, but it may not be appropriate for all datasets.
How does parallel analysis work?
Parallel analysis compares the eigenvalues from your data to eigenvalues from random data with the same dimensions. You generate random data, compute the eigenvalues, and compare them to your observed eigenvalues. Retain components whose eigenvalues exceed the random eigenvalues.
What is the scree plot?
The scree plot displays eigenvalues against the component number. The plot shows the variance explained by each component. The elbow is the point where the curve levels off, and the number of components before the elbow is the number to retain.
How many components should I retain for gene expression data?
The number of components depends on the data and the downstream analysis. Parallel analysis is recommended for gene expression data because it is more robust than the Kaiser criterion or the scree plot. You should also consider the biological interpretation of the components.
What is the cumulative variance rule?
The cumulative variance rule retains components until the cumulative proportion of variance explained reaches a threshold, often 70 to 80 percent. The rule is simple but the threshold is arbitrary and may not align with the biological structure.
Can I use the Kaiser criterion for all data?
The Kaiser criterion is simple and easy to apply, but it may not be accurate for all data. The criterion does not account for the correlation structure among variables. You should use the criterion with caution and compare it to other methods.
What is the difference between the Kaiser criterion and parallel analysis?
The Kaiser criterion uses a fixed threshold of 1 for the eigenvalue. Parallel analysis compares the eigenvalues to those from random data. Parallel analysis is more robust because it uses a statistical comparison.
How do I report the retention decision?
You should report the method you used, the number of components retained, and the justification. You should also report the eigenvalues and the proportion of variance explained. This allows others to understand the decision and to reproduce the analysis.
Using the Evidence
| Source | Best use in this topic | Important limitation |
|---|---|---|
| Research Methods Resources | official guidance | Check the linked page for current local requirements |
| EQUATOR Network | official guidance | Check the linked page for current local requirements |
| Core Practices | official guidance | Check the linked page for current local requirements |
Related Bioinformatics Guides
- Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight
- Metagenomics Data Analysis: From Raw Reads to Biological Insights
- Proteomics Data Analysis Workflow: From Raw Spectra to Biological Insights
- Spatial Omics Data Analysis: From Image Processing to Biological Interpretation
- Single-Cell Sequencing Analysis Pipeline: From Raw Data to Biological Insights
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Research Methods Resources. National Library of Medicine.
- EQUATOR Network. EQUATOR Network.
- Core Practices. Committee on Publication Ethics.
- NIH Grants and Funding. National Institutes of Health.
- ORCID for Researchers. ORCID.
- Data Management and Sharing Policy. National Institutes of Health.
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Thiazide and the Thiazide-Like Diuretics: Review of Hydrochlorothiazide, Chlorthalidone, and Indapamide.. American journal of hypertension, 2022.
- Characterizing CSF inflammatory proteomics in pediatric post-hemorrhagic hydrocephalus and Anti-NMDAR encephalitis.. Journal of neuroinflammation, 2025.
- A posteriori dietary patterns: how many patterns to retain?. The Journal of nutrition, 2014.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.