Shapiro-Wilk vs. Kolmogorov-Smirnov
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- For biological samples with fewer than 50 observations, the Shapiro-Wilk test demonstrates superior statistical power compared to the Kolmogorov-Smirnov test in detecting deviations from normality, making it the preferred choice for small sample sizes common in experiments involving gene expression or enzyme activity assays.
- The Shapiro-Wilk test is particularly effective at identifying departures from normality caused by skewness and kurtosis, which are frequently observed in biological data such as cell counts or hormone levels, thus providing a more reliable basis for deciding between parametric and nonparametric analyses.
- The Kolmogorov-Smirnov test, especially in its standard form, exhibits lower power for small sample sizes and is less sensitive to tail departures and outliers, which are critical considerations for biological data like protein concentrations or patient response measurements.
- Visual inspection of data through histograms and Q-Q plots is crucial and should always accompany formal normality tests, as these graphical methods can reveal patterns like bimodality or extreme outliers that formal tests might miss, especially in very small datasets.
- When normality test results conflict, especially with sample sizes between 10 and 50, the Shapiro-Wilk test's higher power should be given greater weight, and the decision to proceed with parametric tests should be informed by the magnitude of the departure and the robustness of the intended statistical model, such as t-tests or ANOVA.
Quick Answer
- For biological samples with fewer than 50 observations, the Shapiro-Wilk test provides greater statistical power than the Kolmogorov-Smirnov test for detecting departures from normality.
- Use Shapiro-Wilk as your primary normality test for small sample sizes, and report the test statistic along with the p-value in your methods section.
- Both tests can fail to detect non-normality in very small samples, so visual inspection of histograms and Q-Q plots should accompany any formal test.
Normality Testing as a Gatekeeping Step in Biological Analysis
Normality testing serves as a gatekeeping step in many biological analyses. Parametric tests such as t-tests, ANOVA, and linear regression assume that residuals follow a normal distribution. When this assumption is violated, the validity of p-values and confidence intervals becomes questionable. Researchers in the life sciences routinely test their data for normality before deciding whether to apply parametric or nonparametric methods.
The choice of normality test matters because different tests have different sensitivities to various types of departures from normality. Some tests detect skewness more readily, while others are more sensitive to kurtosis or outliers. For small sample sizes, the differences between tests become more pronounced, and selecting the wrong test can lead to incorrect conclusions about whether your data meet the assumptions required for parametric analysis.
The Shapiro-Wilk test and the Kolmogorov-Smirnov test are two of the most commonly used normality tests in biological research. Both are available in major statistical software packages including R, Python, SPSS, and GraphPad Prism. However, they operate on different principles and have different strengths and weaknesses, particularly when applied to small samples typical of many biological experiments.
Biological data frequently violate normality assumptions. Measurements such as enzyme activity, gene expression levels, and cell counts often show skewness or contain outliers. Using a normality test with low power, such as the Kolmogorov-Smirnov test, can lead researchers to incorrectly assume normality and apply parametric tests that are not appropriate. The consequences of this error depend on the robustness of the parametric test being used. The t-test is relatively robust to violations of normality, especially with equal sample sizes. However, ANOVA and regression analyses can be more sensitive to non-normality, particularly when sample sizes are unequal or when the non-normality is severe.
For small samples, the Shapiro-Wilk test provides better protection against these errors. Its higher power means that genuine departures from normality are more likely to be detected, prompting researchers to use appropriate nonparametric alternatives or data transformations.
Statistical Foundations of the Two Tests
How the Shapiro-Wilk Test Works
The Shapiro-Wilk test calculates a statistic, denoted as W, that measures how well the ordered sample values fit a normal distribution. The test compares the order statistics of your sample against the expected values of order statistics from a normal distribution. A W value close to 1 indicates that the data closely match a normal distribution, while values substantially below 1 suggest departure from normality.
The test statistic is computed using coefficients derived from the expected values of order statistics from a standard normal distribution. These coefficients depend on the sample size and must be calculated separately for each n. This sample-size-specific calculation is one reason why the Shapiro-Wilk test performs well across a range of sample sizes.
The Shapiro-Wilk test is particularly effective at detecting departures from normality caused by skewness and kurtosis. It has been shown to have good power even with relatively small samples, which makes it well suited for biological experiments where sample sizes are often constrained by cost, ethics, or practical considerations.
How the Kolmogorov-Smirnov Test Works
The Kolmogorov-Smirnov test, often abbreviated as KS test, takes a different approach. It compares the empirical cumulative distribution function of your sample against the cumulative distribution function of a reference normal distribution. The test statistic, denoted as D, represents the maximum absolute difference between these two cumulative distribution functions.
A key limitation of the standard Kolmogorov-Smirnov test is that it requires the parameters of the reference distribution to be specified in advance. When you estimate the mean and standard deviation from your data, the test becomes conservative, meaning it is less likely to reject the null hypothesis of normality. This conservatism reduces the power of the test, particularly for small sample sizes.
The Lilliefors correction addresses this problem by adjusting the critical values of the Kolmogorov-Smirnov test when parameters are estimated from the data. Many statistical software packages apply this correction automatically, but researchers should verify which version of the test they are using when interpreting results.
Power Analysis for Small Samples
Statistical power refers to the probability that a test will correctly reject the null hypothesis when it is false. For normality testing, power is the probability of detecting non-normality when the data truly come from a non-normal distribution. Low power means that the test often fails to detect real departures from normality, leading researchers to incorrectly assume their data are normal.
For small sample sizes, the Kolmogorov-Smirnov test has notably lower power than the Shapiro-Wilk test. This means that with n below 50, the Kolmogorov-Smirnov test is more likely to miss genuine departures from normality. The practical consequence is that researchers using the Kolmogorov-Smirnov test with small samples may proceed with parametric analyses even when the normality assumption is violated.
The Shapiro-Wilk test maintains better power across a range of non-normal distributions, including distributions with heavy tails, skewness, and mixtures. This makes it the preferred choice for biological data, where sample sizes are frequently small and departures from normality can have meaningful effects on downstream analyses.
Head-to-Head Comparison of Test Characteristics
At a Glance
| Feature | Shapiro-Wilk | Kolmogorov-Smirnov |
|---|---|---|
| Best sample size range | n < 50, works well up to several thousand | n > 50, performs poorly for small samples |
| Statistical power for small n | High, detects skewness and kurtosis effectively | Low, often fails to detect non-normality |
| Parameter estimation | Accounts for estimated mean and variance | Standard version requires known parameters, Lilliefors correction needed |
| Sensitivity to outliers | Moderate, can be influenced by extreme values | Low, less sensitive to individual outliers |
| Software availability | R, Python, SPSS, GraphPad Prism, SAS | R, Python, SPSS, GraphPad Prism, SAS |
| Common use in biology | Recommended for small biological samples | Often used for larger datasets or when comparing distributions |
Sample Size and Test Performance
The performance of normality tests depends heavily on sample size. For very small samples, such as n less than 10, both tests have limited power to detect departures from normality. However, the Shapiro-Wilk test still outperforms the Kolmogorov-Smirnov test in most scenarios.
As sample size increases, the power of both tests improves, but the gap between them narrows. For samples larger than 50, the Kolmogorov-Smirnov test becomes more reliable, though the Shapiro-Wilk test remains a valid choice. Some statisticians recommend using the Shapiro-Wilk test for all sample sizes up to several thousand, after which other considerations may come into play.
For biological research, where sample sizes are often determined by the number of animals, patients, or experimental units that can be reasonably included, the Shapiro-Wilk test is generally the safer choice. The consequences of failing to detect non-normality can be substantial, as subsequent parametric analyses may produce misleading results.
Sensitivity to Different Types of Non-Normality
Non-normal distributions can take many forms. Some are skewed, with a longer tail on one side. Others have heavy tails, meaning more extreme values than expected under normality. Still others are bimodal or have other complex shapes. Different normality tests have different sensitivities to these various departures.
The Shapiro-Wilk test shows good sensitivity to both skewness and kurtosis. It can detect distributions that are moderately skewed or that have heavier tails than the normal distribution. This broad sensitivity makes it useful as a general-purpose normality test.
The Kolmogorov-Smirnov test is more sensitive to differences in the center of the distribution and less sensitive to differences in the tails. This means it may miss distributions with heavy tails or outliers, which are common in biological data. For example, gene expression data often contain outliers, and the Kolmogorov-Smirnov test may fail to flag these data as non-normal.
Practical Workflow for Normality Testing
Step 1: Visual Inspection of Your Data
Before running any formal test, examine your data visually. Create a histogram to see the overall shape of the distribution. Look for obvious skewness, multiple peaks, or extreme outliers. Generate a Q-Q plot, which plots your data quantiles against the quantiles of a normal distribution. Points that fall roughly along a straight line suggest normality, while systematic deviations indicate departures.
Visual inspection provides context that formal tests cannot capture. A formal test may return a non-significant p-value, but if the histogram shows clear bimodality or the Q-Q plot shows strong curvature, you should treat the data as non-normal. Conversely, a significant p-value from a formal test may be driven by a single outlier that has a minimal effect on your actual analysis.
Step 2: Run the Shapiro-Wilk Test
For sample sizes below 50, run the Shapiro-Wilk test as your primary formal test. In R, use the shapiro.test() function. In Python, use scipy.stats.shapiro(). In SPSS, select Analyze, Descriptive Statistics, Explore, and check the Normality plots with tests option. In GraphPad Prism, the normality test is available in the Column Statistics analysis.
Record the test statistic W and the p-value. A p-value below your chosen significance threshold, typically 0.05, indicates that the data depart significantly from normality. A p-value above the threshold does not prove normality, but it suggests that the data are consistent with a normal distribution.
Step 3: Confirm with the Kolmogorov-Smirnov Test When Appropriate
For sample sizes above 50, the Kolmogorov-Smirnov test with the Lilliefors correction can serve as a useful complement to the Shapiro-Wilk test. In R, use the lillie.test() function from the nortest package. In Python, use scipy.stats.kstest() with the appropriate parameters, or use the lilliefors function from the statsmodels package.
When results from the two tests conflict, give more weight to the Shapiro-Wilk result for small samples. The lower power of the Kolmogorov-Smirnov test means that a non-significant result from this test does not provide strong evidence of normality when the Shapiro-Wilk test indicates a departure.
Step 4: Decide on Your Analysis Path
If the Shapiro-Wilk test indicates normality, you can proceed with parametric tests such as t-tests, ANOVA, or linear regression. If the test indicates non-normality, consider the following options:
- Use nonparametric alternatives such as the Mann-Whitney U test, Wilcoxon signed-rank test, or Kruskal-Wallis test
- Apply a data transformation such as log, square root, or Box-Cox transformation and retest the transformed data
- Use robust statistical methods that do not rely on normality assumptions
- Consult a biostatistician if the decision is consequential or the data are complex
The choice among these options depends on your research question, the nature of your data, and the analysis you plan to conduct. For simple comparisons between two groups, nonparametric tests are straightforward alternatives. For regression models, transformations or robust methods may be more appropriate.
Decision Matrix for Test Selection
| Scenario | Recommended Test | Rationale |
|---|---|---|
| n < 50, testing normality of a single sample | Shapiro-Wilk | Higher power detects skewness and kurtosis that KS misses |
| n < 50, comparing two empirical distributions | Kolmogorov-Smirnov (two-sample) | KS is designed for distribution comparison, not normality testing |
| n > 50, testing normality before parametric analysis | Shapiro-Wilk or KS with Lilliefors correction | Both have adequate power, SW remains reliable |
| n < 10, any normality assessment | Visual inspection only | Formal tests lack power, histograms and Q-Q plots are more informative |
| Residuals from regression or ANOVA | Shapiro-Wilk on residuals | Tests the actual assumption required by the model |
| Data with known outliers | Shapiro-Wilk with outlier investigation | SW sensitivity to outliers may flag issues that need examination |
Common Failure Patterns in Normality Testing
Relying Solely on the Kolmogorov-Smirnov Test for Small Samples
A frequent error in biological research is using the Kolmogorov-Smirnov test exclusively for normality assessment, regardless of sample size. With small samples, this test often fails to detect non-normality, leading researchers to proceed with parametric analyses on data that violate the normality assumption.
This failure pattern is particularly problematic because the consequences are invisible. The analysis runs without error, p-values are generated, and conclusions are drawn, but the results may be unreliable. The researcher has no indication that the normality assumption was violated.
Ignoring Visual Evidence in Favor of Test Results
Another common failure is over-relying on formal test results while ignoring visual evidence. A researcher may see a histogram with clear skewness but accept normality because the Shapiro-Wilk test returned a p-value above 0.05. This can happen with small samples where the test lacks power to detect moderate departures from normality.
The reverse error also occurs. A researcher may see a Q-Q plot with a few points slightly off the line and reject normality based on a significant Shapiro-Wilk test, even though the departure is minor and unlikely to affect the results of parametric analyses. Statistical significance does not always imply practical significance.
Using Tests Without Understanding Their Assumptions
Many researchers run normality tests without understanding the assumptions underlying each test. The standard Kolmogorov-Smirnov test assumes that the parameters of the reference distribution are known, not estimated from the data. When researchers estimate the mean and standard deviation from their sample and run the standard test, they get conservative results with reduced power.
The Lilliefors correction addresses this issue, but not all software applies it automatically. Researchers should verify which version of the Kolmogorov-Smirnov test their software implements and report this detail in their methods section.
Testing the Wrong Variable
Normality assumptions apply to the residuals of your statistical model, not necessarily to the raw data. For a t-test comparing two groups, the assumption is that the residuals, or the data within each group, are normally distributed. For regression analysis, the assumption is that the residuals from the model are normally distributed.
Some researchers test the normality of the combined data from both groups, which can mask non-normality within groups or create apparent non-normality when the groups have different means. The correct approach is to test normality within each group or to test the residuals from your fitted model.
Misinterpreting Non-Significant Results as Proof of Normality
A non-significant p-value from any normality test does not establish that data are normal. It only indicates that the test failed to find sufficient evidence against normality. With small samples, this failure may reflect low power instead of true normality. Researchers should phrase results as data being consistent with normality, not as proof of normality.
Records and Measurements for Normality Testing
Documenting Your Testing Procedure
Good research practice requires documenting your normality testing procedure. Your methods section should state which test you used, the version of the test, the software and package used, and the sample size. This documentation allows other researchers to reproduce your analysis and assess the appropriateness of your methods.
For the Shapiro-Wilk test, report the W statistic and the p-value. For the Kolmogorov-Smirnov test, report the D statistic and the p-value, and specify whether the Lilliefors correction was applied. Include this information in your results section or in supplementary materials.
Recording Sample Sizes and Missing Data
Sample size is a critical factor in normality testing. Record the number of observations in each group or condition. Note any missing data and how they were handled. Missing data can affect the distribution of your sample and the power of your normality test.
The National Institutes of Health data management and sharing policy emphasizes the importance of documenting data collection and processing procedures. Clear records of sample sizes and missing data handling support the transparency and reproducibility of your research.
Storing Raw Data and Analysis Scripts
Store your raw data in a format that preserves all original measurements. Keep analysis scripts or syntax files that document exactly how normality tests were run. This allows you to rerun the analysis if needed and provides a complete record for reviewers or collaborators.
The National Library of Medicine research methods resources emphasize the importance of reproducible research practices. Storing raw data and analysis scripts supports reproducibility and allows other researchers to verify your results.
Quality Controls and Verification
Checking Software Outputs
Statistical software can produce unexpected results if used incorrectly. Verify that your software is running the correct version of the normality test. For the Kolmogorov-Smirnov test, confirm whether the Lilliefors correction is applied. For the Shapiro-Wilk test, confirm that the correct sample size is being used.
Run a simple check by generating a random sample from a normal distribution with a known mean and standard deviation. The normality test should return a non-significant p-value. Then generate a sample from a clearly non-normal distribution, such as an exponential distribution, and confirm that the test returns a significant p-value. This verification confirms that your software is functioning correctly.
Cross-Validating with Multiple Methods
When the normality decision is consequential, cross-validate your results using multiple methods. Run both the Shapiro-Wilk and Kolmogorov-Smirnov tests, and compare the results with visual inspection of histograms and Q-Q plots. If all methods agree, you can be more confident in your conclusion.
When methods disagree, investigate the source of the disagreement. Examine the distribution shape, check for outliers, and consider whether the sample size is adequate for reliable testing. The disagreement itself may indicate that the data are borderline, and you should consider the robustness of your planned analysis to violations of normality.
Consulting with a Biostatistician
For complex data structures, unusual distributions, or consequential decisions, consult with a biostatistician. A biostatistician can help you choose the appropriate normality test, interpret conflicting results, and select analysis methods that are robust to violations of normality assumptions.
The National Institutes of Health grants and funding resources emphasize the importance of rigorous statistical methods in funded research. Including a biostatistician in your research team or consulting one during the analysis phase can improve the quality of your research and reduce the risk of statistical errors.
Interpretation Limits and Reporting Standards
What a Non-Significant Result Does Not Mean
A non-significant result from a normality test does not prove that your data are normally distributed. It only indicates that the test did not find sufficient evidence to reject the null hypothesis of normality. With small samples, this result may reflect low statistical power instead of true normality.
Report your normality test results with appropriate caution. State that the data were consistent with a normal distribution, instead of claiming that the data are normal. This phrasing acknowledges the limitations of the test while providing the information needed for readers to assess your methods.
What a Significant Result Does Not Mean
A significant result from a normality test indicates that your data depart from normality in some way, but it does not tell you how severe the departure is or whether it will affect your planned analysis. A large sample can produce a significant result for a trivial departure from normality that has no practical impact on your results.
Consider the magnitude of the departure in addition to the p-value. Examine the Q-Q plot to assess how far the points deviate from the line. Consider whether the departure is likely to affect the specific analysis you plan to conduct. Some parametric tests are robust to minor violations of normality, particularly with balanced designs and equal sample sizes.
Reporting Normality Test Results in Your Manuscript
Report normality test results in your methods section, not in your results section. State which test you used, the test statistic, and the p-value. For example, you might write that the Shapiro-Wilk test indicated that the data did not depart significantly from normality (W = 0.97, p = 0.23).
The EQUATOR Network provides reporting guidelines for various study types, and many journals require adherence to these guidelines. Following reporting standards for statistical methods improves the transparency and reproducibility of your research.
Transparency in Research Reporting
The Committee on Publication Ethics core practices emphasize the importance of accurate and complete reporting of research methods and results. This includes transparent reporting of statistical tests and their results. Failing to report normality testing or reporting it inaccurately can undermine the credibility of your research.
When you use nonparametric tests because your data violated normality assumptions, report this decision and the evidence that supported it. When you use parametric tests despite non-normality, explain why you believe the violation is acceptable, such as citing the robustness of the test or the balanced design.
Data Integrity and Reproducibility Context
Data Management and Sharing Expectations
The National Institutes of Health data management and sharing policy requires researchers to plan for data management and sharing in their grant applications. This includes documenting data collection, processing, and analysis procedures. Normality testing is part of the analysis procedure and should be documented in your data management plan.
Maintaining data integrity requires careful attention to all stages of the research process, from data collection through analysis and interpretation. An overview of quantitative research methods published in Seminars in Oncology Nursing emphasizes the importance of careful data management, including checking for errors and missing values before analysis.
Ethical Reporting of Statistical Results
The Committee on Publication Ethics core practices address issues of data fabrication, falsification, and manipulation. Selecting a normality test that gives you the result you want, or failing to report normality testing when it was conducted, constitutes questionable research practice.
Report your normality testing honestly, even when the results complicate your analysis. If your data are non-normal and you must use nonparametric tests, report this clearly. If your data are borderline and you choose to use parametric tests, explain your reasoning.
Professional Escalation Criteria
Seek professional statistical advice when you encounter any of the following situations:
- Your normality test results conflict with visual inspection of your data
- Your sample size is very small, such as fewer than 10 observations per group
- Your data have a complex structure, such as repeated measures or hierarchical nesting
- Your analysis is consequential, such as a clinical trial or regulatory submission
- You are unsure which normality test is appropriate for your data
A biostatistician can provide guidance on test selection, interpretation of results, and appropriate analysis methods for your specific data structure.
Workflow Diagram for Normality Testing
flowchart TD
A[Collect biological data] --> B[Visual inspection: histogram and Q-Q plot]
B --> C{Clear departure from normality?}
C -->|Yes| D[Use nonparametric methods or transform data]
C -->|No| E{Is n less than 50?}
E -->|Yes| F[Run Shapiro-Wilk test]
E -->|No| G[Run Shapiro-Wilk or KS with Lilliefors correction]
F --> H{Is p less than 0.05?}
G --> H
H -->|Yes| D
H -->|No| I[Proceed with parametric analysis]
D --> J[Document decision and report in methods]
I --> J
A Practical Decision Framework for Conflicting Normality Test Results
When the Shapiro-Wilk and Kolmogorov-Smirnov tests produce conflicting results, researchers need a structured approach to resolve the disagreement instead of relying on intuition or convenience. This section provides a decision framework that integrates test outcomes with visual evidence, sample characteristics, and the specific requirements of your planned analysis. The framework is designed for biological researchers who need practical guidance when formal tests disagree.
The Conflict Resolution Protocol
Conflicting results between the Shapiro-Wilk and Kolmogorov-Smirnov tests occur most frequently with sample sizes between 10 and 50 observations. The protocol below establishes a clear sequence for resolving these conflicts.
Step 1: Verify the test conditions. Confirm that both tests were run correctly. For the Kolmogorov-Smirnov test, verify whether the Lilliefors correction was applied. The standard Kolmogorov-Smirnov test without this correction is conservative and will produce higher p-values than the corrected version. If your software applied the standard version, rerun the test with the Lilliefors correction before interpreting the conflict.
Step 2: Assess the magnitude of disagreement. Record both p-values and calculate the ratio between them. If the Shapiro-Wilk p-value is below 0.05 and the Kolmogorov-Smirnov p-value is above 0.05, the conflict is meaningful. If both p-values are on the same side of the threshold but differ in magnitude, the conflict is less consequential for your decision.
Step 3: Examine the distribution shape. Generate a histogram with a superimposed normal curve and a Q-Q plot. Look for specific features that explain the disagreement. Skewness and heavy tails are more likely to be detected by the Shapiro-Wilk test. The Kolmogorov-Smirnov test is more sensitive to differences in the center of the distribution. If the histogram shows moderate skewness or a few outliers in the tails, the Shapiro-Wilk result is likely the more accurate assessment.
Step 4: Evaluate the impact on your planned analysis. Consider the specific parametric test you intend to use. The t-test is relatively robust to violations of normality, particularly with equal sample sizes and similar variances between groups. ANOVA is moderately robust but less so with unequal group sizes. Linear regression is more sensitive to non-normality in residuals, especially when predicting extreme values. If your planned analysis is robust to minor departures from normality, the conflict may not require changing your analysis plan.
Step 5: Apply the tiebreaker rule. For sample sizes below 50, the Shapiro-Wilk test has greater statistical power and should be given more weight in the decision. This is not a preference but a consequence of the tests' statistical properties. The Kolmogorov-Smirnov test's lower power means that a non-significant result provides weaker evidence of normality. When the tests disagree and your sample is below 50, treat the data as non-normal unless visual inspection strongly supports normality.
A Weighted Scoring System for Normality Decisions
For researchers who prefer a more quantitative approach, the following scoring system integrates multiple sources of evidence into a single decision metric. This system is particularly useful when the normality decision is consequential, such as in clinical research or regulatory submissions.
Assign points based on the following criteria:
| Evidence Source | Condition | Points |
|---|---|---|
| Shapiro-Wilk test | p-value above 0.05 | +2 |
| Shapiro-Wilk test | p-value between 0.01 and 0.05 | 0 |
| Shapiro-Wilk test | p-value below 0.01 | -2 |
| Kolmogorov-Smirnov test with Lilliefors correction | p-value above 0.05 | +1 |
| Kolmogorov-Smirnov test with Lilliefors correction | p-value below 0.05 | -1 |
| Histogram | Symmetric, bell-shaped, no obvious outliers | +1 |
| Histogram | Moderate skewness or a single outlier | 0 |
| Histogram | Strong skewness, multiple peaks, or multiple outliers | -1 |
| Q-Q plot | Points fall approximately along the reference line | +1 |
| Q-Q plot | Systematic curvature or S-shape | -1 |
| Q-Q plot | A few points deviate in the tails | 0 |
| Sample size | n between 10 and 30 | +1 |
| Sample size | n between 31 and 50 | 0 |
| Sample size | n above 50 | -1 |
A total score of +3 or higher supports proceeding with parametric analysis. A score of -3 or lower supports using nonparametric methods or transformations. Scores between -2 and +2 indicate borderline data where the decision should be based on the robustness of your planned analysis and consultation with a biostatistician if the analysis is consequential.
This scoring system is not a substitute for statistical judgment. It provides a structured way to combine evidence from multiple sources and reduces the influence of any single test result on your decision.
Distribution-Specific Decision Rules
Different types of non-normality have different implications for parametric analysis. The following rules address common distribution shapes in biological data.
Skewed distributions. When the Shapiro-Wilk test detects skewness but the Kolmogorov-Smirnov test does not, examine the direction and magnitude of the skewness. Mild positive skewness in biological measurements such as enzyme activity or hormone levels is common and often manageable with a log transformation. If the skewness is moderate to severe, use the nonparametric alternative or transform the data before applying parametric methods.
Heavy-tailed distributions. Distributions with heavy tails produce outliers that can inflate variance estimates and distort means. The Shapiro-Wilk test detects heavy tails more readily than the Kolmogorov-Smirnov test. If your data show heavy tails, consider robust statistical methods that downweight the influence of extreme values, or use nonparametric tests that are based on ranks instead of raw values.
Bimodal distributions. Neither test is particularly effective at detecting bimodality, especially with small samples. Visual inspection of the histogram is the most reliable method for identifying bimodal distributions. If your data show two distinct peaks, the normality assumption is clearly violated regardless of what the formal tests indicate. Investigate whether the bimodality reflects two distinct subpopulations in your sample.
Uniform distributions. Data that are uniformly distributed across a range produce a flat histogram. The Kolmogorov-Smirnov test may detect this departure because the cumulative distribution function differs substantially from the normal cumulative distribution function. The Shapiro-Wilk test may also detect uniformity but with less sensitivity. Both tests should flag uniform data as non-normal.
Implementing the Framework in Practice
The following implementation steps integrate the decision framework into your research workflow.
Step 1: Create a normality assessment record. Before running any tests, create a template that includes fields for sample size, test statistics, p-values, histogram assessment, Q-Q plot assessment, and the weighted score. This record ensures that you collect all necessary information and provides documentation for your methods section.
Step 2: Run the tests in a consistent order. Run the Shapiro-Wilk test first, then the Kolmogorov-Smirnov test with the Lilliefors correction. Record both results before examining the visual evidence. This order prevents visual inspection from biasing your interpretation of the formal test results.
Step 3: Document the visual evidence. Describe the histogram and Q-Q plot in writing instead of relying on memory. Note the presence of skewness, outliers, multiple peaks, or other features. This written description becomes part of your analysis record and supports your final decision.
Step 4: Calculate the weighted score. Apply the scoring system to your evidence. Record the individual points and the total score. This calculation provides a transparent basis for your decision and can be reported in supplementary materials if needed.
Step 5: Make and document your decision. Based on the weighted score and the distribution-specific rules, decide whether to proceed with parametric analysis, use nonparametric methods, or apply a transformation. Document the rationale for your decision, including the evidence that supported it.
Step 6: Verify the decision with sensitivity analysis. If you proceed with parametric analysis despite borderline normality evidence, run the nonparametric alternative as a sensitivity check. If both approaches produce similar conclusions, your results are robust to the normality assumption. If they differ substantially, report both analyses and discuss the discrepancy.
Common Failure Patterns in Applying the Framework
Pattern 1: Ignoring the weighted score when it contradicts expectations. Researchers sometimes have a preferred analysis path and disregard evidence that points in a different direction. If your weighted score indicates non-normality but you proceed with parametric analysis without justification, you are introducing potential bias into your results.
Pattern 2: Applying the framework after the analysis is complete. The decision framework is most useful when applied before running your primary analysis. Applying it retrospectively to justify an already completed analysis undermines its purpose and can lead to confirmation bias.
Pattern 3: Using the framework without understanding the underlying data. The scoring system assumes that your data have been checked for errors and missing values. An overview of quantitative research methods published in Seminars in Oncology Nursing emphasizes the importance of careful data management before analysis. If your data contain entry errors or coding mistakes, the normality assessment will be unreliable regardless of the framework used.
Pattern 4: Overweighting the sample size component. The sample size component in the scoring system is designed to account for the lower reliability of normality tests with very small samples. It should not be used to override strong evidence of non-normality from visual inspection and formal tests.
Records and Measurements for the Decision Framework
Maintain the following records when applying the decision framework:
- The raw data file with all original measurements
- The analysis script or syntax file showing exactly how each test was run
- The normality assessment record with test statistics, p-values, and visual evidence descriptions
- The weighted score calculation with individual component points
- The final decision and the rationale supporting it
- The sensitivity analysis results if a nonparametric alternative was run
These records support the transparency and reproducibility of your analysis. The National Institutes of Health data management and sharing policy emphasizes the importance of documenting data processing and analysis procedures. Complete records allow other researchers to verify your methods and conclusions.
Professional Escalation Criteria for the Decision Framework
Seek professional statistical advice when any of the following situations occur:
- The weighted score falls in the borderline range between -2 and +2 and your analysis is consequential
- The Shapiro-Wilk and Kolmogorov-Smirnov tests disagree and visual inspection does not clearly support either result
- Your data show complex features such as multiple peaks, strong heteroscedasticity, or hierarchical structure
- Your planned analysis is sensitive to normality violations and the consequences of an incorrect decision are severe
- You are preparing results for regulatory submission or a clinical trial where statistical rigor is closely scrutinized
The National Institutes of Health grants and funding resources emphasize the importance of rigorous statistical methods in funded research. Consulting a biostatistician when the framework indicates borderline or complex situations improves the quality of your analysis and reduces the risk of statistical errors.
Reporting the Decision Framework in Your Manuscript
When you use the decision framework, report it transparently in your methods section. State that you used a structured approach to resolve conflicting normality test results, describe the components of the framework, and report the weighted score that supported your decision. This level of detail allows reviewers and readers to assess the appropriateness of your methods.
The EQUATOR Network provides reporting guidelines for various study types, and many journals require adherence to these guidelines. Transparent reporting of your normality assessment process, including how you resolved conflicting test results, supports the reproducibility of your research and aligns with the Committee on Publication Ethics core practices for accurate and complete reporting.
Validation of the Framework with Simulated Data
Before applying the framework to your actual data, validate it with simulated data to confirm that your software and scoring system are working correctly. Generate a sample from a normal distribution with a known mean and standard deviation. Run both normality tests and apply the scoring system. The framework should support proceeding with parametric analysis.
Then generate a sample from a clearly non-normal distribution, such as an exponential or chi-square distribution. Run both tests and apply the scoring system. The framework should support using nonparametric methods or transformations. This validation confirms that your implementation of the framework is functioning correctly and provides a reference for interpreting results from your actual data.
The National Library of Medicine research methods resources provide additional guidance on statistical methods and reproducible research practices. Using these resources alongside the decision framework strengthens your approach to normality testing and supports rigorous analysis of biological data.
Frequently Asked Questions
Why does the Shapiro-Wilk test perform better than the Kolmogorov-Smirnov test for small samples?
The Shapiro-Wilk test is specifically designed to detect departures from normality and uses coefficients that are optimized for each sample size. The Kolmogorov-Smirnov test compares your data against a reference distribution and becomes conservative when parameters are estimated from the data, reducing its power. For small samples, this power difference is substantial, making the Shapiro-Wilk test more likely to detect genuine departures from normality.
Should I use the Kolmogorov-Smirnov test at all for biological data?
The Kolmogorov-Smirnov test can be useful for larger samples, typically above 50 observations, and for comparing two distributions against each other instead of testing normality specifically. For normality testing in biological research, the Shapiro-Wilk test is generally preferred. If you use the Kolmogorov-Smirnov test, ensure that the Lilliefors correction is applied when parameters are estimated from your data.
What should I do if the Shapiro-Wilk test and Kolmogorov-Smirnov test give conflicting results?
For small samples, give more weight to the Shapiro-Wilk test result because it has higher power. Examine your data visually with a histogram and Q-Q plot to understand the nature of the distribution. If the visual evidence supports non-normality, consider using nonparametric methods or transformations regardless of the Kolmogorov-Smirnov result.
Can I use parametric tests if my data fail the Shapiro-Wilk test?
You can use parametric tests if the violation of normality is minor and your test is robust to the violation. The t-test is relatively robust to non-normality, especially with equal sample sizes. However, for ANOVA and regression, non-normality can have greater effects. Consider using nonparametric alternatives or transformations, and consult a biostatistician if the decision is consequential.
How large should my sample be for reliable normality testing?
The Shapiro-Wilk test provides reasonable power for samples as small as 10 observations, though power increases with sample size. For samples below 10, formal normality testing has limited value, and visual inspection should be your primary tool. For samples above 50, both the Shapiro-Wilk and Kolmogorov-Smirnov tests provide adequate power, though the Shapiro-Wilk test remains a reliable choice.
Should I test normality of raw data or residuals?
Test the residuals from your statistical model, not the raw data. For a t-test, test the data within each group separately. For regression or ANOVA, test the residuals from the fitted model. Testing the wrong variable can lead to incorrect conclusions about whether the normality assumption is satisfied.
What is the Lilliefors correction and when do I need it?
The Lilliefors correction adjusts the critical values of the Kolmogorov-Smirnov test when the mean and standard deviation are estimated from the data instead of specified in advance. Without this correction, the test is conservative and has reduced power. Most modern statistical software applies the Lilliefors correction automatically, but you should verify this in your software documentation.
How should I report normality test results in my manuscript?
Report the normality test you used, the test statistic, and the p-value in your methods section. For example, state that the Shapiro-Wilk test indicated no significant departure from normality (W = 0.97, p = 0.23). If you used the Kolmogorov-Smirnov test, specify whether the Lilliefors correction was applied. This information allows readers to assess the appropriateness of your statistical methods.
Related Bioinformatics Guides
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight
- Metagenome Co-Assembly: Strategies for Multi-Sample Data
- Proteomics Mass Spectrometry: From Sample Preparation to Data Analysis
- Spatial Transcriptomics Workflow: From Sample Preparation to Data Analysis
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Research Methods Resources. National Library of Medicine.
- EQUATOR Network. EQUATOR Network.
- Core Practices. Committee on Publication Ethics.
- NIH Grants and Funding. National Institutes of Health.
- ORCID for Researchers. ORCID.
- Data Management and Sharing Policy. National Institutes of Health.
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Interventions for treating cavitated or dentine carious lesions.. The Cochrane database of systematic reviews, 2021.
- Association of Gestational Weight Gain With Adverse Maternal and Infant Outcomes.. JAMA, 2019.
- An Overview of the Fundamentals of Data Management, Analysis, and Interpretation in Quantitative Research.. Seminars in oncology nursing, 2023.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.