How to Check if Your Data Are Normally Distributed: Shapiro-Wilk, Q-Q Plots and When It Matters
By Dr. Zubair Khalid, DVM, MS, PhD ·

A normality test asks a simple question: could this sample have come from a normal (Gaussian) distribution? The Shapiro-Wilk test answers that question with a p value, and a Q-Q plot answers it with a picture. Neither one tells you whether your analysis is valid. That judgment depends on sample size, on what your model actually assumes, and on how badly the data depart from the bell curve.
You will meet this problem constantly. A t test, an ANOVA, a correlation, or a regression all carry an assumption about normality somewhere [4], and reviewers often ask for evidence that you checked it. The trap is that the check itself is easy to misread. Small samples usually pass, large samples fail on trivial deviations, and the assumption in a regression applies to residuals, not to the raw measurements you plotted first.
Quick Answer
- A normality test evaluates the null hypothesis that the sample distribution is normal. A significant result (typically p < 0.05) means the data depart from normality; a non-significant result means the test failed to detect a departure, not that normality is proven [4].
- The Shapiro-Wilk statistic is $W = \left(\sum_{i=1}^{n} a_i x_{(i)}\right)^2 / \sum_{i=1}^{n}(x_i - \bar{x})^2$, where $x_{(i)}$ are the ordered sample values, $\bar{x}$ is the sample mean, and the $a_i$ are constants derived from the means, variances and covariances of the order statistics of a normal sample. Small values of $W$ indicate departure from normality [6].
- Shapiro-Wilk has better power than the Kolmogorov-Smirnov test, even with the Lilliefors correction, and is the test most often recommended for this job [4].
- Always pair the test with a Q-Q plot. Visual inspection alone is unreliable, and a p value alone hides the shape of the problem [4].
- For ANOVA and regression, test the residuals, not the pooled raw data. Two normal groups separated by a constant shift will fail a normality test when pooled and pass when analyzed correctly.
- With samples above roughly 30 to 40, mild non-normality rarely causes major problems, and with hundreds of observations the distribution of the data can often be ignored [4][5].
What a Normality Test Actually Tests
The normal distribution is a symmetric bell-shaped curve defined by a mean and a standard deviation (you can explore it with the Normal Distribution Calculator). Parametric procedures such as correlation, regression, t tests and ANOVA are built on the assumption that the data come from such a distribution [4]. The assumption matters because those methods use the mean and the standard deviation to build confidence intervals and p values, and the arithmetic behind those intervals assumes a particular shape.
A normality test formalizes the check. The null hypothesis is that the sample distribution is normal. A small p value is evidence against that hypothesis, so you reject normality. A large p value means the data are compatible with normality, which is a weaker statement than most people assume [4].
Two families of tests dominate practice. The Shapiro-Wilk test, introduced in 1965, measures how well the ordered data correlate with the scores you would expect from a normal distribution [1][4]. The Kolmogorov-Smirnov test measures the largest vertical gap between the empirical cumulative distribution of your data and a theoretical normal cumulative distribution [7]. The Anderson-Darling test modifies the K-S approach to give more weight to the tails and uses the specific distribution when computing critical values [6].
How the Shapiro-Wilk Test Works
The formula rewards a specific pattern. If your sorted data line up neatly with the sorted values a normal distribution would produce, the weighted sum $\sum a_i x_{(i)}$ grows large relative to the spread of the data, and $W$ approaches 1. If the data are skewed or heavy-tailed, the ordered values drift away from their expected positions, the numerator shrinks, and $W$ falls.
The $a_i$ weights are not arbitrary. They come from the means, variances and covariances of the order statistics of a normal sample, which is why the test performs well in comparison studies against other goodness-of-fit tests [6]. You will rarely compute $W$ by hand. Software returns it along with a p value, and the p value is what you report.
For the Kolmogorov-Smirnov statistic, the calculation is a maximum gap:
$$D = \max_i \left[ F(Y_i) - \frac{i-1}{N},\ \frac{i}{N} - F(Y_i) \right]$$
where $Y_i$ are the ordered data values, $N$ is the sample size, and $F$ is the theoretical cumulative distribution function you are comparing against [7]. The K-S test applies only to continuous distributions, is more sensitive near the center of the distribution than in the tails, and requires the distribution to be fully specified. If you estimate the mean and standard deviation from your own data, the critical region is no longer valid [7]. That last point is the reason the plain K-S test has a poor reputation for normality testing.
Reading the Output: Normality Test Interpretation
Start with the p value, then look at the plot, then think about the sample size. Each step catches a different failure.
A significant Shapiro-Wilk result means the data are not consistent with a normal distribution [4]. A non-significant result means the test did not detect a departure. Those are different claims. An idealized exponential sample, which is strongly right-skewed by construction, gives Shapiro-Wilk p = 0.13 at n = 10. The data are not normal. The test simply lacks the power to say so at that sample size.
Sample size drives everything. For small samples, normality tests have little power to reject the null, so small samples most often pass [4]. For large samples, significant results appear even for trivial deviations from normality, and those small deviations will not affect the results of a parametric test [4]. An idealized sample from a t distribution with 10 degrees of freedom, which has excess kurtosis of 1.0 and represents a mild departure, gives Shapiro-Wilk p = 1.00 at n = 10, p = 0.98 at n = 200, p = 0.0071 at n = 1000, and p = 3.9e-12 at n = 5000. Same distribution, four different verdicts, all technically correct.
This is why the p value cannot be the whole story. Ghasemi and Zahediasl recommend assessing normality both visually and with a test, and they suggest the plain K-S test should no longer be used because of its low power [4].
Q-Q Plots and What They Show
A Q-Q plot places the quantiles of your data against the quantiles of the normal distribution [4]. A normal probability plot is the same idea with ordered response values on the vertical axis and normal order statistic medians on the horizontal axis [8]. If the data are normal, the points fall along an approximately straight line [8].
Departures from that line have characteristic shapes. Right-skewed data produce a quadratic pattern in which all points lie below a reference line drawn between the first and last points, and such data may be better modeled by a lognormal or Weibull distribution [9]. Short tails, heavy tails and outliers each leave their own signature [8]. Ghasemi and Zahediasl note that a Q-Q plot is easier to interpret than a P-P plot for large samples, and that showing plots lets readers judge the assumption themselves [4].
The plot also survives the sample-size problem better than the test. At n = 5000, the t(10) sample fails Shapiro-Wilk decisively, but the Q-Q plot still shows a nearly straight line with slightly heavy tails. That picture tells you the deviation is mild. The p value does not.
Worked Example
Two samples of n = 20, both analyzed the same way.
Sample A (roughly normal): 4.8, 5.1, 5.3, 4.6, 5.0, 5.5, 4.9, 5.2, 4.7, 5.0, 5.4, 4.9, 5.1, 5.3, 4.8, 5.0, 5.2, 4.6, 5.1, 4.9.
Mean 5.020, median 5.00, SD 0.253, skewness 0.055, excess kurtosis -0.584. Shapiro-Wilk W = 0.975, p = 0.862. Lilliefors D = 0.083, p about 0.97. Anderson-Darling A² = 0.175 against a 5% critical value of 0.721. D'Agostino-Pearson p = 0.868. Q-Q plot correlation r = 0.992. Every test agrees, and the plot is a straight line. Conclusion: no evidence against normality.
Sample B (right-skewed, for example a concentration): 1.2, 0.8, 2.5, 1.0, 0.6, 3.9, 1.4, 0.9, 7.8, 1.1, 0.7, 2.0, 1.6, 0.5, 5.2, 1.3, 0.9, 1.8, 12.4, 1.0.
Mean 2.430 against a median of 1.25, SD 2.960, skewness 2.54, excess kurtosis 6.67. Shapiro-Wilk W = 0.637, p = 7.4e-6. Lilliefors D = 0.308, p ≤ 0.001. Anderson-Darling A² = 2.83 against a 1% critical value of 0.992. Q-Q r = 0.787. Conclusion: clearly non-normal.
The naive K-S test with mean and SD estimated from Sample B gives p = 0.035, far less extreme than Lilliefors. That gap is the cost of estimating parameters and then using K-S critical values that assume you did not [7].
After a log transform, log(B) has skewness 1.06, Shapiro-Wilk W = 0.910, p = 0.064, and Anderson-Darling A² = 0.699, just under the 5% critical value of 0.721. The log scale is much closer to normal, which is the standard remedy when data are not normal: transform them, or use a method that does not require normality [5].
Residuals versus raw data. Group 1 is 10.1, 9.6, 10.4, 9.9, 10.0, 10.3, 9.7, 10.2, and group 2 is group 1 plus 5.0. Pooled, the raw data are bimodal and fail Shapiro-Wilk (W = 0.736, p = 0.0004). The residuals from the group means pass (W = 0.931, p = 0.257). For ANOVA and regression, the assumption concerns the errors, not the pooled raw data, and testing the wrong thing produces a false alarm.
| Test | What it measures | Best used when |
|---|---|---|
| Shapiro-Wilk | Correlation between data and normal scores | Default choice; strong power [4][6] |
| Kolmogorov-Smirnov | Maximum gap from the normal CDF | Parameters known in advance [7] |
| Lilliefors | K-S with estimated mean and variance | Parameters estimated from the sample [2] |
| Anderson-Darling | Gap weighted toward the tails | Tail behavior matters [6] |
| Q-Q plot | Visual agreement across the whole range | Always, alongside a test [4] |
Common Mistakes
- Treating a non-significant p value as proof of normality. It means the test failed to detect a departure. At n = 10, an exponential sample gives p = 0.13. Report it as "no evidence against normality."
- Testing raw data when the model assumes normal residuals. Pooled groups separated by a constant shift fail Shapiro-Wilk (W = 0.736, p = 0.0004) while the residuals pass (W = 0.931, p = 0.257). For ANOVA and regression, check the residuals [10].
- Using the plain K-S test with estimated parameters. The critical region is invalid once you estimate location and scale from the data [7]. Use the Lilliefors correction or Shapiro-Wilk instead [4].
- Ignoring sample size when reading the result. Small samples pass because the test has little power [4]. Large samples fail on deviations too small to matter [4].
- Skipping the plot. Visual inspection alone is unreliable, but a p value alone hides whether the problem is skew, heavy tails or a single outlier [4][8].
- Running a normality test and stopping there. Residual plots reveal the error distribution only if the functional part of the model is correct, the SD is constant, there is no drift and the errors are independent [10].
Limitations
The rules of thumb have edges. The "n > 30 or 40" threshold for ignoring non-normality is a rule of thumb quoted by Ghasemi and Zahediasl [4], and how well a test tolerates non-normality depends on how skewed the data are, whether there are outliers, and whether group sizes are equal. Altman and Bland make the underlying point: the means of random samples from any distribution will themselves have a normal distribution, so with samples of hundreds of observations the distribution of the data can often be ignored [5]. What matters is not that the observed data are normal, but that the sample values are compatible with a population having a normal distribution [5].
The sample-size demonstrations above use idealized quantile samples instead of random draws. Real random samples give more variable p values, though the direction holds: low power at small n, high sensitivity at large n. Lilliefors p values from lookup tables are approximate and capped, which is why the worked example reports p ≤ 0.001 and p about 0.97 instead of exact figures.
Whether to run formal normality tests at all is genuinely debated. Many statisticians prefer Q-Q plots of residuals combined with subject knowledge, since the test answers a question nobody asked ("is this exactly normal?") when what matters is "is the deviation large enough to distort my inference?" Software defaults vary too. SPSS, for instance, recommends its K-S (Lilliefors) and Shapiro-Wilk tests only for samples smaller than 50, a software-specific convention, not a statistical law [4]. Check the current documentation for whatever you use.
Frequently Asked Questions
What is a normality test in plain terms?
It is a hypothesis test with the null hypothesis that your sample came from a normal distribution [4]. A significant result means the data are not consistent with normality. A non-significant result means the test found no evidence against it, which is not the same as confirming it.
Should I use Shapiro-Wilk or Kolmogorov-Smirnov?
Shapiro-Wilk, in most cases. It has better power than the K-S test even after the Lilliefors correction, and Ghasemi and Zahediasl suggest the plain K-S test should no longer be used [4]. The K-S test requires a fully specified distribution, so it breaks down when you estimate the mean and SD from your own data [7]. If you need a K-S style test, use the Lilliefors version [2].
Do I need normality for a t test?
The assumption applies to the data, and with samples above roughly 30 to 40 a violation should not cause major problems [4]. With hundreds of observations, the distribution of the data can often be ignored because the sample mean itself is normally distributed [5]. For small samples with clear skew, transform the data or switch to a method that does not require normality [5].
How do I read a Q-Q plot?
If the points fall along an approximately straight line, the data are consistent with normality [8]. Right-skewed data produce a quadratic pattern with points below a reference line drawn between the first and last points [9]. Heavy tails bow outward at both ends, and outliers appear as isolated points far from the line [8].
Can I check normality with skewness and kurtosis instead?
Yes, as a supplement. Both are 0 for a normal distribution. The z statistic, computed as skewness divided by its standard error, is significant at P < 0.05 when it exceeds ±1.96, and in samples of 200 or more the criterion should be ±2.58 [4]. These are descriptive summaries, not a substitute for a Q-Q plot.
References
- Shapiro & Wilk 1965, An analysis of variance test for normality (complete samples), Biometrika 52:591-611
- Lilliefors 1967, On the Kolmogorov-Smirnov test for normality with mean and variance unknown, JASA 62:399-402
- Anderson & Darling 1952, Asymptotic theory of certain goodness of fit criteria, Ann Math Stat 23:193-212
- Ghasemi & Zahediasl 2012, Normality tests for statistical analysis: a guide for non-statisticians, Int J Endocrinol Metab 10:486-489
- Altman & Bland 1995, Statistics notes: the normal distribution, BMJ 310:298
- NIST/SEMATECH e-Handbook 7.2.1.3 Anderson-Darling and Shapiro-Wilk tests
- NIST/SEMATECH e-Handbook 1.3.5.16 Kolmogorov-Smirnov Goodness-of-Fit Test
- NIST/SEMATECH e-Handbook 1.3.3.21 Normal Probability Plot
- NIST/SEMATECH e-Handbook, Normal probability plot: data are skewed right
- NIST/SEMATECH e-Handbook 4.4.4.5 Testing whether random errors are normally distributed
Related Articles
- Cox Proportional Hazards Model Assumptions
- Mann-Whitney U Test: Step-by-Step for Biology
- Log-Transformation and Variance Stabilization in Single-Cell RNA-Seq: A Practical Guide to Choosing the Right Transformation
- T-Test vs Z-Test: Which One to Use (With Examples)
- One-Way vs Two-Way ANOVA: Differences, Assumptions and a Worked Example