Tukey-Kramer Procedure: Definition, Formula and Example

By Dr. Zubair Khalid, DVM, MS, PhD ·

Tukey-Kramer Procedure: Definition, Formula and Example

When your one-way ANOVA is significant, you know at least one group mean differs, but not which one. The Tukey-Kramer procedure answers that question by testing every pair of means at once while holding the family-wise error rate at your chosen alpha. It is the unequal-sample-size version of Tukey's HSD test, and it is the default post-hoc choice in most statistical software.

Quick Answer

  • Use it after a significant one-way ANOVA when you want all pairwise comparisons, not a planned subset.
  • The test statistic is $q = \dfrac{|\bar{y}_i - \bar{y}_j|}{SE}$, where $SE = \sqrt{\dfrac{MS_{within}}{2}\left(\dfrac{1}{n_i} + \dfrac{1}{n_j}\right)}$ [1][2].
  • You compare $q$ to the studentized range critical value $q_{\alpha, k, df_{within}}$, where $k$ is the number of groups [1].
  • With equal sample sizes the formula collapses to $HSD = q \sqrt{MS_{within}/n}$, the classic Tukey HSD [2].
  • Reject the null for a pair when the absolute mean difference exceeds the critical value, or equivalently when the confidence interval excludes 0 [3].

Before You Start

The Tukey-Kramer procedure is a follow-up test, so it assumes the ANOVA itself was appropriate. Check these conditions first.

Independence. Observations must be independent within and across groups. Repeated measures on the same subject violate this and call for a different procedure.

Normality. The residuals should be roughly normal. With small samples, check a normal quantile plot. The test is fairly tolerant of mild non-normality, especially as sample sizes grow.

Equal variances. The groups should share a common variance. The Tukey-Kramer adjustment handles unequal sample sizes, but it does not fix unequal variances. If the largest group SD is several times the smallest, consider a Welch ANOVA with Games-Howell instead.

A significant omnibus test. The usual practice is to run the post-hoc only when the ANOVA F test rejects [4]. Tukey-Kramer controls the family-wise error rate on its own, so this gate is a convention rather than a requirement.

The right kind of question. Tukey-Kramer is for all pairwise comparisons. If you planned two or three specific contrasts in advance, a planned comparison procedure has more power.

Step by Step

  1. Fit the one-way ANOVA. Compute the group means, the within-group sum of squares $SS_{within}$, and the mean square $MS_{within} = SS_{within} / (N - k)$, where $N$ is the total number of observations and $k$ is the number of groups.
  1. Find the studentized range critical value. Look up $q_{\alpha, k, df_{within}}$ for your alpha, the number of groups, and the within-group degrees of freedom. This value comes from the studentized range distribution, not the t distribution [1][5].
  1. Compute the standard error for each pair. For groups $i$ and $j$ with sizes $n_i$ and $n_j$:

$$SE_{ij} = \sqrt{\frac{MS_{within}}{2}\left(\frac{1}{n_i} + \frac{1}{n_j}\right)}$$

When all groups have the same size $n$, this simplifies to $SE = \sqrt{MS_{within}/n}$ [2].

  1. Form the test statistic. For each pair, $q_{ij} = \dfrac{|\bar{y}_i - \bar{y}_j|}{SE_{ij}}$.
  1. Build the confidence interval. The interval for the difference $\bar{y}_i - \bar{y}_j$ is:

$$(\bar{y}_i - \bar{y}_j) \pm q_{\alpha, k, df_{within}} \cdot SE_{ij}$$

  1. Decide. If the interval excludes 0, the pair differs significantly. If it contains 0, you have no evidence of a difference [3].
  1. Report. Give the mean difference, the adjusted confidence interval, and the adjusted p-value for each pair. A compact letter display is a common way to summarize which groups share a letter and therefore do not differ [2].

Worked Example

A greenhouse trial grows plants under three fertilizers, A, B and C, with six plants per fertilizer. Growth is measured in centimeters.

FertilizerGrowth (cm)
A12.1, 13.4, 11.8, 12.9, 13.1, 12.5
B15.2, 14.8, 15.6, 15.1, 14.9, 15.4
C13.0, 12.6, 13.4, 12.8, 13.2, 12.9

Step 1. Group summaries. The means are A: 12.6333, B: 15.1667, C: 12.9833. The sample standard deviations are A: 0.6121, B: 0.3011, C: 0.2858. The grand mean is 13.5944.

Step 2. ANOVA. The between-group sum of squares is 22.6144 and the within-group sum of squares is 2.7350. With $k = 3$ groups and $N = 18$ observations, $df_{between} = 2$ and $df_{within} = 15$. So $MS_{between} = 11.3072$ and $MS_{within} = 0.1823$. The F statistic is $11.3072 / 0.1823 = 62.0140$ with $p = 0.0000$. The omnibus test is significant, so you proceed.

Step 3. Critical value. For $\alpha = 0.05$, $k = 3$ and $df_{within} = 15$, the studentized range critical value is $q = 3.6734$.

Step 4. Standard error. All groups have $n = 6$, so $SE = \sqrt{0.1823/6} = 0.1743$.

Step 5. HSD half-width. $q \cdot SE = 3.6734 \times 0.1743 = 0.6404$. Any absolute mean difference larger than 0.6404 is significant.

Step 6. Pairwise results.

PairMean difference95% CIAdjusted pReject
A vs B2.5333[1.8930, 3.1737]0.0000Yes
A vs C0.3500[-0.2904, 0.9904]0.3562No
B vs C-2.1833[-2.8237, -1.5430]0.0000Yes

Fertilizer B produces significantly more growth than both A and C. Fertilizers A and C do not differ from each other, since their interval contains 0 and the adjusted p-value is 0.3562.

Step 7. Code. The same analysis in Python with statsmodels:

import pandas as pd
from statsmodels.stats.multicomp import pairwise_tukeyhsd
df = pd.DataFrame({'fertilizer': ['A']*6+['B']*6+['C']*6,
                   'growth': [12.1,13.4,11.8,12.9,13.1,12.5,
                              15.2,14.8,15.6,15.1,14.9,15.4,
                              13.0,12.6,13.4,12.8,13.2,12.9]})
print(pairwise_tukeyhsd(df['growth'], df['fertilizer'], alpha=0.05))

Output:

Multiple Comparison of Means - Tukey HSD, FWER=0.05
===================================================
group1 group2 meandiff p-adj   lower  upper  reject
---------------------------------------------------
     A      B   2.5333    0.0   1.893 3.1737   True
     A      C     0.35 0.3562 -0.2904 0.9904  False
     B      C  -2.1833    0.0 -2.8237 -1.543   True
---------------------------------------------------

The output matches the hand calculation exactly. The F statistic is 62.0140, the adjusted p-values are 0.0000, 0.3562 and 0.0000, and the confidence intervals line up to four decimal places.

Other Ways to Do It

R. The base function TukeyHSD() takes an aov fit and returns the same pairwise table. The multcomp package with glht() gives more control over the confidence interval format and works with a wider range of models [3].

SAS. PROC GLM with a MEANS statement and the TUKEY option produces the pairwise comparisons. The Tukey-Kramer adjustment is applied automatically when group sizes differ [2].

Excel. Excel has no built-in Tukey-Kramer function. You can compute it by hand from the ANOVA table: get $MS_{within}$ and $df_{within}$ from the single-factor ANOVA output, look up $q$ from a studentized range table, then apply the formulas above. This is workable for a handful of groups but error-prone for many.

SPSS. The One-Way ANOVA dialog includes a Post Hoc panel where Tukey is one of the listed options. It reports the pairwise table with adjusted significance values.

Troubleshooting

The ANOVA is significant but no pair is significant. This can happen when the omnibus test is borderline and the Tukey adjustment is conservative. It is more common with unequal sample sizes, where the confidence coefficient exceeds $1 - \alpha$ [1][5]. Check whether a planned contrast would answer your actual question.

The adjusted p-values look much larger than the raw ones. That is expected. The adjustment controls the error rate across all pairs at once, so individual p-values rise. With three groups there are three comparisons, with five groups there are ten.

Group sizes differ a lot. Use the per-pair standard error formula, not the equal-n shortcut. Software does this automatically, but hand calculations often get it wrong.

The confidence interval and the p-value disagree. They should not. If they do, you have probably used the wrong $q$ value or mismatched the degrees of freedom.

Variances are clearly unequal. Tukey-Kramer does not correct for heteroscedasticity. Switch to a procedure designed for it, such as a Games-Howell test for all pairs, or a Welch t-test for a single pair of groups.

Common Mistakes

  • Running Tukey-Kramer without a significant ANOVA. The procedure is designed as a follow-up. Tukey-Kramer already controls the family-wise error rate, but running it after a non-significant F test invites conflicting conclusions. Fix: run the omnibus test first and report both results.
  • Using the equal-n formula with unequal groups. The shortcut $SE = \sqrt{MS_{within}/n}$ is only valid when all $n$ are equal [2]. Fix: use the per-pair formula with $1/n_i + 1/n_j$.
  • Reading the raw p-values instead of the adjusted ones. The whole point of the procedure is the multiplicity adjustment [4]. Fix: report the adjusted p-value and the simultaneous confidence interval.
  • Treating a non-significant pair as proof of equality. Failing to reject means the data are consistent with equal means, not that the means are identical. Fix: describe the result as "no evidence of a difference" and report the interval.
  • Confusing Tukey-Kramer with a t-test on each pair. Running separate two-sample t-tests gives a different, higher error rate. Fix: use one simultaneous procedure for all pairs.
  • Ignoring the direction of the difference. A negative mean difference is not an error. Fix: report the sign and interpret it in the units of the outcome.

Limitations

Tukey-Kramer controls the family-wise error rate across all pairwise comparisons, which is exactly what you want when you have no specific hypotheses. That same control makes it conservative. When you only care about a few planned contrasts, or when one group is the natural control, a procedure built for that structure will detect real differences that Tukey-Kramer misses. The test also assumes normally distributed, independent observations with a common variance, and it does not repair violations of those assumptions.

The procedure tells you which pairs differ, not why, how large the effect is in practical terms, or whether the difference matters. A statistically significant difference of 0.5 cm may be meaningless in a greenhouse trial. Always pair the pairwise table with effect sizes and the raw group means before drawing conclusions.

Frequently Asked Questions

What is the difference between Tukey HSD and the Tukey-Kramer procedure?

They are the same test. Tukey's HSD uses a single standard error and is exact when all group sizes are equal. When group sizes differ, the standard error is computed separately for each pair, and that version is called the Tukey-Kramer method [1][2]. Most software applies the Tukey-Kramer version automatically.

When should I use Tukey-Kramer instead of a t-test?

Use Tukey-Kramer whenever you are comparing more than two group means and want to test all pairs. A series of independent t-tests does not control the error rate across the whole set of comparisons, so the chance of at least one false positive grows with each test you run [4].

Does Tukey-Kramer require equal sample sizes?

No. That is its main advantage over the original Tukey HSD. With unequal sample sizes the confidence coefficient is actually greater than $1 - \alpha$, which means the procedure is conservative [1][5]. You lose a little power but the error rate stays controlled.

How do I interpret the adjusted p-value?

The adjusted p-value is the smallest family-wise significance level at which that pair would be declared different. Compare it to your alpha as usual. An adjusted p-value of 0.3562 for the A vs C pair means that pair is not significant at the 0.05 level, while 0.0000 for A vs B means it clearly is.

Can I use Tukey-Kramer after a two-way ANOVA?

Yes, with care. In a two-way design you typically run the pairwise comparisons on the factor of interest, using the residual mean square and residual degrees of freedom from the full model. If an interaction is present, comparing main-effect means can be misleading, so examine the cell means instead.

References

  1. 7.4.7.1. Tukey's method
  2. 2.3: Tukey Test for Pairwise Mean Comparisons - Statistics LibreTexts
  3. Multiple (pair-wise) comparisons using Tukey's HSD and the compact letter display - Statistics with R
  4. Lee S, Lee DK. (2018). What is the proper way to apply the multiple comparison test? Korean journal of anesthesiology
  5. Part 2: Statistical Analysis and Modeling | Introduction to Experimental Design - passel

Further Reading

Related Articles