One-Way vs Two-Way ANOVA: Differences, Assumptions and a Worked Example

By Dr. Zubair Khalid, DVM, MS, PhD ·

One-Way vs Two-Way ANOVA: Differences, Assumptions and a Worked Example

Analysis of variance (ANOVA) is a general technique for testing whether the means of two or more groups are equal [1]. The name sounds backwards at first: you compare means, but the math compares variances. That is the whole trick. ANOVA splits the total variation in your data into the part explained by your grouping factors and the part left over as random error, then asks whether the explained part is large relative to the leftover part [8].

You will meet ANOVA constantly in lab work and clinical research. Comparing a control group against two drug doses, testing whether a coating performs differently across materials, checking whether a genotype changes the response to treatment: all of these are ANOVA questions. The practical decision that trips people up is which version to run. One factor or two? And if two, what does the interaction term actually mean? This article answers that with definitions, the underlying model, and a fully worked example you can reproduce.

Quick Answer

  • One-way ANOVA tests one factor with three or more levels against the null hypothesis $H_0: \mu_1 = \mu_2 = \dots = \mu_k$, where $\mu_i$ are the population means of each group [7]. With only one factor, the analysis is one-way; with two factors, it is two-way [1].
  • Two-way ANOVA tests two factors at once and, critically, their interaction: whether the effect of one factor depends on the level of the other [4].
  • The core statistic is $F = \text{MST} / \text{MSE}$, the ratio of between-group variance to within-group variance. When the null is true, both estimate the same error variance and the ratio sits near 1; when the null is false, the numerator grows [3].
  • The key difference goes beyond "one factor vs two." A two-way design partitions variance across both factors and their interaction, which can remove nuisance variation from the error term and increase sensitivity [9].
  • Interpretation rule: if the interaction is significant, read the simple effects (one factor at each level of the other) before the main effects, because main effects become averages that can hide opposing patterns.

One-Way ANOVA: One Factor, Many Levels

A factor is an independent treatment variable whose settings you control and vary; each setting is a level, and levels can be numbers or simply present/not present [1]. In a one-way design you have exactly one such factor.

The model is:

$$Y_{ij} = \mu + \tau_i + e_{ij}$$

Here $Y_{ij}$ is the $j$-th observation on the $i$-th treatment, $\mu$ is the common effect shared by all observations, $\tau_i$ is the effect of treatment $i$, and $e_{ij}$ is random error [2]. The errors are assumed normally and independently distributed with mean zero and common variance $\sigma^2$ [2].

The intuition: every observation deviates from the grand mean for two reasons. Either its group is genuinely shifted ($\tau_i$), or it is just noisy ($e_{ij}$). ANOVA asks whether the group shifts are big enough to stand out against the noise.

The test statistic comes from the ANOVA table. Treatment degrees of freedom are $k - 1$ (with $k$ groups), error degrees of freedom are $N - k$ (with $N$ total observations), and total degrees of freedom are $N - 1$. Mean squares are sums of squares divided by their degrees of freedom, and $F = \text{MST} / \text{MSE}$ [3].

Why not run t tests on every pair? Because each test carries its own error rate, and they compound. With $k$ comparisons at $\alpha = 0.05$, the familywise error rate is $1 - 0.95^k$ [8]. For three groups that is $1 - 0.95^3 = 0.143$, meaning roughly a 14% chance of at least one false positive even when nothing is going on. ANOVA controls this with a single omnibus test.

Two-Way ANOVA: Two Factors and an Interaction

With two factors, for example temperature and oven position, the analysis is two-way [1]. The model extends to:

$$Y_{ijk} = \mu + \tau_i + \beta_j + \gamma_{ij} + e_{ijk}$$

where $\tau_i$ and $\beta_j$ are the effects of the levels of factors A and B, and $\gamma_{ij}$ is the interaction effect between level $i$ of A and level $j$ of B [4].

Two-way ANOVA answers three questions at once [1]:

  1. Is there a difference in means across levels of factor A?
  2. Is there a difference in means across levels of factor B?
  3. Do A and B interact?

An interaction means the effect of one factor depends on the level of the other [4]. It is tested with $(a-1)(b-1)$ numerator degrees of freedom, where $a$ and $b$ are the numbers of levels of each factor [4]. In a balanced layout, where every cell has the same number of replicates, the sums of squares add up cleanly [4][5]:

$$\text{SS}_{\text{total}} = \text{SS}_A + \text{SS}_B + \text{SS}_{AB} + \text{SSE}$$

The degrees of freedom follow the same structure: $a-1$ for A, $b-1$ for B, $(a-1)(b-1)$ for the interaction, $N - ab$ for error, and $N - 1$ for total [4].

The payoff of adding a second factor is sensitivity. When multiple factors affect a system, allowing for interaction partitions variance between each factor and their interaction, which can sharpen the signal [9]. A factor that looked unimportant in a one-way analysis may become clearly significant once nuisance variation is moved out of the error term.

FeatureOne-way ANOVATwo-way ANOVA
Number of factors12
Hypotheses testedOne (equal means)Three (A, B, interaction)
Error df$N - k$$N - ab$
Interaction termNoneYes, $(a-1)(b-1)$ df
Typical useCompare 3+ groups on one variableTest two variables and whether they combine

Worked Example

This example uses a published dataset from the NIST e-Handbook: a coating tested on 3 materials (factor B) at 2 labs (factor A), with 3 replicates per cell. It is a balanced 2x3 design with $N = 18$.

The data. Lab 1: material 1 = 4.1, 3.9, 4.3; material 2 = 3.1, 2.8, 3.3; material 3 = 3.5, 3.2, 3.6. Lab 2: material 1 = 2.7, 3.1, 2.6; material 2 = 1.9, 2.2, 2.3; material 3 = 2.7, 2.3, 2.5.

Cell means. Lab 1: 4.100, 3.067, 3.433. Lab 2: 2.800, 2.133, 2.500.

Marginal means. Lab 1 = 3.533, lab 2 = 2.478. Material 1 = 3.450, material 2 = 2.600, material 3 = 2.967. Grand mean = 3.006.

Two-way ANOVA table.

SourceSSdfMSFp
Lab (A)5.013915.0139100.283.5e-7
Material (B)2.181121.090621.810.00010
Lab x Material0.134420.06721.340.297
Residual0.6000120.0500
Total7.929417

Because the design is balanced, Type I, Type II and Type III sums of squares are identical here. The reported result:

Lab, $F(1, 12) = 100.28$, $p < .001$, partial $\eta^2 = .89$; material, $F(2, 12) = 21.81$, $p < .001$, partial $\eta^2 = .78$; interaction, $F(2, 12) = 1.34$, $p = .30$.

The interaction is not significant, so the main effects are interpretable directly. Labs differ substantially, and materials differ. Partial eta squared (effect SS divided by effect SS plus error SS) gives .89 for lab, .78 for material and .18 for the interaction, with a model $R^2$ of .924.

Now the one-way comparison. Run a one-way ANOVA on material alone, ignoring lab (you can reproduce this with the ANOVA Calculator). You get $F(2, 15) = 2.85$, $p = 0.090$, not significant. Tukey HSD finds no pairwise material difference (smallest adjusted $p = 0.075$ for material 1 vs 2). The same data that showed a strong material effect in the two-way model now shows nothing.

Why? The lab effect got pooled into the error term. Residual SS jumps from 0.600 in the two-way model to 5.748 in the one-way model. The material signal did not change; the noise floor did.

Post hoc after the two-way model. Using MSE = 0.050 with 12 df, Tukey HSD is $q(0.95; 3, 12) \times \sqrt{0.05/6} = 3.773 \times 0.0913 = 0.344$. All three material differences (0.850, 0.483, 0.367) exceed 0.344, with Tukey $p$ values of 0.00007, 0.0073 and 0.037.

Assumption checks. Shapiro-Wilk on the two-way residuals gives $W = 0.913$, $p = 0.098$. Levene across the 6 cells gives $p = 0.999$. Neither flags a problem.

A case with a real interaction. Consider a second, illustrative dataset: treatment (control vs drug) crossed with genotype (WT, Het, KO), 3 per cell. Cell means are control 10.17 / 10.20 / 10.13 and drug 14.07 / 12.30 / 10.43. The two-way ANOVA gives treatment $F(1, 12) = 121.09$, $p = 1.3e\text{-}7$; genotype $F(2, 12) = 30.79$, $p = 1.9e\text{-}5$; interaction $F(2, 12) = 29.65$, $p = 2.3e\text{-}5$ (SS 19.845, 10.093, 9.720; residual SS 1.967).

The drug raises the mean by 3.90 in WT, 2.10 in Het and 0.30 in KO. The main effect of treatment is an average over genotypes, and that average describes none of them well. When the interaction is significant, interpret simple effects instead.

How to Report and Interpret Results

The convention is to report $F(df_{\text{effect}}, df_{\text{error}}) = \text{value}$, $p = \text{value}$, plus an effect size, for each main effect and the interaction. For the coating example: "material, $F(2, 12) = 21.81$, $p < .001$." Report the overall F test p value and the result of the post hoc comparison [8].

When ANOVA rejects equality, you know the means are not all equal but not which ones differ, so multiple comparison procedures are needed [6]. Repeating pairwise comparisons for all pairs does not work in general because the overall significance level is not what you specified [6]. Tukey's method tests all pairwise differences of means, Scheffe's method tests all possible contrasts, and Bonferroni is used for a pre-selected group of contrasts [6]. Bonferroni divides alpha by the number of comparisons ($0.05/k$) and can be conservative; Tukey HSD, Scheffe and Duncan are frequently preferred [8]. The choice depends on your question, and no single method is universally best.

A quick note on the two-group case: with exactly two groups, one-way ANOVA and the pooled two-sample t test give the same p value, with $F = t^2$. In this dataset, a one-way ANOVA on lab alone gives $F(1, 16) = 27.52$, $p = 0.00008$, and the pooled t test gives $t = 5.245$ with $t^2 = 27.52$. Identical.

Common Mistakes

  • Running multiple t tests instead of ANOVA. Each test inflates the familywise error rate; with three groups it reaches 0.143 [8]. Use ANOVA first, then a controlled post hoc method.
  • Ignoring a significant interaction and reading main effects. Main effects are averages over the other factor's levels and can point the wrong way when effects differ by group. Interpret simple effects instead.
  • Dropping a factor that matters. As the worked example shows, ignoring lab turned a significant material effect ($p = 0.00010$) into a non-significant one ($p = 0.090$) by inflating the error term.
  • Checking normality on raw data instead of residuals. The normality assumption applies to the random errors, which you check through residuals, for example with a normal probability plot that should lie close to a straight line [10].
  • Treating a non-significant interaction as proof of no interaction. A small F with few replicates may simply reflect low power. Report the effect size and the confidence interval alongside the p value.
  • Forgetting that the factor is categorical and the response is numerical. One-way ANOVA requires a categorical factor and a numeric response [7].

Limitations

In unbalanced designs, Type I, Type II and Type III sums of squares differ, and which to use is debated. The examples here are balanced, so all three agree. If your design is unbalanced, check the current documentation for your software and be explicit about which type you used.

The second dataset with the interaction is invented for illustration. Only the NIST coating dataset is published.

ANOVA is generally described as fairly tolerant of mild non-normality in balanced designs, but that tolerance has limits. Do not assume it covers severe skew or outliers.

Post hoc choice is genuinely open. NIST and Kim describe the options but do not prescribe one method as best [6][8]. Pick based on whether you need all pairwise comparisons, all contrasts, or a pre-planned set.

The $F(df_1, df_2) = x$, $p = y$ format is a widespread convention, not a rule quoted from any single source. Follow your target journal's style guide.

Frequently Asked Questions

What is the difference between one-way and two-way ANOVA?

One-way ANOVA tests a single factor with two or more levels, usually three or more [1]. Two-way ANOVA tests two factors simultaneously and adds an interaction term that captures whether the effect of one factor changes across levels of the other [4]. The two-way model also partitions variance more finely, which often increases power.

When should I use two-way ANOVA instead of one-way?

Use two-way ANOVA when you have two categorical factors and you care about both, or when you suspect they interact. It is also worth using when a second factor is a nuisance variable you want to remove from the error term, as the lab factor did in the worked example. If you only have one factor, one-way is correct.

What does a two-way ANOVA interaction mean?

An interaction means the effect of one factor depends on the level of the other [4]. In the illustrative dataset, the drug raised the outcome by 3.90 in wild-type but only 0.30 in knockout. That is an interaction. When it is significant, report simple effects instead of main effects.

Is one-way ANOVA the same as a t test?

With exactly two groups, yes in result: the p values match and $F = t^2$. With three or more groups, no. ANOVA tests all means simultaneously and controls the familywise error rate, while multiple t tests do not [8].

How do I report a two-way ANOVA?

Report $F(df_{\text{effect}}, df_{\text{error}}) = \text{value}$, $p = \text{value}$ for each main effect and the interaction, plus an effect size. For example: "material, $F(2, 12) = 21.81$, $p < .001$." Then report the post hoc result for any factor that reached significance [8].

References

  1. NIST/SEMATECH e-Handbook 7.4.3 Are the means equal?
  2. NIST/SEMATECH e-Handbook 7.4.3.2 The one-way ANOVA model and assumptions
  3. NIST/SEMATECH e-Handbook 7.4.3.3 The ANOVA table and tests of hypotheses about means
  4. NIST/SEMATECH e-Handbook 7.4.3.7 The two-way ANOVA
  5. NIST/SEMATECH e-Handbook 7.4.3.8 Models and calculations for the two-way ANOVA
  6. NIST/SEMATECH e-Handbook 7.4.7 How can we make multiple comparisons?
  7. OpenStax Introductory Statistics 2e, 13.1 One-Way ANOVA
  8. Kim 2014, Analysis of variance (ANOVA) comparing means of more than two groups, Restor Dent Endod 39:74
  9. Krzywinski & Altman 2014, Two-factor designs, Nature Methods 11:1187-1188
  10. NIST/SEMATECH e-Handbook 4.4.4.5 How can I test whether the random errors are distributed normally?

Related Articles