# Two Sample t-Test: Formula, Calculation and Example

The 2 sample t-test compares the means of two independent groups to decide whether the difference you see is real or just sampling noise. You compute a t statistic from the two means, their variances and the sample sizes, then compare it to a t distribution. This article gives you the formula, a full hand calculation, and the interpretation.

## Quick Answer

- The 2 sample t-test tests whether two population means are equal, using one sample from each group [1].
- The test statistic is the difference in sample means divided by the standard error of that difference.
- If you assume equal variances, you pool the two sample variances and use $df = n_1 + n_2 - 2$ [1][2].
- If you do not assume equal variances, use the Welch version, which estimates each variance separately and adjusts the degrees of freedom [3].
- A small p-value (below your chosen alpha, often 0.05) means the group means differ more than chance would explain.

## The Formula

For independent samples with equal variances assumed, the test statistic is:

$$t = \frac{\bar{x}_1 - \bar{x}_2}{s_p \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}}$$

The pooled standard deviation $s_p$ comes from the pooled variance:

$$s_p^2 = \frac{(n_1 - 1)s_1^2 + (n_2 - 1)s_2^2}{n_1 + n_2 - 2}$$

Each symbol means:

| Symbol | Meaning |
|---|---|
| $\bar{x}_1, \bar{x}_2$ | Sample means of group 1 and group 2 |
| $s_1^2, s_2^2$ | Sample variances of group 1 and group 2 |
| $n_1, n_2$ | Sample sizes of group 1 and group 2 |
| $s_p$ | Pooled standard deviation |
| $s_p\sqrt{1/n_1 + 1/n_2}$ | Standard error of the difference in means |
| $df$ | Degrees of freedom, $n_1 + n_2 - 2$ for the pooled version |

The pooled variance weights each group's variance by its degrees of freedom, so the larger sample contributes more to the estimate [2]. The numerator measures how far apart the means are. The denominator measures how much they would bounce around by chance alone. A large ratio means a real difference.

## How to Calculate It Step by Step

1. State the hypotheses. The null hypothesis is $H_0: \mu_1 = \mu_2$. The two-sided alternative is $H_a: \mu_1 \neq \mu_2$ [1].
2. Compute both sample means by summing each group and dividing by its size.
3. Compute both sample variances using the $n-1$ denominator.
4. Compute the pooled variance with the formula above.
5. Take the square root to get the pooled standard deviation $s_p$.
6. Compute the standard error: $s_p \sqrt{1/n_1 + 1/n_2}$.
7. Divide the mean difference by the standard error to get $t$.
8. Find the degrees of freedom, $n_1 + n_2 - 2$.
9. Look up the two-tailed p-value for your $t$ and $df$, or compare $|t|$ to the critical value.

If you want a tool to skip the arithmetic, the [T-Test Calculator](/tools/t-test-calculator) returns the statistic and p-value directly.

## Worked Example

Two teaching methods were used with 8 students each, and every student took the same exam out of 100. The question is whether the two methods produce different average scores.

| Student | Method A | Method B |
|---|---|---|
| 1 | 78 | 72 |
| 2 | 85 | 79 |
| 3 | 92 | 85 |
| 4 | 88 | 80 |
| 5 | 76 | 70 |
| 6 | 90 | 83 |
| 7 | 84 | 77 |
| 8 | 81 | 74 |

Group A scores: 78, 85, 92, 88, 76, 90, 84, 81. Group B scores: 72, 79, 85, 80, 70, 83, 77, 74.

**Step 1. Means.** Group A sums to 674, so the mean is $674/8 = 84.2500$. Group B sums to 620, so the mean is $620/8 = 77.5000$.

**Step 2. Variances.** Using the $n-1$ denominator, variance A is 32.2143 and variance B is 27.7143. If you need a refresher on that step, see [sample variance equation](/blog/data-analysis/sample-variance-equation-formula-examples).

**Step 3. Pooled variance.**

$$s_p^2 = \frac{(8-1)(32.2143) + (8-1)(27.7143)}{8 + 8 - 2} = 29.9643$$

**Step 4. Pooled standard deviation.** $s_p = \sqrt{29.9643} = 5.4740$.

**Step 5. Standard error.** $5.4740 \times \sqrt{1/8 + 1/8} = 2.7370$.

**Step 6. t statistic.** $(84.2500 - 77.5000) / 2.7370 = 2.4662$.

**Step 7. Degrees of freedom.** $8 + 8 - 2 = 14$.

**Step 8. p-value.** The two-tailed p-value is $2 \times P(T > |2.4662|)$ with 14 degrees of freedom, which is 0.0272.

In Python:

```python
from scipy import stats
group_a = [78, 85, 92, 88, 76, 90, 84, 81]
group_b = [72, 79, 85, 80, 70, 83, 77, 74]
t, p = stats.ttest_ind(group_a, group_b, equal_var=True)
```

Output:

```
t = 2.4662, p = 0.0272
```

The dot plot below shows the two score distributions with their means marked.

## How to Interpret the Result

The t statistic of 2.4662 means the observed mean difference is about 2.5 standard errors away from zero. With 14 degrees of freedom, the two-tailed p-value is 0.0272. If the two population means were truly equal, you would see a difference this large or larger about 2.7 percent of the time.

At the 0.05 significance level, you reject the null hypothesis and conclude the two population means are different [1]. Method A's average is higher by 6.75 points.

The p-value tells you how incompatible the data are with equal means. It does not tell you how big the difference is or whether it matters in practice. Report the mean difference and a confidence interval alongside the p-value so readers can judge the size of the effect. The same logic applies to other tests in the inference pillar, such as the [two proportion z-test](/blog/data-analysis/two-proportion-z-test-formula) for comparing proportions.

## Doing It in Software (Excel, R or Python)

**Excel.** Use `T.TEST(array1, array2, tails, type)`. Set `tails` to 2 for a two-sided test and `type` to 2 for two independent samples with equal variances, or 3 for the unequal-variance version. The function returns the p-value only, so compute the t statistic separately if you need it.

**R.** Use `t.test(x, y)`. By default, `var.equal` is `FALSE`, so R estimates the variance separately for each group and applies the Welch modification to the degrees of freedom [3]. Set `var.equal = TRUE` to get the pooled-variance version shown above.

**Python.** Use `scipy.stats.ttest_ind`. The `equal_var` argument defaults to `True`, which gives the pooled-variance test. Set `equal_var=False` for Welch's version.

The choice between pooled and Welch matters when the two groups have very different variances or very different sample sizes. When in doubt, Welch's version is the safer default because it does not assume equal variances [3].

## Common Mistakes

- **Using the paired test on independent groups.** If the two samples have a one-to-one correspondence, such as before-and-after measurements on the same people, the data are paired and the formulas are different [1]. Use a paired t-test instead.
- **Forgetting to check the equal-variance assumption.** The pooled formula assumes both populations share a variance. If the group variances differ a lot, use Welch's version [3].
- **Reporting the p-value without the mean difference.** A significant result says the means differ, not that the difference is large. Always report both means and their difference.
- **Confusing one-tailed and two-tailed p-values.** A two-tailed p-value is roughly double the one-tailed value for the same data. Decide which you need before you look at the output.
- **Treating a non-significant result as proof of equality.** Failing to reject the null hypothesis means the data are consistent with equal means. It does not prove the means are the same.
- **Ignoring the sample size.** With very small samples, the test has low power and will miss real differences. With very large samples, it will flag tiny differences as significant.

## Limitations

The 2 sample t-test assumes independent observations, roughly normal data within each group, and, in the pooled version, equal population variances [1]. It is sensitive to outliers, especially in small samples, because the mean and variance both react strongly to extreme values. If your data are heavily skewed or contain outliers, a rank-based alternative such as the [Mann-Whitney U test](/blog/data-analysis/mann-whitney-u-test-definition-formula-example) may be more appropriate.

The test also says nothing about causality. A significant difference between two groups can come from the treatment, from a confounding variable, or from how the groups were selected. Random assignment is what supports a causal claim, not the t-test itself. For more than two groups, use ANOVA instead of running repeated t-tests, which inflates the chance of a false positive.

## Frequently Asked Questions

### What is the difference between a one-sample and a two-sample t-test?

A one-sample t-test compares a single sample mean to a fixed value. A two-sample t-test compares the means of two independent groups to each other. The formulas differ in the numerator and in how the standard error is built. See the [one-sample t-test guide](/blog/guides/one-sample-t-test-formula-calculation-and-interpretation) for the single-group case.

### When should I use the pooled versus the Welch version?

Use the pooled version when the two groups have similar variances and similar sample sizes. Use Welch's version when variances differ or sample sizes are very unequal [3]. Many statisticians use Welch by default because it performs well even when variances are equal.

### What does a negative t statistic mean?

A negative t statistic means the first group's mean is lower than the second group's mean. The sign depends on the order you entered the groups. For a two-tailed test, only the absolute value matters for the p-value.

### Can I use the 2 sample t-test on small samples?

Yes, but the test assumes the data are approximately normal, and that assumption matters more when samples are small. With fewer than about 30 observations per group, check for outliers and skew before trusting the result. If normality is doubtful, use a nonparametric test.

### How do I report the result?

Report the two means, the mean difference, the t statistic, the degrees of freedom and the p-value. For the example above: method A averaged 84.25 and method B averaged 77.50, a difference of 6.75 points, $t(14) = 2.47$, $p = 0.027$. Include a confidence interval for the difference when you can.

## References

1. [1.3.5.3. Two-Sample <i>t</i>-Test for Equal Means](https://www.itl.nist.gov/div898/handbook/eda/section3/eda353.htm)
2. [](https://people.umass.edu/bwdillon/files/linguist-609-2020/Notes/TwoSampleT-Test.html)
3. [R: Student's t-Test](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/t.test.html)

## Further Reading

- [1.3.5.3.1. Data Used for Two-Sample <i>t</i>-Test](https://www.itl.nist.gov/div898/handbook/eda/section3/eda3531.htm)
- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [Wasserstein RL, Lazar NA (2016). The ASA Statement on p -Values: Context, Process, and Purpose. The American Statistician](https://doi.org/10.1080/00031305.2016.1154108)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)

## Related Articles

- [Two Proportion Z-Test: Formula and Worked Example](/blog/data-analysis/two-proportion-z-test-formula)
- [t Statistic Formula: Definition, Calculation and Examples](/blog/data-analysis/t-statistic-formula)
- [Student's t-Distribution: Definition, Formula and Examples](/blog/data-analysis/students-t-distribution-definition-formula)
- [What Is a Chi-Square Test? Formula and Examples](/blog/data-analysis/chi-square-test-formula-examples)
- [Sample Mean: Definition, Formula and Examples](/blog/data-analysis/sample-mean)
- [One-Sample t-Test: Formula, Calculation, and Interpretation](/blog/guides/one-sample-t-test-formula-calculation-and-interpretation)
- [Understanding the t-Test: Meaning, Assumptions, and Applications](/blog/guides/understanding-the-t-test-meaning-assumptions-and-applications)
- [Test Statistic Formula: How to Calculate and Use It](/blog/guides/test-statistic-formula-how-to-calculate-and-use-it)