T-Test vs Z-Test: Which One to Use (With Examples)
By Dr. Zubair Khalid, DVM, MS, PhD ·

The t-test and the z-test both answer a version of the same question: is the difference between what you observed and what you expected larger than chance alone would explain? Both produce a test statistic, both get compared against a probability distribution, and both return a p-value. The difference between them comes down to one practical detail: whether you truly know the population standard deviation, or whether you had to estimate it from your own data.
You will meet this choice constantly. Comparing a lab batch mean against a reference value, checking whether a treatment group differs from a control, testing whether a response rate beats a historical benchmark, or reading a methods section that says "two-sample t-test" when the authors might have meant something else. Getting the distinction right changes your p-value, your confidence interval, and sometimes your conclusion.
Quick Answer
- Use a t-test when your outcome is numeric, you are testing one or two means, and the population standard deviation (sigma) is unknown and estimated from the sample as s [1][7].
- Use a z-test when sigma is genuinely known from long-run process data or a large reference population, which is uncommon in laboratory and clinical research [1][7].
- The t statistic is $t = (\bar{Y} - \mu_0) / (s / \sqrt{N})$, where $\bar{Y}$ is the sample mean, $\mu_0$ is the hypothesized mean, $s$ is the sample standard deviation, and $N$ is the sample size. It is referred to the t distribution with $N - 1$ degrees of freedom [1][6].
- The z statistic replaces $s$ with the known sigma, $z = (\bar{Y} - \mu_0) / (\sigma / \sqrt{N})$, and is referred to the standard normal distribution [1].
- Modern practice is to use the t distribution whenever s estimates sigma, regardless of sample size [6]. The old "use z when n > 30" rule is a historical shortcut, and the two give nearly identical answers at large n anyway.
- For proportions (yes/no outcomes), the z-test for proportions is standard, and for two groups it is mathematically equivalent to a chi-square test of independence on the 2x2 table [3][8].
What Each Test Actually Does
Both tests start from the same logic. You have a sample mean, you have a hypothesized value, and you want to know how far apart they are in units of standard error. The standard error measures how much a sample mean would bounce around from study to study if you repeated the experiment.
The z-test assumes you know sigma exactly. That means the standard error $\sigma / \sqrt{N}$ is a fixed, known quantity, and the only randomness in the statistic comes from $\bar{Y}$. Under that assumption, the statistic follows a standard normal distribution.
The t-test admits that you do not know sigma. You estimate it with s, and that estimate carries its own uncertainty. With a small sample, s can be too small or too large by a meaningful margin, so the ratio $(\bar{Y} - \mu_0) / (s / \sqrt{N})$ is more variable than a standard normal. William S. Gosset, working at the Guinness brewery in Dublin and publishing under the pen name "Student," worked out the exact distribution of this ratio in 1908 [6][9]. That distribution is the t distribution.
The t distribution has more probability in its tails and less in its center than the standard normal [6]. This is the mathematical price of not knowing sigma. As the sample grows, s becomes a better estimate of sigma, the t distribution tightens up, and it converges on the normal. NIST notes the approximation is quite good for degrees of freedom above 30 [5]. The degrees of freedom are $N - 1$ because the $N$ deviations from the sample mean sum to zero, so only $N - 1$ of them are free to vary [6].
How to Calculate Each Statistic
For a one-sample test of a mean, the two formulas differ only in the denominator:
$$t = \frac{\bar{Y} - \mu_0}{s / \sqrt{N}}, \quad \text{df} = N - 1$$
$$z = \frac{\bar{Y} - \mu_0}{\sigma / \sqrt{N}}$$
Here $\bar{Y}$ is the sample mean, $\mu_0$ is the value you are testing against, $s$ is the sample standard deviation, $\sigma$ is the known population standard deviation, and $N$ is the sample size. The numerator is the same in both. The denominator is where the choice lives.
For two independent samples, the t-test has two common forms. Welch's version, which does not assume equal variances, is:
$$T = \frac{\bar{Y}_1 - \bar{Y}_2}{\sqrt{s_1^2/N_1 + s_2^2/N_2}}$$
with Welch-Satterthwaite degrees of freedom
$$\nu = \frac{(s_1^2/N_1 + s_2^2/N_2)^2}{(s_1^2/N_1)^2/(N_1 - 1) + (s_2^2/N_2)^2/(N_2 - 1)}$$
The equal-variance version pools the standard deviations:
$$T = \frac{\bar{Y}_1 - \bar{Y}_2}{s_p \sqrt{1/N_1 + 1/N_2}}, \quad \text{df} = N_1 + N_2 - 2$$
where $s_p$ is the pooled standard deviation [2]. Welch's test is the adaptation for unequal variances [10], and many analysts use it as the default.
For paired data, where each subject contributes a before and after measurement, compute the differences $d_i = Y_i - Z_i$, then run a one-sample t-test on those differences:
$$t = \frac{\bar{d}}{s_d / \sqrt{N}}, \quad \text{df} = N - 1$$
Pairing removes variability between units and increases power [4].
For a single proportion, the z-test uses:
$$z = \frac{\hat{p} - p_0}{\sqrt{p_0(1 - p_0)/N}}$$
where $\hat{p}$ is the observed proportion, $p_0$ is the hypothesized proportion, and $N$ is the sample size [3]. The normal approximation requires large N, with NIST specifying $N > 30$ and $\min\{Np_0, N(1-p_0)\} \geq 5$ [3], while OpenStax uses $np > 5$ and $nq > 5$ [7].
How to Read the Output
The test statistic itself is not the answer. It is an intermediate number that you convert into a p-value or a confidence interval. A t of 2.092 with 9 degrees of freedom and a z of 2.092 mean different things because they sit at different positions in their respective distributions.
The p-value tells you the probability of seeing a statistic at least as extreme as yours if the null hypothesis were true. It does not tell you the probability that the null is true, and it does not measure effect size. A tiny p-value from a huge sample can accompany a trivial difference.
Confidence intervals are often more informative. A 95% t interval uses a larger critical value than a z interval, so it is wider. With $N = 10$, the t critical value is 2.262 versus 1.960 for z, making the interval about 15% wider. That extra width is the honest cost of estimating sigma from ten observations.
Sign conventions matter. If $\bar{Y} > \mu_0$, the statistic is positive. If $\bar{Y} < \mu_0$, it is negative. For a two-sided test, you compare the absolute value against the critical value. For a one-sided test, the sign tells you which tail you are in, and you must decide the direction before looking at the data.
Worked Example
Take a sample of 10 measurements: 96, 102, 91, 99, 105, 94, 98, 101, 93, 100. The mean is 97.9, the sample standard deviation is $s = 4.383$, and the standard error is $SE = s/\sqrt{10} = 1.386$. Suppose you want to test $H_0: \mu = 95$ against a two-sided alternative.
Correct approach (t-test, sigma unknown). Since sigma is estimated from the data, use the t distribution:
$$t = \frac{97.9 - 95}{1.386} = 2.092, \quad \text{df} = 9$$
The critical t value at the 5% level is 2.262. Since 2.092 falls short, the result is not significant. The p-value is 0.066, and the 95% confidence interval runs from 94.76 to 101.04, which includes 95.
Incorrect approach (z-test using s). If you plug s into the z formula, you get the same statistic, 2.092, but the normal distribution gives p = 0.036. That would falsely claim significance at the 0.05 level. The error is ignoring the extra uncertainty in s when N is only 10.
Legitimate z-test (sigma genuinely known). If sigma were known to be 4.0 from long-run process data, then:
$$z = \frac{97.9 - 95}{4/\sqrt{10}} = 2.293, \quad p = 0.022$$
The 95% confidence interval would be 95.42 to 100.38. If sigma were known to be 6.0, then z = 1.528 and p = 0.126, not significant.
Convergence. The t critical value at the 97.5th percentile is 2.262 at df 9, 2.045 at df 29, 2.000 at df 60, 1.980 at df 120, and 1.960 for the normal. The gap closes quickly.
Proportions. Suppose 58 of 100 patients respond, and you test $H_0: p = 0.5$. Then:
$$z = \frac{0.58 - 0.5}{\sqrt{0.25/100}} = 1.600, \quad p = 0.110$$
The exact binomial test gives p = 0.133. The normal approximation is close but not identical.
Two proportions and chi-square. Comparing 30/50 versus 18/50 gives a pooled proportion of 0.48, z = 2.402, and p = 0.0163. A chi-square test of independence on the table [[30, 20], [18, 32]] without continuity correction gives chi-square = 5.769, which equals $z^2$, with df 1 and the same p = 0.0163. The two tests are equivalent in this setting.
If you want to run these calculations yourself, try the site's T-Test Calculator.
T-Test vs Z-Test vs Chi-Square vs ANOVA
| Situation | Outcome type | Test | Why |
|---|---|---|---|
| One mean, sigma unknown | Numeric | One-sample t-test | s estimates sigma [1] |
| One mean, sigma known | Numeric | z-test | sigma is fixed [1][7] |
| Two independent means | Numeric | Two-sample t-test (Welch or pooled) | sigma estimated from each group [2] |
| Paired before/after | Numeric | Paired t-test | Differences remove between-unit variability [4] |
| One proportion vs reference | Categorical | z-test for proportions | Normal approximation to binomial [3] |
| Two proportions | Categorical | z-test or chi-square | Equivalent on 2x2 without correction |
| Two categorical variables | Categorical | Chi-square test of independence | Tests association in a contingency table [8] |
| Three or more means | Numeric | One-way ANOVA | Multiple t-tests inflate familywise error [10] |
The chi-square test of independence uses $\chi^2 = \sum (O - E)^2 / E$ with df = (rows - 1)(columns - 1) and requires every expected count to be at least 5 [8]. For a 2x2 table, it gives the same p-value as a two-proportion z-test when no continuity correction is applied. Many software packages apply Yates' correction by default, which breaks the exact equivalence.
The ANOVA point deserves emphasis. Running k separate t-tests at alpha = 0.05 gives a familywise error rate of $1 - 0.95^k$ [10]. With three groups, that is about 14%. ANOVA controls this, then post hoc tests identify which groups differ.
Common Mistakes
- Using z when sigma is estimated. This is the most frequent error. Plugging s into the z formula ignores the uncertainty in s and produces p-values that are too small. With N = 10, the difference between p = 0.066 and p = 0.036 can flip a conclusion. Use t whenever s estimates sigma [6].
- Applying the "n > 30" rule mechanically. Some textbooks still teach this. Current practice is to use t regardless of sample size when s estimates sigma, and the difference is negligible at large n anyway [6].
- Ignoring the equal-variance assumption. The pooled two-sample t-test assumes both groups share a common variance. When variances differ, Welch's test is the appropriate choice [10].
- Treating paired data as independent. Before/after measurements on the same subject are correlated. Analyzing them as two independent groups discards the pairing and loses power [4].
- Using a two-proportion z-test when expected counts are small. The normal approximation requires adequate counts [3][7]. When expected counts fall below 5, use Fisher's exact test or an exact binomial approach.
- Running multiple t-tests instead of ANOVA. Comparing three or more group means with repeated t-tests inflates the chance of a false positive [10].
- Confusing statistical significance with practical importance. A small p-value does not mean the effect is large or meaningful.
Limitations
The t-test assumes the data are a simple random sample from an approximately normal population when the sample is small [7]. With larger samples, the central limit theorem makes the test fairly tolerant of non-normality for the mean, but severe skew or outliers can still distort results.
The z-test for a mean requires sigma to be genuinely known. In practice, sigma is almost never known in laboratory or clinical research. Historical process data or a very large reference population can sometimes supply a defensible value, but this is the exception.
The normal approximation for proportions has no single agreed cutoff. NIST uses $\min\{Np_0, N(1-p_0)\} \geq 5$ with $N > 30$ [3], OpenStax uses $np > 5$ and $nq > 5$ [7], and some sources use 10. Check the current documentation for whichever software you use.
The equivalence between the two-proportion z-test and the chi-square test of independence holds only without Yates' continuity correction. Many software defaults apply the correction to 2x2 tables, so the p-values will differ slightly.
The worked example uses invented data, and the "sigma known" scenarios are hypothetical. They illustrate the arithmetic, not a real study.
Frequently Asked Questions
When should I use a t-test instead of a z-test?
Use a t-test whenever the population standard deviation is unknown and you are estimating it from your sample with s. This covers nearly all real research situations [1][6]. Use a z-test only when sigma is genuinely known from external long-run data, which is rare in lab and clinical work [7].
Does the t-test vs z-test choice matter for large samples?
Barely. At df = 120, the t critical value is 1.980 versus 1.960 for z, a difference of about 1%. At df = 1000, it is 1.962. The practical guidance is still to use t whenever s estimates sigma, because it is correct at every sample size and costs nothing [6].
What is the difference between the t distribution and the normal distribution?
The t distribution has thicker tails and a shorter center than the standard normal [6]. This reflects the extra uncertainty from estimating sigma. As degrees of freedom increase, the t distribution converges to the normal, and NIST notes the approximation is quite good above df = 30 [5]. The standard deviation of the t distribution is $\sqrt{\nu/(\nu - 2)}$, which is undefined for $\nu = 1$ or 2 [5].
When do I use a z-test for proportions?
Use it when your outcome is a yes/no proportion, you have a large sample, and the expected counts meet the normal approximation conditions [3][7]. For a single proportion compared with a reference value, the formula is $z = (\hat{p} - p_0)/\sqrt{p_0(1-p_0)/N}$. For two proportions, the z-test is equivalent to a chi-square test of independence on the 2x2 table when no continuity correction is applied.
How does chi-square test vs t-test differ?
They handle different data types. A t-test compares means of numeric outcomes. A chi-square test of independence assesses whether two categorical variables are related in a contingency table [8]. For a 2x2 table, the chi-square statistic equals the square of the two-proportion z statistic, and the p-values match without continuity correction.
References
- NIST/SEMATECH e-Handbook 7.2.2 Are the data consistent with the assumed process mean?
- NIST/SEMATECH e-Handbook 1.3.5.3 Two-Sample t-Test for Equal Means
- NIST/SEMATECH e-Handbook 7.2.4 Does the proportion of defectives meet requirements?
- NIST/SEMATECH e-Handbook 7.3.1.1 Analysis of paired observations
- NIST/SEMATECH e-Handbook 1.3.6.6.4 t Distribution
- OpenStax Introductory Statistics 2e, 8.2 A Single Population Mean using the Student t Distribution
- OpenStax Introductory Statistics 2e, 9.3 Probability Distribution Needed for Hypothesis Testing
- OpenStax Introductory Statistics 2e, 11.3 Test of Independence
- Student 1908, The Probable Error of a Mean, Biometrika 6(1)
- Kim 2014, Analysis of variance (ANOVA) comparing means of more than two groups, Restor Dent Endod 39:74
Related Articles
- Welch's t-Test: When and How to Use It
- ANOVA vs. t-Test: Choosing the Right Statistical Test
- Chi-Square Test of Independence in Genetics
- P-Value Formula and Interpretation: A Researcher's Guide
- One-Way vs Two-Way ANOVA: Differences, Assumptions and a Worked Example
- Standard Deviation vs Variance vs Standard Error: What Each Measures and When to Report It