# What Is a Chi-Square Test? Formula and Examples

A chi square test is a statistical test for categorical data. It compares the counts you actually observed with the counts you would expect if a hypothesis were true, then tells you whether the gap is larger than chance would explain. This article covers the formula, when to use it, and a full worked example.

## Quick Answer

- A chi square test works on counts of categories, not on means or continuous measurements.
- The test statistic is $\chi^2 = \sum \frac{(O - E)^2}{E}$, where $O$ is observed and $E$ is expected [1].
- The two main versions are the goodness of fit test (one variable) and the test of independence (two variables in a contingency table) [1][2].
- Degrees of freedom come from the table size, not the sample size. For independence, $df = (r-1)(c-1)$ [3].
- A large chi square value means observed and expected counts are far apart, so the model fits poorly [4].

## What the Chi Square Test Means

In plain terms, the chi square test asks one question: are the numbers I counted close to the numbers I would expect? If they are close, nothing unusual is happening. If they are far apart, something in the data needs explaining.

The precise definition is narrower. The chi square test is a statistical test that determines whether observed frequencies differ significantly from expected frequencies [5]. You state a null hypothesis of no difference and an alternative hypothesis of a real difference, then reject or fail to reject the null based on the result [5]. The test statistic follows a chi square distribution, which is right skewed for small degrees of freedom and becomes more symmetric as degrees of freedom increase [3]. Because the formula squares each difference, every test statistic is zero or positive, and the test is always right tailed [3].

## How It Works

The test statistic is built from one formula applied to every cell:

$$\chi^2 = \sum \frac{(O - E)^2}{E}$$

Each symbol means the following:

- $\chi^2$ is the chi square test statistic, sometimes written as the chi statistic or chi value.
- $O$ is the observed frequency, the count you actually recorded [3].
- $E$ is the expected frequency, the count predicted under the null hypothesis [3].
- $\sum$ means sum across all cells in the table.

For each cell you subtract the expected count from the observed count, square the difference, and divide by the expected count. Then you add those values across all cells [5]. A large sum means the observed and expected values are not close and the model is a poor fit [4].

Expected counts depend on the test. In a goodness of fit test, expected counts come from a theoretical distribution [1]. In a test of independence, the expected count for a cell is the row total times the column total divided by the grand total, because that is what independence implies [4].

Degrees of freedom also depend on the test. For a two-way table with $r$ rows and $c$ columns, the degrees of freedom are:

$$df = (r-1)(c-1)$$

This follows from the table having $rc$ cells with one determined by the others, and from estimating $(r-1) + (c-1)$ marginal parameters under independence [4][3].

## Worked Example

A survey of 100 respondents recorded gender and product preference. The data are counts of people in each combination.

| Gender | Preference | Count |
|---|---|---|
| Male | Product A | 30 |
| Male | Product B | 20 |
| Female | Product A | 15 |
| Female | Product B | 35 |

Arranged as a 2x2 table with rows for gender and columns for preference, the observed counts are:

| | Product A | Product B | Row total |
|---|---|---|---|
| Male | 30 | 20 | 50 |
| Female | 15 | 35 | 50 |
| Column total | 45 | 55 | 100 |

The row totals are 50 and 50, the column totals are 45 and 55, and the grand total is 100.

Expected counts use row total times column total divided by the grand total. Every cell works out the same way here because both row totals are 50:

| | Product A | Product B |
|---|---|---|
| Male | 22.5000 | 27.5000 |
| Female | 22.5000 | 27.5000 |

Now compute each cell contribution:

- Male, Product A: $(30 - 22.5000)^2 / 22.5000 = 2.5000$
- Male, Product B: $(20 - 27.5000)^2 / 27.5000 = 2.0455$
- Female, Product A: $(15 - 22.5000)^2 / 22.5000 = 2.5000$
- Female, Product B: $(35 - 27.5000)^2 / 27.5000 = 2.0455$

Summing the four contributions gives a chi square statistic of 9.0909. The degrees of freedom are $(2-1)(2-1) = 1$. The p-value is 0.0026.

You can reproduce this in Python:

```python
import numpy as np
from scipy.stats import chi2_contingency
obs = np.array([[30, 20], [15, 35]])
chi2, p, dof, exp = chi2_contingency(obs, correction=False)
print(chi2, p, dof, exp)
```

Output:

```
9.09090909090909 0.0025688315270227164 1 [[22.5 27.5]
 [22.5 27.5]]
```

If you want to run your own numbers without writing code, the [chi square test calculator](/tools/chi-square-calculator) takes a contingency table and returns the statistic, degrees of freedom, and p-value.

## How to Interpret It

Start with the p-value. It answers a specific question: if the null hypothesis were true, how often would you see a chi square value at least this large? A small p-value means the observed counts are unlikely under the null.

In the example, the p-value is 0.0026. At a 0.05 significance level you reject the null hypothesis of independence. Gender and product preference appear to be associated in this sample.

Compare the statistic with a critical value if you prefer that route. The critical value for a chi square distribution with 1 degree of freedom at the 0.05 level is 3.841, so the calculated value of 9.0909 here exceeds it and leads to rejecting the null. You can look up these thresholds in a [chi square table of critical values](/blog/data-analysis/chi-square-table-critical-values).

A significant result tells you that an association exists. It does not tell you which cells drive it. For that you need post-hoc work such as [standardized residuals and partitioning](/knowledge/bioinformatics/post-hoc-tests-after-a-significant-chi-square-standardized-residuals-partitioning-and-pairwise-compa).

## When to Use It (and when not to)

Use a chi square test when your data are counts of categories and you want to compare those counts with expectations. The three standard uses are comparing observed and expected frequency distributions of a nominal variable, testing whether two variables are independent, and using the chi square distribution in work on correlation coefficients [2].

Typical research questions look like this. Does the distribution of eye colors at a university differ from the general population? Do kindergarten students prefer ice cream flavors equally? [6] Both are count comparisons against an expectation.

Do not use a chi square test for continuous measurements such as height or reaction time. Group continuous data into intervals first if you must, though the choice of intervals affects the power of the test [7]. Do not use it when expected counts are very small. A common rule of thumb requires an expected count of at least five in every cell [7]. When a 2x2 table has small expected counts, a [Fisher exact test](/knowledge/bioinformatics/chi-square-test-of-independence-vs-fisher-s-exact-test-a-decision-framework-for-small-sample-sizes-i) is usually the better choice.

## Chi Square Test vs t-Test

The closest related idea is the t-test, which compares means. The two tests answer different questions about different data types.

| Feature | Chi square test | t-test |
|---|---|---|
| Data type | Counts of categories | Continuous measurements |
| Compares | Observed vs expected frequencies | Means of two groups |
| Null hypothesis | No association or no difference from expected | No difference in means |
| Test statistic | $\chi^2 = \sum (O-E)^2/E$ | Based on difference in means over standard error |
| Distribution | Chi square, right tailed | t distribution, symmetric |

If your outcome is a category count, use chi square. If your outcome is a numeric average, use a [two sample t-test](/blog/data-analysis/two-sample-t-test-formula-example). For comparing two proportions directly, a [two proportion z-test](/blog/data-analysis/two-proportion-z-test-formula) is another option.

## Common Mistakes

- **Using raw counts where percentages belong, or the reverse.** The formula needs frequencies, not proportions. Convert percentages back to counts before computing.
- **Computing expected counts incorrectly.** For independence, each expected count is row total times column total divided by grand total [4]. Do not divide by the number of cells.
- **Getting degrees of freedom wrong.** Use $(r-1)(c-1)$ for a contingency table [3], not the number of cells and not the sample size.
- **Ignoring small expected counts.** If any expected count is below about five, the chi square approximation is unreliable [7]. Combine categories or switch to an exact test.
- **Treating a significant result as proof of causation.** The test detects association. It says nothing about which variable causes the other.
- **Running the test on paired data.** Chi square assumes independent observations. Repeated measures on the same subjects need a different method.

## Limitations

The chi square test only detects departures from the null. It cannot tell you the size or direction of an effect, and a significant result in a large sample may reflect a trivial difference. Report the counts alongside the statistic so readers can judge practical relevance.

The test is also sensitive to how you define your categories. For goodness of fit tests on continuous data, the number of groups and how group membership is defined affect the power of the test, and only rules of thumb exist for choosing them [7]. Results can shift when you redraw the boundaries.

## Frequently Asked Questions

### What is the difference between a chi square test and a chi square distribution?

The distribution is the probability curve the statistic follows. The test is the procedure that produces a statistic and compares it against that curve. The distribution is right skewed for small degrees of freedom and becomes more symmetric as degrees of freedom increase [3]. You can read more about the shape and its properties in this guide to the [chi square distribution](/blog/data-analysis/chi-square-distribution).

### What does a chi square value of 0 mean?

A value of 0 means every observed count exactly equals its expected count. There is no deviation at all. In practice this almost never happens with real data, and a value near 0 simply indicates a very close fit.

### Can I use a chi square test with only two categories?

Yes. A 2x2 table is the most common setup, and it gives 1 degree of freedom. The example in this article uses exactly that design. Just check that all expected counts are large enough for the approximation to hold.

### How many observations do I need for a chi square test?

There is no fixed sample size rule. What matters is the expected count in each cell. A common rule of thumb requires an expected count of at least five in every cell [7]. If your expected counts fall below that, combine categories or use an exact test.

### What is the difference between the goodness of fit test and the test of independence?

The goodness of fit test compares one variable's observed frequencies against a theoretical distribution [1]. The test of independence checks whether two categorical variables are associated in a contingency table [1]. Both use the same statistic, but they differ in how expected counts are derived.

## References

1. [11: Chi-Square Test - Statistics LibreTexts](https://stats.libretexts.org/Courses/Citrus_College/Statistics_C1000%3A_Introduction_to_Statistics/11%3A_Chi-Square_Test)
2. [Witt PL, McGrain P. (1986). Nonparametric testing using the chi-square distribution. Physical therapy](https://pubmed.ncbi.nlm.nih.gov/3945680/)
3. [10.1: Chi-Square Test for Independence - Mathematics LibreTexts](https://math.libretexts.org/Courses/Cosumnes_River_College/STAT_300%3A_Introduction_to_Probability_and_Statistics_(Nam_Lam)/10%3A_Goodness_of_Fit_and_Test_for_Independence/10.01%3A_Chi-Square_Test_for_Independence)
4. [Chi-Square Goodness of Fit Test](http://www.stat.yale.edu/Courses/1997-98/101/chigf.htm%25)
5. [CHI-SQUARE](https://lgross.utk.edu/bioed/bealsmodules/chi-square.html)
6. [Chi-Square Tests - Statistics Resources - LibGuides at National University](https://resources.nu.edu/statsresources/Chi-Square)
7. [7.2.1.1. Chi-square goodness-of-fit test](https://www.itl.nist.gov/div898/handbook/prc/section2/prc211.htm)

## Related Articles

- [Chi-Square Distribution: Definition, Formula and Examples](/blog/data-analysis/chi-square-distribution)
- [Two Sample t-Test: Formula, Calculation and Example](/blog/data-analysis/two-sample-t-test-formula-example)
- [Chi-Square Table: How to Read It and Find Critical Values](/blog/data-analysis/chi-square-table-critical-values)
- [Two Proportion Z-Test: Formula and Worked Example](/blog/data-analysis/two-proportion-z-test-formula)
- [Excel ROUND Function: Formula and Examples](/blog/data-analysis/excel-round-function-formula-examples)
- [Genetics lab report example chi-square](/blog/research-skills/genetics-lab-report-example-writing-up-drosophila-crosses-and-chi-square-analysis)
- [Post-Hoc Tests After a Significant Chi-Square](/knowledge/bioinformatics/post-hoc-tests-after-a-significant-chi-square-standardized-residuals-partitioning-and-pairwise-compa)
- [Chi-Square Test of Independence vs. Fisher](/knowledge/bioinformatics/chi-square-test-of-independence-vs-fisher-s-exact-test-a-decision-framework-for-small-sample-sizes-i)