# What Is a Contingency Table? Definition and Examples

A contingency table is a grid that shows how many observations fall into each combination of two categorical variables. Reading across rows and down columns lets you compare groups and spot relationships at a glance. This article explains the structure, the math behind it, and a full worked example.

## Quick Answer

- A contingency table cross-tabulates two categorical variables, with one variable in the rows and the other in the columns.
- Each cell holds a count (a frequency) of observations sharing both characteristics [1].
- Row totals, column totals, and a grand total sit in the margins and let you compute proportions and conditional probabilities [2].
- You compare observed counts against expected counts to test whether the two variables are independent.
- The chi-square test is the standard companion to a contingency table for testing a significant relationship [1][3].

## What a Contingency Table Means

In plain terms, a contingency table is a way to sort your data by two questions at once. If you survey people about their gender and whether they prefer a product, a contingency table shows you the four possible combinations side by side instead of hiding them in a long list.

The precise statistical definition: a contingency table displays sample values for two different variables that may be dependent or contingent on one another [2][4]. The word "contingent" is the key. It asks whether the value of one variable depends on the value of the other. A table with two variables and two categories each is called a 2x2 table, and tables can extend to any number of rows and columns [5].

The table is built from counts, not formulas. At its core it is counts and percentages of observations [1][6]. That simplicity is why it works for almost any pair of categorical variables.

## How It Works

Start with raw observations. Each observation belongs to exactly one row category and one column category, so it lands in exactly one cell. Add up each row and each column to get the marginal totals, then add those to get the grand total [2].

To test whether the two variables are related, you compare what you observed to what you would expect if the variables were independent. The expected count for any cell is:

$$E = \frac{(\text{row total}) \times (\text{column total})}{\text{grand total}}$$

Each symbol means:

- $E$ is the expected count in a cell under the assumption of independence.
- Row total is the sum of counts in that cell's row.
- Column total is the sum of counts in that cell's column.
- Grand total is the total number of observations in the table.

Then you compute the chi-square statistic:

$$\chi^2 = \sum \frac{(O - E)^2}{E}$$

Here $O$ is the observed count and $E$ is the expected count, and you sum over every cell. The degrees of freedom for the test are:

$$df = (\text{rows} - 1)(\text{columns} - 1)$$

A large chi-square value relative to the degrees of freedom suggests the observed counts differ from what independence would predict [5].

## Worked Example

The dataset is a survey of 100 respondents cross-tabulating gender (Male/Female) and preference (Yes/No).

| Gender | Preference | Count |
|---|---|---|
| Male | Yes | 30 |
| Male | No | 20 |
| Female | Yes | 15 |
| Female | No | 35 |

Arranged as a 2x2 table, the observed counts are:

| | Yes | No | Row total |
|---|---|---|---|
| Male | 30 | 20 | 50 |
| Female | 15 | 35 | 50 |
| Column total | 45 | 55 | 100 |

The row totals are Male = 50 and Female = 50. The column totals are Yes = 45 and No = 55. The grand total is 100.

Now compute the expected counts using $E = (\text{row total} \times \text{column total}) / \text{grand total}$:

- E(Male, Yes) = (50 x 45) / 100 = 22.5000
- E(Male, No) = (50 x 55) / 100 = 27.5000
- E(Female, Yes) = (50 x 45) / 100 = 22.5000
- E(Female, No) = (50 x 55) / 100 = 27.5000

The expected counts are:

| | Yes | No |
|---|---|---|
| Male | 22.5 | 27.5 |
| Female | 22.5 | 27.5 |

Summing $(O - E)^2 / E$ across all four cells gives a chi-square statistic of 9.0909 with 1 degree of freedom, and a p-value of 0.0026.

You can reproduce this in Python:

```python
import pandas as pd
from scipy.stats import chi2_contingency
df = pd.DataFrame([[30,20],[15,35]],
                  index=['Male','Female'], columns=['Yes','No'])
chi2, p, dof, exp = chi2_contingency(df, correction=False)
print(f"chi2 = {chi2:.4f}, p = {p:.4f}, dof = {dof}")
print("Expected counts:")
print(exp)
```

Output:

```
chi2 = 9.0909, p = 0.0026, dof = 1
Expected counts:
[[22.5 27.5]
 [22.5 27.5]]
```

The conditional proportions tell the same story. P(Yes | Male) = 30 / 50 = 0.6000, and P(Yes | Female) = 15 / 50 = 0.3000. Men in this sample said yes twice as often as women.

## How to Interpret It

Read the table in three passes.

First, look at the raw counts and the totals. The row and column totals should each add to the grand total, which confirms the table is complete [2].

Second, convert counts to percentages. Raw counts depend on sample size, so percentages make groups comparable. If you are studying cause and effect, compute percentages in the direction of the causal factor, provided the sample is representative in that direction [7]. In the example above, gender is the grouping variable, so you compare P(Yes | Male) against P(Yes | Female).

Third, compare observed to expected. When observed counts sit close to expected counts, the variables look independent. When they diverge sharply, a relationship is likely. In the worked example, men said yes 30 times against an expected 22.5, and women said yes 15 times against an expected 22.5. That gap, combined with p = 0.0026, points to a real association.

## When to Use It (and when not to)

Use a contingency table when both variables are categorical. Examples include survey responses, treatment versus outcome, and any yes/no or category-based split. It is also the natural first step before a chi-square test, and it pairs well with a frequency table when you want to inspect one variable at a time.

Do not use a contingency table when a variable is continuous. Binning a continuous variable into categories throws away information and can create misleading results. For continuous outcomes, use a scatterplot, correlation, or a regression model instead.

Also avoid it when your sample is very small. Sparse cells make the chi-square approximation unreliable, and you may need an exact test instead.

## Contingency Table vs Frequency Table

These two are easy to confuse because both count observations. The difference is how many variables they cover.

| Feature | Contingency Table | Frequency Table |
|---|---|---|
| Number of variables | Two or more | Usually one |
| Layout | Rows and columns crossed | Single list of categories |
| Shows relationships | Yes, between variables | No, only distribution |
| Typical use | Chi-square test, crosstab | Describing one variable |

A frequency table answers "how many of each?" A contingency table answers "how many of each, split by group?" If you only need the first question, a frequency table is enough.

## Common Mistakes

- **Reading cell counts as percentages.** A count of 30 means nothing without the total. Fix: divide by the relevant row or column total before comparing groups.
- **Comparing raw counts across unequal group sizes.** If one group has 200 people and another has 50, raw counts mislead. Fix: compare conditional proportions or row percentages [7].
- **Computing percentages in the wrong direction.** Percentages should run in the direction of the causal factor when you have a cause-and-effect hypothesis [7]. Fix: decide which variable is the grouping variable before you calculate.
- **Ignoring expected counts.** A chi-square result is meaningless if you never check whether expected counts are large enough. Fix: inspect the expected table and treat small expected counts with caution.
- **Treating a small p-value as proof of a strong effect.** Significance and effect size are different things. Fix: report the proportions alongside the p-value.
- **Dropping missing data silently.** Rows with missing values on either variable disappear from the table. Fix: report how many observations were excluded and consider a sensitivity check.

## Limitations

A contingency table shows association, not causation. A relationship between two variables can be produced by a third variable you did not measure, and the table alone cannot rule that out. Percentages can provide key insights, but they cannot take the place of measures of association [7].

The chi-square test also has assumptions. It relies on expected counts being reasonably large, and it becomes unreliable with sparse cells. It tells you whether a relationship exists, not how strong it is or which specific cells drive the result. For those questions you need follow-up analysis, such as standardized residuals or a measure of association.

## Frequently Asked Questions

### What is a contingency table in simple terms?

It is a grid that counts how many observations fall into each combination of two categories. One variable labels the rows, the other labels the columns, and each cell holds a count. The margins hold row totals, column totals, and the grand total [2].

### What is the difference between a 2x2 and a larger contingency table?

A 2x2 table has two rows and two columns, so each variable has two categories. Larger tables simply add more rows or columns. The mechanics are identical, and the degrees of freedom formula $(\text{rows} - 1)(\text{columns} - 1)$ scales with the size [5].

### How do I know if two variables are independent?

Compare observed counts to expected counts. If they are close, the variables look independent. If they diverge, a relationship is likely. The chi-square test formalizes this by producing a statistic and a p-value [1][3].

### Can I use a contingency table with more than two variables?

Yes, but the table becomes harder to read as you add variables. A common approach is to build separate tables for each level of a third variable, or to use a stratified analysis. Software such as SAS and SPSS handles multi-way crosstabs directly [1][6].

### What sample size do I need?

There is no single cutoff, but the chi-square approximation needs adequate expected counts in every cell. Sparse tables with many small expected counts call for an exact test instead. Always check the expected counts before trusting the p-value.

## References

1. [Crosstabs (Contingency Table) - SAS - GSU Library Research Guides at Georgia State University](https://research.library.gsu.edu/sas/crosstabs)
2. [3.5: Contingency Tables - Statistics LibreTexts](https://stats.libretexts.org/Courses/Los_Angeles_City_College/Introductory_Statistics/03%3A_Probability_Topics/3.05%3A_Contingency_Tables)
3. [Crossbars (Contingency Table) - R Studio guide - Research Guides at Franklin University](https://guides.franklin.edu/RStudio/Crossbars)
4. [3.4 Contingency Tables - Introductory Statistics 2e | OpenStax](https://openstax.org/books/introductory-statistics-2e/pages/3-4-contingency-tables)
5. [9.2: Chi-square contingency tables - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Applied_Statistics/Mikes_Biostatistics_Book_(Dohm)/09%3A_Categorical_Data/9.2%3A_Chi-square_contingency_tables)
6. [Crosstabs (Contingency Table) - SPSS - GSU Library Research Guides at Georgia State University](https://research.library.gsu.edu/SPSSHelp/crosstabs)
7. [Contingency Tables](https://pages.uoregon.edu/rgp/PPPM613/class9a.htm)

## Related Articles

- [Frequency Table: Definition, How to Make One, Examples](/blog/data-analysis/frequency-table-definition-and-examples)
- [Contingency Tables in Life Sciences](/knowledge/bioinformatics/contingency-tables-in-life-sciences-a-practical-guide-to-construction-notation-and-common-pitfalls)
- [Handling Missing Data in Contingency Tables](/knowledge/bioinformatics/handling-missing-data-in-contingency-tables-imputation-strategies-and-sensitivity-analysis-for-biolo)
- [How to Calculate a Percentage: Formula and Examples](/blog/data-analysis/how-to-calculate-a-percentage)
- [What Is an Independent Variable? Definition and Examples](/blog/data-analysis/what-is-an-independent-variable)