# Contingency Tables in Life Sciences


## Key Takeaways

- Contingency tables summarize the joint distribution of two categorical variables, essential for assessing associations between factors like treatment groups (e.g., drug vs. placebo) and outcomes (e.g., cured vs. not cured) in life sciences.
- Accurate table construction mandates pre-defining all variable categories in the study protocol to prevent post-hoc adjustments that inflate false-positive rates and compromise reproducibility.
- Statistical validity hinges on verifying expected cell counts; if more than 20% of cells have expected counts below 5, or any cell below 1, the chi-square approximation is unreliable, necessitating the use of Fisher's exact test.
- Handling zero counts and missing categories requires explicit decisions documented in the protocol, as their inclusion or exclusion impacts degrees of freedom and expected counts, thereby influencing statistical test outcomes.
- Contingency tables reveal associations but do not establish causation; results must be interpreted within the study design's limitations, acknowledging potential unmeasured confounders and avoiding causal inference.
- Reproducibility is paramount; meticulously record table construction steps, including software, code, data versions, and decisions on category handling, to enable independent verification and auditing.

---

## Quick Answer

- Build a contingency table by counting how often each combination of categorical variables occurs in your raw data, with one row per category of the primary variable and one column per category of the secondary variable.
- Before analysis, verify that every observed category appears in the table and decide explicitly how to handle zero counts, because missing categories and empty cells change the validity of chi-square and Fisher exact tests.
- A contingency table only summarizes counts, so it cannot reveal causation, measurement error, or the effect of unmeasured confounders, interpret results within the limits of your study design.

## At a Glance

| Decision Point | What to Do | Common Error | Consequence of Error |
| --- | --- | --- | --- |
| Define categorical variables before data collection | Write explicit category definitions and levels in your study protocol | Adding or merging categories after seeing results | Inflates false-positive rate and makes results non-reproducible |
| Count every observed combination | Use a cross-tabulation function or manual tally that includes all category pairs | Dropping rows with zero counts | Biases the table and invalidates the chi-square approximation |
| Handle small expected counts | Check expected cell counts before choosing a test | Running chi-square when expected counts are below the threshold | Produces unreliable p-values and misleading conclusions |
| Record the table construction steps | Document the software, code, and data version used to build the table | No record of how the table was made | Prevents replication and audit by reviewers or regulators |
| Report the table with the study | Include the full contingency table in the manuscript or supplement | Reporting only the p-value | Hides data errors and prevents independent verification |

## What Is a Contingency Table in Life Sciences

A contingency table is a two-way frequency table that shows how often each combination of two categorical variables occurs in a dataset. In life-science research, these tables are used to examine whether an outcome is associated with a treatment, exposure, genotype, or other grouping factor. For example, a table might show the number of mice that developed a tumor or did not develop a tumor, split by whether they received a drug or a placebo.

The table has rows that represent the levels of one variable and columns that represent the levels of the other variable. Each cell contains the count of observations that fall into that specific row and column combination. The row totals and column totals are called marginal totals, and the total number of observations is the grand total. These marginal totals are used to calculate expected counts under the null hypothesis of no association.

Contingency tables are the foundation for several statistical tests used in biology, including the chi-square test of independence, Fisher exact test, and McNemar test for paired data. The structure of the table determines which test is appropriate and whether the test assumptions are met. A table with two rows and two columns is called a 2x2 table and is common in clinical and laboratory studies. Larger tables, such as 3x3 or 2x4, are used when a variable has more than two categories.

The key distinction from other data summaries is that a contingency table preserves the joint distribution of two variables. A simple frequency table shows the distribution of one variable alone. A contingency table shows how the distribution of one variable changes across the levels of the other variable. This joint structure is what allows a researcher to test for association.

## Why Correct Table Structure Matters

The structure of a contingency table determines which statistical test can be applied and whether the test result is valid. A table that is missing a category, has merged categories, or includes zero counts in the wrong place can produce a misleading p-value. The researcher may then draw a biological conclusion that is not supported by the data.

One common problem is the inclusion of a category with zero observations in both groups. This can happen when a researcher defines a category in the protocol but no animal or sample falls into it. The zero row or column changes the degrees of freedom and the expected counts, which can alter the test result. The researcher must decide whether to keep the zero category in the table or remove it, and that decision must be made before the analysis.

Another problem is the merging of categories after the data are collected. If a researcher sees that two categories have similar counts and merges them to achieve a significant result, the analysis is no longer confirmatory. This practice is a form of p-hacking and violates the principle of pre-specified analysis. The Committee on Publication Ethics Core Practices state that researchers should report their methods honestly and avoid selective reporting of results. The full table and the analysis plan should be documented before the test is run.

The table structure also affects the interpretation of the odds ratio or relative risk. In a 2x2 table, the odds ratio is calculated from the four cells. If the rows and columns are swapped, the odds ratio is inverted. If a category is collapsed, the odds ratio changes. The researcher must define the reference category and the direction of the comparison in the protocol.

## Core Principles of Contingency Table Construction

### Define Variables and Categories Before Data Collection

The first principle is to define the categorical variables and their levels in the study protocol before any data are collected. This includes the name of the variable, the type of variable (nominal or ordinal), and the exact list of categories. For example, a study on bacterial resistance might define the variable "resistance status" with two levels: resistant and susceptible. A study on disease severity might define four levels: none, mild, moderate, and severe.

The definition must be specific enough that two different researchers would classify the same observation into the same category. This is called inter-rater reliability. The protocol should include the criteria for each category, such as a threshold value for a continuous measurement or a standard laboratory result. The National Library of Medicine provides access to research-methods references that describe how to define and measure variables in biomedical studies.

The categories should be mutually exclusive and exhaustive. Mutually exclusive means that each observation can belong to only one category. Exhaustive means that every observation can be placed into one of the listed categories. If a category is missing, the researcher must add an "other" category or a "missing" category to the protocol. This decision must be made before the data are collected, not after.

### Count Observations, Not Measurements

A contingency table contains counts of observations, not measurements or averages. Each observation is a single experimental unit, such as a mouse, a cell culture, a patient, or a field plot. The researcher must decide what the experimental unit is before building the table. If a mouse is measured multiple times, the researcher must decide whether to use the mouse as the unit or the measurement as the unit. This decision affects the sample size and the validity of the test.

For example, if a researcher measures the blood pressure of a mouse three times, the three measurements are not independent. The contingency table should contain one count per mouse, not three counts. The researcher must aggregate the measurements into a single category for each mouse, such as "hypertensive" or "normal," based on a pre-defined threshold.

### Use a Cross-Tabulation Function

The most reliable way to build a contingency table is to use a statistical software function that cross-tabulates two variables. In R, the `table()` function or the `xtabs()` function. In Python, the `crosstab()` function in pandas. In a spreadsheet, a pivot table. These functions automatically count the number of observations for each combination of categories and produce a table with the correct dimensions.

The researcher should not build the table by hand in a spreadsheet unless the dataset is very small. Hand-built tables are prone to errors in counting and in placing counts in the wrong cell. The software function also provides the marginal totals, which are needed for the expected count calculation.

### Verify the Table Against the Raw Data

After the table is built, the researcher must verify that the table matches the raw data. The sum of all cells in the table must equal the total number of observations in the dataset. The sum of each row must equal the number of observations in that category. The sum of each column must equal the number of observations in that category.

The researcher should also check that no observation is counted twice and that no observation is missing. A simple way to verify is to compare the marginal totals from the table with the frequency counts of each variable alone. If the marginal totals do not match, there is an error in the table construction.

## Step-by-Step Framework for Building a Contingency Table

The following workflow provides a step-by-step framework for building a contingency table from raw experimental data. The workflow is designed for a life-science researcher who has a dataset with two categorical variables.

### Step 1: Identify the Two Categorical Variables

The first step is to identify the two categorical variables that will form the rows and columns of the table. The researcher must decide which variable is the primary variable and which is the secondary variable. The primary variable is often the exposure or treatment, and the secondary variable is the outcome. The primary variable is placed in the rows, and the secondary variable is placed in the columns.

For example, in a study of a new antibiotic, the primary variable is "treatment" with levels "drug" and "placebo." The secondary variable is "outcome" with levels "cured" and "not cured." The table has two rows and two columns.

### Step 2: List All Categories for Each Variable

The researcher must list all categories for each variable. This list must match the protocol. If the protocol defined three categories for the outcome, the table must have three columns. If the protocol defined four categories for the treatment, the table must have four rows.

The researcher should write the list of categories in the same order as in the protocol. For ordinal variables, the categories should be in natural order, such as "none, mild, moderate, severe." For nominal variables, the order does not matter, but the researcher should be consistent.

### Step 3: Count the Observations for Each Combination

The researcher must count the number of observations that fall into each combination of a row category and a column category. This is the cross-tabulation step. The researcher can use a software function or a manual tally for a small dataset.

For each observation in the raw data, the researcher determines the row category and the column category and adds one to the corresponding cell. The researcher must ensure that every observation is counted exactly once.

### Step 4: Fill in the Table with Counts

The researcher fills in the table with the counts. The table has a row for each category of the primary variable and a column for each category of the secondary variable. The cell at the intersection of row i and column j contains the count of observations that have the row category i and the column category j.

The researcher also calculates the row totals, the column totals, and the grand total. The row totals are the sum of the counts in each row. The column totals are the sum of the counts in each column. The grand total is the sum of all cells.

### Step 5: Handle Zero Counts and Missing Categories

The researcher must decide how to handle zero counts and missing categories. A zero count is a cell with no observations. A missing category is a category that was defined in the protocol but has no observations in the data.

For a zero count, the researcher must decide whether to keep the cell in the table or to remove the entire row or column. The decision depends on the statistical test. For a chi-square test, a zero cell is acceptable if the expected count is not too low. For a Fisher exact test, a zero cell is acceptable. The researcher must document the decision.

For a missing category, the researcher must decide whether to add a row or column with all zeros or to remove the category from the table. Adding a row with all zeros changes the degrees of freedom and the expected counts. Removing the category changes the interpretation of the table. The decision must be made before the analysis and documented in the protocol.

### Step 6: Verify the Table

The researcher verifies the table by checking the marginal totals against the raw data. The sum of the row totals must equal the grand total. The sum of the column totals must equal the grand total. The researcher should also check that the table is square or rectangular and that the number of rows and columns matches the number of categories.

### Step 7: Record the Table and the Process

The researcher records the table and the process used to build it. The record should include the software version, the code or the function used, the date, and the version of the raw data. This record is essential for reproducibility. The National Institutes of Health Data Management and Sharing Policy requires that researchers plan for the sharing of data and metadata. The table construction process is part of the metadata.

## Handling Missing Categories and Zero Counts

Missing categories and zero counts are common in life-science data. A category may have no observations because the condition is rare, because the sample size is small, or because the category was defined but not observed. The researcher must handle these cases correctly to avoid invalid statistical tests.

### Zero Counts in a Cell

A zero count in a cell means that no observation has that combination of categories. For example, in a study of a new drug, the cell for "drug" and "cured" might have a zero count if the drug is completely ineffective. The zero count is a valid data point. It does not mean that the cell is missing.

The researcher must keep the zero count in the table. Removing the zero count would change the marginal totals and the expected counts. The chi-square test can handle zero counts if the expected counts are not too low. The Fisher exact test can handle zero counts without any problem.

### Missing Categories

A missing category is different from a zero count. A missing category is a category that was defined in the protocol but has no observations in the entire dataset. For example, if the protocol defined "severe" as a category of the outcome, but no patient had a severe outcome, then the "severe" category is missing.

The researcher must decide whether to include the missing category in the table. If the category is included, the table has a row or column with all zeros. This row or column changes the degrees of freedom and the expected counts. If the category is excluded, the table has fewer rows or columns, and the interpretation of the table changes.

The decision must be based on the research question. If the research question is about the association between the treatment and the outcome, and the "severe" category is not observed, the researcher might exclude the category. If the research question is about the distribution of the outcome, the researcher might include the category with a zero count.

The researcher must document the decision in the protocol and in the manuscript. The decision should be made before the analysis, not after seeing the results.

### Small Counts and the Chi-Square Test

The chi-square test of independence is based on the expected counts. The expected count for each cell is calculated as the row total times the column total divided by the grand total. The chi-square test is valid when the expected counts are sufficiently large. The rule of thumb is that no more than 20% of the cells should have an expected count below 5, and no cell should have an expected count below 1.

If the expected counts are too low, the chi-square test is not valid. The researcher should use the Fisher exact test instead. The Fisher exact test calculates the exact probability of the observed table under the null hypothesis, without relying on the approximation. The Fisher exact test is valid for small sample sizes and for tables with zero counts.

The researcher must check the expected counts before choosing the test. The expected counts are calculated from the table. The researcher should not run the chi-square test and then check the expected counts after the fact. The check should be done before the test.

## Choosing the Right Statistical Test

The contingency table is the input for the statistical test. The choice of test depends on the structure of the table and the research question. The researcher must choose the test before running the analysis.

### Chi-Square Test of Independence

The chi-square test of independence is used to test whether two categorical variables are associated. The null hypothesis is that the variables are independent. The test statistic is calculated from the observed counts and the expected counts. The p-value is the probability of observing a test statistic as extreme as the one observed, assuming the null hypothesis is true.

The chi-square test is appropriate for tables with two or more rows and two or more columns. The test is valid when the expected counts are sufficiently large. The researcher must check the expected counts before running the test.

### Fisher Exact Test

The Fisher exact test is used for 2x2 tables and for larger tables when the expected counts are too low for the chi-square test. The test calculates the exact probability of the observed table under the null hypothesis. The test is valid for any sample size and for tables with zero counts.

The Fisher exact test is computationally intensive for large tables. For a 2x2 table, the test is fast. For a 3x3 table, the test is slower. The researcher should use the Fisher exact test when the chi-square test is not valid.

### McNemar Test for Paired Data

The McNemar test is used for paired data, such as before and after measurements on the same subject. The test is used for 2x2 tables where the rows and columns represent the same variable measured at two time points. The test compares the discordant pairs, which are the cells where the row and column categories differ.

The McNemar test is appropriate when the data are paired. The researcher must not use the chi-square test for paired data because the observations are not independent.

### Test Selection Criteria

The researcher must select the test based on the table structure and the data. The criteria are the number of rows and columns, the expected counts, and the pairing of the data. The researcher must document the test selection in the protocol.

The researcher should also consider the direction of the hypothesis. If the research question is about a one-sided alternative, the researcher must specify the direction before the analysis. The p-value for a one-sided test is different from the p-value for a two-sided test.

## Common Failure Patterns in Contingency Table Analysis

The following are common failure patterns that lead to errors in contingency table analysis. The researcher should be aware of these patterns and avoid them.

### Failure to Pre-Specify Categories

The most common failure is to define the categories after the data are collected. The researcher sees the data and decides to merge or split categories to achieve a significant result. This practice is a form of p-hacking and is not acceptable. The Committee on Publication Ethics Core Practices require that researchers report their methods honestly and avoid selective reporting. The categories must be defined in the protocol before the data are collected.

### Failure to Check Expected Counts

The researcher runs a chi-square test without checking the expected counts. The expected counts are too low, and the p-value is unreliable. The researcher must check the expected counts before running the test. If the expected counts are too low, the researcher must use the Fisher exact test.

### Failure to Handle Zero Counts

The researcher removes a row or column with a zero count without documenting the decision. The removal changes the degrees of freedom and the expected counts. The researcher must document the decision and the reason for the decision.

### Failure to Verify the Table

The researcher builds the table by hand and makes a counting error. The table does not match the raw data. The researcher must verify the table against the raw data by checking the marginal totals.

### Failure to Document the Process

The researcher does not record the table construction process. The table cannot be reproduced by another researcher. The researcher must record the table, the code, the data version, and the date.

### Failure to Report the Table

The researcher reports only the p-value and not the table. The reader cannot verify the analysis or interpret the results. The researcher must report the full contingency table in the manuscript or the supplementary material.

## Records and Measurements for Reproducibility

Reproducibility requires that the researcher records the table construction process. The record should include the raw data, the code used to build the table, the version of the software, and the date. The record should also include the decisions made about missing categories and zero counts.

The NIH Data Management and Sharing Policy requires that researchers plan for the sharing of data and supporting the documentation. The policy applies to research funded by the NIH. The researcher should include the contingency table and the construction process in the data management plan.

The researcher should also use a unique identifier for the data and the table. The ORCID for Researchers provides a unique identifier for the researcher, which can be linked to the data and the table. The identifier helps to ensure that the data and the table are attributed to the correct researcher.

The researcher should also record the version of the raw data. If the raw data are updated, the table must be rebuilt. The researcher should not use an old table with new data.

## Common Failure Patterns in Analysis

The following are common failure patterns in the analysis of contingency tables. The researcher should be aware of these patterns and avoid them.

### Using the Wrong Test

The researcher uses the chi-square test when the expected counts are too low. The researcher uses the chi-square test for paired data. The researcher must select the test based on the table structure and the data.

### Interpreting the P-Value Incorrectly

The p-value is the probability of the observed data under the null hypothesis. The p-value is not the probability that the null hypothesis is true. The p-value is not the probability that the alternative hypothesis is true. The researcher must interpret the p-value correctly.

### Ignoring the Effect Size

The p-value indicates whether the association is statistically significant. The p-value does not indicate the strength of the association. The researcher must report the effect size, such as the odds ratio or the relative risk, to describe the strength of the association.

### Ignoring the Confidence Interval

The confidence interval provides a range of plausible values for the effect size. The researcher must report the confidence interval along with the p-value. The confidence interval is more informative than the p-value alone.

### Ignoring the Sample Size

The p-value depends on the sample size. A small sample size may not have enough power to detect a real association. A large sample size may detect a small association that is not biologically meaningful. The researcher must consider the sample size when interpreting the results.

## Reporting the Contingency Table in a Manuscript

The contingency table must be reported in the manuscript. The table should be labeled with a number and a title. The table should include the row and column categories, the counts, and the marginal totals. The table should also include the p-value and the effect size.

The researcher should follow the reporting guidelines for the study design. The EQUATOR Network provides a list of reporting guidelines for different study designs. The researcher should select the appropriate guideline and follow it. The guideline will specify the information that must be reported in the manuscript.

The researcher should also report the test used, the expected counts, and the decision to handle missing categories. The researcher should report the software and the version used for the analysis.

The researcher should also report the data-sharing plan. The NIH Data Management and Sharing Policy requires that the researcher plan for the sharing of data. The plan should include the data, the code, and the documentation.

## Limitations of Contingency Table Analysis

The contingency table analysis has several limitations. The researcher must be aware of these limitations and interpret the results within the bounds of the design.

### No Causation

A contingency table can show an association between two variables. The table cannot show causation. The association may be due to a confounder, a third variable that is associated with both the exposure and the outcome. The researcher must not conclude causation from a contingency table.

### No Measurement of the Strength of the Association

The p-value indicates whether the association is statistically significant. The p-value does not indicate the strength of the association. The researcher must report the effect size, such as the odds ratio or the relative risk, to describe the strength of the association.

### No Handling of Confounders

The contingency table does not control for confounders. The researcher must use a more advanced method, such as logistic regression, to control for confounders. The researcher must not use a contingency table to control for confounders.

### No Handling of Continuous Variables

The contingency table requires categorical variables. If the variable is continuous, the researcher must categorize the variable. The categorization can lose information and can affect the results. The researcher must choose the categories carefully.

### No Handling of the Missing Data

The contingency table does not handle missing data. If the data are missing, the researcher must decide how to handle the missing data. The researcher must not ignore the missing data.

## Professional Escalation Criteria

The researcher should escalate the analysis to a professional statistician when the analysis is complex or when the researcher is not sure about the analysis. The following are the criteria for escalation:

- The table has more than two rows or two columns.
- The expected counts are too low for the chi-square test.
- The data are paired.
- The data are missing.
- The researcher is not sure about the test selection.
- The researcher is not sure about the interpretation of the results.

The researcher should consult a statistician before the analysis, not after the analysis. The statistician can help the researcher to select the test, to check the assumptions, and to interpret the results.

## A Practical Decision Framework for Structuring Contingency Tables

Building a contingency table correctly requires more than following a generic workflow. Researchers need a concrete decision framework they can apply at each step, especially when the data do not fit neatly into the expected structure. The framework below gives you a sequence of checkpoints, each with a specific question to answer and a rule to apply. This approach reduces the risk of introducing structural errors that invalidate downstream tests.

### Checkpoint 1: Confirm the Experimental Unit

Before any counting begins, confirm what constitutes a single observation. The experimental unit is the smallest entity that is independently assigned to a treatment or exposure group. In a mouse study, the unit is the mouse, not a single tumor or a single blood measurement. In a field trial, the unit is the plot, not the individual plant. In a clinical study, the unit is the patient, not the visit.

Ask this question: does each row in the raw dataset represent one experimental unit? If the dataset contains repeated measurements on the same unit, the table will overcount and inflate the sample size. The decision is to aggregate repeated measurements into a single category per unit before counting. For example, if a mouse has three blood pressure readings, classify the mouse as hypertensive or normal based on a pre-defined threshold, then count the mouse once.

### Checkpoint 2: Verify the Category Lists Against the Protocol

The next checkpoint is to compare the categories present in the raw data against the categories defined in the study protocol. Write the protocol-defined categories for each variable in a separate column next to the observed categories. Any category in the protocol that does not appear in the data is a missing category. Any value in the data that does not match a protocol category is an unclassified observation.

For each discrepancy, make one of three decisions. First, if the protocol category is missing because the condition is rare, decide whether to keep the zero row or column or to collapse it with an adjacent category. Second, if the data contain an unclassified value, decide whether to assign it to an existing category based on the protocol criteria or to create an "other" category. Third, if the protocol itself is ambiguous, document the ambiguity and resolve it before proceeding. The Committee on Publication Ethics Core Practices require honest reporting of methods, so the resolution of each discrepancy must be recorded.

### Checkpoint 3: Apply the Expected Count Rule Before Choosing a Test

The choice between the chi-square test and the Fisher exact test depends on the expected counts, not the observed counts. The expected count for each cell is the row total multiplied by the column total divided by the grand total. Calculate these expected counts from the table before running any test.

Apply this rule: if more than 20% of the cells have an expected count below 5, or if any cell has an expected count below 1, the chi-square approximation is not reliable. In that case, use the Fisher exact test. This rule is a standard threshold in biostatistics practice. The decision must be made before the test is run, not after seeing the p-value.

### Checkpoint 4: Decide on Zero Rows and Columns

A zero row or column is a category with no observations in the entire dataset. The decision to keep or remove it changes the degrees of freedom and the expected counts. The decision depends on the research question.

If the research question is about the association between two variables, and the zero category is a level of the outcome that never occurred, removing the category is often appropriate because it does not contribute information about association. If the research question is about the distribution of the outcome, the zero category should be kept to show the full distribution. The decision must be documented in the protocol and in the manuscript.

### Checkpoint 5: Cross-Check Marginal Totals

After the table is built, verify that the marginal totals match the raw data. The row total for each category must equal the number of observations in that category in the raw data. The column total must equal the number of observations in that category. The grand total must equal the total number of experimental units.

A mismatch indicates a counting error, a duplicate count, or a missing observation. The table must be rebuilt before any test is run. This verification step is the final gate before analysis.

### A Record Sheet for Table Construction

A structured record sheet helps to document the decisions made at each checkpoint. The record should include the following fields for each variable: the protocol-defined categories, the observed categories, the number of observations per category, and the decision made for any discrepancy. The record should also include the expected count check result, the test selected, and the date of the analysis.

This record serves as the metadata for the table. The National Institutes of Health Data Management and Sharing Policy requires that researchers plan for the sharing of data and supporting documentation. The record sheet is part of that documentation and should be included in the data management plan.

### Common Failure Patterns in the Decision Process

The most common failure is to skip the category verification step and proceed directly to counting. This leads to a table that does not match the protocol and cannot be reproduced. A second failure is to choose the test based on the observed counts instead of the expected counts. A third failure is to remove a zero row or column without documenting the decision, which makes the analysis non-reproducible.

A fourth failure is to use the wrong experimental unit. This happens when the researcher counts measurements instead of units, inflating the sample size and producing a p-value that is too small. The researcher must confirm the unit of analysis before building the table.

### When to Escalate to a Statistician

The decision framework is not a substitute for professional statistical advice. Escalate to a statistician when the table has more than two rows or two columns, when the expected counts are too low for the chi-square test, when the data are paired, when the data are missing, or when the researcher is uncertain about the test selection. The statistician should be consulted before the analysis, not after the results are known.

## Frequently Asked Questions

### What is the difference between a frequency table and a contingency table?

A frequency table shows the distribution of one categorical variable. A contingency table shows the joint distribution of two categorical variables. The contingency table has rows for one variable and columns for the other variable, and each cell contains the count of observations that have the specific combination of categories.

### How do I decide which variable goes in the rows and which goes in the columns?

The primary variable, usually the exposure or treatment, goes in the rows. The secondary variable, usually the outcome, goes in the columns. The choice does not affect the chi-square test of independence, but it affects the interpretation of the odds ratio and the relative risk.

### What should I do if a category has zero observations?

If a category has zero observations in a cell, keep the zero count in the table. If a category has zero observations in the entire table, decide whether to include the category with a zero row or column or to exclude the category. Document the decision in the protocol.

### When should I use the Fisher exact test instead of the chi-square test?

Use the Fisher exact test when the expected counts are too low for the chi-square test. The rule of thumb is that no more than 20% of the cells should have an expected count below 5, and no cell should have an expected count below 1. The Fisher exact test is valid for any table size and for tables with zero counts.

### Can I use a contingency table to prove causation?

No. A contingency table can show an association between two variables, but it cannot prove causation. The association may be due to a confounder or to chance. The researcher must not conclude causation from a contingency table.

### How do I report the contingency table in my manuscript?

Report the full contingency table with the row and column categories, the counts, and the marginal totals. Report the test used, the p-value, and the effect size. Follow the reporting guidelines for your study design, which are available from the EQUATOR Network.

### What is the role of the NIH Data Management and Sharing Policy in contingency table analysis?

The NIH Data Management and Sharing Policy requires that researchers plan for the sharing of data and supporting documentation. The contingency table and the code used to build it are part of the supporting documentation. The researcher should include the table and the code in the data-sharing plan.

### How do I ensure that my contingency table is reproducible?

Record the raw data, the code used to build the table, the version of the software, and the date. Record the decisions made for missing categories and zero counts. Use a unique identifier for the data and the table, such as an ORCID for the researcher.

## Using the Evidence

| Source | Best use in this topic | Important limitation |
|---|---|---|
| [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books) | official guidance | Check the linked page for current local requirements |
| [EQUATOR Network](https://www.equator-network.org/) | official guidance | Check the linked page for current local requirements |
| [Core Practices](https://publicationethics.org/core-practices) | official guidance | Check the linked page for current local requirements |

## Related Bioinformatics Guides

- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Microbiome Data Analysis in R: A Practical Guide for Compositional Data](/knowledge/bioinformatics/microbiome-data-analysis-in-r-a-practical-guide-for-compositional-data)
- [What Is a Data Warehouse? A Practical Guide for Life Science Organizations](/knowledge/bioinformatics/what-is-a-data-warehouse-a-practical-guide-for-life-science-organizations)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Data Science and AI in Life Sciences: Applications and Emerging Trends](/knowledge/bioinformatics/data-science-and-ai-in-life-sciences-applications-and-emerging-trends)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Diagnostic accuracy of 16S rDNA PCR, multiplex PCR and metagenomic next-generation sequencing in periprosthetic joint infections: a systematic review and meta-analysis.](https://pubmed.ncbi.nlm.nih.gov/40023316). Clinical microbiology and infection : the official publication of the European Society of Clinical Microbiology and Infectious Diseases, 2025.
- [Histological differential diagnostics of renal oncocytoma.](https://pubmed.ncbi.nlm.nih.gov/32996740). Experimental oncology, 2020.
- [The role of correspondence analysis in medical research.](https://pubmed.ncbi.nlm.nih.gov/38584915). Frontiers in public health, 2024.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.