Post-Hoc Tests After a Significant Chi-Square
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- After a significant chi-square test of independence, standardized residuals are the primary diagnostic tool to identify specific cells deviating from expected frequencies, with absolute values greater than approximately 2 (or 1.96/3) indicating meaningful contributions to the overall association.
- Table partitioning involves dividing the original contingency table into smaller subtables to test specific hypotheses about group differences, but requires corrections for multiple comparisons due to the non-independence of these subtables.
- Pairwise comparisons, analogous to post-hoc tests in ANOVA, are crucial for identifying which specific groups or categories differ, and necessitate corrections like Bonferroni or Holm-Bonferroni to control the family-wise error rate.
- The choice of post-hoc method (standardized residuals, partitioning, or pairwise comparisons) is dictated by the research question's structure (exploratory vs. planned comparisons) and the contingency table's dimensions (2xk vs. rx c).
- A structured workflow, including checking chi-square assumptions (expected cell counts ≥ 5), examining residuals, selecting an appropriate post-hoc strategy with multiple testing correction, and documenting the process, ensures reproducible and defensible analysis.
- Common failures include omitting post-hoc analysis, treating standardized residuals as formal tests without correction, performing pairwise comparisons without correction, or using overly conservative corrections like Bonferroni when less conservative methods like Holm-Bonferroni or False Discovery Rate control are more appropriate.
Quick Answer
- After a significant chi-square test, examine standardized residuals to identify which cells contribute most to the overall association.
- Partition the contingency table into smaller subtables and test each with a chi-square, applying a correction for multiple comparisons.
- Standardized residuals are descriptive diagnostics, not formal significance tests, so pair them with adjusted p-values from pairwise comparisons.
Understanding the Post-Hoc Problem in Chi-Square Analysis
A significant chi-square test of independence tells you that the observed frequencies in your contingency table are unlikely to have arisen by chance under the null hypothesis of independence. It does not tell you which specific cells, categories, or combinations of categories are driving that departure from expectation. This distinction is the core problem that post-hoc procedures address.
Consider a typical biological experiment. You might compare the distribution of a categorical outcome across three treatment groups. The overall chi-square test returns a p-value below your chosen alpha threshold. You now know that the groups are not all identical in their outcome distribution, but you do not know whether all three groups differ from each other, whether one group differs from the other two, or whether the difference is concentrated in a particular outcome category. Reporting only the omnibus test leaves your conclusion vague and limits the practical value of the analysis.
The post-hoc analysis after a significant chi-square test is analogous to post-hoc testing after a significant ANOVA. In ANOVA, a significant F-test leads to pairwise comparisons among group means. In chi-square analysis, the significant test leads to an examination of which cells in the table deviate from expected counts, which categories drive the association, or which pairs of groups differ in their distribution. The statistical logic is the same: the omnibus test establishes that something is happening, and the post-hoc procedures identify where it is happening.
The choice of post-hoc method depends on the structure of your data and the question you are asking. If you have a single categorical variable with more than two levels and you are comparing observed frequencies to a theoretical distribution, you would use a goodness-of-fit approach. If you have two categorical variables and you want to know which cells deviate from independence, standardized residuals are the primary diagnostic tool. If you have a table with more than two rows or columns and you want to compare specific groups, you would partition the table and run pairwise comparisons.
The statistical logic here is straightforward. The chi-square statistic is a sum of contributions from each cell in the table. Each cell contributes a quantity calculated as the squared difference between the observed and expected frequency divided by the expected frequency. The cells with the largest contributions are the ones most responsible for the significant result. Examining these contributions is the first step in any post-hoc analysis.
Standardized Residuals as the Primary Diagnostic
The most direct way to identify which cells drive a significant chi-square result is to examine the standardized residuals for each cell. The standardized residual for a cell is calculated by taking the difference between the observed and expected frequency and dividing by the standard error of that difference. This produces a value that can be interpreted in relation to the standard normal distribution.
A standardized residual with an absolute value greater than approximately 2 indicates that the cell contributes meaningfully to the overall chi-square statistic. Some researchers use a threshold of 1.96, which corresponds to the 5 percent significance level in a two-tailed test. Others use a more conservative threshold of 3 to reduce the risk of false positives when many cells are examined. The choice of threshold should be made before examining the data and should be reported in the methods section.
The standardized residual is preferred over the raw contribution to the chi-square statistic because it is scaled to account for the expected frequency. A cell with a large expected count can produce a large contribution to the chi-square even when the proportional deviation is small. The standardized residual corrects for this by dividing by the square root of the expected count, making the values comparable across cells with different expected frequencies.
Consider a concrete example from a biological laboratory. Suppose you are studying the distribution of a genetic marker across four treatment conditions. The overall chi-square test is significant. The standardized residuals show that one treatment condition has a residual of 2.8 for the presence of the marker, while the other three conditions have residuals below 1.5. This pattern suggests that the significant result is driven by the elevated frequency of the marker in that one treatment condition. The other conditions do not appear to differ meaningfully from expectation.
The standardized residual approach has an important limitation. It is a descriptive diagnostic tool, not a formal hypothesis test. The thresholds used to identify meaningful residuals are heuristic and do not carry the same inferential guarantees as a properly corrected pairwise comparison. When you examine many cells, the probability of observing a large residual by chance increases. This is the same multiple-comparison problem that affects all post-hoc analyses.
For this reason, standardized residuals are best used as a screening tool to identify candidate cells for further investigation. The formal confirmation of differences should come from pairwise comparisons or partitioning procedures that incorporate corrections for multiple testing. The residuals tell you where to look, and the corrected tests tell you whether the differences are statistically defensible.
Partitioning the Contingency Table
Partitioning is a method for breaking down a significant chi-square result into smaller, more interpretable components. The approach involves dividing the original contingency table into smaller subtables and testing each subtable separately. The partitions are designed to answer specific questions about the structure of the association.
The most common partitioning strategy is to compare each group to the remaining groups. For a table with three groups, you would create three subtables. The first subtable compares group one to groups two and three combined. The second compares group two to groups one and three combined. The third compares group three to groups one and two combined. Each subtable is tested with its own chi-square test.
Another partitioning strategy is to compare groups in a hierarchical manner. If you have an ordered categorical variable, you might partition the table to test for trends across the ordered categories. This approach is more specialized and requires that the categories have a meaningful order.
The partitioning approach has a clear advantage. It produces results that are directly interpretable in terms of the research question. Instead of saying that the overall table shows a significant association, you can say that group 1 differs from groups 2 and 3 combined, while groups 2 and 3 do not differ from each other. This level of specificity is valuable for biological interpretation.
The partitioning approach also has a limitation. The subtables are not independent of each other because the same data are used in multiple partitions. This non-independence means that the tests are not truly independent, and the overall error rate across the set of tests is higher than the nominal level for each individual test. A correction for multiple testing is needed to control the family-wise error rate.
The degrees of freedom for the partitions should sum to the degrees of freedom of the original table. This property allows you to verify that the partitioning is complete and that no information has been lost. For a table with r rows and c columns, the degrees of freedom are calculated as the product of the number of rows minus one and the number of columns minus one. The degrees of freedom of the partitions should sum to this total.
Pairwise Comparisons with Corrections
The most common approach to post-hoc analysis after a significant chi-square test is to perform pairwise comparisons between the groups and apply a correction for multiple testing. This approach is analogous to the post-hoc procedures used after ANOVA.
For a table with k groups, there are k choose 2 possible pairwise comparisons. Each comparison is a 2 by c subtable that tests whether the distribution of the outcome variable differs between the two groups. The chi-square test for each subtable is calculated in the standard way.
The multiple testing problem is immediate. If you have five groups, you have ten pairwise comparisons. If each comparison is tested at the 0.05 level, the probability of finding at least one significant result by chance is substantially higher than 0.05. The correction procedures are designed to control this inflated error rate.
The Bonferroni correction is the simplest approach. The alpha level for each individual test is divided by the number of comparisons. For ten comparisons, each test would be evaluated at the 0.005 level. The Bonferroni correction is conservative, meaning it reduces the power to detect true differences, but it is simple to apply and easy to explain.
The Holm-Bonferroni method is a stepwise improvement over the standard Bonferroni correction. The p-values from the pairwise comparisons are ordered from smallest to largest. The smallest p-value is compared to the adjusted threshold of alpha divided by the number of comparisons. If it is significant, the next p-value is compared to alpha divided by the number of comparisons minus one. The process continues until a non-significant result is found. The Holm method is more powerful than the standard Bonferroni correction while maintaining the same family-wise error rate.
The false discovery rate approach is another option. The Benjamini-Hochberg procedure controls the expected proportion of false discoveries among the rejected hypotheses. This approach is less conservative than the Bonferroni methods and is often preferred in exploratory analyses where the goal is to identify candidate differences for further study.
The choice of correction method should be made before the analysis and reported in the methods section. The choice depends on the research context. If the analysis is confirmatory and the cost of a false positive is high, the Bonferroni or Holm correction is appropriate. If the analysis is exploratory and the goal is to generate hypotheses, the false discovery rate approach may be more suitable.
The At a Glance Table
| Method | What It Identifies | When to Use | Key Limitation |
|---|---|---|---|
| Standardized residuals | Cells with observed counts far from expected | First step after significant chi-square | Descriptive only, no formal inference |
| Table partitioning | Which groups or categories differ from the rest | When you have a specific comparison structure | Subtables are not independent |
| Pairwise comparisons with Bonferroni | Which pairs of groups differ | Confirmatory analysis with few comparisons | Conservative, may miss true differences |
| Pairwise comparisons with Holm | Which pairs of groups differ | Confirmatory analysis with moderate comparisons | Less conservative than Bonferroni |
| False discovery rate | Which pairs of groups differ | Exploratory analysis with many comparisons | Controls proportion of false positives |
Practical Workflow for Post-Hoc Analysis
The post-hoc analysis should follow a structured workflow that begins with the overall chi-square test and proceeds through the diagnostic and confirmatory stages. The workflow ensures that the analysis is complete, reproducible, and defensible.
The first step is to verify that the overall chi-square test is appropriate for the data. The assumptions of the chi-square test must be checked before the test is interpreted. The expected frequency in each cell should be at least 5 for the standard chi-square test to be valid. If the expected frequencies are lower, Fisher exact test or a different approach should be used.
The second step is to examine the standardized residuals for each cell. This examination identifies the cells that contribute most to the overall significance. The residuals are calculated as part of the chi-square output in most statistical software packages. The cells with absolute standardized residuals above the chosen threshold are flagged for further investigation.
The third step is to decide on the post-hoc strategy. The choice depends on the research question and the structure of the table. If the goal is to identify which groups differ from each other, pairwise comparisons are appropriate. If the goal is to identify which cells deviate from expectation, the standardized residuals are the primary tool. If the goal is to test a specific structure, partitioning is appropriate.
The fourth step is to perform the chosen post-hoc tests with the appropriate correction for multiple testing. The results should be reported with the corrected p-values and the effect sizes for each comparison.
The fifth step is to interpret the results in the context of the biological question. The post-hoc analysis identifies the specific differences, but the biological significance of those differences must be evaluated by the researcher.
The sixth step is to document the entire procedure in the methods section of the report. The documentation should include the overall chi-square result, the post-hoc method chosen, the correction applied, and the thresholds used.
Software Implementation
The post-hoc procedures described in this article are implemented in most standard statistical software packages. The specific commands and output formats vary by software, but the underlying calculations are the same.
In R, the chisq.test function provides the overall chi-square test and the standardized residuals. The residuals are stored in the residuals component of the test result object. The pairwise comparisons can be performed using the chisq.test function on subsets of the data, with the p-values adjusted using the p.adjust function.
In Python, the scipy.stats.chi2_contingency function provides the overall chi-square test. The standardized residuals can be calculated manually from the observed and expected frequencies. The pairwise comparisons can be performed using the same function on subsets of the data, with the p-values adjusted using the statsmodels multipletests function.
In SPSS, the crosstabs procedure provides the chi-square test and the standardized residuals. The pairwise comparisons can be performed by selecting subsets of the data and running the crosstabs procedure on each subset. The p-values must be adjusted manually or using a custom script.
In GraphPad Prism, the contingency analysis provides the chi-square test and the standardized residuals. The pairwise comparisons can be performed using the multiple comparisons option in the contingency analysis dialog.
The choice of software does not affect the underlying statistical logic. The same calculations are performed regardless of the software package. The researcher should choose the software that is most familiar and document the specific commands used in the analysis.
Common Failure Patterns
The most common failure in post-hoc analysis after a significant chi-square test is the failure to perform any post-hoc analysis at all. Many researchers report only the overall chi-square result and do not identify the specific differences. This leaves the reader with an incomplete understanding of the data.
The second common failure is the use of the standardized residuals as formal tests without any correction for multiple testing. The residuals are descriptive diagnostics, and the thresholds used to identify meaningful residuals are heuristic. Treating the residuals as formal tests inflates the probability of false positives.
The third common failure is the application of pairwise comparisons without any correction for multiple testing. This is particularly problematic when the number of groups is large. The probability of finding at least one false positive increases rapidly with the number of comparisons.
The fourth common failure is the use of the Bonferroni correction in situations where it is too conservative. The Bonferroni correction assumes that the comparisons are independent, which is not the case when the same data are used in multiple comparisons. The Holm-Bonferroni method is a better choice in most situations.
The fifth common failure is the failure to check the assumptions of the chi-square test before performing the post-hoc analysis. If the expected frequencies are too low, the chi-square test is not valid, and the post-hoc results are also not valid.
The sixth common failure is the failure to report the post-hoc procedure in sufficient detail. The report should include the specific method used, the correction applied, and the number of comparisons performed. Without this information, the reader cannot evaluate the validity of the post-hoc results.
Limitations of Post-Hoc Tests
The post-hoc tests described in this article have several limitations that should be acknowledged in the research report.
The first limitation is the loss of statistical power. The correction for multiple testing reduces the power to detect true differences. This is the trade-off for controlling the family-wise error rate. The researcher should be aware of this trade-off and plan the sample size accordingly.
The second limitation is the dependence of the pairwise comparisons. The pairwise comparisons are not independent because the same data are used in multiple comparisons. This dependence is not fully accounted for by the standard corrections.
The third limitation is the descriptive nature of the standardized residuals. The residuals are useful for identifying the cells that contribute most to the overall significance, but they do not provide a formal test of the difference in any particular cell.
The fourth limitation is the sensitivity of the chi-square test to sample size. With a large sample, the chi-square test can be significant even when the differences are small and biologically unimportant. The post-hoc tests will also be significant in this situation, but the practical significance of the differences must be assessed by the researcher.
The fifth limitation is the difficulty of interpreting the results when the table is large. With many rows and columns, the number of pairwise comparisons becomes very large, and the results become difficult to interpret. In this situation, a different approach, such as correspondence analysis, may be more appropriate.
Reporting the Post-Hoc Results
The reporting of post-hoc results should follow the established guidelines for statistical reporting in the biological sciences. The EQUATOR Network provides a collection of reporting guidelines for different types of research studies. The guidelines emphasize the importance of reporting the statistical methods in sufficient detail to allow the reader to evaluate the validity of the results.
The report should include the overall chi-square test result, including the test statistic, the degrees of freedom, and the p-value. The report should also include the post-hoc method used, the correction applied, and the number of comparisons performed. The results of the post-hoc tests should be presented in a table that shows the comparison, the test statistic, the p-value, and the adjusted p-value.
The report should also include the standardized residuals for the cells that are identified as contributing to the overall significance. The residuals should be presented in a table or a figure that shows the observed and expected frequencies for each cell.
The report should describe the biological interpretation of the post-hoc results. The statistical significance of the differences should be interpreted in the context of the research question and the existing literature.
The report should also acknowledge the limitations of the post-hoc analysis. The limitations should be described in the discussion section of the report.
Records and Reproducibility
The post-hoc analysis should be reproducible. The researcher should document the specific commands used in the analysis, the version of the software, and the data file. This documentation allows another researcher to reproduce the analysis and verify the results.
The data management and sharing policy of the National Institutes of Health requires that research data be shared in a timely manner. The policy applies to research funded by the NIH and requires that the data be made available to other researchers. The post-hoc analysis should be documented in a way that allows other researchers to reproduce the results from the shared data.
The researcher should also document the decisions made during the analysis. The choice of the post-hoc method, the correction, and the threshold should be documented in the methods section of the report. The documentation should be sufficient to allow another researcher to understand the rationale for the decisions.
The researcher should also document the version of the software used for the analysis. The results of the analysis can vary between software versions, and the documentation should allow the reader to identify the specific version used.
Professional Escalation Criteria
The post-hoc analysis may reveal results that require professional escalation. The following situations should be discussed with a statistician or a more experienced researcher.
The first situation is when the expected frequencies are too low for the chi-square test to be valid. The chi-square test requires that the expected frequency in each cell be at least 5. If the expected frequencies are lower, the Fisher exact test should be used instead. The Fisher exact test is more computationally intensive but does not have the same assumption.
The second situation is when the table is large and the number of pairwise comparisons is very large. The correction for multiple testing becomes very conservative, and the results may be difficult to interpret. A statistician can help to choose a more appropriate approach.
The third situation is when the results of the post-hoc analysis are inconsistent with the overall chi-square test. This can happen when the overall test is significant but none of the pairwise comparisons are significant. This situation can occur when the overall test is driven by a combination of small differences across multiple cells. A statistician can help to interpret this result.
The fourth situation is when the data are not independent. The chi-square test assumes that the observations are independent. If the data are clustered or repeated measures, the chi-square test is not valid, and a different approach is needed.
The fifth situation is when the research involves human subjects or animal subjects. The research must be conducted in accordance with the relevant ethical guidelines. The Committee on Publication Ethics provides guidance on the ethical conduct of research and publication.
A Decision Framework for Selecting the Post-Hoc Method
The choice among standardized residuals, table partitioning, and pairwise comparisons is often presented as a matter of statistical preference, but it is better understood as a decision that follows from the structure of your research question and the design of your study. A practical decision framework helps you select the appropriate method before you run the analysis, document that choice in your methods section, and avoid the common failure of switching methods after seeing the results.
Step 1: Define the Comparison Structure
The first decision point is whether your research question specifies a comparison structure in advance. A planned comparison structure exists when you have a hypothesis about which groups or categories should differ before you collect the data. For example, if you are studying a dose response and you expect the highest dose to differ from the control and the intermediate doses, you have a planned structure. In this case, table partitioning is the appropriate method because it tests the specific contrasts you predicted.
If you do not have a planned structure and are exploring which groups differ after seeing a significant overall test, pairwise comparisons with a correction for multiple testing are appropriate. This is the exploratory scenario that most researchers face after a significant chi-square test.
The distinction matters because planned comparisons do not require the same correction for multiple testing as post-hoc comparisons. The correction is needed when the comparisons are data-driven because the probability of finding a false positive increases with the number of comparisons examined. When the comparisons are specified in advance, the multiple testing problem is reduced because you are not searching over all possible comparisons.
Step 2: Assess the Table Dimensions
The dimensions of your contingency table constrain the available post-hoc methods. A 2 by 2 table has only one degree of freedom, and a significant chi-square test means the two variables are associated. There are no meaningful pairwise comparisons to perform because the table has only two rows and two columns. The standardized residuals for the four cells describe the direction of the association, and the analysis is complete.
A 2 by k table, where k is greater than 2, has one categorical variable with two levels and another with k levels. The post-hoc question is which of the k levels differ from each other in their distribution across the two levels of the other variable. Pairwise comparisons among the k levels are appropriate, with the correction applied to the number of comparisons.
An r by c table, where both r and c are greater than 2, has the most complex post-hoc structure. The pairwise comparisons can be made among the rows, among the columns, or both. The number of comparisons grows quickly. For a 4 by 4 table, there are six pairwise comparisons among the rows and six among the columns, for a total of twelve comparisons. The correction for multiple testing becomes more conservative as the number of comparisons increases.
Step 4: Evaluate the Expected Frequencies
The expected frequencies in the table determine whether the chi-square test and the post-hoc procedures are valid. The standard chi-square test requires that the expected frequency in each cell be at least 5. If the expected frequencies are lower, the chi-square test is not valid, and the post-hoc procedures based on the chi-square distribution are also not valid.
When the expected frequencies are low, the Fisher exact test is the appropriate alternative. The Fisher exact test calculates the exact probability of the observed table under the null hypothesis of independence, without relying on the chi-square approximation. The Fisher exact test can be extended to tables larger than 2 by 2, but the computation becomes intensive for large tables.
The post-hoc analysis after a Fisher exact test follows the same logic as the post-hoc analysis after a chi-square test. The pairwise comparisons are performed using the Fisher exact test on the subtables, and the p-values are adjusted for multiple testing. The standardized residuals are still useful for identifying the cells that contribute to the overall association.
Step 5: Determine the Number of Comparisons
The number of comparisons is the key factor in choosing the correction method. The Bonferroni correction divides the alpha level by the number of comparisons. For a small number of comparisons, such as three or four, the Bonferroni correction is not too conservative, and the loss of power is acceptable.
For a moderate number of comparisons, such as six to ten, the Holm-Bonferroni method is preferred. The Holm method is more powerful than the standard Bonferroni correction because it uses a stepwise approach that adjusts the alpha level based on the ordered p-values. The Holm method maintains the same family-wise error rate as the Bonferroni correction but with greater power.
For a large number of comparisons, such as more than ten, the false discovery rate approach may be more appropriate. The Benjamini-Hochberg procedure controls the expected proportion of false discoveries among the rejected hypotheses. This approach is less conservative than the Bonferroni methods and is often preferred in exploratory analyses where the goal is to identify candidate differences for further study.
The choice of correction method should be documented in the methods section of the report. The documentation should include the number of comparisons performed and the rationale for the correction method chosen.
Step 6: Document the Decision
The decision framework should be documented in the methods section of the research report. The documentation should include the comparison question, the table dimensions, the expected frequencies, and the number of comparisons. This documentation allows the reader to evaluate the validity of the post-hoc analysis and to reproduce the analysis from the shared data.
The documentation should also include the specific commands used in the analysis and the version of the software. The National Institutes of Health data management and sharing policy requires that research data be shared in a timely manner, and the documentation should allow other researchers to reproduce the results from the shared data.
A Record System for Post-Hoc Decisions
A record system for post-hoc decisions helps you track the choices made during the analysis and provides a basis for the methods section of the report. The record should include the following elements for each analysis:
The first element is the research question and the comparison structure. This entry records whether the comparisons were planned or exploratory and what specific differences were of interest.
The second element is the table dimensions. This entry records the number of rows and columns in the contingency table and the degrees of freedom for the overall test.
The third element is the expected frequency check. This entry records the minimum expected frequency in the table and whether the chi-square test was valid or whether the Fisher exact test was used.
The fourth element is the post-hoc method chosen. This entry records whether the analysis used standardized residuals, table partitioning, pairwise comparisons, or a combination of these methods.
The fifth element is the correction method. This entry records the correction applied to the p-values and the number of comparisons performed.
The sixth element is the threshold for the standardized residuals. This entry records the threshold used to identify meaningful residuals and the rationale for the choice.
The seventh element is the software and version. This entry records the software package and version used for the analysis.
The record system should be maintained in a laboratory notebook or an electronic file that is accessible to the research team. The record provides the basis for the methods section of the report and allows the analysis to be reproduced by other researchers.
Troubleshooting the Post-Hoc Analysis
The decision framework and record system help you avoid common failure patterns, but problems can still arise during the analysis. The following troubleshooting steps address the most common issues.
The first issue is a significant overall test with no significant pairwise comparisons. This situation can occur when the overall test is driven by a combination of small differences across multiple cells. The overall test has more power than the individual pairwise comparisons because it uses all of the data. The result should be interpreted as evidence of an overall association, but the specific differences cannot be identified with the available sample size. The standardized residuals can help identify the cells with the largest deviations, even if the pairwise comparisons are not significant.
The second issue is a significant pairwise comparison that is not supported by the standardized residuals. This situation can occur when the pairwise comparison is driven by a combination of small deviations across multiple cells in the subtable. The standardized residuals for the individual cells may be below the threshold, but the combined effect across the cells is significant. The pairwise comparison is the more reliable result because it is a formal test with a corrected p-value.
The third issue is a large number of pairwise comparisons that makes the correction very conservative. The Bonferroni correction for a table with many groups can reduce the power to detect true differences. The Holm-Bonferroni method is more powerful and should be used in this situation. If the number of comparisons is very large, the false discovery rate control may be more appropriate.
The fourth issue is the presence of zero cells in the table. A zero cell can occur when a particular combination of categories is not observed in the sample. The zero cell can cause the expected frequency to be zero, which makes the chi-square test invalid. The Fisher exact test should be used when the expected frequencies are low or zero.
The fifth issue is the presence of a structural zero in the table. A structural zero occurs when a particular combination of categories is impossible by design. For example, if you are studying the distribution of a treatment outcome and one treatment is not administered to one group, the cell for that combination is a structural zero. The structural zero should be excluded from the analysis, and the table should be redefined to remove the impossible combination.
The Role of the Decision Framework in Reporting
The decision framework provides the structure for the methods section of the report. The methods section should describe the comparison question, the table dimensions, the expected frequency check, the post-hoc method chosen, the correction method, and the number of comparisons. This description allows the reader to evaluate the validity of the post-hoc analysis and to understand the rationale for the choices made.
The reporting should follow the established guidelines for statistical reporting in the biological sciences. The EQUATOR Network provides a collection of reporting guidelines for different types of research studies. The guidelines emphasize the importance of reporting the statistical methods in sufficient detail to allow the reader to evaluate the validity of the results.
The report should also include the standardized residuals for the cells that are identified as contributing to the overall significance. The residuals should be presented in a table or a figure that shows the observed and expected frequencies for each cell.
The report should describe the biological interpretation of the post-hoc results. The statistical significance of the differences should be interpreted in the context of the research question and the existing literature.
The report should also acknowledge the limitations of the post-hoc analysis. The limitations should be described in the discussion section of the report.
The Decision Framework in Practice
The decision framework is best understood through a concrete example. Suppose you are studying the distribution of a categorical outcome across four treatment groups. The overall chi-square test is significant. The table has 4 rows and 2 columns, so the degrees of freedom are 3.
The first step is to define the comparison question. You do not have a planned structure, so the comparisons are exploratory. The second step is to assess the table dimensions. The table has 4 rows and 2 columns, so there are six pairwise comparisons among the four groups. The third step is to check the expected frequencies. The minimum expected frequency is 6, which is above the threshold of 5, so the chi-square test is valid.
The fourth step is to choose the post-hoc method. Because the comparisons are exploratory, pairwise comparisons are appropriate. The fifth step is to choose the correction method. The number of comparisons is six, which is moderate, so the Holm-Bonferroni method is appropriate. The sixth step is to document the decision in the methods section.
The analysis produces six pairwise comparisons with adjusted p-values. The results show that group 1 differs from group 2 and group 3, but not from group 4. The standardized residuals show that the difference is driven by the elevated frequency of the outcome in group 1. The results are reported in a table that shows the comparison, the test statistic, the p-value, and the adjusted p-value.
The decision framework ensures that the post-hoc analysis is systematic, documented, and reproducible. The framework prevents the common failure of switching methods after seeing the results and provides a clear record for the methods section of the report.
Frequently Asked Questions
What is the difference between a standardized residual and a chi-square contribution?
The chi-square contribution for a cell is the squared difference between the observed and expected frequencies divided by the expected frequency. The standardized residual is the square root of the chi-square contribution, with the sign indicating whether the observed frequency is higher or lower than expected. The standardized residual is easier to interpret because it is on a scale that can be compared to the standard normal distribution.
When should I use the Bonferroni correction versus the Holm-Bonferroni method?
The Holm-Bonferroni method is generally preferred because it is more powerful than the standard Bonferroni correction while maintaining the same family-wise error rate. The standard Bonferroni correction is simpler to explain and is appropriate when the number of comparisons is small. The Holm-Bonferroni method should be used when the number of comparisons is moderate to large.
Can I use the standardized residuals as a formal test of significance?
The standardized residuals are descriptive diagnostics, not formal tests. The threshold of 2 or 3 is a heuristic that identifies cells with large deviations from expectation. The residuals do not provide a formal p-value for the difference in any particular cell. Formal testing requires pairwise comparisons or partitioning with corrections for multiple testing.
What should I do if the overall chi-square test is significant but none of the pairwise comparisons are significant?
This situation can occur when the overall test is driven by a combination of small effects across multiple cells. The overall test has more power than the individual pairwise comparisons because it uses all of the data. The result should be interpreted as evidence of an overall association, but the specific differences cannot be identified with the available sample size.
How do I choose the threshold for the standardized residuals?
The threshold should be chosen before the analysis and reported in the methods section. The threshold of 2 is commonly used and corresponds to the 95 percent confidence level for a single cell. The threshold of 3 is more conservative and reduces the risk of false positives when many cells are examined. The choice of threshold should be based on the number of cells and the cost of false positives.
What is the difference between a post-hoc test and a planned comparison?
A post-hoc test is performed after the overall test is significant and is not specified in advance. A planned comparison is specified before the data is collected and is based on the research question. The planned comparisons do not require the same correction for multiple testing because they are specified in advance. The post-hoc tests require the correction because they are data-driven.
Can I use the post-hoc tests for a chi-square goodness-of-fit test?
The post-hoc procedures for a chi-square test of independence can be adapted for a goodness-of-fit test. The goodness-of-fit test compares the observed distribution to a theoretical distribution. The post-hoc analysis identifies the categories that deviate from the theoretical distribution. The standardized residuals are the primary diagnostic tool for the goodness-of-fit test.
How do I report the post-hoc results in a research paper?
The report should include the overall chi-square test result, the post-hoc method used, the correction applied, and the number of comparisons performed. The results should be presented in a table that shows the standardized residuals and the adjusted p-values for the pairwise comparisons. The interpretation of the results should be described in the discussion section.
Using the Evidence
| Source | Best use in this topic | Important limitation |
|---|---|---|
| Research Methods Resources | official guidance | Check the linked page for current local requirements |
| EQUATOR Network | official guidance | Check the linked page for current local requirements |
| Core Practices | official guidance | Check the linked page for current local requirements |
Related Bioinformatics Guides
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Foundation Models for Genomics: From Single Cells to Health Trajectories
- Genomic Diagnostics: Choosing the Right Test for Clinical and Veterinary Applications
- Spatial Transcriptomics vs. Single-Cell RNA Sequencing: Which Approach Fits Your Research?
- The Complete Guide to Sequence File Formats and Pairwise Alignment in Bioinformatics
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Research Methods Resources. National Library of Medicine.
- EQUATOR Network. EQUATOR Network.
- Core Practices. Committee on Publication Ethics.
- NIH Grants and Funding. National Institutes of Health.
- ORCID for Researchers. ORCID.
- Data Management and Sharing Policy. National Institutes of Health.
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Effectiveness of oral health education intervention among 12-15-year-old school children in Dharan, Nepal: a randomized controlled trial.. BMC oral health, 2021.
- Fisher's exact approach for post hoc analysis of a chi-squared test.. PloS one, 2017.
- [Outcomes for [(177)Lu]Lu-PSMA-617 with and Without Concurrent Use of Androgen Receptor Pathway Inhibitors in Patients with Metastatic Castration-resistant Prostate Cancer.](https://pubmed.ncbi.nlm.nih.gov/40752988). European urology oncology, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.