Choosing Between Parametric and Nonparametric Tests for Repeated Measures
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Repeated Measures ANOVA is indicated for continuous outcomes exhibiting approximate normality and satisfying the sphericity assumption, which posits equal variances of differences between all pairs of conditions. Violation of sphericity necessitates corrections (e.g., Greenhouse-Geisser) or an alternative approach.
- The Friedman test serves as a nonparametric alternative for ordinal or non-normally distributed continuous data, eschewing normality and sphericity assumptions by ranking observations within subjects. However, this ranking discards magnitude information, potentially reducing statistical power for subtle biological effects.
- Mixed-effects models offer superior flexibility for repeated measures data, adeptly handling unbalanced designs, missing time points, and complex correlation structures without the stringent sphericity assumption. They are often the most defensible choice for complex biological datasets.
- Data structure and distribution shape are primary determinants for test selection: complete, balanced, and normally distributed data favor ANOVA; ordinal or skewed continuous data suggest the Friedman test; while unbalanced or incomplete data mandate mixed-effects models.
- Sphericity assessment via Mauchly's test is critical before interpreting repeated measures ANOVA; however, its sensitivity to sample size warrants consideration of direct covariance matrix inspection. Significant violations necessitate corrections or a shift to models not reliant on this assumption.
- Power considerations are paramount: while ANOVA offers higher power for normally distributed data, the Friedman test's robustness to non-normality and outliers can sometimes render it more powerful in specific biological contexts, albeit at the cost of magnitude information.
Quick Answer
- Choose repeated measures ANOVA when your outcome is continuous, approximately normal, and sphericity holds, otherwise consider Friedman test or mixed-effects models.
- Test sphericity with Mauchly's test before interpreting repeated measures ANOVA results, and report the correction when sphericity is violated.
- The Friedman test does not require normality or sphericity but discards magnitude information by ranking data, which can reduce power for subtle biological effects.
Understanding Repeated Measures Designs in Biology
Repeated measures designs are common across the life sciences. A researcher measures the same biological unit multiple times under different conditions or across time points. Examples include tracking cell proliferation across several days, measuring plant height at weekly intervals, recording enzyme activity before and after treatment, or monitoring gene expression across developmental stages. The defining feature is that observations within each experimental unit are correlated, because they come from the same subject, culture, or field plot.
This correlation violates the independence assumption of ordinary ANOVA. When you analyze repeated measures data with methods designed for independent groups, you treat each time point as if it came from a separate subject. That approach inflates the error degrees of freedom, increases the risk of false positives, and produces confidence intervals that are too narrow. The statistical framework must account for the within-subject correlation structure to produce valid inference.
The two most frequently taught approaches are repeated measures ANOVA and the Friedman test. Both are extensions of simpler methods. Repeated measures ANOVA extends the paired t-test to more than two time points or conditions. The Friedman test extends the Wilcoxon signed-rank test to more than two related groups. Each method answers the same broad question, whether the population distributions differ across conditions, but they operate under different assumptions and use different information from the data.
A third option, mixed-effects models, has become increasingly accessible in modern statistical software. These models handle unbalanced designs, missing time points, and complex correlation structures more flexibly than traditional repeated measures ANOVA. For many biological datasets, a linear mixed-effects model is the most defensible choice, especially when the study design is not perfectly balanced or when you need to model the correlation structure explicitly.
The decision framework in this article helps you evaluate your data structure, check assumptions, and choose a defensible analysis path. The goal is not to memorize a flowchart but to understand why each method works and when it fails.
At a Glance
| Method | Data type | Key assumptions | Strength | Limitation |
|---|---|---|---|---|
| Repeated measures ANOVA | Continuous, approximately normal | Sphericity, normality of residuals | High power for detecting mean differences | Sensitive to sphericity violations and outliers |
| Friedman test | Ordinal or continuous | Related groups, no normality required | Robust to non-normal distributions | Ranks discard magnitude, lower power for subtle effects |
| Mixed-effects model | Continuous or non-normal with link function | Correct random effect structure | Handles missing data and unbalanced designs | Requires more statistical expertise and careful model checking |
Core Principles of Repeated Measures Analysis
The Nature of Correlated Observations
In a repeated measures design, each experimental unit contributes multiple observations. A mouse measured at five time points yields five data points, but those five points are not independent. The mouse's baseline size, health status, and genetic background influence all five measurements. The correlation between measurements taken close in time is often stronger than the correlation between measurements taken far apart.
This correlation structure is the reason why standard ANOVA is inappropriate. Standard ANOVA assumes that every observation is independent, meaning the outcome for one unit tells you nothing about the outcome for another unit. In repeated measures data, the outcome for one time point tells you a great deal about the outcome for the next time point on the same unit. Ignoring this structure treats the data as if you had five times as many independent samples, which is not true.
The repeated measures ANOVA handles this by partitioning the total variance into between-subjects variance and within-subjects variance. The within-subjects variance is further partitioned into the effect of time or condition and the residual error. The test statistic compares the variance explained by the time effect to the residual error variance. This approach accounts for the fact that the same subjects are measured repeatedly.
Sphericity and the Compound Symmetry Assumption
Repeated measures ANOVA requires a specific assumption about the covariance structure, called sphericity. Sphericity means that the variances of the differences between all pairs of conditions are equal. In practical terms, the correlation between any two time points should be roughly the same, and the variance of the differences between any two conditions should be similar.
When sphericity is violated, the F-test for the within-subjects effect becomes too liberal. The actual Type I error rate exceeds the nominal alpha level, meaning you are more likely to declare a significant effect when none exists. Mauchly's test is commonly used to assess sphericity, but it is sensitive to sample size. With small samples, Mauchly's test may fail to detect a real violation. With large samples, it may detect trivial violations that have little practical impact.
When sphericity is violated, you have several options. You can apply a correction to the degrees of freedom, such as the Greenhouse-Geisser or Huynh-Feldt corrections. These corrections adjust the p-value to account for the degree of sphericity violation. Alternatively, you can switch to a mixed-effects model that does not require sphericity, or use the Friedman test if the data are not normally distributed.
Normality of Residuals
Repeated measures ANOVA assumes that the residuals are normally distributed. This assumption applies to the error terms, not necessarily to the raw data. In practice, researchers often check the distribution of the dependent variable at each time point, but the formal assumption is about the residuals after accounting for the model effects.
Normality violations can arise from skewed biological measurements, such as cell counts, enzyme activities, or gene expression levels. When the data are heavily skewed, the mean may not be a good summary, and the ANOVA may be misleading. In such cases, you may consider transforming the data, using a nonparametric approach, or using a generalized linear mixed model with an appropriate distribution.
The Friedman Test
How the Friedman Test Works
The Friedman test is a nonparametric alternative to repeated measures ANOVA. It ranks the observations within each subject across the conditions. For each subject, the smallest value gets a rank of 1, the next gets a rank of 2, and so on. The test then compares the sum of ranks for each condition across all subjects.
The null hypothesis is that the distributions of the conditions are identical. The test statistic follows a chi-square distribution with k-1 degrees of freedom, where k is the number of conditions. If the test is significant, you conclude that at least one condition differs from the others.
The Friedman test does not require normality or sphericity. It only requires that the data are measured on at least an ordinal scale and that the observations are related. This makes it a robust choice for data that are clearly non-normal or when the assumptions of repeated measures ANOVA are seriously violated.
Strengths and Limitations of the Friedman Test
The main strength of the Friedman test is its robustness. It does not rely on distributional assumptions, so it is safe to use with skewed data, ordinal data, or data with outliers. It is also relatively simple to compute and interpret.
The main limitation is that ranking discards information about the magnitude of differences. If the differences between conditions are large but the ranks are similar, the test may not detect the effect. For example, if one condition produces values that are consistently higher but the differences are small, the ranks may not reflect the magnitude of the difference. This can reduce the power of the test, especially when the data are approximately normal and the effect is subtle.
Another limitation is that the Friedman test does not provide estimates of the effect size in the original units. It tells you whether the distributions differ, but it does not tell you how much they differ. For biological research, where the magnitude of an effect is often important, this can be a drawback.
When to Use the Friedman Test
Use the Friedman test when the data are clearly non-normal, when the data are ordinal, or when the assumptions of repeated measures ANOVA are not met. It is also a reasonable choice when the sample size is small and you cannot reliably assess normality or sphericity.
For example, if you are measuring a behavioral score that is an ordinal rating, the Friedman test is appropriate. If you are measuring a biomarker that is heavily skewed, the Friedman test may be more robust than ANOVA. If you have a small sample size and cannot verify the assumptions of ANOVA, the Friedman test is a conservative choice.
Mixed-Effects Models as an Alternative
Why Mixed-Effects Models Are Useful
Mixed-effects models, also called linear mixed models or hierarchical models, provide a flexible framework for repeated measures data. They model the fixed effects of the conditions and the random effects of the subjects. The random effects account for the within-subject correlation without requiring sphericity.
A mixed-effects model can handle unbalanced data, where subjects have different numbers of observations. It can also handle missing data, as long as the missingness is not informative. This is a major advantage over repeated measures ANOVA, which requires complete data for all subjects.
Mixed-effects models also allow you to model the correlation structure explicitly. You can choose an unstructured covariance matrix, a compound symmetry structure, or an autoregressive structure, depending on the pattern of correlation in your data. This flexibility makes mixed-effects models the preferred choice for many biological datasets.
When to Use Mixed-Effects Models
Use a mixed-effects model when you have missing data, unbalanced designs, or when you want to model the correlation structure explicitly. They are also useful when you have multiple levels of nesting, such as measurements within animals and animals within treatment groups.
Mixed-effects models require more statistical expertise than repeated measures ANOVA or the Friedman test. You need to specify the random effects structure, choose a covariance structure, and check the model assumptions. However, modern statistical software makes this more accessible, and the benefits often outweigh the additional complexity.
Model Checking and Interpretation
After fitting a mixed-effects model, you should check the residuals for normality and homoscedasticity. You should also check the random effects for normality. If the residuals are not normal, you may need to transform the data or use a generalized linear mixed model with a different distribution.
The interpretation of a mixed-effects model is similar to that of repeated measures ANOVA. The fixed effects for the conditions tell you whether the conditions differ. The random effects tell you how much variability is due to the subjects. The model also provides estimates of the effect sizes and confidence intervals.
Practical Workflow for Choosing a Test
Step 1: Examine the Data Structure
Start by examining the structure of your data. How many subjects do you have? How many time points or conditions? Are the data balanced, meaning each subject has the same number of observations? Are there missing values?
If the data are balanced and complete, repeated measures ANOVA or the Friedman test may be appropriate. If the data are unbalanced or have missing values, a mixed-effects model is a better choice.
Step 2: Check the Distribution
Examine the distribution of the dependent variable at each time point. Create histograms or Q-Q plots to assess normality. If the data are clearly skewed, consider a transformation or a nonparametric test.
If the data are approximately normal, you can proceed with repeated measures ANOVA. If the data are not normal, the Friedman test is a safer choice.
Step 3: Test Sphericity
If you plan to use repeated measures ANOVA, test sphericity with Mauchly's test. If the test is significant, apply a correction such as Greenhouse-Geisser or Huynh-Feldt. Alternatively, switch to a mixed-effects model.
Step 4: Choose the Test
Based on the data structure and assumptions, choose the appropriate test. If the data are normal and sphericity holds, use repeated measures ANOVA. If the data are not normal, use the Friedman test. If the data are unbalanced or have missing values, use a mixed-effects model.
Step 5: Perform the Test and Report
Perform the test and report the results, including the test statistic, degrees of freedom, and p-value. For repeated measures ANOVA, report the sphericity correction if applied. For the Friedman test, report the chi-square statistic and p-value. For mixed-effects models, report the fixed effects and random effects.
Records and Measurements
Data Collection Records
Maintain a clear record of your data collection process. Record the date and time of each measurement, the subject identifier, the condition, and the measured value. This documentation is essential for reproducibility and for verifying the assumptions of your analysis.
Assumption Checks
Document the results of your assumption checks. Record the p-value of Mauchly's test for sphericity, the results of normality tests, and any transformations applied. This documentation helps you justify your choice of test and allows others to evaluate the validity of your analysis.
Analysis Records
Record the statistical software and version you used, the specific test and options, and the output. This documentation is important for reproducibility and for reporting in your research paper.
Common Failure Patterns
Ignoring Sphericity
A common failure is to run repeated measures ANOVA without checking sphericity. This can lead to inflated Type I error rates and false conclusions. Always test sphericity and apply corrections if needed.
Using the Friedman Test with Normal Data
Another failure is to use the Friedman test when the data are normal and the assumptions of ANOVA are met. This reduces power and discards information about the magnitude of the effects. Use the Friedman test only when the data are not normal or the assumptions are not met.
Overlooking Missing Data
Missing data can be a problem for repeated measures ANOVA, which requires complete data. If you have missing values, you may need to use a mixed-effects model or to impute the missing values. Ignoring missing data can bias the results.
Misinterpreting the Friedman Test
The Friedman test only tells you whether the conditions differ. It does not tell you which conditions differ or how much they differ. You need to perform post-hoc tests to identify the specific differences.
Limitations and Interpretation
The Friedman Test and Effect Size
The Friedman test does not provide a direct estimate of the effect size. You can calculate a measure of effect size, such as Kendall's W, but it is not as intuitive as the effect size from ANOVA. This limitation is important when you need to report the magnitude of the effect.
The Repeated Measures ANOVA and Missing Data
Repeated measures ANOVA requires complete data for all subjects. If you have missing data, you may need to drop subjects or impute the missing values. This can reduce the sample size and the power of the test. Mixed-effects models are more flexible in this regard.
The Assumption of Sphericity
The sphericity assumption is often violated in biological data, especially when the time points are not equally spaced. The Greenhouse-Geisser correction is a conservative approach, but it can reduce the power of the test. The mixed-effects model is a more flexible alternative.
Safety and Regulatory Context
Data Management and Sharing
When you conduct research, you must follow the data management and sharing policies of your funding agency. The National Institutes of Health (NIH) requires that you plan for data management and sharing in your grant application. This includes describing how you will handle the data, including the statistical analysis. The NIH policy is described in the Data Management and Sharing Policy. You must also follow the NIH Grants and Funding guidelines for the application and review process.
Reporting Guidelines
When you report your research, you should follow the reporting guidelines that are appropriate for your study design. The EQUATOR Network provides a list of reporting guidelines for different study types. These guidelines help you report your methods and results transparently, which is important for the reproducibility of your research.
Publication Ethics
When you publish your research, you must follow the publication ethics that are described in the Core Practices of the Committee on Publication Ethics. This includes the proper attribution of authorship, the disclosure of conflicts of interest, and the responsible reporting of the data.
Researcher Identity
Maintain your researcher identity by registering for an ORCID identifier. This identifier links your research outputs, including your data and publications, to your professional record. This is important for the transparency and the reproducibility of your research.
A Field Decision Protocol for Repeated Measures Analysis in Growth and Longitudinal Studies
The Core Decision Problem in Practice
Researchers collecting repeated measures data face a practical problem that textbooks rarely address directly. The statistical choice between repeated measures ANOVA and the Friedman test is not a single decision made once at the start of analysis. It is a sequence of decisions that must be revisited as data accumulate, as assumptions are checked, and as the research context evolves. The most common failure pattern is not choosing the wrong test. It is choosing a test and then failing to verify that the choice remains defensible after the data are fully collected and examined.
This section provides a structured decision protocol that you can apply to your own data. The protocol is built around observable data characteristics, not abstract statistical theory. It gives you concrete thresholds, checkpoints, and escalation criteria. The goal is to help you make a defensible choice that you can explain to a reviewer, a collaborator, or a regulatory body.
The Four Gate Decision Framework
The protocol uses four sequential gates. Each gate requires you to examine a specific aspect of your data and make a documented decision. The gates are ordered so that the most fundamental data characteristics are examined first. This ordering prevents you from spending time on detailed assumption checks for a test that is already inappropriate for your data structure.
Gate 1: Data Completeness and Balance
The first gate examines the structure of your dataset. You need to answer three questions before any statistical testing begins.
The first question is whether every subject has a measurement at every time point. In a balanced repeated measures design, each subject contributes exactly one observation per condition. If any subject has a missing time point, the design is unbalanced. Repeated measures ANOVA in most standard software packages requires complete data. When a subject has a missing observation, the entire subject is often dropped from the analysis. This is called listwise deletion. If you have a small sample size, dropping even one subject can substantially reduce your power.
The second question is whether the time points are equally spaced. Some biological measurements are taken at regular intervals, such as daily weights or weekly cell counts. Other measurements are taken at irregular intervals, such as blood samples collected at clinically determined times. Equally spaced time points are compatible with the standard repeated measures ANOVA. Unequally spaced time points create a more complex correlation structure that the standard ANOVA may not handle well.
The third question is whether the number of time points is the same for all subjects. In some longitudinal studies, subjects enter the study at different times or leave early. This creates a design where some subjects have more measurements than others. This is a form of imbalance that repeated measures ANOVA cannot accommodate.
If your data are complete, balanced, and equally spaced, you can proceed to Gate 2. If any of these conditions fail, you should move directly to a mixed-effects model. The Friedman test also requires complete data, so it does not solve the missing data problem. A mixed-effects model can handle missing data and unbalanced designs without dropping subjects, provided the missingness is not informative.
Gate 2: Measurement Scale and Distribution Shape
Gate 2 examines the nature of your outcome variable. The first question is whether the measurement is continuous or ordinal. Continuous measurements can take any value within a range, such as plant height in centimeters, body weight in grams, or enzyme activity in units per milligram. Ordinal measurements are categorical with a natural order, such as a disease severity score from 1 to 5 or a behavioral rating scale.
If your outcome is ordinal, the Friedman test is the appropriate choice. Repeated measures ANOVA assumes a continuous outcome. Applying ANOVA to ordinal data treats the categories as if they were equally spaced, which is rarely justified. The Friedman test ranks the data within each subject, which is the correct way to handle ordinal measurements.
If your outcome is continuous, you need to examine the distribution shape. Create a histogram or a Q-Q plot of the residuals from a preliminary model, or examine the distribution of the outcome at each time point. The key question is whether the distribution is approximately symmetric and bell-shaped, or whether it is heavily skewed.
For continuous data that are approximately normal, repeated measures ANOVA is a reasonable choice. For continuous data that are heavily skewed, you have three options. You can transform the data, such as a log transformation for count data or a square root transformation for variance that increases with the mean. You can use the Friedman test, which does not require normality. Or you can use a generalized linear mixed model with an appropriate distribution, such as a negative binomial for count data or a gamma distribution for continuous positive data.
The choice between these three options depends on your research question. If you need to estimate the magnitude of the effect in the original units, a transformation or a generalized linear mixed model is preferable. The Friedman test only tells you whether the distributions differ, not how much they differ.
Gate 3: Sphericity Assessment for Parametric Path
Gate 3 applies only if you have passed Gate 1 and Gate 2 and are considering repeated measures ANOVA. The critical assumption to check is sphericity. Sphericity means that the variances of the differences between all pairs of conditions are equal. In practical terms, the correlation between any two time points should be roughly the same.
Mauchly's test is the standard method for assessing sphericity. The test has a known limitation. It is sensitive to sample size. With a small sample, Mauchly's test may fail to detect a real violation. With a large sample, it may detect a trivial violation that has no practical impact on the results.
A more practical approach is to examine the covariance matrix directly. If you have a reasonable number of subjects, you can estimate the covariance matrix and inspect the variances of the differences between pairs of conditions. If the variances are roughly similar, sphericity is likely to hold. If some variances are much larger than others, sphericity is violated.
When sphericity is violated, you have three options. The first is to apply a correction to the degrees of freedom, such as the Greenhouse-Geisser or Huynh-Feldt correction. These corrections adjust the p-value to account for the degree of sphericity violation. The Greenhouse-Geisser correction is more conservative, meaning it is less likely to declare a significant effect. The Huynh-Feldt correction is less conservative and is often preferred when the Greenhouse-Geisser correction is too severe.
The second option is to switch to a mixed-effects model. A mixed-effects model does not require sphericity because it models the correlation structure explicitly. You can choose an unstructured covariance matrix, which makes no assumptions about the pattern of correlation, or a structured covariance matrix, such as an autoregressive structure that assumes correlations decay with time.
The third option is to use the Friedman test. This is appropriate if the data are also non-normal or if you are concerned about the robustness of the ANOVA. The Friedman test does not require sphericity or normality.
Gate 4: Sample Size and Power Considerations
Gate 4 addresses the sample size and the power of the test. The choice between parametric and nonparametric methods is also about assumptions. It is also about the ability to detect the effect you are studying.
The Friedman test ranks the data within each subject. This ranking discards information about the magnitude of the differences. If the differences between conditions are large but the ranks are similar, the test may not detect the effect. For example, if one condition produces values that are consistently higher but the differences are small, the ranks may not reflect the magnitude of the difference. This can reduce the power of the test, especially when the data are approximately normal and the effect is subtle.
The power of the Friedman test relative to repeated measures ANOVA depends on the distribution of the data. When the data are approximately normal, the Friedman test is less powerful than the ANOVA. The loss of power can be substantial, sometimes requiring a 10 to 15 percent larger sample size to achieve the same power. When the data are heavily skewed or have outliers, the Friedman test can be more powerful than the ANOVA, because the ANOVA is sensitive to outliers and the Friedman test is not.
A practical approach is to consider the sample size in relation to the number of conditions. The Friedman test requires a minimum number of subjects to produce a valid p-value. With a small sample size, the chi-square approximation may not be accurate. In such cases, you should use an exact version of the Friedman test or a permutation-based approach.
A Record System for the Decision Protocol
The decision protocol is only useful if you document your choices. A clear record of the decision process serves several purposes. It helps you justify your choice of test to a reviewer. It allows you to revisit the decision if the data change. It provides a basis for reporting the analysis in your research paper.
The record should include the following elements. First, record the data structure, including the number of subjects, the number of time points, and the number of missing values. Second, record the measurement type and the distribution shape, including the results of any normality tests and the histograms or Q-Q plots. Third, record the sphericity assessment, including the p-value of Mauchly's test and the variances of the differences. Fourth, record the chosen test and the reason for the choice. Fifth, record the results of the test, including the test statistic, the degrees of freedom, and the p-value.
This record can be kept in a simple spreadsheet or a laboratory notebook. The goal is to create a transparent trail that shows how you arrived at your analysis choice. This transparency is important for the reproducibility of your research. The EQUATOR Network provides reporting guidelines that emphasize the importance of documenting the statistical methods and the assumptions.
Troubleshooting Common Decision Failures
The decision protocol is designed to prevent common failures, but some failures still occur. The following troubleshooting guide addresses the most frequent problems.
Failure 1: The data are normal but the Friedman test was used. This is a common failure when researchers default to the nonparametric test without checking the distribution. The result is a loss of power and a failure to detect a real effect. The solution is to check the distribution before choosing the test. If the data are approximately normal, use repeated measures ANOVA or a mixed-effects model.
Failure 2: The data are skewed but the ANOVA is used. This occurs when researchers assume that the central limit theorem will protect them. The central limit theorem applies to the sampling distribution of the mean, but it does not protect against the effect of skew on the F-test. The result can be an inflated Type I error rate. The solution is to examine the distribution and use a transformation or a nonparametric test.
Failure 3: The sphericity violation is ignored. This is the most common failure in repeated measures ANOVA. The researcher runs the ANOVA without checking sphericity, and the result is a false positive. The solution is to always test sphericity and apply a correction when the test is significant.
Failure 4: The missing data are handled by dropping subjects. This is a common failure when the researcher uses repeated measures ANOVA with missing data. The result is a reduced sample size and a loss of power. The solution is to use a mixed-effects model that can handle the missing data.
Failure 5: The Friedman test is used with a small sample size. The chi-square approximation of the Friedman test is not accurate with a small sample size. The result is a p-value that is not reliable. The solution is to use an exact version of the test or a permutation test.
The Role of the Decision Protocol in Research Reporting
The decision protocol is beyond an internal tool. It is also a reporting tool. When you write your research paper, you should describe the decision process that led to your choice of test. This description should include the data structure, the distribution shape, the sphericity assessment, and the sample size considerations.
The EQUATOR Network provides reporting guidelines that can help you structure this description. The guidelines for observational studies and experimental studies include specific items about the statistical methods. You should report the test used, the assumptions checked, and the corrections applied. This transparency is important for the reproducibility of your research.
The Committee on Publication Ethics also emphasizes the importance of transparent reporting. The core practices include the responsible reporting of research. This includes the accurate description of the statistical methods. A clear decision protocol helps you meet this standard.
The Decision Protocol in the Context of Data Management
The decision protocol is also relevant to data management. The NIH Data Management and Sharing Policy requires that you plan for data management and sharing in your grant application. This includes describing how you will handle the data, including the statistical analysis.
The decision protocol can be part of your data management plan. You can describe the steps you will take to choose the appropriate test. This description shows that you have a plan for the analysis and that you are aware of the assumptions. This is important for the review of your grant application. The NIH Grants and Funding guidelines describe the review process and the criteria for a strong application.
The decision protocol also supports the reproducibility of your research. When you share your data, you should also share the analysis code and the decision record. This allows other researchers to verify your analysis and to understand the choices you made. The ORCID identifier can link your data and your publications to your professional record, which supports the transparency of your research.
A Worked Example of the Protocol in Action
To illustrate the protocol, consider a growth study with 12 mice measured at four time points. The outcome is body weight in grams. The data are complete, with no missing values. The time points are equally spaced at weekly intervals.
Gate 1 is passed. The data are complete, balanced, and equally spaced. The researcher can consider repeated measures ANOVA or the Friedman test.
Gate 2 is the distribution shape. The researcher creates a histogram of the body weight at each time point. The distributions are approximately normal, with no severe skew. The researcher can consider repeated measures ANOVA.
Gate 3 is the sphericity assessment. The researcher runs Mauchly's test. The p-value is 0.03, which is significant at the 0.05 level. The sphericity assumption is violated. The researcher applies the Greenhouse-Geisser correction to the degrees of freedom. The corrected p-value is reported.
Gate 4 is the sample size. The researcher has 24 subjects, which is a reasonable sample size for a repeated measures design. The power is adequate for the effect size of interest.
The researcher reports the repeated measures ANOVA with the Greenhouse-Geisser correction. The decision record includes the histogram, the Mauchly test result, and the correction applied. This record is included in the supplementary material of the paper.
When to Escalate to a Statistical Consultant
The decision protocol is designed to be practical, but it has limits. Some situations require the expertise of a statistician. You should escalate to a statistician when you encounter any of the following situations.
First, when the data have a complex correlation structure. If the correlation between time points is not constant and does not follow a simple pattern, a mixed-effects model with a structured covariance matrix may be needed. This requires statistical expertise.
Second, when the data are not normal and a transformation does not help. If the data are heavily skewed and no transformation produces a normal distribution, a generalized linear mixed model may be needed. This requires statistical expertise.
Third, when the sample size is very small. With a small sample size, the assumptions of the tests are difficult to verify. A statistician can help you choose a test that is appropriate for the sample size.
Fourth, when the missing data are not random. If the missingness is related to the outcome, the analysis is more complex. A statistician can help you handle the missing data appropriately.
The decision protocol is a starting point, not a substitute for statistical expertise. When in doubt, consult a statistician. The cost of a consultation is small compared to the cost of an incorrect analysis.
Frequently Asked Questions
What is the difference between repeated measures ANOVA and the Friedman test?
Repeated measures ANOVA is a parametric test that assumes normality and sphericity. The Friedman test is a nonparametric test that does not require these assumptions. The Friedman test ranks the data within each subject, while the ANOVA uses the raw values.
When should I use the Friedman test instead of repeated measures ANOVA?
Use the Friedman test when the data are not normally distributed, when the data are ordinal, or when the assumptions of the ANOVA are not met. The Friedman test is also a good choice when the sample size is small and you cannot verify the assumptions.
What is sphericity and why does it matter?
Sphericity is the assumption that the variances of the differences between all pairs of conditions are equal. It matters because the repeated measures ANOVA is sensitive to violations of this assumption. If sphericity is violated, the test can produce false positives.
How do I test for sphericity?
You can test for sphericity with Mauchly's test. If the test is significant, you should apply a correction to the degrees of freedom, such as the Greenhouse-Geisser or Huynh-Feldt correction.
Can I use the Friedman test for data with missing values?
The Friedman test requires complete data for all subjects. If you have missing values, you may need to drop the subjects or use a mixed-effects model, which can handle missing data.
What is a mixed-effects model and when should I use it?
A mixed-effects model is a statistical model that includes both fixed effects and random effects. It is useful for repeated measures data when the data are unbalanced, have missing values, or have a complex correlation structure. It is more flexible than repeated measures ANOVA.
How do I report the results of a repeated measures analysis?
Report the test statistic, the degrees of freedom, and the p-value. For repeated measures ANOVA, report the sphericity correction if applied. For the Friedman test, report the chi-square statistic and the p-value. For mixed-effects models, report the fixed effects and the confidence intervals.
What are the limitations of the Friedman test?
The Friedman test discards information about the magnitude of the differences, which can reduce the power to detect subtle effects. It also does not provide an estimate of the effect size in the original units. The test only tells you whether the conditions differ, not how much they differ.
Using the Evidence
| Source | Best use in this topic | Important limitation |
|---|---|---|
| Research Methods Resources | official guidance | Check the linked page for current local requirements |
| EQUATOR Network | official guidance | Check the linked page for current local requirements |
| Core Practices | official guidance | Check the linked page for current local requirements |
Related Bioinformatics Guides
- Longitudinal Microbiome Data Analysis: Methods and Best Practices
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Genomic Data Visualization Tools: Choosing and Using Them Effectively
- Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight
- Genomic Diagnostics: Choosing the Right Test for Clinical and Veterinary Applications
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Research Methods Resources. National Library of Medicine.
- EQUATOR Network. EQUATOR Network.
- Core Practices. Committee on Publication Ethics.
- NIH Grants and Funding. National Institutes of Health.
- ORCID for Researchers. ORCID.
- Data Management and Sharing Policy. National Institutes of Health.
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- A Fasting-Mimicking Diet Affects the Inflammatory Response Following Periodontal Treatment: A Multi-centre Feasibility Randomised Controlled Pilot Trial.. Journal of clinical periodontology, 2026.
- Parasympathetic Tone Activity Evaluation to Discriminate Ketorolac and Ketorolac/Tramadol Analgesia Level in Swine.. Anesthesia and analgesia, 2019.
- The effect of ischaemic postconditioning on mucosal integrity and function in equine jejunal ischaemia.. Equine veterinary journal, 2022.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.