# How to Choose the Correlation Structure for Your Repeated Measures Analysis


## Key Takeaways

- The selection of a correlation structure for repeated measures analysis hinges on the temporal spacing of observations and the strength of within-subject correlations, akin to how viral load dynamics (e.g., influenza A) are influenced by sampling intervals and host immune response.
- Model comparison using information criteria like AIC and BIC is crucial, analogous to choosing between diagnostic tests (e.g., ELISA vs. Western blot) based on their sensitivity, specificity, and cost-effectiveness for detecting specific biomarkers.
- Compound Symmetry (CS) is suitable for balanced designs with equal time spacing, mirroring scenarios where all animals in a treatment group receive the same vaccination protocol at fixed intervals, assuming uniform immune response.
- Autoregressive (AR1) structures are appropriate for equally spaced time points where correlations decay geometrically with time distance, comparable to tracking antibody titers after a single vaccination dose, where the immune response wanes predictably over time.
- The Unstructured (UN) correlation structure, while flexible, is best reserved for a small number of time points due to its high parameter burden, similar to how a comprehensive genomic sequencing of a pathogen is feasible for a few isolates but impractical for large-scale surveillance.
- Unequal time spacing necessitates careful consideration, potentially favoring Compound Symmetry or Unstructured models, much like monitoring disease progression in a herd where sampling occurs at irregular intervals dictated by clinical signs or intervention points.

---

## Quick Answer

- Select the correlation structure that matches how observations are spaced in time and how strongly nearby measurements relate to each other.
- Fit candidate models with different structures and compare them using information criteria such as AIC or BIC.
- No single structure works for all datasets, and choosing the wrong one can bias standard errors and change your conclusions.

## At a Glance

| Correlation Structure | Best For | Key Assumption | When to Avoid |
| --- | --- | --- | --- |
| Compound Symmetry (CS) | Balanced designs with equal time spacing | Equal correlation between all pairs of time points | Unequal spacing or when correlations decay with time distance |
| Autoregressive (AR1) | Equally spaced time points | Correlation decays geometrically as time distance increases | Unequal spacing or when correlations do not decay smoothly |
| Unstructured (UN) | Small numbers of time points | No pattern assumed, all correlations estimated freely | Many time points or small sample sizes due to parameter burden |
| Toeplitz | Equally spaced time points | Correlations depend on time lag but not on absolute time | Unequal spacing or when correlations vary by absolute time |

## Understanding Correlation Structures in Repeated Measures

Repeated measures data arise when you observe the same experimental unit multiple times. In animal farming research, this could mean weighing the same pen of pigs at weekly intervals, measuring milk yield from the same cow across lactation months, or tracking egg production from the same flock over several cycles. The defining feature of repeated measures is that observations from the same unit are correlated, meaning they share something in common that makes them more similar to each other than to observations from different units.

The correlation structure is a mathematical description of how those within-unit observations relate to each other. It is a set of assumptions about the pattern of correlation between pairs of measurements taken on the same subject. When you fit a mixed model or a generalized estimating equation to repeated measures data, the correlation structure is a critical component that determines how the model accounts for the non-independence of observations.

The choice of correlation structure matters because it directly affects the standard errors of your fixed-effect estimates. Standard errors are used to calculate p-values and confidence intervals. If the correlation structure is wrong, the standard errors can be too small or too large, leading to conclusions that do not reflect the true uncertainty in your data. In some cases, the point estimates themselves can shift, though the primary impact is on inference.

The National Library of Medicine provides access to authoritative biomedical books and research-method references that cover the foundations of statistical modeling in biological research. These resources can help you understand the broader context of longitudinal data analysis and the role of correlation structures within it.

## Core Principles of Correlation Structure Selection

### The Nature of Within-Subject Correlation

The first principle is that correlation in repeated measures data is a property of the data, not a choice you make arbitrarily. The correlation structure should reflect the actual relationship between observations on the same subject. In animal science, this relationship is often driven by biology. An animal's weight at one time point is correlated with its weight at the next time point because the animal is the same individual with a consistent genetic background, nutritional history, and health status.

The pattern of correlation can take several forms. In some cases, the correlation between any two time points is roughly the same, regardless of how far apart they are. This is the compound symmetry assumption. In other cases, the correlation is stronger for observations that are closer in time and weaker for observations that are farther apart. This is the autoregressive pattern. In still other cases, the correlation pattern is complex and does not follow a simple mathematical rule.

### The Relationship Between Time Spacing and Correlation

The spacing of your time points is a critical determinant of which correlation structure is appropriate. Autoregressive structures assume that the correlation between two observations depends on the time lag between them. This assumption works well when the time points are equally spaced, such as weekly measurements or monthly measurements. When the time points are equally spaced, the lag between any two consecutive observations is the same, and the autoregressive model can be applied directly.

When time points are unequally spaced, the autoregressive structure becomes more complicated. Some software packages can handle continuous-time autoregressive structures that account for the actual time distances between observations. However, these models are more complex and require careful specification. Compound symmetry does not depend on time spacing at all, because it assumes the correlation is the same for all pairs of observations. This can be an advantage when time points are irregular, but it is also a strong assumption that may not hold.

### The Number of Time Points and Parameter Burden

The number of time points in your study affects which correlation structures are feasible. The unstructured correlation structure estimates a separate correlation parameter for every pair of time points. If you have t time points, the unstructured structure requires t times (t minus 1) divided by 2 correlation parameters. For a study with 3 time points, this is 3 parameters. For a study with 10 time points, this is 45 parameters. Estimating many parameters requires a large sample size to achieve stable estimates.

When the number of time points is small, the unstructured structure is often a good choice because it makes no assumptions about the pattern of correlation. When the number of time points is large, the unstructured structure becomes impractical, and you need to use a more parsimonious structure that imposes some assumptions.

## Practical Workflow for Selecting a Correlation Structure

### Step 1: Examine Your Study Design

Start by documenting the structure of your repeated measures data. Record the number of time points, the spacing between time points, and the number of subjects or experimental units. Determine whether the time points are equally spaced or unequally spaced. This information will narrow down the set of correlation structures that are appropriate for your data.

For equally spaced time points, you can consider compound symmetry, autoregressive, Toeplitz, or unstructured structures. For unequally spaced time points, compound symmetry and unstructured structures are more straightforward, while autoregressive structures require special handling.

### Step 2: Plot the Data and Examine Correlation Patterns

Before fitting any models, plot your data to visualize the pattern of correlation. Create spaghetti plots that show the trajectory of each subject over time. This will help you see whether subjects tend to maintain their relative positions over time, which would suggest a strong correlation structure.

You can also compute the empirical correlation matrix from your data. This involves calculating the correlation between observations at each pair of time points. If the correlations are roughly equal across all pairs, compound symmetry may be appropriate. If the correlations decrease as the time lag increases, an autoregressive structure may be appropriate. If the correlations do not follow a clear pattern, an unstructured structure may be needed.

### Step 3: Fit Candidate Models

Fit the same fixed-effects model with different correlation structures. The fixed effects should be the same across all candidate models so that the comparison is only about the correlation structure. Common candidates include compound symmetry, autoregressive, Toeplitz, and unstructured.

Use restricted maximum likelihood (REML) for linear mixed models when comparing correlation structures. REML produces less biased estimates of variance components than maximum likelihood, and it is the recommended method for comparing nested models with the same fixed effects.

### Step 4: Compare Models Using Information Criteria

Information criteria such as Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) provide a way to compare models with different correlation structures. These criteria balance the fit of the model against the number of parameters. Lower values indicate a better trade-off between fit and complexity.

AIC is more suitable when the goal is prediction or when you want to avoid overfitting. BIC penalizes additional parameters more heavily than AIC, so it tends to select simpler models. In practice, you should compare both criteria and look for consistency. If AIC and BIC agree on the best structure, you can be more confident in the choice. If they disagree, you need to consider the context of your study and the consequences of choosing a simpler or more complex structure.

### Step 5: Check the Sensitivity of Your Conclusions

After selecting a correlation structure, check whether your conclusions about the fixed effects are sensitive to the choice of structure. Fit the model with several plausible structures and compare the estimates, standard errors, and p-values for your primary fixed effects. If the conclusions are the same across structures, your results are robust. If the conclusions change, you need to be cautious in your interpretation and report the results from the structure that best fits the data.

## Common Correlation Structures and Their Properties

### Compound Symmetry

Compound symmetry assumes that the variance is the same at each time point and that the correlation between any two observations on the same subject is the same. This means that the correlation between observations at time 1 and time 2 is the same as the correlation between observations at time 1 and time 5. This structure has two parameters: the common variance and the common correlation.

Compound symmetry is a reasonable assumption when the correlation between observations does not depend on the time lag. This can happen when the repeated measurements are exchangeable, meaning that the order of the measurements does not matter. In animal science, this might be the case when the measurements are taken under similar conditions at each time point and the biological process does not have a strong temporal trend.

The main limitation of compound symmetry is that it does not account for the decay of correlation over time. If the correlation between observations decreases as the time between them increases, compound symmetry will be too restrictive and may lead to incorrect standard errors.

### Autoregressive Structure

The autoregressive structure of order 1, often denoted AR(1), assumes that the correlation between two observations on the same subject is a function of the time lag between them. Specifically, the correlation between observations at time t and time t+k is equal to the correlation raised to the power k. This means that the correlation decays geometrically as the time lag increases.

The AR(1) structure has two parameters: the variance and the correlation parameter. It is appropriate when the time points are equally spaced and when the correlation between observations decreases as the time between them increases. This is a common pattern in growth studies, where the weight of an animal at one time point is more strongly correlated with its weight at the next time point than with its weight several time points later.

The main limitation of the AR(1) structure is that it requires equally spaced time points. If the time points are unequally spaced, the AR(1) structure cannot be applied directly. Some software packages offer a continuous-time autoregressive structure that can handle unequally spaced time points, but this is more complex and requires careful specification.

### Toeplitz Structure

The Toeplitz structure assumes that the correlation between observations depends only on the time lag between them, but it does not impose a specific functional form on the correlation. This means that the correlation between observations at lag 1 is a parameter, the correlation at lag 2 is another parameter, and so on. This structure is more flexible than the AR(1) structure because it does not assume that the correlation decays geometrically.

The Toeplitz structure requires the estimation of as many correlation parameters as there are time lags. For a study with t time points, this requires t minus 1 correlation parameters. This is more than the AR(1) structure but less than the unstructured structure. The Toeplitz structure is appropriate when the time points are equally spaced and when the correlation pattern is not known in advance.

### Unstructured Correlation

The unstructured correlation structure makes no assumptions about the pattern of correlation. It allows each pair of time points to have its own correlation parameter. This is the most flexible structure, but it also requires the estimation of the most parameters.

The unstructured structure is appropriate when the number of time points is small and when the pattern of correlation is not known in advance. It is often used as a starting point for analysis because it does not impose any assumptions. However, when the number of time points is large, the unstructured structure can lead to unstable estimates and convergence problems.

### Choosing Between Structures

The choice between these structures depends on the characteristics of your data. If the time points are equally spaced and the correlation decays with time, the AR(1) structure is a good choice. If the time points are equally spaced but the correlation pattern is unknown, the Toeplitz structure provides more flexibility. If the time points are unequally spaced or the number of time points is small, the unstructured structure may be appropriate. If the correlation between observations is constant across time, the compound symmetry structure is the simplest choice.

## Information Criteria for Model Comparison

### Akaike Information Criterion

The Akaike Information Criterion (AIC) is a measure of the relative quality of a statistical model. It is calculated as the negative of twice the log-likelihood plus twice the number of parameters. The AIC provides a trade-off between the fit of the model and the complexity of the model. Lower AIC values indicate a better model.

AIC is useful for comparing models with different correlation structures because it penalizes models with more parameters. This helps to avoid selecting a model that overfits the data. When comparing models with the same fixed effects but different correlation structures, the model with the lowest AIC is preferred.

### Bayesian Information Criterion

The Bayesian Information Criterion (BIC) is similar to AIC but with a stronger penalty for the number of parameters. BIC is calculated as the negative of twice the log-likelihood plus the number of parameters times the natural logarithm of the sample size. The penalty for the number of parameters increases with the sample size, so BIC tends to select simpler models than AIC.

BIC is useful when you want to select a model that is parsimonious and when the sample size is large. When comparing correlation structures, the model with the lowest BIC is preferred. However, BIC can sometimes select a model that is too simple, especially when the sample size is small.

### Using Information Criteria in Practice

When comparing correlation structures, you should fit the same fixed-effects model with each candidate structure and record the AIC and BIC values. The structure with the lowest AIC and BIC is the best among the candidates. If AIC and BIC select different structures, you should consider the context of your study and the consequences of choosing a simpler or more complex model.

Information criteria are relative measures. They can only tell you which model is better among the models you have fitted. They cannot tell you whether the best model is actually a good model. You should also examine the residuals and the fit of the model to ensure that the model is adequate.

## Model Fitting and Estimation Methods

### Restricted Maximum Likelihood

Restricted maximum likelihood (REML) is a method for estimating variance components in mixed-effects models. Unlike maximum likelihood, REML accounts for the loss of degrees of freedom from estimating the fixed effects. This produces less biased estimates of the variance components, especially when the number of subjects is small.

REML is the recommended method for comparing correlation structures when the fixed effects are the same across models. When you use REML, the likelihood values from different models can be compared directly, provided that the fixed effects are identical. If the fixed effects differ, you should use maximum likelihood for model comparison.

### Maximum Likelihood

Maximum likelihood (ML) is a method for estimating model parameters by maximizing the likelihood function. In mixed-effects models, ML estimates of variance components are biased downward because they do not account for the uncertainty in the fixed effects. However, ML is useful for comparing models with different fixed effects.

When you are comparing models with different fixed effects and different correlation structures, you should use ML. When you are comparing models with the same fixed effects but different correlation structures, you should use REML.

### Convergence and Numerical Issues

Some correlation structures can cause convergence problems during model fitting. The unstructured structure is particularly prone to convergence issues when the number of time points is large or when the sample size is small. If the model does not converge, you may need to try a simpler structure or increase the number of iterations.

You should also check for boundary estimates. A correlation parameter that is estimated at the boundary of its allowed range, such as a correlation of 1 or negative 1, may indicate that the model is not appropriate for the data. In such cases, you should consider a different correlation structure.

## Records and Measurements for Correlation Selection

### Data Requirements

To select a correlation structure, you need data that includes repeated measurements on the same subjects. The data should include a subject identifier, a time variable, and the outcome variable. You also need to know the fixed effects that you want to include in the model.

The quality of the correlation structure selection depends on the quality of the data. Missing data can affect the estimation of the correlation structure. If some subjects have missing observations at some time points, the correlation structure may be estimated from a subset of the data. You should examine the pattern of missing data and consider whether it is missing at random or related to the outcome.

### Recording Your Model Selection Process

It is important to record the process of selecting the correlation structure. This includes the candidate structures that you considered, the information criteria for each model, and the final choice. This documentation is important for reproducibility and for reporting in your research.

The EQUATOR Network provides reporting guidelines for various types of research studies. These guidelines can help you report your statistical methods in a transparent and complete way. Following reporting guidelines ensures that other researchers can understand and replicate your analysis.

### Reproducibility and Data Management

Reproducibility is a key principle in scientific research. To make your analysis reproducible, you should document the software and version used for the analysis, the code used to fit the models, and the data used for the analysis. This documentation allows others to verify your results and to apply the same methods to their own data.

The National Institutes of Health Data Management and Sharing Policy describes expectations for managing and sharing research data. This policy applies to NIH-funded research and emphasizes the importance of data management planning and data sharing. Even if your research is not funded by NIH, following these principles can improve the quality and transparency of your work.

## Common Failure Patterns in Correlation Structure Selection

### Using the Same Structure for Every Dataset

A common mistake is to use the same correlation structure for every repeated measures analysis without examining the data. This can lead to incorrect standard errors and misleading conclusions. The correlation structure should be selected based on the characteristics of the data, not on habit or convention.

### Ignoring Unequal Time Spacing

Another common mistake is to use an autoregressive structure when the time points are unequally spaced. The standard AR(1) structure assumes equal spacing, and applying it to unequally spaced data can produce incorrect results. If your time points are unequally spaced, you should use a structure that accounts for the actual time of observations or use a structure that does not depend on time spacing.

### Overfitting with Unstructured Correlation

The unstructured correlation structure is flexible, but it requires many parameters. When the number of time points is large, the unstructured structure can overfit the data and produce unstable estimates. This can lead to standard errors that are too small and p-values that are too significant. You should consider the number of time points and the sample size when choosing between the unstructured structure and a more parsimonious structure.

### Ignoring the Sensitivity of Results

A common failure is to select a correlation structure and then report the results without checking whether the conclusions are sensitive to the choice of structure. If the conclusions change when you use a different structure, you need to report this and interpret your results with caution. The sensitivity of the results to the correlation structure is an important piece of information for the reader.

## Limitations and Interpretation

### The Correlation Structure Is an Assumption

The correlation structure is an assumption about the pattern of correlation in the data. It is not a fact that can be proven. Even when you select a structure using information criteria, the structure is still an approximation of the true correlation pattern. The information criteria can help you choose the best structure among the candidates, but they cannot guarantee that the chosen structure is correct.

### The Impact of the Correlation Structure on Inference

The correlation structure affects the standard errors of the fixed effects. If the correlation structure is misspecified, the standard errors can be biased. This can lead to p-values that are too small or too large. The direction of the bias depends on the nature of the misspecification. In some cases, the bias can be substantial, leading to conclusions that are not supported by the data.

### The Role of the Sample Size

The sample size affects the ability to estimate the correlation structure. With a small sample size, the correlation parameters may be estimated with low precision. This can make it difficult to distinguish between different correlation structures. With a large sample size, the correlation parameters can be estimated more precisely, and the information criteria can be more reliable.

### The Need for Domain Knowledge

The selection of the correlation structure should be informed by domain knowledge. In animal breeding, you know the biology of the animals and the expected pattern of correlation. For example, you may know that the weight of an animal at one time is strongly correlated with its weight at the next time, but the correlation decreases as the time between measurements increases. This knowledge can help you choose a plausible correlation structure and interpret the results of the information criteria.

## Reporting Your Analysis

### Reporting the Correlation Structure

When you report the results of a repeated measures analysis, you should describe the correlation structure that you used. This includes the type of structure, the estimated parameters, and the method used to select the structure. This information is important for the reader to understand the analysis and to assess the validity of the results.

### Reporting the Model Selection Process

You should also report the process of selecting the correlation structure. This includes the candidate structures that you considered, the information criteria for each model, and the final choice. This information allows the reader to see how the model was selected and to evaluate the robustness of the results.

### Reporting the Sensitivity Analysis

If you conducted a sensitivity analysis to check whether the results are robust to the choice of correlation structure, you should report the results of this analysis. This can include the estimates and standard errors for the fixed effects under different correlation structures. This information helps the reader understand the uncertainty in the results.

### The Role of Reporting Guidelines

The EQUATOR Network provides a collection of reporting guidelines for various types of research studies. These guidelines can help you report your analysis in a complete and transparent way. Following the appropriate reporting guideline can improve the quality of your research and make it easier for other researchers to understand and replicate your work.

## Professional Escalation Criteria

### When to Seek Statistical Advice

If you are uncertain about the choice of correlation structure, you should seek advice from a statistician. This is particularly important when the results of your analysis are sensitive to the choice of structure or when the model does not converge. A statistician can help you explore the data, fit the models, and interpret the results.

### When to Consider a Different Modeling Approach

If the correlation structure is difficult to select or if the models do not fit the data well, you may need to consider a different modeling approach. For example, you might consider a generalized estimating equation (GEE) approach, which does not require the specification of the full correlation structure. GEE models are robust to misspecification of the correlation structure, but they require a larger sample size.

### When to Escalate to a Senior Researcher

If the results of your analysis have important implications for animal welfare or for the management of the animals, you should discuss the results with a senior researcher or a veterinarian. The choice of correlation structure can affect the conclusions of the study, and it is important to ensure that the conclusions are supported by the data.

## A Field Decision Framework for Correlation Structure Selection

### The Five Question Screening Method

Before fitting any models, work through a structured set of questions that narrows your candidate correlation structures from the full menu to a shortlist. This framework is designed for farm researchers who need a repeatable process that does not depend on statistical software defaults.

**Question 1: How many time points do you have?**

Count the number of repeated measurements per subject. If you have three or fewer time points, the unstructured correlation structure is often the most defensible choice because it requires only three correlation parameters. If you have four to six time points, you can consider structured options. If you have more than six time points, the unstructured structure becomes impractical because it requires 15 or more correlation parameters, and you should focus on parsimonious structures.

**Question 2: Are the time points equally spaced?**

Measure the actual time intervals between consecutive observations. If the intervals are equal, such as weekly weighings every seven days, then autoregressive and Toeplitz structures are viable candidates. If the intervals are unequal, such as weighing at birth, weaning, and market, you should eliminate the standard autoregressive structure from consideration. Compound symmetry and unstructured structures do not depend on time spacing and remain available.

**Question 3: Does the correlation between observations decay with time distance?**

Plot the empirical correlation between observations at different time lags. Calculate the correlation between observations one time point apart, two time points apart, and so on. If the correlations decrease as the lag increases, an autoregressive or Toeplitz structure is plausible. If the correlations stay roughly constant across all lags, compound symmetry is a better match. If the correlations show no clear pattern, the unstructured structure is the safest starting point.

**Question 4: How many subjects or experimental units do you have?**

The number of animals, pens, or flocks determines how many correlation parameters you can estimate reliably. A common rule is that you need at least 10 subjects per estimated parameter. The unstructured structure with four time points requires six correlation parameters plus variance components, so you would want at least 60 subjects. If you have fewer subjects, you should use a more parsimonious structure.

**Question 5: What does the biology of the animal suggest?**

Use your knowledge of the production system. In growth studies, weight at one time is strongly correlated with weight at the next time because the animal is the same individual. This suggests an autoregressive pattern. In studies where measurements are taken under similar conditions at each time point, such as repeated milk samples under the same diet, compound symmetry may be more plausible. The biology of the animal should guide your candidate set.

### A Practical Decision Table

| Question | Answer | Recommended Structure |
| --- | --- | --- |
| Number of time points | 3 or fewer | Unstructured |
| Number of time points | 4 to 8 | Compare AR1, Toeplitz, and unstructured |
| Number of time points | More than 8 | AR1 or Toeplitz |
| Time spacing | Unequal | Compound symmetry or unstructured |
| Correlation decay | Strong decay with lag | AR1 or Toeplitz |
| Correlation constant | Constant across lags | Compound symmetry |
| Sample size | Fewer than 10 subjects per parameter | Compound symmetry or AR1 |
| Sample size | More than 10 subjects per parameter | Unstructured possible |

### Implementing the Decision Framework

**Step 1: Create a design summary table**

For each dataset, record the number of time points, the actual time intervals, the number of subjects, and the number of observations per subject. This table becomes the basis for your decision.

**Step 2: Compute the empirical correlation matrix**

Use your statistical software to calculate the correlation between observations at each pair of time points. This is a descriptive step that does not require fitting any models. The matrix gives you a visual and numeric view of the correlation pattern.

**Step 3: Apply the five questions**

Work through the questions in order. The answers will narrow your candidate set to two or three structures. For example, a study with five weekly weighings on 50 pigs would allow you to compare AR1, Toeplitz, and unstructured. A study with three measurements at birth, weaning, and market on 30 calves would point to unstructured as the primary candidate.

**Step 4: Fit the candidate models and compare**

Fit the same fixed-effects model with each candidate structure. Record the AIC and BIC values. The structure with the lowest values is the best among the candidates. If AIC and BIC disagree, consider the sample size and the consequences of a more complex model.

**Step 5: Verify with a sensitivity check**

Fit the model with the selected structure and with one alternative structure. Compare the estimates, standard errors, and p-values for the primary fixed effects. If the conclusions are the same, your choice is robust. If they differ, report the results from the best-fitting structure and note the sensitivity.

### Records and Measurements for the Decision Process

Maintain a written record of the decision process for each analysis. This record should include the design summary, the empirical correlation matrix, the candidate structures, the AIC and BIC values for each model, and the final selection. This documentation is essential for reproducibility and for reporting in research papers.

The National Library of Medicine provides access to authoritative biomedical books and research-method references that cover the foundations of statistical modeling in biological research. These resources can help you understand the broader context of longitudinal data analysis and the role of correlation structures within it.

The EQUATOR Network provides reporting guidelines for various types of research studies. These guidelines can help you report your statistical methods in a transparent and complete way. Following reporting guidelines ensures that other researchers can understand and replicate your analysis.

### Troubleshooting the Decision Framework

**Problem: The empirical correlation matrix shows a pattern that does not match any candidate structure**

If the correlations are high at lag 1, drop at lag 2, and then rise again at lag 3, no simple structure will fit well. In this case, the unstructured structure is the safest choice if the sample size allows. If the sample size is too small, consider whether the time points can be combined or whether a different modeling approach such as generalized estimating equations is more appropriate.

**Problem: The information criteria select different structures**

When AIC and BIC disagree, the difference is usually due to the penalty for the number of parameters. BIC penalizes more heavily and tends to select simpler models. If the sample size is large, BIC is more reliable. If the sample size is small, AIC may be more appropriate. You should also consider the consequences of choosing a simpler model. If the simpler model produces standard errors that are very different from the more complex model, you should report the results from the more complex model.

**Problem: The selected structure does not converge**

If the model with the selected structure does not converge, check the data for issues such as missing values or extreme outliers. Try a simpler structure that is still plausible based on the decision framework. If the unstructured structure does not converge, use the Toeplitz or AR1 structure instead.

**Problem: The conclusions change across structures**

This is the most important signal that the correlation structure matters for your study. When the conclusions change, you need to report the results from the best-fitting structure and describe the sensitivity. You should also consider whether the study design is adequate for the research question. If the conclusions are not robust, the study may need more subjects or a different design.

### Welfare and Safety Context

The choice of correlation structure can affect the conclusions of a study that has implications for animal welfare. For example, if you are comparing two feeding treatments and the conclusion about which treatment is better changes with the correlation structure, you need to be cautious before making management decisions. The Committee on Publication Ethics Core Practices emphasize the importance of accurate reporting and the avoidance of misleading conclusions. You should report the sensitivity of your results so that readers can assess the strength of the evidence.

### Professional Escalation Criteria

If the decision framework leads to a structure that produces conclusions that are not robust, or if the model does not converge, you should seek advice from a statistician. This is particularly important when the results have implications for animal welfare or for the management of the animals. A statistician can help you explore the data, fit the models, and interpret the results. If the results are sensitive to the correlation structure, you should discuss the findings with a senior researcher or a veterinarian before making any management decisions.

## Frequently Asked Questions

### What is the difference between compound symmetry and autoregressive correlation?

Compound symmetry assumes that the correlation between any two observations on the same subject is the same, regardless of the time between them. Autoregressive assumes that the correlation decreases as the time between observations increases. The choice between them depends on the pattern of correlation in your data.

### How do I know if my time points are equally spaced?

Time points are equally spaced if the time between consecutive observations is the same for all pairs of consecutive observations. For example, measurements taken at weeks 1, 2, 3, and 4 are equally spaced. Measurements taken at weeks 1, 2, 4, and 8 are not equally spaced.

### Can I use an autoregressive structure with unequally spaced time points?

The standard autoregressive structure assumes equally spaced time points. If your time points are unequally spaced, you can use a continuous-time autoregressive structure, which accounts for the actual time of observations. This structure is more complex and may require specialized software.

### What is the unstructured correlation structure?

The unstructured correlation structure allows each pair of time points to have its own correlation parameter. It makes no assumptions about the pattern of correlation. This structure is flexible but requires many parameters, so it is best used when the number of time points is small.

### How do I compare models with different correlation structures?

Fit the same fixed-effects model with each candidate correlation structure and compare the information criteria. The model with the lowest AIC or BIC is the best among the candidates. You should also check whether the conclusions about the fixed effects are sensitive to the choice of structure.

### What should I do if the model does not converge?

If the model does not converge, you can try a simpler correlation structure, increase the number of iterations, or check the data for issues. The unstructured structure is often the cause of convergence problems when the number of time points is large. You may need to use a more parsimonious structure.

### How does the correlation structure affect my p-values?

The correlation structure affects the standard errors of the fixed effects. If the correlation structure is misspecified, the standard errors can be too small or too large, which can change the p-values. This is why it is important to select a correlation structure that fits the data well.

### Should I always use the unstructured correlation structure?

No. The unstructured structure requires many parameters and can be unstable when the number of time points is large. It is a good choice when the number of time points is small and the pattern of correlation is unknown. For larger numbers of time points, a more structured approach such as autoregressive or Toeplitz may be more appropriate.

## Related Bioinformatics Guides

- [Longitudinal Microbiome Data Analysis: Methods and Best Practices](/knowledge/bioinformatics/longitudinal-microbiome-data-analysis-methods-and-best-practices)
- [Selecting Persistent Identifiers for Research Data: A Decision Framework](/knowledge/bioinformatics/selecting-persistent-identifiers-for-research-data-a-decision-framework)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Microbiome Data Analysis in R: A Practical Guide for Compositional Data](/knowledge/bioinformatics/microbiome-data-analysis-in-r-a-practical-guide-for-compositional-data)
- [Foundation Models in Genetics: Opportunities and Challenges](/knowledge/bioinformatics/foundation-models-in-genetics-opportunities-and-challenges)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Correlation Coefficients for a Study with Repeated Measures.](https://pubmed.ncbi.nlm.nih.gov/32300374). Computational and mathematical methods in medicine, 2020.
- [Physical fitness assessment: an update.](https://pubmed.ncbi.nlm.nih.gov/16700660). Journal of long-term effects of medical implants, 2006.
- [Repeated measures discriminant analysis using multivariate generalized estimation equations.](https://pubmed.ncbi.nlm.nih.gov/34898331). Statistical methods in medical research, 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.