# Mixed-Effects Models for Longitudinal Data: A Practical Guide for Biologists


## Key Takeaways

- Linear mixed-effects models, particularly using `lmer()` in R, are crucial for analyzing longitudinal biological data where subjects are repeatedly measured, as they explicitly account for the non-independence of observations within subjects by modeling both fixed (population-level) and random (subject-specific) effects.
- The choice between random intercepts (modeling varying baseline levels) and random slopes (modeling varying trajectories over time) is a critical modeling decision, best informed by likelihood ratio tests comparing nested models and biological plausibility regarding individual variability in response.
- Mixed-effects models offer robust handling of unbalanced data and missing observations, a significant advantage over traditional methods, but require rigorous assumption checking, including normality of residuals and random effects, and examination of the missing data mechanism (MCAR, MAR, MNAR).
- Data preparation is paramount, necessitating a "long format" structure where each row represents a single observation, with a correctly specified subject identifier as a factor variable, and careful outlier detection through plotting individual trajectories and residual analysis.
- Model selection between competing random effects structures should be guided by likelihood ratio tests for nested models and information criteria like AIC or BIC for non-nested models, balancing model fit with parsimony to avoid overfitting.
- Reporting should meticulously detail the model structure (fixed and random effects formula), fixed effect estimates with standard errors and p-values, random effects variances, and model fit statistics (e.g., AIC, BIC) to ensure reproducibility and interpretability.

---

## Quick Answer

- Fit a linear mixed-effects model with `lmer()` from the lme4 package in R when your biological data contains repeated measurements from the same subjects, specifying subject as a random intercept.
- Decide between random intercepts and random slopes by comparing model fit with likelihood ratio tests and examining whether individual trajectories vary meaningfully across subjects.
- Mixed-effects models handle unbalanced data and missing observations gracefully, but they require careful checking of assumptions including normality of residuals and random effects.

## Understanding Longitudinal Data in Biology

Longitudinal studies track the same biological subjects across multiple time points. This design appears throughout the life sciences, from plant growth experiments measuring height weekly to clinical studies tracking biomarker levels monthly. The defining feature is that each subject contributes multiple observations, creating a hierarchical structure where measurements are nested within individuals.

This nesting violates the independence assumption of ordinary linear regression. When you measure the same mouse five times, those five measurements are correlated because they come from the same animal. Ignoring this correlation leads to standard errors that are too small, which means you may declare effects significant when they are not. Mixed-effects models address this by explicitly modeling both the fixed effects that describe population-level patterns and the random effects that capture subject-specific deviations.

The terminology can be confusing at first. Fixed effects are the parameters you want to estimate and interpret, such as the effect of a drug treatment or the rate of growth over time. Random effects are the subject-specific adjustments that account for the fact that some individuals start higher or grow faster than others. The model treats these individual deviations as coming from a distribution, typically a normal distribution centered at zero.

### Why Ordinary Regression Fails for Repeated Measurements

Consider a simple experiment where you measure the body weight of 20 rats at 5 time points. If you treat all 100 observations as independent, you are assuming that the weight of rat 3 at week 2 tells you nothing about the weight of rat 3 at week 4. This is false. Rat 3 is likely to be consistently heavier or lighter than the group average across all time points.

The consequence of ignoring this correlation is that your standard errors are underestimated. The model thinks it has more independent information than it actually does. This inflates the test statistics and produces p-values that are too small. You may conclude that a treatment works when the evidence is not there.

Mixed-effects models solve this by partitioning the variance into two components. The residual variance captures the measurement error and the within-subject fluctuation around each subject's trajectory. The random effect variance captures the between-subject variability, the degree to which individuals differ from each other in their baseline levels or their trajectories.

### The Hierarchical Structure of Biological Data

Biological data often has multiple levels of nesting. Consider a study of gene expression in plants grown in different greenhouses. Leaves are nested within plants, plants are nested within greenhouses, and greenhouses may be nested within sites. Each level of nesting introduces a potential source of correlation.

Mixed-effects models can accommodate this hierarchy through nested random effects. You can specify random intercepts for greenhouse and for plant within greenhouse. This allows the model to account for the fact that plants in the same greenhouse share environmental conditions, and leaves from the same plant share genetic and developmental factors.

The key decision is which levels to treat as random and which to treat as fixed. The general rule is that the levels of a factor are random if you are interested in the population of possible levels, and fixed if you are interested in the specific levels in your study. For example, if you chose your three greenhouses because they represent the range of conditions in your region, greenhouse is a random effect. If you chose them because they are the only three greenhouses you have access to and you want to compare those specific greenhouses, greenhouse is a fixed effect.

## Core Principles of Linear Mixed-Effects Models

The linear mixed-effects model extends the ordinary linear model by adding random effects to the linear predictor. The general form is written as y = Xβ + Zb + ε, where y is the vector of observations, X is the design matrix for the fixed effects, β is the vector of fixed effect coefficients, Z is the design matrix for the random effects, b is the vector of random effects, and ε is the vector of residuals.

The fixed effects β are the population-level parameters you want to estimate and interpret. The random effects b are assumed to follow a normal distribution with mean zero and covariance matrix G. The residuals ε are assumed to follow a normal distribution with mean zero and covariance matrix R. The total variance of the observations is then ZGZ' + R.

### Fixed Effects and Random Effects

Fixed effects are the parameters that describe the average relationship between predictors and the outcome. In a growth study, the fixed effect for time tells you the average rate of growth across all subjects. The fixed effect for treatment tells you the average difference between the treatment and control groups.

Random effects describe how individual subjects deviate from the population average. A random intercept allows each subject to have their own baseline level. A random slope allows each subject to have its own trajectory over time. The model estimates the variance of these random effects, which tells you how much subjects vary in their baselines and trajectories.

The distinction matters for interpretation. Fixed effects are what you report in your results section. Random effects are the variance components that you report to show how much individual variation exists. The random effects also affect the standard errors of the fixed effects, because they account for the correlation among repeated measurements.

### Random Intercepts

The random intercept model is the simplest mixed-effects model. It assumes that each subject has its own baseline level, but the effect of time is the same for all subjects. The model is written as y_ij = β0 + β1 * time_ij + b_i + ε_ij, where b_i is the random intercept for subject i.

This model is appropriate when you believe that subjects differ in their starting points but respond to time in the same way. For example, if you are measuring the effect of a fertilizer on plant height, you might expect that all plants grow at the same rate but start at different heights due to differences in seed size or initial conditions.

The random intercept variance tells you how much subjects differ in their baseline levels. A large variance means that subjects are quite different from each other at the start of the study. A small variance means that subjects are similar in their baseline levels.

### Random Slopes

The random slope model allows each subject to have its own trajectory. The model is written as:

y_ij = β + β1 * time_ij + b_i + b_i1 * time_ij + ε_ij

where b_i is the random intercept and b_i1 is the random slope for subject i. This model allows each subject to have its own baseline and its own rate of change over time.

The random slope model is appropriate when you believe that subjects respond differently to the passage of time. For example, in a study of tumor growth, some mice may grow tumors quickly while others grow slowly. A random slope captures this variation in growth rates.

The random slope model has more parameters than the random intercept model. It estimates the variance of the random intercepts, the variance of the random slopes, and the covariance between them. The covariance is important because it tells you whether subjects with higher baselines tend to have steeper or shallower slopes.

### Fixed Effects Structure

The fixed effects structure includes all the predictors you are interested in. This can include time, treatment, and their interaction. The interaction term is important because it tells you whether the effect of treatment changes over time.

For example, in a study of a drug treatment, you might have a fixed effect for treatment, a fixed effect for time, and a fixed effect for the treatment by time interaction. The interaction term tells you whether the difference between the treatment and control groups changes over time. If the interaction is significant, the treatment effect is not constant.

The fixed effects structure should be based on your research question and your biological understanding of the system. You should not include every possible interaction just because you can. Each additional parameter reduces the power of your analysis and increases the complexity of interpretation.

## Preparing Your Data for Mixed-Effects Modeling

The quality of your mixed-effects model depends on the quality of your data preparation. Mixed-effects models require data in a specific format, and they are sensitive to missing data and outliers. Careful data preparation will save you time and prevent errors.

### Long Format Data Structure

Mixed-effects models require data in long format, where each row represents one observation. This means that if you have 10 subjects measured at 5 time points, you need 50 rows of data. Each row contains the subject identifier, the time point, the outcome value, and any other predictors.

The subject identifier is critical. It tells the model which observations belong to the same subject. Without a proper subject identifier, the model cannot estimate the random effects. The subject identifier should be a factor variable, not a numeric variable, to prevent the model from treating it as a continuous predictor.

The time variable should be numeric if you are modeling time as a continuous predictor. If you are modeling time as a categorical predictor, it should be a factor. The choice depends on your research question. Continuous time is more powerful if the relationship is linear. Categorical time is more flexible if the relationship is nonlinear.

### Handling Missing Data

Missing data is common in longitudinal studies. Animals die, samples are lost, and equipment fails. Mixed-effects models have a major advantage over repeated measures ANOVA in this regard. They can handle missing data without dropping the entire subject.

The mixed-effects model uses all available data. If a subject has 8 of 10 observations, the model uses those 8 observations. The subject is not dropped from the analysis. This is because the model estimates the random effects based on the available data.

However, the missing data mechanism matters. If the missing data is missing completely at random, the model estimates are unbiased. If the missing data is missing at random, meaning the missingness depends on observed variables, the model estimates are also unbiased. If the missing data is missing not at random, meaning the missingness depends on the unobserved outcome, the model estimates may be biased.

You should examine the pattern of missing data in your study. If subjects with more severe disease are more likely to drop out, the missing data may be missing not at random. In this case, you should consider more advanced methods such as joint modeling or multiple imputation.

### Outliers and Data Quality Checks

Outliers can have a strong influence on mixed-effects model estimates. A single extreme observation can pull the fixed effect estimates and inflate the random effect variances. You should examine the data for outliers before fitting the model.

Plot the data by subject to identify unusual trajectories. A subject with a sudden spike or drop in the outcome may have a measurement error or a genuine biological event. You should investigate the source of the outlier before deciding whether to exclude it.

The model residuals can also be used to identify outliers. After fitting the model, you can examine the residuals for each observation. Observations with large residuals may be outliers. You should investigate these observations and decide whether they are data errors or genuine biological variation.

## Fitting a Linear Mixed-Effects Model in R

The lme4 package in R is the standard tool for fitting linear mixed-effects models. The primary function is lmer(), which fits the model using restricted maximum likelihood or maximum likelihood.

### Installing and Loading Required Packages

You need to install the lme4 package and the lmerTest package. The lmerTest package provides p-values for the fixed effects, which are not provided by default in lme4.

```r
install.packages("lme4")
install.packages("lmerTest")
library(lme4)
library(lmerTest)
```

The lmerTest package uses the Satterthwaite approximation to compute p-values for the fixed effects. This is the standard approach for mixed-effects models.

### The lmer() Function Syntax

The lmer() function takes a formula and a data frame. The formula specifies the fixed effects and the random effects. The random effects are specified in parentheses with a vertical bar separating the random effect from the grouping variable.

```r
model <- lmer(outcome ~ time + treatment + (1 | subject), data = mydata)
```

This formula specifies that the outcome is predicted by time and treatment as fixed effects, and the random intercept for subject. The (1 | subject) term specifies a random intercept for each subject.

To add a random slope for time, you would use:

```r
model <- lmer(outcome ~ time + treatment + (1 + time | subject), data = mydata)
```

This specifies that both the intercept and the slope for time vary by subject.

### Choosing the Random Effects Structure

The choice of random effects structure is a modeling decision. You should start with the maximal random effects structure that is justified by your design. This means including random slopes for all within-subject predictors.

The maximal structure is the random intercept plus random slopes for all within-subject predictors. In a study with time as the only within-subject predictor, the maximal structure is (1 + time | subject). In a study with time and treatment as within-subject predictors, the maximal structure is (1 + time + treatment | subject).

You can then compare the maximal model to simpler models using the likelihood ratio test. The likelihood ratio test compares the fit of two nested models. The test statistic is the difference in the deviance, which follows a chi-square distribution with degrees of freedom equal to the difference in the number of parameters.

```r
model_max <- lmer(outcome ~ time + treatment + (1 + time | subject), data = data)
model_simple <- lmer(outcome ~ time + treatment + (1 | subject), data = data)
anova(model_max, model_simple)
```

If the likelihood ratio test is not significant, the simpler model is preferred. If the test is significant, the more complex model is preferred.

### Fitting the Model

Once you have decided on the random effects structure, you can fit the model. The lmer() function uses maximum likelihood or restricted maximum likelihood. The default is restricted maximum likelihood, which is preferred for estimating the random effects variances.

```r
model <- lmer(outcome ~ time + treatment + (1 + time | subject), data = data)
summary(model)
```

The summary output shows the fixed effects estimates, the standard errors, the t-values, and the p-values. It also shows the random effects variances and the residual variance.

### Checking Model Assumptions

The linear mixed-effects model makes several assumptions. You should check these assumptions before interpreting the results.

The residuals should be normally distributed. You can check this with a histogram or a Q-Q plot of the residuals. The residuals should also be homoscedastic, meaning the variance of the residuals is constant across the fitted values. You can check this with a plot of the residuals against the fitted values.

The random effects should also be normally distributed. You can check this with a Q-Q plot of the random effects. The random effects should be centered at zero and approximately normally distributed.

The model also assumes that the random effects and the residuals are independent. This is a structural assumption that is difficult to check directly.

## Interpreting the Output

### Fixed Effects Estimates

The fixed effects estimates are the population-level effects. The intercept is the expected outcome when all predictors are zero. The coefficient for time is the expected change in the outcome for a one-unit increase in time, holding the other predictors constant. The coefficient for treatment is the expected difference between the treatment and control groups, holding the other predictors constant.

The standard errors of the fixed effects account for the correlation in the data. They are larger than the standard errors from an ordinary linear model that ignores the correlation. This is because the model recognizes that the effective sample size is smaller than the total number of observations.

The p-values from the summary output are computed using the Satterthwaite approximation. This is the recommended approach for mixed-effects models. The p-values are approximate, but they are more accurate than the p-values from an ordinary linear model.

### Random Effects Variances

The random effects variances describe the between-subject variability. The random intercept variance is the variance of the subject-specific baselines. The random slope variance is the variance of the subject-specific slopes.

The random effects variances are important for understanding the data. If the random intercept variance is large, the subjects are quite different in their baseline levels. If the random slope variance is large, the subjects are quite different in their trajectories.

The correlation between the random intercept and the random slope is also informative. A positive correlation means that subjects with higher baselines tend to have steeper slopes. A negative correlation means that subjects with higher baselines tend to have flatter slopes.

### Intraclass Correlation Coefficient

The intraclass correlation coefficient is the proportion of the total variance that is due to the between-subject variation. It is computed as the random intercept variance divided by the sum of the random intercept variance and the residual variance.

The ICC ranges from 0 to 1. An ICC of 0 means that there is no between-subject variation, and all the variation is within-subject. An ICC of 1 means that there is no within-subject variation, and all the variation is between-subject.

The ICC is a useful measure of the degree of correlation in the data. A high ICC means that the repeated measurements are highly correlated, and the mixed-effects model is necessary. A low ICC means that the repeated measurements are not highly correlated, and the mixed-effects model may not be necessary.

## Worked Example: Plant Growth Over Time

To illustrate the process, consider a study of plant growth. You have 20 plants, and you measure the height of each plant at 5 time points. The plants are randomly assigned to a treatment or control group. The data is in long format with columns for plant ID, time, treatment, and height.

### Data Preparation

The data is already in long format. Each row represents one measurement of one plant. The plant ID is a factor variable. The time is a numeric variable. The treatment is a factor variable with two levels.

```r
head(data)
```

The data should be checked for missing values and outliers. The summary of the data can be used to check the range of the variables.

### Fitting the Random Intercept Model

The first model is a random intercept model. This model assumes that all plants have the same growth rate, but they have different baseline heights.

```r
model_ri <- lmer(height ~ time + treatment + (1 | plant), data = data)
summary(model_ri)
```

The summary shows the fixed effects estimates. The intercept is the expected height at time zero for the control group. The coefficient for time is the expected change in height per unit time. The coefficient for treatment is the expected difference between the treatment and control groups.

The random effects section shows the variance of the random intercepts and the residual variance. The ICC is the random intercept variance divided by the total variance.

### Fitting the Random Slope Model

The second model is a random slope model. This allows each plant to have its own growth rate.

```r
model_rs <- lmer(height ~ time + treatment + (1 + time | plant), data = data)
summary(model_rs)
```

The summary shows the fixed effects estimates and the random effects. The random effects section shows the variance of the random intercepts, the variance of the random slopes, and the correlation between them.

### Comparing the Models

The two models are compared with the likelihood ratio test.

```r
anova(model_ri, model_rs)
```

The output shows the log-likelihood of each model and the likelihood ratio test statistic. If the p-value is less than 0.05, the random slope model is preferred.

### Interpreting the Results

The fixed effects estimates are interpreted in the context of the biological question. The treatment effect is the difference between the treatment and control groups. The time effect is the growth rate. The interaction between treatment and time, if included, is the difference in growth rates between the treatment and control groups.

The random effects variances describe the between-plant variation. The random intercept variance is the variation in baseline heights. The random slope variance is the variation in growth rates.

## Model Selection and Comparison

### Likelihood Ratio Tests

The likelihood ratio test is the primary tool for comparing nested models. A model is nested in another model if the simpler model is a special case of the more complex model. For example, the random intercept model is nested in the random slope model because the random slope model reduces to the random intercept model when the random slope variance is zero.

The likelihood ratio test statistic is the difference in the log-likelihoods of the two models. The test statistic follows a chi-squared distribution with degrees of freedom equal to the difference in the number of parameters. The p-value is the probability of observing a test statistic as large as the one observed if the simpler model is true.

The likelihood ratio test is valid when the models are fitted with maximum likelihood. When the models are fitted with restricted maximum likelihood, the likelihood ratio test is only valid for comparing models with the same fixed effects structure. To compare models with different fixed effects structures, you must fit the models with maximum likelihood.

### Akaike Information Criterion

The Akaike Information Criterion is a model selection criterion that balances the fit of the model with the complexity of the model. The AIC is calculated as -2 times the log-likelihood plus 2 times the number of parameters. The model with the lower AIC is preferred.

The AIC is useful for comparing models that are not nested. The AIC can be used to compare models with different fixed effects structures and different random effects structures. The AIC is a relative measure, so it is only useful for comparing models fitted to the same data.

### Bayesian Information Criterion

The Bayesian Information Criterion is similar to the AIC, but it penalizes the number of parameters more heavily. The BIC is calculated as -2 times the log-likelihood plus the number of parameters times the natural logarithm of the sample size. The model with the lower BIC is preferred.

The BIC tends to select simpler models than the AIC. The BIC is useful when you want to avoid overfitting the data. The BIC is also useful for comparing models that are not nested.

## Common Failure Patterns and How to Avoid Them

### Singular Fit

A singular fit occurs when the model estimates a variance component to be zero. This can happen when the random effect variance is very small or when the model is overparameterized. The model output will show a warning that the model is singular.

A singular fit can be caused by having too few subjects or too few observations per subject. It can also be caused by having a random effect that does not vary. If the random slope variance is estimated to be zero, the random slope is not needed.

The solution is to simplify the random effects structure. Remove the random effect that is estimated to be zero. The model should be refitted without the problematic random effect.

### Convergence Warnings

Convergence warnings occur when the model does not converge to a stable solution. The model may be overparameterized, or the data may not support the model. The warning indicates that the model estimates may not be reliable.

The first step is to try to refit the model with a different optimizer. The lme4 package provides several optimizers. The model can be refit with the optimizer argument.

If the model still does not converge, the random effects structure should be simplified. The model may have too many random effects for the data. The model should be refit with a simpler random effects structure.

### Overparameterization

Overparameterization occurs when the model has more parameters than the data can support. This can lead to unstable estimates and convergence problems. The model should be simplified by removing unnecessary fixed effects or random effects.

The fixed effects structure should be based on the biological hypothesis. The random effects structure should be based on the design of the study. The model should not include every possible interaction or every possible random effect.

## Reporting Results from Mixed-Effects Models

### Reporting the Model Structure

The model structure should be reported in the methods section. The fixed effects and the random effects should be described. The model should be described in a way that allows the reader to understand the analysis.

The model can be described in words or with the model formula. The model formula is the most precise way to describe the model. The formula should be included in the methods section.

### Reporting the Fixed Effects

The fixed effects estimates should be reported with their standard errors and p-values. The estimates should be reported in the context of the biological question. The treatment effect should be reported as the difference between the treatment and control groups.

The confidence intervals for the fixed effects should also be reported. The confidence intervals can be computed with the confint() function in R. The confidence intervals provide a range of plausible values for the fixed effects.

### Reporting the Random Effects

The random effects variances should be reported. The random intercept variance and the random slope variance should be reported. The correlation between the random intercepts and the random slopes should also be reported.

The ICC should be reported if it is relevant. The ICC provides a measure of the degree of correlation in the data.

### Reporting the Model Fit

The model fit should be reported. The log-likelihood, the AIC, and the BIC should be reported. The model fit statistics allow the reader to compare the model with other models.

The residuals should be checked and reported if there are any issues. The assumptions of the model should be checked and reported.

## At a Glance

| Model Type | Random Structure | When to Use | Key Output |
| --- | --- | --- | --- |
| Random Intercept | (1 \| subject) | Subjects differ in baseline but respond to time similarly | Fixed effects for population average, random intercept variance for between-subject baseline variation |
| Random Slope | (1 + time \| subject) | Subjects differ in both baseline and trajectory over time | Fixed effects for population average, random slope variance for between-subject trajectory variation |
| Nested Random Effects | (1 \| group/subject) | Subjects are nested within higher-level groups | Fixed effects for population average, variance components for each level of nesting |

## Practical Implementation Steps

### Step 1: Examine the Data Structure

Before fitting any model, examine the data. Check the number of subjects, the number of observations per subject, and the pattern of missing data. Plot the data to see the trajectories of individual subjects.

### Step 2: Decide on the Fixed Effects

Decide on the fixed effects based on your biological hypothesis. Include the predictors that are relevant to your research question. Include the interaction terms that are relevant.

### Step 3: Decide on the Random Effects Structure

Decide on the random effects structure based on the design of the study. Include random intercepts for the subjects. Include random slopes for the within-subject predictors if the design supports them.

### Step 4: Fit the Model

Fit the model using the lmer() function. Check the summary output for convergence warnings and singular fit warnings.

### Step 5: Check the Model Assumptions

Check the assumptions of the model. Examine the residuals for normality and homoscedasticity. Examine the random effects for normality.

### Step 6: Interpret the Results

Interpret the fixed effects estimates in the context of the biological question. Interpret the random effects variances in the context of the between-subject variation.

### Step 7: Report the Results

Report the model structure, the fixed effects estimates, the random effects variances, and the model fit. Report the results in a way that allows the reader to reproduce the analysis.

## Records and Measurements

### Data Documentation

Document the data collection process. Record the number of subjects, the number of time points, and the number of observations per subject. Record the pattern of missing data.

### Model Documentation

Document the model fitting process. Record the model formula, the optimizer, and the convergence status. Record the model selection process.

### Analysis Documentation

Document the analysis process. Record the software version, the package version, and the date of the analysis. Record the random seed if the analysis involves any random processes.

## Common Failure Patterns

### Failure to Account for Correlation

The most common failure is to ignore the correlation in the data. This leads to standard errors that are too small and p-values that are too low. The mixed-effects model is the solution to this problem.

### Failure to Check Assumptions

The second common failure is to ignore the assumptions of the model. The model assumes normality of the residuals and the random effects. The model assumes homoscedasticity of the residuals. These assumptions should be checked.

### Failure to Report the Model Structure

The third common failure is to report the model structure. The reader cannot evaluate the analysis if the model structure is not reported. The model formula should be reported in the methods section.

## Limitations of Mixed-Effects Models

### Assumption of Normality

The linear mixed-effects model assumes that the random effects and the residuals are normally distributed. This assumption may not hold for all biological data. If the data is not normal, the model estimates may be biased.

### Assumption of Linearity

The linear mixed-effects model assumes that the relationship between the predictors and the outcome is linear. If the relationship is nonlinear, the model may not fit the data well. The model can be extended to include nonlinear terms.

### The Complexity of the Model

The mixed-effects model is more complex than the ordinary linear model. The model requires careful specification of the random effects structure. The model can be difficult to fit and interpret.

## Professional Escalation Criteria

### When to Consult a Statistician

You should consult a statistician if you are unsure about the random effects structure. You should consult a statistician if the model does not converge. You should consult a statistician if the model assumptions are not met.

### When to Consider a Different Model

You should consider a different model if the data is not normal. You should consider a different model if the relationship is not linear. You should consider a different model if the missing data is not missing at random.

## Frequently Asked Questions

### What is the difference between a fixed effect and a random effect?

A fixed effect is a parameter that describes the population-level relationship between a predictor and the outcome. A random effect is a parameter that describes the individual-level deviation from the population average. Fixed effects are the parameters you want to estimate and interpret. Random effects are the parameters that account for the correlation in the data.

### How do I decide whether to include a random slope?

You should include a random slope if you believe that the effect of the within-subject predictor varies across subjects. You can test this by comparing the model with the random slope to the model without the random slope using the likelihood ratio test. If the test is significant, the random slope is supported.

### What is the intraclass correlation coefficient?

The intraclass correlation coefficient is the proportion of the total variance that is due to the between-subject variation. It is calculated as the random intercept variance divided by the sum of the random intercept variance and the residual variance. The ICC ranges from 0 to 1.

### How do I handle missing data in a mixed-effects model?

The mixed-effects model can handle missing data without dropping the entire subject. The model uses all the available data. However, the missing data mechanism matters. If the missing data is missing not at random, the model estimates may be biased.

### What is the difference between maximum likelihood and restricted maximum likelihood?

Maximum likelihood estimates the fixed effects and the random effects variances simultaneously. Restricted maximum likelihood estimates the random effects variances after accounting for the fixed effects. Restricted maximum likelihood is the default in the lme4 package because it produces less biased estimates of the random effects variances.

### How do I report the results of a mixed-effects model?

You should report the model structure, the fixed effects estimates, the standard errors, the p-values, and the confidence intervals. You should also report the random effects variances and the model fit statistics. The model formula should be included in the methods section.

### What should I do if the model does not converge?

If the model does not converge, you should try to refit the model with a different optimizer. If the model still does not converge, you should simplify the random effects structure. The model may have too many random effects for the data.

### What is the difference between a mixed-effects model and a repeated measures ANOVA?

A mixed-effects model is more flexible than a repeated measures ANOVA. The mixed-effects model can handle unbalanced data, missing data, and continuous predictors. The repeated measures ANOVA requires balanced data and categorical time points.

## Using the Evidence

| Source | Best use in this topic | Important limitation |
|---|---|---|
| [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books) | official guidance | Check the linked page for current local requirements |
| [EQUATOR Network](https://www.equator-network.org/) | official guidance | Check the linked page for current local requirements |
| [Core Practices](https://publicationethics.org/core-practices) | official guidance | Check the linked page for current local requirements |

## Related Bioinformatics Guides

- [Gene Set Enrichment Analysis in R: A Practical Tutorial for Interpreting Omics Data](/knowledge/bioinformatics/gene-set-enrichment-analysis-in-r-a-practical-tutorial-for-interpreting-omics-data)
- [FAIR Data Maturity Model: A Practical Assessment Framework for Bioinformatics Workflows](/knowledge/bioinformatics/fair-data-maturity-model-a-practical-assessment-framework-for-bioinformatics-workflows)
- [Metabolomics Data Analysis in R: A Practical Workflow](/knowledge/bioinformatics/metabolomics-data-analysis-in-r-a-practical-workflow)
- [Microbiome Data Analysis in R: A Practical Guide for Compositional Data](/knowledge/bioinformatics/microbiome-data-analysis-in-r-a-practical-guide-for-compositional-data)
- [Longitudinal Microbiome Data Analysis: Methods and Best Practices](/knowledge/bioinformatics/longitudinal-microbiome-data-analysis-methods-and-best-practices)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [Research Methods Resources](https://www.ncbi.nlm.nih.gov/books). National Library of Medicine.
- [EQUATOR Network](https://www.equator-network.org/). EQUATOR Network.
- [Core Practices](https://publicationethics.org/core-practices). Committee on Publication Ethics.
- [NIH Grants and Funding](https://grants.nih.gov/). National Institutes of Health.
- [ORCID for Researchers](https://info.orcid.org/researchers). ORCID.
- [Data Management and Sharing Policy](https://sharing.nih.gov/data-management-and-sharing-policy). National Institutes of Health.
- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Tutorial on Biostatistics: Longitudinal Analysis of Correlated Continuous Eye Data.](https://pubmed.ncbi.nlm.nih.gov/32744149). Ophthalmic epidemiology, 2021.
- [Application of mixed-effects models to study the country-specific outpatient antibiotic use in Europe: a tutorial on longitudinal data analysis.](https://pubmed.ncbi.nlm.nih.gov/22096069). The Journal of antimicrobial chemotherapy, 2011.
- [Avoiding bias in mixed model inference for fixed effects.](https://pubmed.ncbi.nlm.nih.gov/21751227). Statistics in medicine, 2011.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.