# Likelihood Ratio Test: Definition, Formula and Examples

The likelihood ratio test, often written as the l r test, compares two nested statistical models and asks whether the extra parameters in the larger model earn their keep. You compute the log-likelihood of each model, double the difference, and compare the result to a chi-square distribution. If the gap is large enough, the simpler model is rejected.

## Quick Answer

- The l r test statistic is $LR = 2(\ln L_{full} - \ln L_{reduced})$, where $L_{full}$ is the maximum likelihood of the model with more parameters and $L_{reduced}$ is the maximum likelihood of the nested model with fewer [1].
- The two models must be nested, meaning the reduced model is the full model with some parameters fixed or removed.
- Degrees of freedom equal the difference in the number of estimated parameters between the two models.
- Under the null hypothesis that the reduced model is adequate, the statistic follows a chi-square distribution with those degrees of freedom [1].
- A small p-value means the extra parameters improve fit enough to reject the simpler model.

## What the Likelihood Ratio Test Means

In plain terms, the test asks how much better one model explains your data than a simpler version of itself. If the improvement is tiny, you keep the simpler model. If the improvement is large, the extra terms are doing real work.

The precise definition: the likelihood ratio test evaluates a null hypothesis that restricts some parameters of a model, usually by setting them to zero or to fixed values. You fit the unrestricted model and the restricted model, take the maximum likelihood of each, and form the ratio $\lambda = L_{reduced} / L_{full}$. Because the reduced model cannot fit better than the full model, $\lambda$ lies between 0 and 1. The smaller $\lambda$ is, the less plausible the restriction becomes [1].

The test is general. Many model assumptions, from equal variances to dropped predictors, can be written as restrictions on parameters, and the likelihood ratio test checks whether those restrictions are consistent with the data [1].

## How It Works

The test statistic is built from the ratio of the two likelihoods. Working on the log scale keeps the numbers manageable.

$$\lambda = \frac{L_{reduced}}{L_{full}}$$

$$LR = -2 \ln \lambda = 2(\ln L_{full} - \ln L_{reduced})$$

Each symbol:

- $L_{full}$ is the maximum value of the likelihood when all parameters are free and the maximum likelihood estimates are substituted in [1].
- $L_{reduced}$ is the maximum value of the likelihood when parameters are restricted and reduced in number [1].
- $\lambda$ is the likelihood ratio, bounded between 0 and 1.
- $LR$ is the test statistic, also written as $-2 \ln \lambda$ [1].
- $df$ is the number of parameters lost, equal to $k_{full} - k_{reduced}$ [1].

The NIST handbook states the procedure directly: calculate $\chi^2 = -2 \ln \lambda$ and compare it to a chi-square distribution [1]. The log-likelihood values themselves come from fitting each model, which is the same quantity you would inspect when reading a regression summary. If you want the background on that quantity first, see [What Is Log-Likelihood?](/blog/data-analysis/what-is-log-likelihood).

## Worked Example

The dataset is clinic systolic blood pressure in mmHg for 30 patients, with age and a treatment indicator as predictors. We compare a full model using age and treatment against a reduced model using age alone.

| age | trt | bp | | age | trt | bp |
|---|---|---|---|---|---|---|
| 25 | 0 | 105.75 | | 66 | 1 | 132.30 |
| 32 | 1 | 112.60 | | 67 | 0 | 125.85 |
| 38 | 0 | 111.90 | | 69 | 1 | 135.95 |
| 41 | 1 | 116.55 | | 70 | 0 | 126.50 |
| 45 | 0 | 114.75 | | 72 | 1 | 136.60 |
| 47 | 1 | 122.85 | | 35 | 0 | 110.25 |
| 50 | 0 | 116.50 | | 44 | 1 | 119.20 |
| 52 | 1 | 126.60 | | 49 | 0 | 116.95 |
| 54 | 0 | 117.70 | | 55 | 1 | 128.25 |
| 56 | 1 | 127.80 | | 59 | 0 | 120.45 |
| 58 | 0 | 121.90 | | 62 | 1 | 131.10 |
| 60 | 1 | 128.00 | | 65 | 0 | 124.75 |
| 61 | 0 | 125.55 | | 68 | 1 | 133.40 |
| 63 | 1 | 128.65 | | 71 | 0 | 131.05 |
| 64 | 0 | 126.20 | | 74 | 1 | 134.70 |

Steps with the computed values:

1. Full model log-likelihood for bp ~ age + trt: $ll_{full} = -53.0238$.
2. Reduced model log-likelihood for bp ~ age: $ll_{red} = -78.7672$.
3. LR statistic: $2(-53.0238 - (-78.7672)) = 51.4869$.
4. Degrees of freedom: $3 - 2 = 1$.
5. p-value: $P(\chi^2_{1} > 51.4869) = 0.0000$.

The full model estimates three parameters (intercept, age, treatment). The reduced model estimates two (intercept, age). Dropping treatment costs one degree of freedom.

```python
import statsmodels.api as sm, scipy.stats as st
m_full = sm.OLS(y, sm.add_constant(X_full)).fit()
m_red  = sm.OLS(y, sm.add_constant(X_red)).fit()
LR = 2*(m_full.llf - m_red.llf)
df = X_full.shape[1] - X_red.shape[1]
p  = st.chi2.sf(LR, df)  # LR=51.4869, p=0.0000
```

Output:

```
LR = 51.4869, df = 1, p-value = 0.0000
```

The chi-square distribution with 1 degree of freedom places almost no probability mass beyond 51.4869, so the tail area is effectively zero. Treatment adds explanatory power that age alone does not capture.

## How to Interpret It

The null hypothesis says the reduced model is adequate, meaning the dropped parameters are zero or the restriction holds. A large LR statistic relative to its degrees of freedom produces a small p-value, which is evidence against the null.

Read the result in three parts. First, the size of the statistic. Second, the degrees of freedom, which set the reference distribution. Third, the p-value, which converts the statistic into a probability. In the example, LR = 51.4869 with 1 degree of freedom gives p = 0.0000, so you reject the reduced model.

The direction matters. You are not testing whether the full model is good in absolute terms. You are testing whether the extra parameters improve fit beyond what chance would produce. A model can be a poor fit overall and still beat a nested competitor. For a broader view of how test statistics convert to decisions, see [Test Statistic Formula: How to Calculate and Use It](/blog/guides/test-statistic-formula-how-to-calculate-and-use-it).

## When to Use It (and when not to)

Use the likelihood ratio test when you have two nested models fit by maximum likelihood to the same data, and you want to know whether the larger one is justified. Typical cases include dropping predictors from a regression, testing whether a variance component is zero, and checking constraints in demand systems, where software such as the micEconAids package reports the likelihood ratio chi-squared statistic and its p-value directly [2].

Do not use it when the models are not nested. Comparing two models with different predictors that neither contains the other falls outside this test. Do not use it when the models are fit to different datasets or different sample sizes, because the likelihoods are then not comparable. Do not use it when the fitting method is not maximum likelihood, since the statistic depends on maximum likelihood estimates [1].

The test also requires care with boundary cases. When a restriction sits on the edge of the parameter space, such as testing whether a variance is zero, the usual chi-square reference distribution can be wrong. The NIST handbook notes that the calculations are difficult in general and typically need to be built into software [1].

## Likelihood Ratio Test vs Wald Test

The Wald test is the closest alternative. Both compare a full and a reduced model, but they get there differently. The likelihood ratio test refits the reduced model and compares likelihoods. The Wald test uses only the full model and checks whether the restricted estimates are far from the unrestricted ones relative to their standard errors.

| Feature | Likelihood Ratio Test | Wald Test |
|---|---|---|
| Models fit | Full and reduced | Full only |
| Basis | Difference in log-likelihoods | Distance of estimates from restriction |
| Statistic | $2(\ln L_{full} - \ln L_{reduced})$ | Quadratic form using standard errors |
| Reference distribution | Chi-square, df = parameters lost | Chi-square, same df |
| Behavior near boundary | More stable | Can misbehave |

For large samples the two usually agree. For small samples, or when a parameter estimate is near a boundary, the likelihood ratio test tends to be the safer choice.

## Common Mistakes

- Comparing non-nested models. The test requires that the reduced model be a special case of the full one. Fix: confirm one model is the other with parameters removed or fixed.
- Using the wrong degrees of freedom. The df is the difference in parameter counts, not the number of parameters in either model. Fix: count estimated parameters in each and subtract.
- Forgetting the factor of 2. The statistic is $2(\ln L_{full} - \ln L_{reduced})$, not the raw difference. Fix: double the log-likelihood gap.
- Mixing log-likelihood and likelihood. The formula uses log-likelihoods. Fix: take logs before subtracting, or use the $-2 \ln \lambda$ form [1].
- Comparing models on different samples. Dropped rows change the likelihood scale. Fix: fit both models to the identical dataset.
- Treating a tiny p-value as proof the full model is correct. It only shows the restriction is inconsistent with the data. Fix: check fit diagnostics separately, such as [R and R-Squared](/blog/data-analysis/r-and-r-squared-interpretation).

## Limitations

The test cannot tell you whether either model is a good description of reality. It only ranks two nested candidates. A rejected reduced model does not mean the full model is right, and a retained reduced model does not mean it fits well.

The chi-square approximation relies on large samples. With small samples the p-value can be inaccurate, and the test can be unreliable when parameters sit near the boundary of their allowed range [1]. The procedure also requires maximum likelihood fitting, which needs software support and is not always readily available [1]. If your models are not nested, compare them with an information criterion such as AIC or BIC instead. If you only need a distribution-free comparison of groups rather than of models, consider the [Mann-Whitney U Test](/blog/data-analysis/mann-whitney-u-test-definition-formula-example) for two groups or the [Kruskal-Wallis Test](/blog/data-analysis/kruskal-wallis-test-definition-formula-examples) for more than two groups.

## Frequently Asked Questions

### What is the likelihood ratio test in simple terms?

It compares a complex model to a simpler version of itself. You fit both, measure how much better the complex one explains the data, and decide whether that improvement is larger than chance would produce. The statistic is twice the difference in log-likelihoods.

### What is the formula for the l r test?

The formula is $LR = 2(\ln L_{full} - \ln L_{reduced})$, where the two terms are the maximum log-likelihoods of the full and reduced models [1]. You compare the result to a chi-square distribution with degrees of freedom equal to the number of parameters lost.

### How do I find the degrees of freedom?

Subtract the number of estimated parameters in the reduced model from the number in the full model. In the blood pressure example, the full model has three parameters and the reduced model has two, so df = 1.

### What does a significant result mean?

It means the data are inconsistent with the restriction imposed by the reduced model. You reject the simpler model in favor of the one with more parameters. It does not confirm that the larger model is correct, only that the restriction is not supported.

### Can I use the likelihood ratio test for non-nested models?

No. The test requires nesting, meaning the reduced model must be the full model with some parameters fixed or removed. For non-nested comparisons you need a different criterion, such as an information criterion that does not depend on nesting.

## References

1. [8.2.3.3. Likelihood ratio tests](https://www.itl.nist.gov/div898/handbook/apr/section2/apr233.htm)
2. [R: Likelihood Ratio test for Almost Ideal Demand Systems](https://search.r-project.org/CRAN/refmans/micEconAids/html/lrtest.aidsEst.html)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [Wasserstein RL, Lazar NA (2016). The ASA Statement on p -Values: Context, Process, and Purpose. The American Statistician](https://doi.org/10.1080/00031305.2016.1154108)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)

## Related Articles

- [What Is Log-Likelihood? Definition, Formula and Examples](/blog/data-analysis/what-is-log-likelihood)
- [Maximum Likelihood Estimation: Definition and Example](/blog/data-analysis/maximum-likelihood-estimation)
- [Two Sample t-Test: Formula, Calculation and Example](/blog/data-analysis/two-sample-t-test-formula-example)
- [Mann-Whitney U Test: Definition, Formula and Example](/blog/data-analysis/mann-whitney-u-test-definition-formula-example)
- [Example of a Ratio: Definition and Real Data Examples](/blog/data-analysis/example-of-a-ratio-definition-examples)
- [Statistical Parameter: Definition, Types, and Estimation](/blog/guides/statistical-parameter-definition-types-and-estimation)
- [Statistical Tests: Choosing the Right One for Your Data](/blog/guides/statistical-tests-choosing-the-right-one-for-your-data)
- [Test Statistic Formula: How to Calculate and Use It](/blog/guides/test-statistic-formula-how-to-calculate-and-use-it)