# Generalized Linear Models: Definition and Examples

Generalized linear models (GLMs) are a family of regression models that extend ordinary linear regression to outcomes that are not continuous and normally distributed, such as counts, proportions and binary results. Instead of modeling the mean directly, a GLM models a function of the mean as a linear combination of predictors. That function is the link function, and it is the single idea that makes the whole family work.

## Quick Answer

- A generalized linear model has three parts: a random component (the distribution of the outcome), a systematic component (the linear predictor $\eta = X\beta$) and a link function $g(\mu) = \eta$ that connects them [1].
- The link function lets the expected mean be non-linearly related to the predictors while the coefficients stay linear in the model [2].
- Common examples are logistic regression for binary outcomes, Poisson regression for counts and gamma regression for positive skewed data [2].
- GLMs are usually estimated by maximum likelihood, not by ordinary least squares [2].
- Use a GLM when your outcome is bounded, a count, or has a variance that changes with its mean. Use ordinary linear regression when the outcome is continuous and roughly normal with constant variance.

## What Generalized Linear Models Mean

In plain terms, a generalized linear model is a regression model that keeps the familiar "predictors times coefficients" structure but changes two things: the distribution it assumes for the outcome, and the scale on which the mean is predicted.

The precise definition is this. A GLM specifies that each observation $y_i$ is drawn independently from a distribution in the exponential family, with mean $\mu_i$. A linear predictor $\eta_i = \beta_0 + \beta_1 x_{i1} + \dots + \beta_p x_{ip}$ is formed from the predictors. A monotonic, differentiable link function $g$ connects the two so that $g(\mu_i) = \eta_i$ [1]. The exponential family includes the normal, binomial, Poisson, gamma and inverse Gaussian distributions, which is why one framework covers so many outcome types [1].

This matters because ordinary least squares restricts the coefficients to have a constant effect on the outcome, while a GLM allows that effect to vary across the range of the predictors [1]. A Poisson model with a log link, for instance, produces a multiplicative effect on the count scale even though the model is additive on the log scale.

## How It Works

Every GLM is built from three components [1]:

1. **Random component.** The conditional distribution of the outcome $Y$, such as Poisson or binomial.
2. **Systematic component.** The linear predictor $\eta = X\beta$.
3. **Link function.** The function $g$ such that $g(\mu) = \eta$, where $\mu = E(Y)$.

The general equation is:

$$g(\mu_i) = \beta_0 + \beta_1 x_{i1} + \dots + \beta_p x_{ip}$$

Each symbol means the following. $\mu_i$ is the expected value of the outcome for observation $i$. $g(\cdot)$ is the link function. $\beta_0$ is the intercept and $\beta_1 \dots \beta_p$ are the coefficients. $x_{ij}$ are the predictor values.

Because the link is invertible, you can always move back to the mean scale:

$$\mu_i = g^{-1}(\eta_i)$$

The choice of link depends on the outcome. For counts, the log link gives $\mu = \exp(\eta)$, which keeps predicted means positive. For binary outcomes, the logit link gives $\mu = 1/(1+e^{-\eta})$, which keeps predicted probabilities between 0 and 1. For positive continuous data, the inverse or log link is common.

Estimation is by maximum likelihood, which finds the coefficient values that make the observed data most probable under the assumed distribution [2]. This is a different criterion from least squares, and it is why GLM outputs include deviance and log-likelihood instead of sums of squared residuals. If you want the underlying machinery, the article on [maximum likelihood estimation](/blog/data-analysis/maximum-likelihood-estimation) walks through it step by step.

## Worked Example

The dataset below records 20 manufacturing batches, each with a temperature in degrees Celsius and the number of defects found in that batch. Defects are counts, so a Poisson GLM with a log link is a natural choice.

| temperature_C | defects | temperature_C | defects |
|---|---|---|---|
| 60 | 2 | 80 | 8 |
| 62 | 3 | 82 | 11 |
| 64 | 2 | 84 | 10 |
| 66 | 4 | 86 | 13 |
| 68 | 5 | 88 | 12 |
| 70 | 4 | 90 | 15 |
| 72 | 6 | 92 | 14 |
| 74 | 7 | 94 | 18 |
| 76 | 6 | 96 | 17 |
| 78 | 9 | 98 | 21 |

**Step 1: Model form.** We assume $\log(\mu_i) = \beta_0 + \beta_1 T_i$ with $y_i \sim \text{Poisson}(\mu_i)$.

**Step 2: Link function.** The link is $g(\mu) = \log(\mu)$, and its inverse is $\mu = \exp(\eta)$.

**Step 3: Fitted coefficients.** Fitting the Poisson GLM gives $\beta_0 = -2.1331$ and $\beta_1 = 0.0530$.

**Step 4: Prediction at 80 °C.** The linear predictor is $\eta = -2.13314 + 0.053015 \times 80 = 2.1080$. Converting back to the count scale gives $\mu = \exp(2.1080) = 8.2321$ expected defects.

**Step 5: Comparison with ordinary least squares.** Fitting the same data with OLS gives $\beta_0 = -27.1504$ and $\beta_1 = 0.4620$, with $R^2 = 0.9455$. The OLS line predicts a constant increase of 0.462 defects per degree, while the Poisson model predicts a constant 5.4 percent increase per degree on the count scale.

**Step 6: Fit statistics.** The Poisson deviance is $D = 3.1050$ and the log-likelihood is $\ell = -40.4827$, giving an AIC of 84.9654. The OLS Gaussian log-likelihood is $\ell = -33.2998$, giving an AIC of 70.5996. The two AICs are not directly comparable here because they come from different likelihoods for different assumed distributions, which is a common trap discussed below.

Here is the code that produced the Poisson fit:

```python
import statsmodels.api as sm
X = sm.add_constant(temp)
pois = sm.GLM(defects, X, family=sm.families.Poisson()).fit()
print(pois.params)  # [-2.1331, 0.0530]
```

Output:

```text
[-2.13314301  0.05301488]
```

## How to Interpret It

Coefficients in a GLM live on the link scale, so their interpretation depends on the link.

For a log link, $\exp(\beta_1)$ is a multiplicative effect. Here $\exp(0.0530) = 1.0544$, so each additional degree of temperature multiplies the expected defect count by about 1.054, a 5.4 percent increase. This interpretation stays valid across the whole range of temperature, which is the practical payoff of the link function.

For a logit link, $\exp(\beta_1)$ is an odds ratio. A coefficient of 0.7 means each one-unit increase in the predictor multiplies the odds of the outcome by about 2.01.

For an identity link, the coefficient is a plain additive change, exactly as in linear regression.

Predicted values should always be reported on the original scale, not the link scale, because that is what a reader can act on. A predicted mean of 8.23 defects is meaningful. A linear predictor of 2.108 is not.

## When to Use It (and when not to)

Use a generalized linear model when the outcome has a distribution that ordinary linear regression handles poorly. The clearest cases are:

- **Binary or proportion outcomes.** Logistic regression is the standard GLM here. See [logistic regression](/blog/data-analysis/logistic-regression-definition-formula) for the full treatment.
- **Count outcomes.** Poisson or negative binomial regression, especially when counts are small or include many zeros [1].
- **Positive skewed continuous outcomes.** Gamma or inverse Gaussian models.
- **Outcomes where the variance grows with the mean.** Poisson variance equals the mean, so a single constant error term is wrong.

Do not reach for a GLM when the outcome is continuous, unbounded and roughly symmetric with constant variance. Ordinary linear regression is simpler, faster and easier to explain, and it will usually be the better model. Also avoid a GLM when your real problem is a nonlinear relationship between a continuous predictor and a continuous outcome. A [quadratic regression](/blog/data-analysis/quadratic-regression-analysis) or another transformed linear model may fit better and stay within the least squares framework.

Before committing to any model, it is worth checking the assumptions you are making. The guide to [model assumptions in statistical analysis](/blog/guides/understanding-model-assumptions-in-statistical-analysis) covers the diagnostics that apply across both frameworks.

## Generalized Linear Models vs Linear Regression

| Feature | Linear regression (OLS) | Generalized linear model |
|---|---|---|
| Outcome distribution | Normal | Any exponential family member [1] |
| Link function | Identity, $g(\mu) = \mu$ | Log, logit, probit, inverse and others |
| Mean-variance link | Constant variance | Variance depends on the mean |
| Estimation | Least squares | Maximum likelihood [2] |
| Typical use | Continuous, unbounded outcomes | Counts, binary, proportions, skewed data |
| Coefficient meaning | Additive change in the mean | Change on the link scale |

The key structural difference is that OLS is a special case of a GLM with a normal distribution and an identity link. Everything else in the GLM family relaxes one or both of those choices [3].

## Common Mistakes

- **Comparing AIC across different likelihoods.** A Poisson AIC comes from probabilities and a Gaussian AIC comes from densities, so they are not on the same scale. Only compare AIC values between models fitted to the same outcome data whose likelihoods are of the same kind and include all constants.
- **Interpreting coefficients on the original scale.** A coefficient of 0.0530 is not "0.053 more defects." It is a change in the log of the mean. Exponentiate before you interpret.
- **Ignoring overdispersion in count models.** If the residual deviance is much larger than the residual degrees of freedom, the Poisson variance assumption is failing. Switch to a negative binomial or a quasi-Poisson model [1].
- **Using a GLM when OLS is fine.** If the outcome is continuous and roughly normal, the extra complexity buys you nothing and costs interpretability.
- **Forgetting that prediction intervals differ from confidence intervals.** GLMs predict the mean on the link scale. A prediction interval for a single new observation needs the full distribution, not just the fitted mean.
- **Treating prediction and causal explanation as the same task.** The covariates you include, the functional forms you allow and the way you interpret results differ between a model built for prognosis and one built for causal inference [4].

## Limitations

A GLM fixes the relationship between the mean and the variance through the chosen distribution. If the real data have variance that does not follow that rule, the standard errors will be wrong even if the coefficients look reasonable. Overdispersion and underdispersion are the usual culprits in count data.

The link function also imposes a specific shape on the fitted curve. A log link forces the mean to grow multiplicatively, which may not match the data at the extremes. GLMs handle nonlinearity in the mean but they do not automatically handle nonlinearity in the predictors, interactions or unmodeled clustering. For clustered or longitudinal data, methods such as [generalized estimating equations](/knowledge/bioinformatics/generalized-estimating-equations-gee-in-life-sciences-when-and-how-to-use-them) extend the GLM framework to account for within-group correlation. Variable selection in GLMs is also its own problem, and penalized likelihood approaches have been developed specifically for it [5].

## Frequently Asked Questions

### What is the difference between a GLM and a general linear model?

A general linear model is just linear regression written in matrix form, with a normal outcome and an identity link. A generalized linear model allows other outcome distributions and a nonlinear link function. The names sound similar but the scope is very different.

### Do I always need a link function?

Yes. Every GLM has one. In ordinary linear regression the link is the identity function, $g(\mu) = \mu$, which is why it often goes unmentioned. The link is what lets the linear predictor map onto a valid range for the mean.

### How do I choose the right link function?

Start from the range of your outcome. Counts need a link that keeps the mean positive, so log is standard. Binary outcomes need a link that keeps the mean between 0 and 1, so logit or probit. Positive continuous outcomes often use log or inverse. The canonical link for each distribution is a reasonable default.

### Can I use a GLM for count data with many zeros?

You can, but a plain Poisson model may fit poorly if zeros dominate. Options include a negative binomial model for overdispersion or a zero-inflated model when there is a separate process generating the zeros [1]. Check the residual deviance first.

### Is a GLM a machine learning method?

It is a statistical model that can be used for prediction, but the goals differ. A GLM built for prediction and a GLM built for causal explanation require different covariate choices, different evaluation and different interpretation [4]. Treating them as interchangeable leads to errors in both directions.

## References

1. [Topic 4](https://web.pdx.edu/~crkl/ec575/ec575-4.htm)
2. [Generalized Linear Models Spring 2026 | JPSM | Joint Program in Survey Methodology | University of Maryland](https://jpsm.umd.edu/node/4894)
3. [14.1: The CLM and the GLM - Statistics LibreTexts](https://stats.libretexts.org/Courses/Knox_College/Linear_Models_and_Rurita_Kralovstvi/14%3A_Generalized_Linear_Models/14.01%3A_The_CLM_and_the_GLM)
4. [Arnold KF, Davies V, de Kamps M, Tennant PWG, Mbotwa J, Gilthorpe MS. (2021). Reflection on modern methods: generalized linear models for prognosis and intervention-theory, practice and implications for machine learning. International journal of epidemiology](https://pubmed.ncbi.nlm.nih.gov/32380551/)
5. [Li Z, Wang S, Lin X. (2012). Variable selection and estimation in generalized linear models with the smooth <i>L</i><sub>0</sub> penalty. The Canadian journal of statistics = Revue canadienne de statistique](https://pmc.ncbi.nlm.nih.gov/articles/PMC3600656/)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)

## Related Articles

- [Assumptions of Linear Regression: Definition and Examples](/blog/data-analysis/assumptions-of-linear-regression)
- [Quadratic Regression Analysis: Equation and Example](/blog/data-analysis/quadratic-regression-analysis)
- [Bivariate Data: Definition, Examples and Analysis](/blog/data-analysis/bivariate-data-definition-examples)
- [Multivariate Analysis: Definition, Methods and Examples](/blog/data-analysis/multivariate-analysis)
- [Maximum Likelihood Estimation: Definition and Example](/blog/data-analysis/maximum-likelihood-estimation)
- [Understanding Model Assumptions in Statistical Analysis](/blog/guides/understanding-model-assumptions-in-statistical-analysis)
- [Generalized Estimating Equations (GEE) in Life Sciences](/knowledge/bioinformatics/generalized-estimating-equations-gee-in-life-sciences-when-and-how-to-use-them)
- [Simple Linear Regression in Biology: A Step-by-Step Guide with Worked Examples](/knowledge/bioinformatics/simple-linear-regression-in-biology-a-step-by-step-guide-with-worked-examples)