# What Are Residuals in Statistics? Definition and Formula

The meaning of residuals in statistics is simple: a residual is the difference between an observed value and the value your model predicts for it. If your regression line predicts sales of 10.0 and the actual figure is 10.1, the residual is 0.1. Residuals are the raw material for judging whether a model fits well or badly.

## Quick Answer

- A residual is $e_i = y_i - \hat{y}_i$, the observed value minus the fitted value.
- Positive residuals mean the model under-predicted. Negative residuals mean it over-predicted.
- Residuals are computed from your sample. Errors are the unobservable deviations from the true population relationship [1].
- Squaring residuals and summing them gives the residual sum of squares (RSS), the quantity least squares regression minimizes.
- A good model leaves residuals scattered randomly around zero with no pattern and roughly constant spread [2].

## What Residuals Mean

In plain terms, a residual tells you how far off your model was for one specific data point. It is the leftover part of the observation that the model did not explain. When you fit a line through a scatterplot, most points do not sit exactly on the line. The vertical gap between a point and the line is its residual.

The precise statistical definition: given a fitted model, the residual for observation $i$ is the difference between the observed response $y_i$ and the fitted (predicted) value $\hat{y}_i$ produced by the model. The National Institute of Standards and Technology describes residuals as estimates of experimental error obtained by subtracting the predicted responses from the observed responses, and as elements of variation unexplained by the fitted model [2].

One distinction matters from the start. The true error is the deviation of an observation from the real population function, and you can never observe it because you do not know the true function. The residual is the deviation from your estimated function, and you can always compute it [1]. People often use "error" and "residual" interchangeably in casual speech, but they are different objects.

## How It Works

For a simple linear regression with one predictor, the fitted line is:

$$\hat{y}_i = b_0 + b_1 x_i$$

The residual is then:

$$e_i = y_i - \hat{y}_i$$

Each symbol means the following.

| Symbol | Meaning |
|---|---|
| $y_i$ | The observed value of the response for observation $i$ |
| $\hat{y}_i$ | The value the fitted model predicts for observation $i$ |
| $e_i$ | The residual for observation $i$ |
| $b_0$ | The estimated intercept of the regression line |
| $b_1$ | The estimated slope of the regression line |
| $x_i$ | The value of the predictor for observation $i$ |

The slope and intercept come from the least squares formulas:

$$b_1 = \frac{S_{xy}}{S_{xx}}, \qquad b_0 = \bar{y} - b_1\bar{x}$$

where $S_{xx} = \sum (x_i - \bar{x})^2$ and $S_{xy} = \sum (x_i - \bar{x})(y_i - \bar{y})$.

Two properties always hold for a model fitted with an intercept by ordinary least squares. The residuals sum to zero, and they are uncorrelated with the fitted values. That is why a residual plot should look like random noise around a horizontal line at zero [1].

To summarize the overall size of the residuals, you square each one and add them up. That gives the residual sum of squares:

$$RSS = \sum_{i=1}^{n} e_i^2$$

Squaring removes the sign, so large positive and large negative residuals both count as bad fits. The mean squared error divides the sum of squared residuals by $n - p - 1$, where $p$ is the number of predictors excluding the intercept, which produces an unbiased estimate of the error variance [1].

## Worked Example

The dataset below records advertising spend and sales, both in thousands of dollars, across 8 months.

| advertising_spend_k | sales_k |
|---|---|
| 1 | 2.1 |
| 2 | 3.9 |
| 3 | 6.2 |
| 4 | 7.8 |
| 5 | 10.1 |
| 6 | 12.2 |
| 7 | 13.8 |
| 8 | 16.1 |

**Step 1: Summary statistics.** With $n = 8$, the mean of $x$ is 4.5000 and the mean of $y$ is 9.0250. The sums of squares are $S_{xx} = 42.0000$ and $S_{xy} = 83.9000$.

**Step 2: Slope.** $b_1 = 83.9000 / 42.0000 = 1.9976$.

**Step 3: Intercept.** $b_0 = 9.0250 - 1.9976 \times 4.5000 = 0.0357$.

**Step 4: Fitted line.** $\hat{y} = 0.0357 + 1.9976x$.

**Step 5: Residuals.** Subtract each fitted value from the observed value.

| $x$ | $y$ | $\hat{y}$ | $e = y - \hat{y}$ |
|---|---|---|---|
| 1 | 2.1000 | 2.0333 | 0.0667 |
| 2 | 3.9000 | 4.0310 | -0.1310 |
| 3 | 6.2000 | 6.0286 | 0.1714 |
| 4 | 7.8000 | 8.0262 | -0.2262 |
| 5 | 10.1000 | 10.0238 | 0.0762 |
| 6 | 12.2000 | 12.0214 | 0.1786 |
| 7 | 13.8000 | 14.0190 | -0.2190 |
| 8 | 16.1000 | 16.0167 | 0.0833 |

The residuals switch sign with no clear pattern and stay small, which is what you want to see.

**Step 6: Residual sum of squares.** $RSS = 0.1948$.

**Step 7: R-squared.** The total sum of squares is $TSS = 167.7950$, so $R^2 = 1 - 0.1948/167.7950 = 0.9988$. About 99.88% of the variation in sales is explained by advertising spend.

Here is the same calculation in Python.

```python
import numpy as np
x = np.array([1,2,3,4,5,6,7,8], dtype=float)
y = np.array([2.1,3.9,6.2,7.8,10.1,12.2,13.8,16.1])
b1, b0 = np.polyfit(x, y, 1)
yhat = b0 + b1 * x
resid = y - yhat
RSS = float(np.sum(resid**2))
print(round(b0, 4), round(b1, 4), round(RSS, 4))  # intercept, slope, residual sum of squares
```

Output:

```text
0.0357 1.9976 0.1948
```

## How to Interpret It

Start with the sign. A positive residual means the observed value came in above the prediction, so the model under-predicted that point. A negative residual means the model over-predicted it.

Then look at the size. A residual of 0.0667 on a response measured in thousands of dollars is about $67 of unexplained variation, which is tiny next to sales values in the thousands. The same residual on a response measured in single dollars would be enormous. Always judge residuals against the scale of the response.

Then look at the pattern. Plot the residuals against the fitted values or against the predictor. If the linear model is appropriate, the points should scatter randomly about zero with no trend [1]. A curved pattern suggests the true relationship is quadratic or higher order. A funnel shape, where the spread grows as fitted values grow, indicates heteroscedasticity. If the spread stays constant, you have homoscedasticity [1]. A single point far from the rest is a candidate outlier worth investigating. For a fuller treatment of these patterns, see [Residual Plots: How to Interpret Them with Examples](/blog/data-analysis/residual-plots-how-to-interpret).

## When to Use It (and when not to)

Use residuals whenever you fit a model and want to check whether it is trustworthy. They are the standard diagnostic for linear regression, analysis of variance, and design of experiments. NIST calls examining residuals a key part of all statistical modeling, because careful inspection tells you whether your assumptions are reasonable and your model choice is sound [2].

Do not use residuals as a substitute for thinking about the problem. A model can have small residuals and still be wrong if you extrapolate far outside the observed range of the predictor. Do not compare residuals across models fitted to different response scales without standardizing them first. And do not treat the residual for one point as evidence about the population error at that point. It is a single sample deviation, not the true error [1].

## Residual vs Error

These two terms get mixed up constantly. The table below separates them.

| Aspect | Residual | Error |
|---|---|---|
| Definition | Observed value minus fitted value | Observed value minus the true population value |
| Observable? | Yes, you compute it from your data | No, the true function is unknown [1] |
| Depends on | Your sample and your fitted model | The population relationship |
| Notation | Usually $e_i$ | Usually $\varepsilon_i$ |
| Use | Model diagnostics, RSS, $R^2$ | Theoretical properties of estimators |

The mean squared error of a regression is computed from the sum of squares of the residuals, not of the unobservable errors. Dividing that sum by $n - p - 1$ instead of $n$ removes the bias and gives an unbiased estimate of the variance of the unobserved errors [1].

## Common Mistakes

- **Subtracting in the wrong order.** The formula is observed minus predicted, not the reverse. Flipping the order flips every sign and reverses your reading of under- and over-prediction. Fix: write $e_i = y_i - \hat{y}_i$ at the top of your work and check one point by hand.
- **Reporting residuals without squaring when summarizing.** Positive and negative residuals cancel to zero by construction, so the raw sum is always zero and tells you nothing. Fix: use the residual sum of squares or the mean squared error. See [Residual Sum of Squares: Formula and Example](/blog/research-skills/residual-sum-of-squares-formula-and-example) for the full breakdown.
- **Judging fit from $R^2$ alone.** A high $R^2$ does not prove the model is correct. Fix: always inspect a residual plot alongside $R^2$.
- **Ignoring the scale of the response.** A residual of 5 is trivial for sales in millions and severe for sales in units. Fix: compare residuals to the standard deviation of the response, which you can get from [Mean and Standard Deviation: Definition, Formula and Examples](/blog/data-analysis/mean-and-standard-deviation).
- **Treating every large residual as a data entry error.** Some large residuals are real and informative. Fix: investigate before deleting, and report what you removed.
- **Confusing residuals with errors.** They are different quantities with different properties. Fix: use "residual" for anything computed from your fitted model.

## Limitations

Residuals only tell you about the data you have. They cannot detect a problem that your sample does not cover, such as a nonlinear relationship that only appears at predictor values you never observed. They also cannot tell you which of several competing models is scientifically correct. They only show which one fits this sample better.

Residual analysis is also partly subjective. Deciding whether a pattern is "random enough" depends on judgment and on the sample size. With very few observations, residual plots are hard to read and formal tests have little power. With very many, trivial departures from assumptions can look statistically significant while having no practical effect on your conclusions. The residual sum of squares alone is not sufficient for choosing the most appropriate model, because it never increases as you add more predictors [3].

## Frequently Asked Questions

### What is the meaning of residuals in simple terms?

A residual is how far off a prediction was for one data point. You take the actual value, subtract the value your model predicted, and the difference is the residual. Positive means the model guessed too low, negative means it guessed too high.

### How do you calculate a residual?

Fit your model first to get the predicted values. Then subtract each predicted value from the matching observed value using $e_i = y_i - \hat{y}_i$. In the worked example above, the first observation had $y = 2.1000$ and $\hat{y} = 2.0333$, giving a residual of 0.0667.

### What does a residual squared tell you?

Squaring a residual removes its sign and gives more weight to large misses. Summing the squared residuals produces the residual sum of squares, the quantity that least squares regression minimizes. In the example, the eight squared residuals add to 0.1948.

### What is a good residual value?

There is no universal threshold. A good residual is small relative to the scale of the response and shows no pattern when plotted. Residuals should scatter randomly around zero with roughly constant spread across all fitted values [2].

### Can residuals be negative?

Yes. A negative residual simply means the observed value fell below the model's prediction. In the example, three of the eight residuals are negative. The signs are expected to balance out, since residuals from a least squares fit with an intercept always sum to zero [1].

## References

1. [Errors and residuals - Wikipedia](https://en.wikipedia.org/wiki/Errors_and_residuals)
2. [5.2.4. Are the model residuals well-behaved?](https://www.itl.nist.gov/div898/handbook/pri/section2/pri24.htm)
3. [1.3.5.18.1. Defining Models and Prediction Equations](https://www.itl.nist.gov/div898/handbook/eda/section3/eda35i1.htm)

## Further Reading

- [R: Compute Weighted Residuals](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/weighted.residuals.html)
- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [Altman N, Krzywinski M (2015). Simple linear regression. Nature Methods](https://doi.org/10.1038/nmeth.3627)

## Related Articles

- [Residual Plots: How to Interpret Them with Examples](/blog/data-analysis/residual-plots-how-to-interpret)
- [Sample Mean: Definition, Formula and Examples](/blog/data-analysis/sample-mean)
- [Mean and Standard Deviation: Definition, Formula and Examples](/blog/data-analysis/mean-and-standard-deviation)
- [Formula for Range: Definition, Formula and Examples](/blog/data-analysis/formula-for-range)
- [Parameter Definition in Statistics: Meaning and Examples](/blog/data-analysis/parameter-definition-statistics)
- [Residual Sum of Squares: Formula and Example](/blog/research-skills/residual-sum-of-squares-formula-and-example)