# OLS Regression: What Ordinary Least Squares Means

Ordinary least squares (OLS) is the standard method for fitting a straight line to data. It chooses the slope and intercept that make the sum of the squared vertical distances between the observed points and the line as small as possible. That single rule is why OLS regression is the default in most statistical software, and why its output looks the way it does.

## Quick Answer

- OLS picks the one line that minimizes the sum of squared residuals, where a residual is the vertical gap between a data point and the fitted line [1].
- The fitted line is written $\hat{y} = b_0 + b_1 x$, with $b_0$ the intercept and $b_1$ the slope.
- Squaring the residuals does two things: it makes positive and negative gaps count equally, and it punishes large misses far more than small ones.
- The slope tells you how much the predicted outcome changes for a one-unit increase in the predictor, holding everything else in the model fixed [2].
- $R^2$ reports the share of the outcome's variation the model explains, and it is not a measure of whether the model is appropriate [3].

## What OLS Means

In plain terms, OLS is a rule for drawing the best-fitting straight line through a cloud of points. "Least squares" names the rule: among all possible lines, pick the one whose squared errors are smallest.

The precise definition is a minimization problem. For each observation $i$, define the residual as the difference between the observed value $y_i$ and the value the line predicts, $\hat{y}_i$. OLS chooses $b_0$ and $b_1$ to minimize

$$SSE = \sum_{i=1}^{n} (y_i - \hat{y}_i)^2 = \sum_{i=1}^{n} \left(y_i - b_0 - b_1 x_i\right)^2$$

The word "ordinary" distinguishes this from other least-squares variants. Weighted least squares gives some points more influence than others, and nonlinear least squares fits curves instead of straight lines [4]. Ordinary least squares treats every observation equally and fits a linear model.

## How It Works

Minimizing SSE has a closed-form solution, so no search is needed. The two formulas are built from sums of deviations:

$$b_1 = \frac{S_{xy}}{S_{xx}}, \qquad b_0 = \bar{y} - b_1 \bar{x}$$

Each symbol means the following.

| Symbol | Meaning |
|---|---|
| $n$ | Number of observations |
| $x_i, y_i$ | The predictor and outcome values for observation $i$ |
| $\bar{x}, \bar{y}$ | The sample means of the predictor and outcome |
| $S_{xx}$ | $\sum (x_i - \bar{x})^2$, the total squared spread of the predictor |
| $S_{xy}$ | $\sum (x_i - \bar{x})(y_i - \bar{y})$, how the two variables move together |
| $b_1$ | Slope, the change in predicted $y$ per one-unit change in $x$ |
| $b_0$ | Intercept, the predicted $y$ when $x = 0$ |
| $SSE$ | Sum of squared residuals, the quantity OLS minimizes |
| $SST$ | $\sum (y_i - \bar{y})^2$, total variation in the outcome |
| $SSR$ | $SST - SSE$, variation explained by the line |

The slope formula has a useful reading. $S_{xy}$ measures how much $x$ and $y$ rise and fall together, and $S_{xx}$ measures how much $x$ moves at all. Dividing one by the other converts that shared movement into units of $y$ per unit of $x$. The intercept then shifts the line so it passes through the point $(\bar{x}, \bar{y})$, which every OLS line does.

## Worked Example

The dataset is monthly advertising spend and sales, in thousands of dollars, for 8 months.

| ad_spend | sales |
|---|---|
| 1 | 14 |
| 2 | 17 |
| 3 | 19 |
| 4 | 23 |
| 5 | 24 |
| 6 | 28 |
| 7 | 30 |
| 8 | 33 |

Working through the steps:

- $n = 8$
- $\bar{x} = 36.0 / 8 = 4.5000$
- $\bar{y} = 188.0 / 8 = 23.5000$
- $S_{xx} = 42.0000$
- $S_{xy} = 113.0000$
- $b_1 = 113.0000 / 42.0000 = 2.6905$
- $b_0 = 23.5000 - 2.6905 \times 4.5000 = 11.3929$
- $SSE = 1.9762$
- $SST = 306.0000$
- $SSR = 306.0000 - 1.9762 = 304.0238$
- $R^2 = 304.0238 / 306.0000 = 0.9935$
- Pearson $r = 0.9968$

The fitted line is $\hat{y} = 11.3929 + 2.6905x$. Each extra thousand dollars of ad spend is associated with about 2.69 thousand dollars more in sales. The line explains 99.35% of the variation in sales.

Here is the same fit in Python with statsmodels:

```python
import statsmodels.api as sm
X = sm.add_constant(ad_spend)
model = sm.OLS(sales, X).fit()
print(model.params, model.rsquared)
```

Output:

```
intercept=11.3929, slope=2.6905, R2=0.9935
```

Excel returns the same values with `SLOPE=2.6905`, `INTERCEPT=11.3929`, and `RSQ=0.9935`. If you want to reproduce this without writing code, the [linear regression calculator](/tools/linear-regression-calculator) takes a pair of columns and returns the coefficients.

## How to Interpret It

Read the slope first. It is a rate of change in the outcome's units per one unit of the predictor. A slope of 2.6905 means predicted sales rise by 2.6905 thousand dollars for each additional thousand dollars of ad spend. The sign tells you the direction of the relationship.

Read the intercept second, and carefully. It is the predicted outcome when the predictor equals zero. That value is only meaningful if zero is inside the range of real data. Here, an ad spend of zero is plausible, so 11.3929 thousand dollars is a reasonable baseline prediction. In many datasets zero is impossible, and the intercept is just a positioning constant.

Read the fit statistics third. $R^2$ of 0.9935 means the line accounts for 99.35% of the variation in sales around its mean. The Pearson correlation of 0.9968 measures the strength of the linear association on its own scale. These two numbers are closely related in simple regression, and the [difference between R and R-squared](/blog/data-analysis/r-and-r-squared-interpretation) matters when you report them. If you add predictors, compare [adjusted R-squared against R-squared](/blog/data-analysis/adjusted-r-squared-vs-r-squared) instead of trusting the raw value.

Finally, check the residuals. Plot them against the fitted values and look for curvature, fanning, or a few points far from the rest. Those patterns are what the [assumptions of linear regression](/blog/data-analysis/assumptions-of-linear-regression) are about, and they decide whether the coefficients mean what you think they mean [3].

## When to Use It (and when not to)

Use OLS when the outcome is continuous, the relationship is roughly linear, and the residuals behave reasonably well. It is the right default for a first pass at almost any numeric outcome, and it stays useful when you add predictors, because the interpretation of each coefficient as a partial effect is straightforward [2]. For a broader orientation, see [what regression analysis is](/blog/guides/what-is-regression-analysis-a-practical-introduction).

Do not use OLS when the outcome is a count, a proportion bounded between 0 and 1, or a category. Those outcomes violate the constant-variance and normality ideas that make OLS inference work, and a [generalized linear model](/blog/data-analysis/generalized-linear-models-explained) is usually the better fit. Do not use OLS when the relationship is clearly curved, since a straight line will leave systematic structure in the residuals. A [quadratic regression](/blog/data-analysis/quadratic-regression-analysis) or another curved form handles that case.

## OLS vs Correlation

Correlation and OLS answer different questions about the same two variables.

| Aspect | OLS regression | Pearson correlation |
|---|---|---|
| Question answered | How much does $y$ change when $x$ changes? | How strongly do $x$ and $y$ move together? |
| Output | Slope and intercept with units | A single number from -1 to 1 |
| Direction | Directional, $x$ predicts $y$ | Symmetric, order does not matter |
| Units | Depends on the units of $x$ and $y$ | Unit-free |
| Prediction | Produces predicted values | Produces no predictions |

In simple regression the two are linked: $r^2$ equals $R^2$, and the slope equals $r$ times the ratio of the standard deviations. The correlation tells you the strength, and the regression tells you the size and the prediction.

## Common Mistakes

- **Reading the intercept as a real-world baseline when zero is out of range.** If your predictor never approaches zero, the intercept is an extrapolation. Report it as a positioning constant, or center the predictor so the intercept becomes the predicted value at the mean.
- **Treating a high $R^2$ as proof the model is correct.** A high value says the line tracks the data, not that the right variables are in the model. Always inspect residuals before trusting the fit [3].
- **Ignoring outliers.** Least squares squares every residual, so one or two extreme points can pull the line noticeably [4]. Fit the model with and without the suspect points and compare the slopes.
- **Confusing correlation with causation.** A slope of 2.6905 does not mean ad spend caused the sales increase. It describes an association in these 8 months.
- **Extrapolating beyond the observed range.** The line was fit between ad spend of 1 and 8. Predicting sales at an ad spend of 20 assumes the straight-line pattern continues, which nothing in the data confirms.
- **Dropping the units from your write-up.** A slope of 2.6905 is meaningless without "thousands of dollars of sales per thousand dollars of ad spend."

## Limitations

OLS describes a linear relationship in the data you gave it. It cannot tell you whether the predictor causes the outcome, whether an omitted variable is doing the real work, or whether the pattern holds outside the observed range. It also gives every observation equal weight, so a single unusual point can shift the coefficients more than you expect [4].

The method is also sensitive to how the model is specified. If the true relationship is curved, OLS will still return a slope, and that slope will be a misleading average of a pattern it cannot capture. Alternatives such as symmetric or rank-based regression exist precisely because OLS can be thrown off by outliers and heavy-tailed errors [5]. Treat the coefficients as a summary of the data in front of you, and check the diagnostics before you treat them as a description of the world.

## Frequently Asked Questions

### What does "least squares" actually minimize?

It minimizes the sum of the squared vertical distances between each observed point and the fitted line. Those distances are called residuals. Squaring removes the sign and gives large misses more weight, which is why a single far-off point can move the line.

### Why square the residuals instead of taking absolute values?

Squaring produces a smooth function with a closed-form minimum, so the slope and intercept can be written as formulas rather than found by search. It also matches the normal-distribution assumption used for standard errors and p-values. Absolute-value methods exist and are more resistant to outliers, but they need iterative fitting.

### Is OLS the same as linear regression?

In practice, yes, when people say "linear regression" they usually mean ordinary least squares. Strictly, linear regression names the model form, a straight-line relationship, and OLS names the estimation method used to fit it. Other methods can fit the same linear model.

### What is a good R-squared value?

It depends entirely on the field and the question. In controlled physical measurements, values above 0.99 are common, as in the example here. In social science data with human behavior, 0.30 can be a meaningful result. A low $R^2$ does not automatically mean the model is useless, and a high one does not mean it is correct [3].

### Can OLS handle more than one predictor?

Yes. Multiple linear regression extends the same least-squares rule to several predictors at once, and each coefficient then reports the effect of that predictor while the others are held fixed [2]. The geometry changes, but the objective, minimizing squared residuals, does not.

## References

1. [Altman N, Krzywinski M (2015). Simple linear regression. Nature Methods](https://doi.org/10.1038/nmeth.3627)
2. [Krzywinski M, Altman N (2015). Multiple linear regression. Nature Methods](https://doi.org/10.1038/nmeth.3665)
3. [Altman N, Krzywinski M (2016). Regression diagnostics. Nature Methods](https://doi.org/10.1038/nmeth.3854)
4. [4.1.4.2. Nonlinear Least Squares Regression](https://www.itl.nist.gov/div898/handbook/pmd/section1/pmd142.htm)
5. [Greco L, Luta G, Krzywinski M et al. (2025). Symmetric alternatives to the ordinary least squares regression. Nature Methods](https://doi.org/10.1038/s41592-025-02760-w)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)

## Related Articles

- [Assumptions of Linear Regression: Definition and Examples](/blog/data-analysis/assumptions-of-linear-regression)
- [Quadratic Regression Analysis: Equation and Example](/blog/data-analysis/quadratic-regression-analysis)
- [Explanatory Variable: Definition, Examples and Role in Regression](/blog/data-analysis/explanatory-variable)
- [What Is Y-Hat? Predicted Values in Regression Explained](/blog/data-analysis/y-hat-predicted-values)
- [R and R-Squared: What They Mean and How to Interpret Them](/blog/data-analysis/r-and-r-squared-interpretation)
- [What is Regression Analysis? A Practical Introduction](/blog/guides/what-is-regression-analysis-a-practical-introduction)
- [Regression Analysis Biostatistics](/blog/guides/regression-analysis-biostatistics)
- [Multiple Linear Regression for Biologists](/knowledge/bioinformatics/multiple-linear-regression-for-biologists-handling-confounders-and-interactions)