# Adjusted R-Squared vs R-Squared: Differences and Examples

Adjusted R-squared and R-squared are both measures of how well a regression model fits your data. R-squared tells you the proportion of variance in the outcome explained by the predictors, while adjusted R-squared applies a penalty for the number of predictors in the model. The two values are close when you have many observations and few predictors, and they diverge as you add more predictors or work with smaller samples.

## Quick Answer

- R-squared ($R^2$) always increases or stays the same when you add a predictor, even a useless one.
- Adjusted R-squared ($\bar{R}^2$) can decrease when a new predictor does not pull its weight.
- Use $R^2$ to describe how much variance your final model explains.
- Use adjusted $R^2$ to compare models with different numbers of predictors.
- With 30 homes and 3 predictors, $R^2 = 0.9939$ and adjusted $R^2 = 0.9932$. With 6 predictors, $R^2 = 0.9979$ and adjusted $R^2 = 0.9973$.

## Key Differences

| Feature | R-squared | Adjusted R-squared |
|---|---|---|
| What it measures | Share of outcome variance explained | Share of variance explained, penalized for predictor count |
| Effect of adding a predictor | Never decreases | Can decrease |
| Formula | $1 - \frac{SS_{res}}{SS_{tot}}$ | $1 - \frac{(1-R^2)(n-1)}{n-k-1}$ |
| Range | 0 to 1 (for models with an intercept) | Can be negative |
| Best used for | Describing a single model's fit | Comparing models with different predictor counts |
| Sensitive to sample size | Indirectly | Directly, through $n$ and $k$ |
| Typical reporting | "The model explains 99.4% of variance" | "Model 2 has the higher adjusted R-squared" |

## R-Squared Explained

R-squared, also called the coefficient of determination, is the fraction of total variation in the outcome that your model accounts for. If $R^2 = 0.80$, the predictors explain 80% of the variation in the response, and 20% remains unexplained.

The formula is:

$$R^2 = 1 - \frac{SS_{res}}{SS_{tot}}$$

where $SS_{res}$ is the sum of squared residuals and $SS_{tot}$ is the total sum of squares around the mean of the outcome. For a deeper treatment of the underlying idea, see coefficient of determination (R-squared).

The key property is monotonicity. Every time you add a predictor to an ordinary least squares model, $SS_{res}$ can only shrink or stay the same, so $R^2$ can only rise or stay the same. Add a column of random noise and $R^2$ will still tick upward, usually by a tiny amount. That makes raw $R^2$ a poor tool for deciding whether a new variable belongs in the model.

## Adjusted R-Squared Explained

Adjusted R-squared corrects for that inflation by dividing the unexplained variance by the residual degrees of freedom. The formula is:

$$\bar{R}^2 = 1 - \frac{(1 - R^2)(n - 1)}{n - k - 1}$$

where $n$ is the sample size and $k$ is the number of predictors (not counting the intercept). The term $n - k - 1$ is the residual degrees of freedom.

The mechanics are simple. Adding a predictor raises $k$, which lowers $n - k - 1$, which enlarges the penalty fraction. If the new predictor reduces the residual sum of squares by more than the penalty costs, adjusted $R^2$ rises. If it does not, adjusted $R^2$ falls. That is the behavior you want when comparing candidate models.

Adjusted $R^2$ is always less than or equal to $R^2$ for the same model, and the gap widens as $k$ grows relative to $n$. It can even go negative when $R^2$ is smaller than $k/(n-1)$, meaning the predictors explain little relative to their number, which is a useful warning sign. If you want the broader context of how these fit statistics relate to correlation, see R and R-squared interpretation.

## Worked Example

The dataset contains 30 single-family homes with size, bedroom, bathroom, age, garage, lot, and sale price in thousands of dollars. Here are the first ten rows.

| sqft | beds | baths | age | garage | lot | price |
|---|---|---|---|---|---|---|
| 1400 | 2 | 1 | 30 | 1 | 5000 | 210 |
| 1600 | 3 | 2 | 25 | 1 | 5500 | 240 |
| 1700 | 3 | 2 | 20 | 2 | 6000 | 255 |
| 1800 | 3 | 2 | 18 | 2 | 6500 | 270 |
| 1900 | 4 | 2 | 15 | 2 | 7000 | 290 |
| 2000 | 4 | 3 | 12 | 2 | 7500 | 310 |
| 2100 | 4 | 3 | 10 | 2 | 8000 | 330 |
| 2200 | 4 | 3 | 8 | 2 | 8500 | 350 |
| 2300 | 4 | 3 | 6 | 3 | 9000 | 375 |
| 2400 | 5 | 3 | 5 | 3 | 9500 | 400 |

The remaining 20 rows follow the same pattern, with prices rising as square footage rises and age falls.

We fit two models. Model 1 uses three predictors: square footage, bedrooms, and bathrooms. Model 2 adds age, garage, and lot, for six predictors total. The sample size is $n = 30$ in both cases.

**Model 1 (k = 3):**

$$R^2 = 0.9939$$

$$\bar{R}^2 = 1 - \frac{(1 - 0.9939)(30 - 1)}{30 - 3 - 1} = 0.9932$$

**Model 2 (k = 6):**

$$R^2 = 0.9979$$

$$\bar{R}^2 = 1 - \frac{(1 - 0.9979)(30 - 1)}{30 - 6 - 1} = 0.9973$$

**Change from Model 1 to Model 2:**

- $\Delta R^2 = 0.9979 - 0.9939 = 0.0039$
- $\Delta \bar{R}^2 = 0.9973 - 0.9932 = 0.0041$

Both statistics rise, and adjusted $R^2$ rises slightly more. That happens because the three added predictors (age, garage, lot) carry real signal in this dataset, so the reduction in residual variance outweighs the degrees-of-freedom penalty. If those three columns had been random noise, $R^2$ would still have risen slightly, but adjusted $R^2$ would usually have fallen.

Here is the Python code that produced these values.

```python
import statsmodels.api as sm
X1 = sm.add_constant(df[['sqft','beds','baths']])
m1 = sm.OLS(df['price'], X1).fit()
print(f"{m1.rsquared:.4f} {m1.rsquared_adj:.4f}")
X2 = sm.add_constant(df[['sqft','beds','baths','age','garage','lot']])
m2 = sm.OLS(df['price'], X2).fit()
print(f"{m2.rsquared:.4f} {m2.rsquared_adj:.4f}")
```

Output:

```
0.9939 0.9932
0.9979 0.9973
```

The bar chart below compares the two fit statistics across the two models.

Bars: R2 0.9939 vs adj 0.9932 (3 preds), R2 0.9979 vs adj 0.9973 (6 preds).

## Which One Should You Use?

Use $R^2$ when you are describing a single finished model to an audience that wants to know how much variance the model explains. It is intuitive and easy to communicate. "The model explains 99.4% of the variation in sale price" is a sentence most readers understand immediately.

Use adjusted $R^2$ when you are choosing between models that differ in the number of predictors. It is the fairer comparison because it accounts for the cost of complexity. If Model A has 3 predictors and Model B has 6, comparing raw $R^2$ tilts the contest toward Model B before you even look at the data.

In practice, report both. Give $R^2$ so readers know the absolute fit, and give adjusted $R^2$ so they can judge whether the extra predictors earned their place. If you are working across languages and want to know which environment makes this easier, see R vs Python for data analysis.

## Common Mistakes

- **Assuming a high $R^2$ means a good model.** A model can fit the sample well and still predict poorly on new data. Check residual plots and consider cross-validation.
- **Comparing raw $R^2$ across models with different predictor counts.** This always favors the larger model. Use adjusted $R^2$ or another penalized criterion instead.
- **Forgetting that adjusted $R^2$ can be negative.** A negative value means the predictors explain too little variance to justify their number. Do not treat it as an error.
- **Using $k$ to mean the total number of coefficients including the intercept.** The formula uses predictors only. Counting the intercept inflates the penalty and gives the wrong answer.
- **Treating adjusted $R^2$ as a hypothesis test.** It is a descriptive fit measure. It does not tell you whether a predictor is statistically significant.
- **Reporting many decimal places.** Four decimals on a fit statistic implies precision the data rarely supports. Two or three is usually enough.

## Limitations

Neither statistic tells you whether the model is correctly specified. A high $R^2$ or adjusted $R^2$ can coexist with omitted variables, wrong functional form, or heteroscedasticity. Both are also sensitive to outliers, since squared residuals give large errors outsized influence. A single extreme point can move either value substantially.

Adjusted $R^2$ penalizes predictor count but not predictor quality. It does not detect multicollinearity, and it does not tell you which predictors matter. It also assumes you are comparing models on the same outcome variable and the same set of observations. Comparing adjusted $R^2$ across models fit to different samples or different response transformations is meaningless.

## Frequently Asked Questions

### Can adjusted R-squared be higher than R-squared?

No. For the same model, adjusted $R^2$ is always less than or equal to $R^2$. The penalty term $(n-1)/(n-k-1)$ is greater than 1 whenever $k > 0$, so it shrinks the explained fraction. The gap grows as the number of predictors rises relative to the sample size.

### What does a negative adjusted R-squared mean?

It means $R^2$ is smaller than $k/(n-1)$, so the predictors explain less variation than you would expect from that many useless predictors. This typically happens with small samples and many predictors, or with a model that fits very poorly. A negative adjusted $R^2$ is a signal to simplify the model or reconsider the predictors.

### Should I report R-squared or adjusted R-squared in a paper?

Report both when space allows. Give $R^2$ for the reader who wants the absolute share of variance explained, and adjusted $R^2$ for the reader comparing your model to alternatives. If you must pick one for model selection, use adjusted $R^2$.

### Does adjusted R-squared fix overfitting?

It reduces the incentive to add predictors that do not help, but it does not eliminate overfitting. Adjusted $R^2$ is still computed on the same data used to fit the model. For a real check on overfitting, use cross-validation or an independent test set.

### Why did my adjusted R-squared go down when I added a predictor?

The new predictor did not reduce the residual sum of squares enough to offset the degrees-of-freedom penalty. In other words, it added complexity without adding explanatory power. Consider dropping it, or check whether it is correlated with predictors already in the model.

## References

This article draws on the standard references listed under Further Reading.

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [Altman N, Krzywinski M (2015). Simple linear regression. Nature Methods](https://doi.org/10.1038/nmeth.3627)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)
- [Wasserstein RL, Lazar NA (2016). The ASA Statement on p -Values: Context, Process, and Purpose. The American Statistician](https://doi.org/10.1080/00031305.2016.1154108)

## Related Articles

- [R and R-Squared: What They Mean and How to Interpret Them](/blog/data-analysis/r-and-r-squared-interpretation)
- [Coefficient of Determination (R-Squared): Definition and Examples](/blog/data-analysis/coefficient-of-determination-r-squared)
- [Quadratic Regression Analysis: Equation and Example](/blog/data-analysis/quadratic-regression-analysis)
- [Correlation vs Covariance: Differences and When to Use Each](/blog/data-analysis/correlation-vs-covariance)
- [Measures of Variability: Range, Variance and Standard Deviation](/blog/data-analysis/measures-of-variability)
- [Which R-Value Represents the Strongest Correlation? A Guide to Interpreting Correlation Coefficients](/blog/guides/which-r-value-represents-the-strongest-correlation-a-guide-to-interpreting-correlation-coefficie)
- [One-Way vs Two-Way ANOVA: Differences, Assumptions and a Worked Example](/blog/research-skills/one-way-vs-two-way-anova)
- [Standard Deviation vs Variance vs Standard Error: What Each Measures and When to Report It](/blog/research-skills/standard-deviation-vs-variance-vs-standard-error)