# Coefficient of Determination (R-Squared): Definition and Examples

The coefficient of determination, written R-squared or $R^2$, is a number between 0 and 1 that tells you how much of the variation in your outcome variable a regression model explains. If $R^2 = 0.85$, the model accounts for 85% of the variation in the outcome. The remaining 15% is left unexplained by the predictors you used.

## Quick Answer

- The coefficient of determination is the proportion of total variation in the dependent variable that the regression model explains.
- It is computed as $R^2 = 1 - \frac{SS_{res}}{SS_{tot}}$, where $SS_{res}$ is the residual sum of squares and $SS_{tot}$ is the total sum of squares.
- Values run from 0 (the model explains nothing) to 1 (the model explains everything).
- In simple linear regression, $R^2$ equals the square of the Pearson correlation $r$ between the two variables.
- A high $R^2$ means the model fits the data you have. It does not prove the model is correct or that the predictors cause the outcome.

## What the Coefficient of Determination Means

In plain terms, the coefficient of determination answers one question: how much better does my model describe the data than just using the average? If you had to predict every outcome value without any predictors, your best guess would be the mean of the outcome. A regression model tries to do better by using the predictor variables. $R^2$ measures how much of that improvement you actually captured.

The precise statistical definition: $R^2$ is the fraction of the total sum of squares (the squared deviations of the observed values from their mean) that is accounted for by the regression sum of squares. Equivalently, it is one minus the fraction of total variation that remains in the residuals. Linear least squares regression fits parameters by minimizing the sum of squared residuals, so $R^2$ is a natural summary of how well that fit performs [1].

People also search for this as the "determination coefficient" or the "correlation of determination." Both refer to the same quantity. The standard name is the coefficient of determination.

## How It Works

The formula is built from three sums of squares.

$$R^2 = 1 - \frac{SS_{res}}{SS_{tot}} = \frac{SS_{reg}}{SS_{tot}}$$

Each symbol means the following.

- $SS_{tot} = \sum (y_i - \bar{y})^2$ is the total sum of squares. It measures how much the observed values $y_i$ vary around their mean $\bar{y}$.
- $SS_{res} = \sum (y_i - \hat{y}_i)^2$ is the residual sum of squares. It measures the variation left over after fitting, where $\hat{y}_i$ is the predicted value.
- $SS_{reg} = SS_{tot} - SS_{res}$ is the regression sum of squares, the part of the variation the model explains.
- $R^2$ is the ratio of explained variation to total variation.

Because $SS_{res}$ can never be negative and can never exceed $SS_{tot}$ when the model includes an intercept, $R^2$ falls between 0 and 1. In simple linear regression with one predictor, $R^2$ is exactly $r^2$, the square of the Pearson correlation coefficient between the predictor and the outcome.

## Worked Example

The dataset below records study hours and exam scores for 8 students.

| Hours | Score |
|-------|-------|
| 1 | 52 |
| 2 | 58 |
| 3 | 63 |
| 4 | 68 |
| 5 | 72 |
| 6 | 78 |
| 7 | 85 |
| 8 | 90 |

The steps below use these values.

- Sample size: $n = 8$
- Mean study hours: $\bar{x} = 4.5000$
- Mean exam score: $\bar{y} = 70.7500$
- Fitted intercept: $b_0 = 46.6429$
- Fitted slope: $b_1 = 5.3571$
- Regression equation: $\hat{y} = 46.6429 + 5.3571 \cdot x$
- Total sum of squares: $SS_{tot} = 1209.5000$
- Residual sum of squares: $SS_{res} = 4.1429$
- Regression sum of squares: $SS_{reg} = SS_{tot} - SS_{res} = 1205.3571$
- Coefficient of determination: $R^2 = 1 - \frac{4.1429}{1209.5000} = 0.9966$
- Correlation between hours and scores: $r = 0.9983$
- Check: $r^2 = 0.9966$, which matches $R^2$

So the model explains 99.66% of the variation in exam scores. Only about 0.34% of the variation is left in the residuals.

You can reproduce this in Python with statsmodels or in Excel with the RSQ function.

```python
import pandas as pd
import statsmodels.api as sm
df = pd.DataFrame({'hours': [1, 2, 3, 4, 5, 6, 7, 8],
                   'scores': [52, 58, 63, 68, 72, 78, 85, 90]})
X = sm.add_constant(df['hours'])
model = sm.OLS(df['scores'], X).fit()
print(f"OLS R-squared = {model.rsquared:.4f}")  # 0.9966
```

Output:

```
OLS R-squared = 0.9966
```

In Excel, `=RSQ(B2:B9, A2:A9)` with scores in B2:B9 and hours in A2:A9 also returns 0.9966, so the manual, Python and Excel routes agree. If you want to see how the slope and intercept are estimated before $R^2$ is computed, the article on [OLS regression](/blog/data-analysis/ols-regression-ordinary-least-squares) walks through that step.

## How to Interpret It

Read $R^2$ as a percentage of explained variation, not as a grade for your model.

- $R^2 = 0.9966$ means the predictors account for 99.66% of the variation in the outcome. The fit is very tight.
- $R^2 = 0.50$ means half the variation is explained and half is not. Whether that is good depends on the field.
- $R^2 = 0$ means the model does no better than predicting the mean every time.

Context sets the bar. In a controlled physics experiment, you might expect $R^2$ above 0.99. In social science data with noisy human behavior, an $R^2$ of 0.30 can be a meaningful result. The number alone does not tell you if the model is useful.

When you add predictors to a model, $R^2$ can only stay the same or rise, even if the new predictors are useless. That is why analysts often report the [adjusted R-squared](/blog/data-analysis/adjusted-r-squared-vs-r-squared), which penalizes the count of predictors. For a single predictor, the two are close, and the gap grows as you add more.

## When to Use It (and when not to)

Use $R^2$ when you want a single summary of how well a linear model fits the outcome, and when you are comparing models that predict the same outcome on the same data.

Do not use it as your only measure. A high $R^2$ does not confirm that the [assumptions of linear regression](/blog/data-analysis/assumptions-of-linear-regression) hold. Residual plots, influence checks, and domain reasoning still matter [2]. Do not compare $R^2$ across datasets with different outcomes or different scales. Do not treat a low $R^2$ as proof that the predictors are unrelated to the outcome, since a nonlinear relationship can produce a weak linear fit. If the relationship curves, a model such as [quadratic regression](/blog/data-analysis/quadratic-regression-analysis) may describe it better.

## Coefficient of Determination vs Correlation Coefficient

The two are closely related but not the same thing.

| Feature | Coefficient of determination ($R^2$) | Correlation coefficient ($r$) |
|---------|--------------------------------------|-------------------------------|
| Range | 0 to 1 | -1 to 1 |
| Sign | Always non-negative | Positive or negative |
| Meaning | Share of outcome variation explained | Strength and direction of a linear relationship |
| Relationship | $R^2 = r^2$ in simple linear regression | $r = \pm\sqrt{R^2}$ |
| Use | Regression fit quality | Association between two variables |

In the worked example, $r = 0.9983$ and $R^2 = 0.9966$. The correlation tells you the relationship is strong and positive. The coefficient of determination tells you the model explains 99.66% of the variation. If you need to describe the spread of a variable on its own, the [coefficient of variation](/blog/data-analysis/coefficient-of-variation-formula) is a better tool, and you can compute it with the [Coefficient of Variation Calculator](/tools/coefficient-of-variation-calculator).

## Common Mistakes

- **Treating a high $R^2$ as proof of a good model.** A model can fit well and still violate assumptions or be misspecified. Fix: check residual plots and diagnostics alongside $R^2$ [2].
- **Comparing $R^2$ across different outcomes.** A model predicting income and a model predicting test scores cannot be ranked by $R^2$. Fix: compare only models with the same outcome and the same observations.
- **Assuming a low $R^2$ means no relationship.** A curved relationship can give a low linear $R^2$ even when the variables are strongly linked. Fix: plot the data before trusting the number.
- **Reading causation into $R^2$.** A high $R^2$ says nothing about cause and effect. Fix: keep causal claims tied to study design, not fit statistics.
- **Forgetting that $R^2$ never falls when you add predictors.** Adding noise variables can nudge $R^2$ upward. Fix: use adjusted $R^2$ when comparing models with different numbers of predictors.
- **Confusing $R^2$ with $r$.** The correlation can be negative, but $R^2$ cannot. Fix: report the one that matches your question.

## Limitations

$R^2$ summarizes fit for the data at hand. It says nothing about whether the model will predict new data well, and it does not detect overfitting. A model with many predictors can post a high $R^2$ on the sample it was fit to and still generalize poorly.

It also cannot tell you which predictors matter, whether the relationship is linear, or whether the errors behave as assumed. Two models with identical $R^2$ can differ completely in their residuals and their usefulness. Treat $R^2$ as one number in a larger diagnostic picture, not as a verdict.

## Frequently Asked Questions

### What is a good R-squared value?

There is no universal threshold. What counts as good depends on the field and the noise in the data. In tightly controlled experiments, values above 0.95 are common. In social or behavioral data, 0.20 to 0.50 can be meaningful. Judge $R^2$ against the context and against competing models, not against a fixed cutoff.

### Can R-squared be negative?

For an ordinary least squares model that includes an intercept and is fit to the same data used to compute $R^2$, the value cannot be negative. It can turn negative if you evaluate a model on new data, or if you fit a model without an intercept and use the standard formula. In those cases the model performs worse than simply predicting the mean.

### Is R-squared the same as correlation?

No. The correlation coefficient $r$ measures the strength and direction of a linear relationship and ranges from -1 to 1. The coefficient of determination is $r^2$ in simple linear regression, so it ranges from 0 to 1 and measures explained variation. They carry related but different information.

### Does a high R-squared mean the model is correct?

No. A high $R^2$ means the model fits the observed data closely. It does not confirm that the model is correctly specified, that assumptions hold, or that the predictors cause the outcome. Always pair $R^2$ with diagnostic checks [2].

### How is R-squared different from adjusted R-squared?

$R^2$ can only stay the same or increase when you add predictors. Adjusted $R^2$ applies a penalty for the number of predictors, so it can fall when a new predictor adds little. Use adjusted $R^2$ when comparing models with different numbers of predictors. The article on [R and R-squared interpretation](/blog/data-analysis/r-and-r-squared-interpretation) covers how the two behave in practice.

## References

1. [4.1.4.1. Linear Least Squares Regression](https://www.itl.nist.gov/div898/handbook/pmd/section1/pmd141.htm)
2. [Altman N, Krzywinski M (2016). Regression diagnostics. Nature Methods](https://doi.org/10.1038/nmeth.3854)

## Further Reading

- [Altman N, Krzywinski M (2015). Simple linear regression. Nature Methods](https://doi.org/10.1038/nmeth.3627)
- [Krzywinski M, Altman N (2015). Multiple linear regression. Nature Methods](https://doi.org/10.1038/nmeth.3665)
- [1.4.2.1.3. Quantitative Output and Interpretation](https://www.itl.nist.gov/div898/handbook/eda/section4/eda4213.htm)
- [1.4.2.8.3. Quantitative Output and Interpretation](https://www.itl.nist.gov/div898/handbook/eda/section4/eda4283.htm)

## Related Articles

- [R and R-Squared: What They Mean and How to Interpret Them](/blog/data-analysis/r-and-r-squared-interpretation)
- [Adjusted R-Squared vs R-Squared: Differences and Examples](/blog/data-analysis/adjusted-r-squared-vs-r-squared)
- [Coefficient of Variation: Formula and Examples](/blog/data-analysis/coefficient-of-variation-formula)
- [OLS Regression: What Ordinary Least Squares Means](/blog/data-analysis/ols-regression-ordinary-least-squares)
- [Covariance Formula: Definition and Calculation Examples](/blog/data-analysis/covariance-formula-definition)