# Correlation vs Covariance: Differences and When to Use Each

Covariance and correlation both measure how two variables move together, but they answer slightly different questions. Covariance tells you the direction of a linear relationship in the original units of your data, while correlation rescales that same relationship onto a fixed range from -1 to 1. If you want a number you can compare across studies, report correlation. If you want a number that feeds into a formula, such as a regression slope or a covariance matrix, report covariance.

## Quick Answer

- **Covariance** measures the direction of a linear relationship and keeps the units of both variables, so its size depends on the scale you measured on [1].
- **Correlation** is covariance divided by the product of the two standard deviations, which makes it dimensionless and bounded between -1 and 1 [1].
- The **sign** means the same thing in both. Positive means the variables tend to move in the same direction, negative means they move in opposite directions.
- The **magnitude** of covariance is not interpretable on its own. A covariance of 16 could be strong or weak depending on the units.
- The **magnitude** of correlation is interpretable directly. Values near -1 or 1 indicate a strong linear relationship, values near 0 indicate a weak one.

## Key Differences

| Property | Covariance | Correlation |
|---|---|---|
| Formula | $\text{cov}(X,Y) = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{n-1}$ | $r = \frac{\text{cov}(X,Y)}{s_x s_y}$ |
| Units | Product of the units of X and Y | None (dimensionless) [1] |
| Range | Unbounded, any real number | -1 to 1 |
| Effect of rescaling | Changes when you change units | Unchanged by linear rescaling |
| Interpretation | Direction only | Direction and strength |
| Typical use | Covariance matrices, regression math | Reporting and comparing relationships |

The core difference is that correlation is a standardized version of covariance. Dividing by the two standard deviations removes the units and fixes the scale, so the same underlying relationship always produces the same correlation no matter how you measured the variables [1].

## Covariance Explained

Covariance measures whether two variables tend to be above or below their means at the same time. When both are above their means, the product $(x_i - \bar{x})(y_i - \bar{y})$ is positive. When one is above and the other below, the product is negative. Averaging those products gives the covariance [2].

$$\text{cov}(X,Y) = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{n-1}$$

The $n-1$ denominator gives the unbiased sample estimate, which is the default in most statistical software [2]. A positive covariance means the variables move together on average. A negative covariance means they move in opposite directions. A covariance near zero means there is little linear co-movement.

The problem with covariance is that its size depends on the units. Measure height in centimeters instead of meters and the covariance changes by a factor of 100. That makes it hard to compare covariances across datasets or judge whether a value is large. Covariance still matters because it is the building block for other quantities. Regression slopes, covariance matrices, and multivariate methods all work directly with covariances [3]. If you want to see the formula applied step by step, the [covariance formula guide](/blog/data-analysis/covariance-formula-definition) walks through the arithmetic.

## Correlation Explained

Correlation fixes the scale problem by dividing covariance by the product of the two standard deviations [1].

$$r = \frac{\text{cov}(X,Y)}{s_x s_y}$$

Because both the numerator and denominator carry the same units, they cancel and the result is dimensionless [1]. The value always falls between -1 and 1. A correlation of 1 means one variable is an exact positive linear function of the other, and -1 means an exact negative linear function [1]. A value of 0 means there is no linear relationship, though the variables could still be related in a curved way.

This standardization is why correlation is the number people usually report. You can compare a correlation of 0.8 between height and weight with a correlation of 0.8 between study time and test scores, even though the variables have completely different units. The [correlation coefficient calculator](/tools/correlation-coefficient-calculator) lets you check these values quickly on your own data. For guidance on reading the strength of a coefficient, see [which R-value represents the strongest correlation](/blog/guides/which-r-value-represents-the-strongest-correlation-a-guide-to-interpreting-correlation-coefficie).

## Worked Example

Take a small dataset of hours studied and exam scores out of 50 for 8 students.

| hours_studied | exam_score |
|---|---|
| 2 | 22 |
| 3 | 25 |
| 4 | 28 |
| 5 | 30 |
| 6 | 33 |
| 7 | 36 |
| 8 | 38 |
| 9 | 41 |

The mean hours studied is 5.5000 and the mean exam score is 31.6250. Using the sample formula, the covariance is:

$$\text{cov} = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{n-1} = 16.0714$$

The standard deviation of hours is 2.4495 and the standard deviation of scores is 6.5670. Dividing gives the correlation:

$$r = \frac{16.0714}{2.4495 \times 6.5670} = 0.9991$$

Now rescale the scores to a 0 to 100 range by multiplying each by 2. The new mean is 63.2500 and the new standard deviation is 13.1339. The covariance doubles to 32.1429 because the scores now carry twice the scale. The correlation stays at 0.9991 because the rescaling cancels out.

Here is the same calculation in Python:

```python
import statistics
hours = [2,3,4,5,6,7,8,9]
scores = [22,25,28,30,33,36,38,41]
n = len(hours)
mh, ms = statistics.mean(hours), statistics.mean(scores)
cov = sum((h-mh)*(s-ms) for h,s in zip(hours,scores))/(n-1)
r = cov/(statistics.stdev(hours)*statistics.stdev(scores))
print(round(cov, 4), round(r, 4))  # 16.0714 0.9991
```

Output:

```
16.0714 0.9991
```

In Excel, `=COVARIANCE.S(B2:B9,A2:A9)` returns 16.0714 and `=CORREL(A2:A9,B2:B9)` returns 0.9991. The covariance changed when you rescaled the scores, the correlation did not.

## Which One Should You Use?

Use **correlation** when you are reporting a relationship to an audience. It is scale-free, bounded, and comparable across variables and studies [1]. If you are describing how strongly two things move together, correlation is the number to put in a table or a sentence.

Use **covariance** when you are doing math that needs it. Regression coefficients, covariance matrices, and multivariate methods such as principal component analysis all operate on covariances [3]. A covariance matrix holds the variances on the diagonal and the covariances off the diagonal, and it captures the full spread of several variables at once [2]. If you are working with several variables together, see [multivariate analysis](/blog/data-analysis/multivariate-analysis) for how these matrices are used.

A practical rule: report correlation, compute with covariance. When you present results, give the correlation and its interpretation. When you build a model or run a decomposition, keep the covariance.

## Common Mistakes

- **Reading covariance magnitude as strength.** A covariance of 50 is not "stronger" than a covariance of 5 unless the units are identical. Fix: convert to correlation before judging strength.
- **Comparing covariances across different units.** Covariances in centimeters and meters are not comparable. Fix: standardize to correlation first.
- **Assuming correlation of 0 means no relationship.** Correlation only captures linear relationships. Fix: plot the data before concluding there is no association, as covered in correlation vs causation.
- **Using the wrong denominator.** Dividing by $n$ instead of $n-1$ gives a biased sample covariance. Fix: use the $n-1$ sample formula unless you specifically want the population value [2].
- **Ignoring outliers.** Both covariance and correlation are sensitive to extreme points [4]. Fix: inspect the scatter plot and consider a robust estimator when outliers are present [4].
- **Mixing up the sign convention.** A negative covariance and a negative correlation both mean the variables move in opposite directions. Fix: check the sign, not just the size.

## Limitations

Neither measure captures a curved relationship. Two variables can follow a perfect U-shape and still produce a correlation near zero, because correlation only detects linear co-movement. Always plot your data before trusting a single number.

Covariance and correlation are also sensitive to outliers, and a single extreme point can pull either value far from the rest of the data [4]. Sample estimates of covariance and correlation are not linked in any simple way, so they generally need to be treated separately when you assess their uncertainty [1]. Correlation does not tell you which variable causes the other, and it does not tell you the size of the effect in practical terms.

## Frequently Asked Questions

### What is the main difference between covariance and correlation?

Covariance measures the direction of a linear relationship in the original units of the variables. Correlation divides that covariance by the product of the two standard deviations, which removes the units and bounds the result between -1 and 1 [1]. The sign is the same in both, but only correlation tells you strength directly.

### Why is correlation scale-free?

Correlation is scale-free because the units in the numerator and denominator cancel. When you divide covariance by the product of the standard deviations, any change in measurement units affects both parts equally, so the ratio stays the same [1]. That is why rescaling scores from 50 points to 100 points left the correlation at 0.9991 in the example above.

### Can covariance be larger than 1?

Yes. Covariance has no upper or lower bound because it keeps the units of both variables [1]. A covariance can be 0.001 or 10,000 depending on the scale. Correlation is the bounded version, so if you need a value between -1 and 1, use correlation.

### Which should I report in a paper?

Report correlation when you are describing the strength and direction of a relationship, because readers can interpret it without knowing your units. Report covariance when the value feeds into a model or a covariance matrix [3]. Many papers report both, giving the correlation for interpretation and the covariance for the underlying math.

### Does a correlation of 0.9991 mean the variables are almost identical?

No. It means the two variables have an almost perfect positive linear relationship in this sample. A correlation near 1 says the points fall close to a straight line, not that the variables measure the same thing. With only 8 observations, the estimate is also uncertain, and a larger sample would give a more stable value.

## References

1. [Covariance and correlation - Wikipedia](https://en.wikipedia.org/wiki/Covariance_and_correlation)
2. [6.5.4.1. Mean Vector and Covariance Matrix](https://www.itl.nist.gov/div898/handbook/pmc/section5/pmc541.htm)
3. [Covariance, SciPy v1.18.0 Manual](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.Covariance.html)
4. [2.6. Covariance estimation, scikit-learn 1.9.1 documentation](https://scikit-learn.org/stable/modules/covariance.html)

## Further Reading

- [R: Correlation and Covariance Matrices](https://www.math.ucla.edu/~anderson/rw1001/library/base/html/cor.html)
- [Altman N, Krzywinski M (2015). Association, correlation and causation. Nature Methods](https://doi.org/10.1038/nmeth.3587)

## Related Articles

- [Nominal vs Ordinal Variables: Differences and Examples](/blog/data-analysis/nominal-vs-ordinal-variables)
- [Covariance Formula: Definition and Calculation Examples](/blog/data-analysis/covariance-formula-definition)
- [Multivariate Analysis: Definition, Methods and Examples](/blog/data-analysis/multivariate-analysis)
- [Bivariate Data: Definition, Examples and Analysis](/blog/data-analysis/bivariate-data-definition-examples)
- [Predictor vs. Covariate: Clarifying Terminology in Research](/blog/guides/predictor-vs-covariate-clarifying-terminology-in-research)
- [Which R-Value Represents the Strongest Correlation? A Guide to Interpreting Correlation Coefficients](/blog/guides/which-r-value-represents-the-strongest-correlation-a-guide-to-interpreting-correlation-coefficie)
- [Correlational Research](/blog/guides/correlational-research)