# What Is Zero Correlation? Definition and Examples

Zero correlation means there is no linear relationship between two variables. When you compute a Pearson correlation coefficient, a value near 0 tells you that changes in one variable do not move the other in any consistent straight-line way. This article explains what zero correlation means, how to compute it, and how to recognize it in real data.

## Quick Answer

- Zero correlation means the Pearson correlation coefficient $r$ is 0 or close to 0, so the two variables have no linear relationship.
- A correlation of zero does not mean the variables are unrelated in every way. They can still have a strong curved or non-linear relationship.
- The sign of $r$ shows direction and the size shows strength. At $r = 0$ there is no direction to report.
- Zero correlation is about linear association only. Always plot your data before trusting a single number [1].
- A value near 0 can also appear when you restrict the range of your data, even if the full dataset shows a real relationship [2].

## What Zero Correlation Means

In plain terms, zero correlation means that knowing the value of one variable tells you nothing about the value of the other, as long as you are only looking for a straight-line pattern. If study hours go up, the quiz score is just as likely to go down as up. There is no consistent trend.

The precise statistical definition uses the Pearson correlation coefficient, usually written $r$. It measures the strength and direction of the linear relationship between two quantitative variables. The value of $r$ always falls between $-1$ and $+1$. A value of $+1$ is a perfect positive line, $-1$ is a perfect negative line, and $0$ means no linear association at all.

So a correlation of zero is the midpoint of that scale. It is the point where the best-fitting straight line is flat, and the two variables carry no shared linear signal.

## How It Works

The Pearson correlation coefficient is built from how each pair of values deviates from its own mean. The formula is:

$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}$$

Each symbol means the following:

- $x_i$ and $y_i$ are the two values for one observation, such as one student's hours and score.
- $\bar{x}$ and $\bar{y}$ are the means of each variable.
- $(x_i - \bar{x})$ is how far one $x$ value sits from the mean of $x$.
- $(y_i - \bar{y})$ is how far the matching $y$ value sits from the mean of $y$.
- The numerator multiplies those two deviations for every pair and adds them up. This is the covariance part.
- The denominator rescales that sum by the spread of each variable, which forces $r$ into the range from $-1$ to $+1$.

When the deviations tend to have the same sign, the products are mostly positive and $r$ is positive. When they tend to have opposite signs, $r$ is negative. When the signs are mixed and roughly cancel out, the numerator is near 0 and you get zero correlation.

## Worked Example

The dataset below has 12 students. The first pair is study hours against a quiz score that was randomly assigned, so the true relationship is near zero. The second pair, x2 and y2, is included as a contrast with a strong positive relationship.

| study_hours | quiz_score | x2 | y2 |
|---|---|---|---|
| 1 | 72 | 1 | 2 |
| 2 | 55 | 2 | 4 |
| 3 | 68 | 3 | 5 |
| 4 | 81 | 4 | 4 |
| 5 | 60 | 5 | 5 |
| 6 | 74 | 6 | 7 |
| 7 | 66 | 7 | 8 |
| 8 | 79 | 8 | 9 |
| 9 | 58 | 9 | 10 |
| 10 | 70 | 10 | 12 |
| 11 | 63 | 11 | 13 |
| 12 | 77 | 12 | 15 |

Here are the steps for the study-hours and quiz-score pair.

- Number of pairs, $n$: 12
- Mean study hours, $\bar{x}$: 6.5000
- Mean quiz score, $\bar{y}$: 68.5833
- Sample standard deviation of hours: 3.6056
- Sample standard deviation of scores: 8.4473
- Numerator, $\sum (x_i - \bar{x})(y_i - \bar{y})$: 37.5000
- Denominator, $\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}$: 335.0270
- Pearson $r$: $37.5000 / 335.0270 = 0.1119$

The fitted line has a slope of 0.2622 and an intercept of 66.8788. That slope is small, and the correlation of 0.1119 is close to zero. The second pair gives $r = 158.0000 / 161.1780 = 0.9803$, which is a strong positive relationship. The contrast shows how differently a near-zero value and a near-one value behave.

You can reproduce the first result with this code.

```python
import numpy as np
hours = [1,2,3,4,5,6,7,8,9,10,11,12]
scores = [72,55,68,81,60,74,66,79,58,70,63,77]
r = np.corrcoef(hours, scores)[0,1]
print(f"r = {r:.4f}")  # r = 0.1119
```

Output:

```
r = 0.1119
```

You can check your own numbers with the [Correlation Coefficient Calculator](/tools/correlation-coefficient-calculator).

## How to Interpret It

A correlation of zero tells you the linear trend is flat. The fitted line is essentially horizontal, so a one-unit increase in $x$ predicts almost no change in $y$. In the example, the slope of 0.2622 means each extra study hour predicts about a quarter of a point on the quiz, which is tiny next to the score spread of 8.4473.

Interpretation depends on context. In some fields a value of 0.1119 is treated as effectively no relationship. In others, even a small $r$ can matter if the sample is large. The number alone is not enough. You also want the scatter plot, the sample size, and the units of the variables.

A near-zero $r$ also does not prove the variables are independent. It only says the linear component is missing. Two variables can be tightly linked through a curve and still produce $r$ near 0. For a fuller picture of how direction and strength work across signs, see [correlation examples with positive, negative and zero relationships](/blog/data-analysis/correlation-examples-positive-negative).

## When to Use It (and when not to)

Use the Pearson correlation coefficient when both variables are quantitative, the relationship looks roughly linear, and you want a single number for strength and direction. It is a good first summary when you have already looked at a scatter plot.

Do not rely on it when the relationship is clearly curved, when there are extreme outliers, or when one variable is categorical. In those cases the coefficient can hide the real pattern. If your data is curved, a value near zero may be misleading, and you should describe the shape instead.

Be careful with restricted ranges. If you only sample part of the range of a variable, the correlation can drop toward zero even when the full dataset shows a clear relationship [2]. This is a common trap in subgroup analysis.

## Zero Correlation vs No Relationship

These two ideas sound the same but are not. Zero correlation is a statement about linear association only. No relationship is a stronger claim that the variables are independent in every way.

| Feature | Zero correlation | No relationship |
|---|---|---|
| What it measures | Linear association only | Any form of association |
| Typical value | $r$ near 0 | No dependence at all |
| Curved patterns | Can still exist | Excluded by definition |
| Example | $r = 0.1119$ with a flat line | Independent random values |
| Safe conclusion | No straight-line trend | No link of any kind |

A dataset can have zero correlation and still show a perfect U-shape, where $y$ rises, falls, and rises again as $x$ increases. The linear coefficient misses it. For more on this distinction, read [no correlation: definition, graphs and examples](/blog/data-analysis/no-correlation-definition-graphs-examples).

## Common Mistakes

- Treating zero correlation as proof of independence. Fix: plot the data and check for curves before concluding there is no link.
- Ignoring the scatter plot. Fix: always draw the points, since the same $r$ can come from very different shapes [1].
- Confusing a small $r$ with a meaningless one. Fix: judge the size against your field and sample size, not a fixed cutoff.
- Forgetting range restriction. Fix: check whether you trimmed the data, because a narrow range can push $r$ toward zero [2].
- Assuming a flat line means no effect. Fix: look at the slope and units, since a small slope can still matter in context.
- Reading causation into any correlation, including zero. Fix: remember that a coefficient describes association, never cause.

## Limitations

The Pearson coefficient captures linear association only. It cannot detect a strong curved relationship, and it can report a value near zero when the variables are actually tightly linked through a curve. It is also sensitive to outliers, since a single extreme point can pull $r$ far from its true value.

A single number also hides the shape of the data. Two datasets with the same $r$ can look completely different. Range restriction can shrink $r$ toward zero even when the full data shows a real trend [2]. For all these reasons, treat zero correlation as one piece of evidence, not a final verdict.

## Frequently Asked Questions

### Does zero correlation mean the variables are unrelated?

No. Zero correlation means there is no linear relationship. The variables can still be related through a curve, a threshold, or some other non-linear pattern. You need a scatter plot and possibly other methods to rule out those shapes.

### Can two variables have zero correlation but still be dependent?

Yes. A classic case is a symmetric U-shape, where $y$ depends strongly on $x$ but the positive and negative parts cancel in the linear formula. The coefficient lands near zero while the variables are clearly dependent.

### What is the difference between zero correlation and no correlation?

They usually mean the same thing in practice, both pointing to $r$ near 0. The phrase zero correlation is more precise, because it names the linear coefficient. No correlation is looser and can wrongly suggest the variables are fully independent.

### Is a correlation of 0.1119 close enough to zero?

It is close to zero and shows almost no linear trend. The fitted slope of 0.2622 is small relative to the score spread. Whether you call it zero depends on context, but the practical message is that study hours give little linear information about these quiz scores.

### Can zero correlation still be useful?

Yes. A near-zero result can rule out a simple linear effect and point you toward other explanations. It can also flag a restricted range or a non-linear pattern worth investigating. For related traps, see [spurious correlation](/blog/data-analysis/spurious-correlation-definition-examples) and correlation vs causation.

## References

1. [Altman N, Krzywinski M (2015). Association, correlation and causation. Nature Methods](https://doi.org/10.1038/nmeth.3587)
2. [Bland JM, Altman DG (2011). Correlation in restricted ranges of data. BMJ](https://doi.org/10.1136/bmj.d556)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [Altman N, Krzywinski M (2015). Simple linear regression. Nature Methods](https://doi.org/10.1038/nmeth.3627)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)

## Related Articles

- [Negative Correlation Examples: Definition and Real Data Cases](/blog/data-analysis/negative-correlation-examples-definition)
- [Correlation Examples: Positive, Negative and Zero Relationships](/blog/data-analysis/correlation-examples-positive-negative)
- [No Correlation: Definition, Graphs and Examples](/blog/data-analysis/no-correlation-definition-graphs-examples)
- [What Is Correlation? Definition, Formula and Examples](/blog/data-analysis/what-is-correlation)
- [Spurious Correlation: Definition, Examples and How to Spot It](/blog/data-analysis/spurious-correlation-definition-examples)