Negative Correlation Examples: Definition and Real Data Cases

By Dr. Zubair Khalid, DVM, MS, PhD ·

Negative Correlation Examples: Definition and Real Data Cases

Negative correlations examples are cases where one variable goes up while the other goes down. If you plot them, the points trend downward from left to right. The Pearson correlation coefficient $r$ is negative, and the closer it gets to $-1$, the tighter that downward pattern is.

Quick Answer

  • A negative correlation means higher values of one variable pair with lower values of the other.
  • The sign of $r$ gives the direction. The size of $|r|$ gives the strength.
  • $r = -1$ is a perfect downward line. $r = 0$ means no linear relationship. Values near $-0.7$ to $-1$ are usually called strong.
  • A negative slope on a scatterplot is the visual signal of a negative correlation [1].
  • Correlation measures association, not cause. A negative $r$ never proves that one variable drives the other [2].

What Negative Correlation Means

A negative correlation is a relationship in which the two variables move in opposite directions. When one increases, the other tends to decrease. When one decreases, the other tends to increase.

The precise statistical definition uses the Pearson correlation coefficient, usually written $r$. It measures the strength and direction of a linear relationship between two numeric variables. A negative value of $r$ means the linear trend has a negative slope. A positive value means the trend has a positive slope. A value near zero means there is little or no linear trend [1].

The coefficient always falls between $-1$ and $+1$. The sign tells you the direction. The absolute value tells you how closely the points follow a straight line. So $r = -0.9$ is a stronger negative relationship than $r = -0.3$.

How It Works

The Pearson correlation compares how two variables vary together against how much each varies on its own. The formula is:

$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2} \cdot \sqrt{\sum (y_i - \bar{y})^2}}$$

Each symbol means the following.

  • $x_i$ and $y_i$ are the individual paired values.
  • $\bar{x}$ and $\bar{y}$ are the means of each variable.
  • $(x_i - \bar{x})$ is how far a point sits from the mean of $x$.
  • $(y_i - \bar{y})$ is how far the same point sits from the mean of $y$.
  • The numerator sums the products of those deviations. When high $x$ pairs with low $y$, those products are negative, so the numerator turns negative.
  • The denominator rescales the result so it lands between $-1$ and $+1$.

The key idea is the numerator. If large values of $x$ consistently sit with small values of $y$, the products are mostly negative and $r$ comes out negative. That is the arithmetic behind every negative correlation.

Worked Example

This dataset covers advertising spend and sales for 8 stores, in thousands.

ad_spend_ksales_k
1095
1290
1488
1682
1878
2072
2268
2460

Here are the steps with the computed values.

  • $n = 8$ stores.
  • Mean advertising spend: $(10+12+14+16+18+20+22+24) / 8 = 17.0000$.
  • Mean sales: $(95+90+88+82+78+72+68+60) / 8 = 79.1250$.
  • Numerator $\sum (x-\bar{x})(y-\bar{y}) = -407.0000$.
  • Denominator $\sqrt{\sum (x-\bar{x})^2} = 12.9615$.
  • Denominator $\sqrt{\sum (y-\bar{y})^2} = 31.6050$.
  • Pearson $r = -407.0000 / (12.9615 \times 31.6050) = -0.9935$.
  • Slope $b_1 = -407.0000 / 168.0000 = -2.4226$.
  • Intercept $b_0 = 79.1250 - (-2.4226 \times 17.0000) = 120.3095$.
  • $r^2 = (-0.9935)^2 = 0.9871$.

The negative numerator drives the negative $r$. The slope is also negative, which matches the direction. The value $r = -0.9935$ is close to $-1$, so the points sit almost on a downward line.

You can reproduce this in Python.

import numpy as np
from scipy import stats
ad_spend = [10, 12, 14, 16, 18, 20, 22, 24]
sales    = [95, 90, 88, 82, 78, 72, 68, 60]
r, p = stats.pearsonr(ad_spend, sales)
print(round(r, 4))  # -0.9935

Output:

r = -0.9935

In Excel, =CORREL(ad_spend_range, sales_range) returns the same value, $-0.9935$. If you want to check your own numbers quickly, the correlation coefficient calculator handles the arithmetic for you.

How to Interpret It

Direction and strength are two separate readings.

Direction comes from the sign. A negative sign means the variables move in opposite directions. In the example, stores that spent more on advertising had lower sales in this dataset.

Strength comes from the absolute value. A rough guide many analysts use:

$r$ valueReading
$-1.0$Perfect negative line
$-0.7$ to $-1.0$Strong negative
$-0.4$ to $-0.7$Moderate negative
$-0.1$ to $-0.4$Weak negative
$0$No linear relationship

The $r^2$ value tells you the share of variation explained by the linear relationship. Here $r^2 = 0.9871$, so about 98.7 percent of the variation in sales is explained by the linear trend with advertising spend. That is a very tight fit for this small dataset.

Always look at the scatterplot alongside the number. A single $r$ value can hide curves, clusters and outliers. The NIST handbook recommends plotting all response variables in pairwise scatterplots and checking the slope before trusting any correlation figure [1]. A downward slope confirms a negative relationship, and a random cloud with no slope suggests the variables are probably not correlated.

When to Use It (and when not to)

Use the Pearson correlation when both variables are numeric, the relationship looks roughly linear, and you have no extreme outliers. It is a fast way to screen for relationships before building a regression model. In the semiconductor example from NIST, checking pairwise correlations between yield and device measurements helped guide later yield-improvement models [1].

Do not use it when the relationship is clearly curved. A strong U-shaped pattern can produce an $r$ near zero even though the two variables are tightly linked. Do not use it when one variable is categorical. Do not use it when outliers dominate the data, since a single extreme point can flip the sign of $r$.

One more caution applies to restricted data ranges. If you only sample a narrow slice of a wider range, the correlation you measure can be much weaker than the true relationship, or even reversed [3]. This is a common trap when you filter data before analysis.

Negative Correlation vs Positive Correlation

Both describe linear direction. The difference is the sign and what it means in practice.

FeatureNegative correlationPositive correlation
DirectionOne up, other downBoth move together
Sign of $r$NegativePositive
Scatterplot slopeDownwardUpward
ExampleAd spend up, sales down (this dataset)Height up, weight up
InterpretationInverse relationshipDirect relationship

If you want to compare both patterns side by side, see these correlation examples with positive, negative and zero relationships. When the points show no trend at all, that is a separate case covered in no correlation graphs and examples.

Common Mistakes

  • Reading the sign as strength. A large negative number is not "more negative" in a way that means stronger than a smaller one. Compare absolute values. The fix is to read $|r|$ for strength and the sign only for direction.
  • Assuming causation. A negative $r$ does not mean one variable causes the other to fall. The fix is to treat correlation as a screening tool and test causation separately. See correlation vs causation.
  • Ignoring the scatterplot. A single number can hide a curve or an outlier. The fix is to plot the data and check the slope before interpreting $r$ [1].
  • Using Pearson on curved data. A curved relationship can give a misleading $r$. The fix is to inspect the shape first and consider a different measure if the pattern is not linear.
  • Filtering to a narrow range. Restricting the data range can weaken or reverse the correlation [3]. The fix is to check whether your range covers enough of the variable.
  • Mixing up $r$ and $r^2$. $r^2$ is always positive and reports explained variation. The fix is to report both, with $r$ for direction and $r^2$ for fit.

Limitations

The Pearson correlation only captures linear relationships. Two variables can be strongly related in a curved way and still produce an $r$ near zero. It is also sensitive to outliers, since a single distant point can pull the coefficient toward or away from zero. It says nothing about cause, and it does not tell you which variable, if any, drives the other [2].

A negative correlation can also appear for reasons that have nothing to do with a real link. Both variables might be influenced by a third factor, or the pattern might be a coincidence in a small sample. With only 8 stores, the example above is a clean illustration, but small samples are unstable. Always check sample size and context before acting on a correlation.

Frequently Asked Questions

What is a simple example of a negative correlation?

A simple example is the one in this article: as advertising spend rises from 10 to 24 thousand, sales fall from 95 to 60 thousand, giving $r = -0.9935$. Other everyday cases include outdoor temperature and heating bills, or hours of exercise and resting heart rate.

Is a negative correlation always strong?

No. The sign only tells you the direction. Strength depends on the absolute value. An $r$ of $-0.2$ is a weak negative relationship, while $-0.9$ is strong. Always read the size of the number, not just the minus sign.

Can a negative correlation mean one thing causes the other to decrease?

Not on its own. Correlation measures association, not cause [2]. A negative relationship can come from a third factor, from reverse causation, or from coincidence. You need a controlled design or further analysis to support a causal claim.

What does an r value of -1 mean?

An $r$ of $-1$ means a perfect negative linear relationship. Every point falls exactly on a downward straight line. In real data this is rare, since measurement noise and other factors usually keep $r$ away from the extremes.

Why did my negative correlation disappear after filtering the data?

Restricting the range of your data can change the correlation, sometimes weakening it or reversing the sign [3]. If you filter to a narrow band of values, you remove the variation that produced the original relationship. Check the full range before drawing conclusions.

References

  1. 3.4.2.1. Response Correlations
  2. Altman N, Krzywinski M (2015). Association, correlation and causation. Nature Methods
  3. Bland JM, Altman DG (2011). Correlation in restricted ranges of data. BMJ

Further Reading

Related Articles