# Spurious Correlation: Definition, Examples and How to Spot It

Spurious statistics describe relationships that look real in the numbers but have no causal basis. A spurious correlation is a strong statistical association between two variables that is produced by a third factor, by the way the data were built, or by chance, not by any direct link between the two variables. This article defines the concept, shows a worked example with real computed values, and gives you practical checks to run before you trust a relationship.

## Quick Answer

- A spurious correlation is a strong measured association between two variables that does not reflect a direct causal relationship.
- It usually appears because a third variable (a confounder) drives both, because the data are ratios or indices, or because outliers and skew distort the coefficient [1][2].
- A high Pearson $r$ and a small p-value do not rule out a spurious relationship. They only say the sample association is unlikely under a null of zero correlation.
- Detection relies on domain reasoning, checking for a common cause, plotting the data, and testing whether the relationship survives controls.
- Machine learning models learn spurious correlations too, and they can fail badly when those patterns change in the real world [3][4].

## What Spurious Correlation Means

In plain terms, a spurious correlation is a relationship that exists in your dataset but not in the world. Two variables move together in the numbers, yet changing one would not change the other.

The precise statistical definition is narrower. A correlation is spurious when the observed association between $X$ and $Y$ is not caused by $X$ affecting $Y$ or $Y$ affecting $X$, but is instead induced by a common cause, by selection, by the construction of the variables, or by sampling noise. The coefficient itself is computed correctly. The interpretation is what fails.

This distinction matters because the arithmetic is not wrong. If you feed two columns into a correlation function, you get a number. The problem is that the number answers "do these move together in this sample" and not "does one cause the other." That gap is where spurious statistics cause real damage.

## How It Works

The Pearson correlation coefficient is the usual starting point:

$$r = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_{i=1}^{n}(x_i - \bar{x})^2}\sqrt{\sum_{i=1}^{n}(y_i - \bar{y})^2}}$$

Each symbol means the following.

- $x_i$ and $y_i$ are the paired values for observation $i$.
- $\bar{x}$ and $\bar{y}$ are the sample means of each variable.
- $n$ is the number of paired observations.
- The numerator is the sum of products of deviations, which is large when high $x$ values pair with high $y$ values.
- The denominator rescales that sum by the spread of each variable, so $r$ always falls between $-1$ and $1$.

The formula has no term for cause. It measures co-movement only. That is why a confounder such as time, temperature, population size or income can push $r$ toward $1$ without any direct link between $X$ and $Y$. When a third variable drives both, the correlation is real in the data and spurious in meaning. If you want the mechanics of how the numerator and denominator relate, see [correlation vs covariance](/blog/data-analysis/correlation-vs-covariance).

## Worked Example

The dataset below is a simulated 24-month series of ice cream sales (units) and drowning incidents, built to illustrate a spurious correlation. Both variables rise in summer and fall in winter, so a common seasonal driver moves them together.

| Month | Ice cream sales | Drownings | Month | Ice cream sales | Drownings |
|---|---|---|---|---|---|
| M01 | 120 | 3 | M13 | 138 | 4 |
| M02 | 128 | 4 | M14 | 145 | 5 |
| M03 | 135 | 4 | M15 | 155 | 6 |
| M04 | 150 | 5 | M16 | 172 | 7 |
| M05 | 178 | 7 | M17 | 200 | 9 |
| M06 | 210 | 10 | M18 | 235 | 12 |
| M07 | 245 | 14 | M19 | 270 | 15 |
| M08 | 260 | 16 | M20 | 285 | 17 |
| M09 | 240 | 13 | M21 | 262 | 14 |
| M10 | 195 | 9 | M22 | 215 | 10 |
| M11 | 150 | 5 | M23 | 168 | 6 |
| M12 | 130 | 4 | M24 | 145 | 5 |

Here are the steps with the computed values.

- Sample size: $n = 24$.
- Mean ice cream sales: $\bar{x} = 188.7917$.
- Mean drowning incidents: $\bar{y} = 8.5000$.
- Sample standard deviation of ice cream sales: $s_x = 51.7830$.
- Sample standard deviation of drownings: $s_y = 4.4036$.
- Covariance: $\text{cov} = 226.3696$.
- Pearson $r$ from covariance: $r = 226.3696 / (51.7830 \times 4.4036) = 0.9927$.
- Manual numerator: $\sum (x_i - \bar{x})(y_i - \bar{y}) = 5206.5000$.
- Manual denominator: $\sqrt{\sum (x_i - \bar{x})^2 \times \sum (y_i - \bar{y})^2} = 5244.6721$.
- Manual $r$: $5206.5000 / 5244.6721 = 0.9927$.
- `scipy.stats.pearsonr`: $r = 0.9927$, $p = 0.0000$.
- Correlation of ice cream sales with time: $r = 0.3498$.
- Correlation of drownings with time: $r = 0.2988$.
- Partial correlation controlling for time: $r_{\text{partial}} = 0.9935$.
- Excel `=CORREL(A2:A25,B2:B25)` returns 0.9927.

```python
import numpy as np
from scipy import stats
x = np.array([120,128,135,150,178,210,245,260,240,195,150,130,
              138,145,155,172,200,235,270,285,262,215,168,145])
y = np.array([3,4,4,5,7,10,14,16,13,9,5,4,
              4,5,6,7,9,12,15,17,14,10,6,5])
r, p = stats.pearsonr(x, y)
print(f"r = {r:.4f}, p = {p:.4f}")  # r = 0.9927, p = 0.0000
```

Output:

```
r = 0.9927, p = 0.0000
```

The correlation is 0.9927 and the p-value rounds to 0.0000. Neither number tells you that ice cream does not cause drowning. The partial correlation controlling for time is 0.9935, which is even higher, so a simple "control for time" step does not remove the association here. That is a useful lesson: partial correlation only helps when the control variable captures the shared driver, and a linear time trend does not capture a repeating seasonal cycle.

## How to Interpret It

Read a correlation as a description of co-movement in your sample, nothing more. A value near $1$ or $-1$ means the points fall close to a straight line. It does not mean one variable moves the other.

Ask three questions before you trust a relationship.

1. Is there a plausible mechanism? If you cannot describe how $X$ would change $Y$, treat the correlation as unexplained.
2. Is there a common cause? A third variable that drives both is the most common source of spurious statistics.
3. Does the relationship hold when the conditions change? Split the data by season, region or subgroup and recompute.

For the ice cream example, the mechanism is missing and the common cause is obvious: warm weather increases both ice cream purchases and swimming. The correlation is genuine in the data and meaningless as a causal claim. For more on this reasoning, see correlation vs causation.

## When to Use It (and when not to)

Use correlation when you want to summarize whether two variables move together, when you are screening many variables for patterns worth investigating, or when you need a compact input for a model. It is a fast, interpretable first look at [bivariate data](/blog/data-analysis/bivariate-data-definition-examples).

Do not use it as evidence of causation, as a stand-alone basis for a policy or business decision, or as a substitute for an experiment or a properly specified model. If your goal is to estimate the effect of changing $X$ on $Y$, correlation is the wrong tool. You need a design or a model that handles confounders.

## Spurious Correlation vs Confounding

Confounding is one mechanism that produces spurious correlations, so the two ideas overlap but are not identical. Confounding is a property of the data-generating process. Spurious correlation is a property of the estimate you report.

| Aspect | Spurious correlation | Confounding |
|---|---|---|
| What it is | A reported association with no direct causal link | A third variable that distorts an association |
| Where it lives | In your result | In the data-generating process |
| Typical sign | High $r$, small p-value, no mechanism | Association changes when the confounder is controlled |
| Fix | Investigate mechanism, plot, split, control | Measure and adjust for the confounder, or randomize |
| Can exist without the other | Yes, for example from ratio variables or outliers [1][2] | Yes, confounding can bias an effect estimate without making $r$ look "spurious" |

## Common Mistakes

- **Treating a small p-value as proof of a real effect.** The p-value tests whether the sample correlation differs from zero. Fix: report the effect size and the mechanism alongside the p-value.
- **Ignoring a common cause.** Two variables can both track a third, such as time, temperature or population. Fix: name candidate confounders and test whether the association survives controlling for them.
- **Correlating ratios or indices without thinking.** Ratios share components and can produce surprising, spurious results [2]. Fix: use randomization tests and define the correct null correlation for ratio data.
- **Letting outliers and skew drive the coefficient.** Skewed or heavy-tailed variables can make correlations spuriously large or small [1]. Fix: run exploratory plots first, then transform or handle discordant observations.
- **Assuming a linear control removes a seasonal pattern.** In the worked example, controlling for a linear time trend left $r_{\text{partial}} = 0.9935$. Fix: model the actual shared driver, such as season, not a proxy.
- **Trusting a model that learned a shortcut.** Neural networks can rely on spurious features and fail when those features change [3][4]. Fix: test on data where the spurious feature is absent or altered.

## Limitations

Correlation and partial correlation cannot prove or disprove causation on their own. They are descriptive tools. A partial correlation only adjusts for the variables you measured and included, so an unmeasured confounder can keep an association looking real. In the worked example, the partial correlation controlling for time was 0.9935, higher than the raw 0.9927, which shows that a poorly chosen control can leave the spurious pattern intact.

Detection also depends on context you may not have. You can compute $r$ for any two columns, but judging whether the relationship is meaningful requires knowing how the data were collected, what else varies with them, and whether the pattern holds outside the sample. When those facts are missing, the honest answer is that the correlation is unexplained, not that it is real. For a broader list of traps, see [common correlation and regression mistakes](/knowledge/bioinformatics/common-mistakes-in-correlation-and-regression-analysis-in-life-sciences-and-how-to-avoid-them).

## Frequently Asked Questions

### What is a spurious correlation in simple terms?

It is a strong statistical relationship between two variables that has no direct causal link. The numbers move together, but changing one would not change the other. A hidden third factor, the way the variables were built, or chance usually explains it.

### Can a spurious correlation have a very small p-value?

Yes. A p-value near zero only says the sample association is unlikely if the true correlation were zero. It says nothing about cause. The ice cream and drowning example has $r = 0.9927$ and $p = 0.0000$, yet no causal link exists.

### How do I detect a spurious correlation?

Start with a plausible mechanism. If you cannot explain how one variable would affect the other, look for a common cause. Plot the data, check for outliers and skew, split the sample by subgroup or season, and test whether the association survives controls for the variables you suspect [1].

### Does controlling for a third variable always remove a spurious correlation?

No. It only works when the control captures the shared driver. In the worked example, controlling for a linear time trend gave a partial correlation of 0.9935, which is higher than the raw value. The real driver was a repeating seasonal cycle, which a straight trend line does not represent.

### Do machine learning models suffer from spurious correlations too?

Yes. Deep networks can learn and rely on spurious patterns in training data, then fail when those patterns change in deployment [3]. Some spurious signals are weak yet still harmful, and they can be learned from only a handful of samples [3][4]. Testing on shifted data helps expose them.

If you want to compute a coefficient on your own data, the [correlation coefficient calculator](/tools/correlation-coefficient-calculator) returns $r$ from two columns. Pair it with a scatter plot and a clear statement of what the number does and does not mean.

## References

1. [Halperin S. (1986). Spurious correlations--causes and cures. Psychoneuroendocrinology](https://pubmed.ncbi.nlm.nih.gov/3704066/)
2. [Jackson DA, Somers KM. (1991). The spectre of 'spurious' correlations. Oecologia](https://pubmed.ncbi.nlm.nih.gov/28313173/)
3. [New Technique Overcomes Spurious Correlations Problem in AI | NC State News](https://news.ncsu.edu/2025/03/ai-spurious-correlations/)
4. [Uncovering memorization effect in the presence of spurious correlations - PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC12216586/)

## Further Reading

- [Wilson G, Bryan J, Cranston K et al. (2017). Good enough practices in scientific computing. PLOS Computational Biology](https://doi.org/10.1371/journal.pcbi.1005510)
- [Wilkinson MD, Dumontier M, Aalbersberg IJ et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data](https://doi.org/10.1038/sdata.2016.18)

## Related Articles

- [Negative Correlation Examples: Definition and Real Data Cases](/blog/data-analysis/negative-correlation-examples-definition)
- [Correlation Examples: Positive, Negative and Zero Relationships](/blog/data-analysis/correlation-examples-positive-negative)
- [Correlation vs Covariance: Differences and When to Use Each](/blog/data-analysis/correlation-vs-covariance)
- [No Correlation: Definition, Graphs and Examples](/blog/data-analysis/no-correlation-definition-graphs-examples)
- [Statistical Synonyms: A Guide to Terminology in Statistics](/blog/guides/statistical-synonyms-a-guide-to-terminology-in-statistics)
- [Fundamental Statistics: Core Concepts Explained](/blog/research-skills/fundamental-statistics-core-concepts-explained)
- [Correlation Analysis in SPSS](/knowledge/bioinformatics/correlation-analysis-in-spss-how-to-run-interpret-and-report-pearson-and-spearman-coefficients)