What Is Correlation? Definition, Formula and Examples

By Dr. Zubair Khalid, DVM, MS, PhD ·

What Is Correlation? Definition, Formula and Examples

What is correlation? It is a statistical measure of how strongly two variables move together in a straight-line pattern. The most common version, the Pearson correlation coefficient, is written as $r$ and runs from -1 to +1. This article explains the definition, the formula, how to compute $r$ by hand, and how to read the result without overclaiming.

Quick Answer

  • Correlation measures the direction and strength of a linear relationship between two variables [1].
  • The Pearson correlation coefficient $r$ is a unit-free number between -1 and +1 [2].
  • A positive $r$ means both variables tend to rise together. A negative $r$ means one tends to rise as the other falls [2].
  • The closer $r$ is to zero, the weaker the linear relationship [2].
  • Correlation does not prove causation, and it cannot describe curved relationships [1].

What Correlation Means

In plain terms, correlation tells you whether two measurements tend to move together and how tightly they track each other. If you record two numbers for each subject, correlation summarizes whether high values on one tend to pair with high values on the other, with low values, or with no pattern at all.

The precise statistical definition is narrower. Correlation is a statistical method used to assess a possible linear association between two continuous variables [3]. The word "linear" matters. Correlation asks how well the points fit an imaginary straight line drawn through the data [2]. A perfect curved relationship can still produce a correlation near zero because the pattern is not a straight line.

The sample correlation coefficient, $r$, quantifies the strength of the relationship, and correlations are also tested for statistical significance [1]. That is why a correlation is usually reported with two numbers, $r$ and a p-value [2].

How It Works

The Pearson correlation coefficient compares how far each data point sits from the mean of each variable. The formula is:

$$r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{(n-1)\, s_x \, s_y}$$

Each symbol has a specific job:

SymbolMeaning
$x_i, y_i$The two measurements for one subject
$\bar{x}, \bar{y}$The mean of each variable
$n$The number of paired observations
$s_x, s_y$The standard deviation of each variable
$\sum$Sum across all subjects

The numerator is the covariance numerator, the sum of the paired deviations. Dividing by $n-1$ gives the covariance. Dividing the covariance by the product of the two standard deviations removes the units and rescales the result to the -1 to +1 range. That is why $r$ is unit-free: it does not matter whether you measure hours, dollars, or centimeters.

A closely related variant is the Spearman correlation, which is similar in usage but applies to ranked data [2]. Rank correlations measure a different type of association and do not require the increase to be linear [4].

Worked Example

The dataset below contains 12 students, with their weekly study hours and their exam scores.

Study hoursExam score
252
358
461
566
670
774
878
982
1085
1188
1291
1394

Here are the steps with the computed values.

  1. Count the pairs: $n = 12$.
  2. Mean study hours: $\bar{x} = 7.5000$.
  3. Mean exam score: $\bar{y} = 74.9167$.
  4. Population standard deviation of hours: $s_x = 3.4521$.
  5. Population standard deviation of scores: $s_y = 13.1178$.
  6. Covariance numerator sum: $541.5000$.
  7. Population covariance, to match the population standard deviations: $541.5000 / 12 = 45.1250$.
  8. Correlation: $r = 45.1250 / (3.4521 \times 13.1178) = 0.9965$.

The result is $r = 0.9965$, an almost perfect positive linear relationship. The same data also produce a regression slope of 3.7867 and an intercept of 46.5163, so each extra hour of study is associated with about 3.79 more points on the exam.

You can reproduce the coefficient in Python:

import numpy as np
hours = [2,3,4,5,6,7,8,9,10,11,12,13]
scores = [52,58,61,66,70,74,78,82,85,88,91,94]
r = np.corrcoef(hours, scores)[0,1]
print(round(r, 4))  # 0.9965

Output:

0.9965

If you want to run your own numbers, the Correlation Coefficient Calculator takes two columns of data and returns $r$ directly.

How to Interpret It

Two things matter when you read a correlation: the sign and the size.

The sign tells you the direction. Positive $r$ values indicate a positive correlation, where the values of both variables tend to increase together. Negative $r$ values indicate a negative correlation, where the values of one variable tend to increase when the values of the other variable decrease [2]. A classic example is elevation and summer temperature at campsites: as elevation increases, the temperature drops, so they are negatively correlated [1]. For more real cases, see negative correlation examples.

The size tells you the strength. The closer $r$ is to zero, the weaker the linear relationship [2]. A common rule of thumb treats values near 0.1 to 0.3 as weak, 0.3 to 0.5 as moderate, and above 0.5 as strong, though the cutoff depends on your field [3].

You can also convert $r$ into an effect size. The coefficient of determination, or R-squared, is found by squaring $r$. If $r = .3$, the effect size is .09, which means 9% of the variability in one variable is explained by the other variable [5]. For the study hours example, $0.9965^2$ is about 0.993, so roughly 99% of the variation in scores is explained by the linear relationship with hours.

A correlation is also tested for statistical significance, reported as a p-value alongside $r$ [1]. A large $r$ from a tiny sample can still be statistically uncertain.

When to Use It (and when not to)

Use correlation when you have two continuous variables measured on the same subjects and you want a quick, unit-free summary of their linear association [3]. It is a good first step before building a regression model, and it pairs naturally with the covariance formula, which is the same calculation before rescaling.

Do not use Pearson correlation when the relationship is clearly curved, when either variable is a ranking rather than a measurement, or when outliers dominate the data. Correlation cannot accurately describe curvilinear relationships [1], and it will not detect outliers, so those points can skew the result [2]. It also cannot look at the presence or effect of other variables outside of the two being explored [1].

If you have several variables at once, a covariance matrix gives you all the pairwise relationships in one table.

Correlation vs Covariance

Covariance and correlation use the same numerator. The difference is scaling.

FeatureCovarianceCorrelation
UnitsCarries the units of both variablesUnit-free
RangeUnbounded-1 to +1
DirectionSign shows directionSign shows direction
StrengthHard to judgeEasy to judge
Formula link$\text{cov} = \frac{\sum (x_i-\bar{x})(y_i-\bar{y})}{n-1}$$r = \text{cov} / (s_x s_y)$

In the worked example, the population covariance is 45.1250, and the sample covariance, dividing by $n-1$, is 49.2273. On its own, that number is hard to interpret because it depends on the units of hours and points. Dividing by the two standard deviations turns it into 0.9965, which you can compare against any other correlation on the same scale.

Common Mistakes

  • Treating correlation as causation. A relationship may be observed, but you cannot say one variable caused or affected the other. The relationship may be due to other variables not accounted for in the model [5]. Fix: describe the association, then test causation with a controlled design.
  • Ignoring outliers. Pearson correlation will not detect outliers and can be skewed by them [2]. Fix: plot the data before trusting $r$.
  • Applying it to curved data. Correlation cannot accurately describe curvilinear relationships [1]. Fix: inspect the scatterplot, or use a method built for nonlinear patterns.
  • Reading a small $r$ as "no relationship." A weak linear $r$ can hide a strong curved pattern. Fix: check the shape of the data, not just the number.
  • Comparing $r$ across very different ranges. The coefficient is sensitive to the range of observations [6]. Fix: compare correlations only when the data cover a similar spread.
  • Using $r$ to check agreement. The coefficient is invalid when used to assess agreement of two methods aiming to measure a certain value [6]. Fix: use an intraclass coefficient or Bland-Altman limits of agreement instead [6].

Limitations

Correlation only looks at the two variables at hand and will not give insight into relationships beyond the bivariate data [2]. It cannot account for confounders, so a third variable can drive both measurements and inflate or hide the association. It also says nothing about cause and effect [1].

The coefficient is sensitive to the range of observations and assumes a linear association [6]. Restrict your sample to a narrow band of values and $r$ can shrink toward zero even when a real relationship exists. Include a few extreme points and $r$ can jump toward 1. Always pair the number with a scatterplot, and treat the p-value as a separate piece of evidence about whether the pattern could be chance.

Frequently Asked Questions

What is the correlation coefficient range?

The correlation coefficient $r$ is a unit-free value between -1 and 1 [2]. A value of +1 means a perfect positive linear relationship, -1 means a perfect negative one, and 0 means no linear relationship. Values in between describe partial relationships.

What does a negative correlation mean?

Negative $r$ values indicate a negative correlation, where the values of one variable tend to increase when the values of the other variable decrease [2]. Elevation and summer temperature are a standard example, since temperature drops as elevation rises [1]. The strength still depends on how close $r$ is to -1.

Does correlation mean causation?

No. Correlation is a common tool for describing simple relationships without making a statement about cause and effect [1]. A third variable you did not measure could explain both. Only a controlled experiment or a carefully designed causal study can support a causation claim.

What is the difference between correlation and regression?

Correlation summarizes the strength and direction of a linear relationship on a -1 to +1 scale. Regression fits an equation that predicts one variable from another, producing a slope and an intercept. In the worked example, $r = 0.9965$ while the regression line has a slope of 3.7867 and an intercept of 46.5163.

Can correlation be used with ranked data?

Pearson correlation is built for continuous measurements. For ranked data, the Spearman correlation is a closely related variant that is similar in usage but applicable to ranked data [2]. Rank coefficients measure a different type of association and do not require the increase to be linear [4].

References

  1. Correlation
  2. Correlation Coefficient
  3. Mukaka MM. (2012). Statistics corner: A guide to appropriate use of correlation coefficient in medical research. Malawi medical journal : the journal of Medical Association of Malawi
  4. Correlation - Wikipedia
  5. Correlation - Statistics Resources - LibGuides at National University
  6. Janse RJ, Hoekstra T, Jager KJ, Zoccali C, Tripepi G, Dekker FW, van Diepen M. (2021). Conducting correlation analysis: important limitations and pitfalls. Clinical kidney journal

Related Articles