# Cramer's V: Definition, Formula and Examples

Cramer's V is a measure of association between two categorical variables. It takes a chi-square statistic and converts it into a number between 0 and 1, where 0 means no association and 1 means a perfect association. It works for tables of any size, which makes it the default strength measure after a significant chi-square test [1][2].

## Quick Answer

- Cramer's V measures how strongly two categorical variables are related, on a scale from 0 to 1.
- The formula is $V = \sqrt{\dfrac{\chi^2}{n \cdot \min(r-1,\ c-1)}}$, where $\chi^2$ is the chi-square statistic, $n$ is the sample size, and $r$ and $c$ are the table's row and column counts.
- It is symmetric, so it does not assume one variable causes the other [1].
- It applies to nominal, ordinal, or binary variables, though Phi is preferred for two binary variables [1].
- A common rough guide treats values below 0.10 as weak, 0.10 to 0.25 as moderate, and above 0.25 as strong in social science research [3].

## What Cramer's V Means

In plain terms, Cramer's V tells you how much knowing one variable's category helps you predict the other variable's category. If every group has the same mix of outcomes, V is near 0. If each category of one variable lines up almost perfectly with one category of the other, V is near 1.

The precise definition: Cramer's V is the square root of the chi-square statistic divided by the product of the sample size and the smaller of (rows minus one) or (columns minus one) [1]. It is a symmetric, non-directional measure, meaning it treats both variables the same way and does not point to a cause [1]. The size of the table and the number of categories in each variable do not change the scale, so you can compare V values across tables with different dimensions [1].

## How It Works

The formula is:

$$V = \sqrt{\frac{\chi^2}{n \cdot \min(r-1,\ c-1)}}$$

Each symbol means:

- $\chi^2$ is the chi-square statistic from the test of independence, computed as $\sum \frac{(O-E)^2}{E}$, where $O$ is an observed cell count and $E$ is the expected count under independence.
- $n$ is the total number of observations in the table.
- $r$ is the number of rows.
- $c$ is the number of columns.
- $\min(r-1,\ c-1)$ is the smaller of the two values, which corrects for the table's shape.

The chi-square statistic alone tells you whether an association exists, but its size grows with the sample size, so it cannot measure strength on its own [2]. Dividing by $n$ removes that dependence, and dividing by $\min(r-1,\ c-1)$ rescales the result so the maximum possible value is 1. The square root brings the value back to a linear scale.

## Worked Example

A survey records responses (Low, Medium, High) for a Control group and a Treatment group. The contingency table is:

| Group | Low | Medium | High | Row total |
|---|---|---|---|---|
| Control | 18 | 22 | 10 | 50 |
| Treatment | 12 | 28 | 30 | 70 |
| Column total | 30 | 50 | 40 | 120 |

The grand total is $n = 120$. Expected counts come from $E = \frac{\text{row total} \times \text{column total}}{n}$. For the Control/Low cell, $E = \frac{50 \times 30}{120} = 12.5000$, and the cell contribution is $\frac{(18-12.5000)^2}{12.5000} = 2.4200$.

Repeating this for every cell gives:

| Cell | Observed | Expected | $(O-E)^2/E$ |
|---|---|---|---|
| Control, Low | 18 | 12.5000 | 2.4200 |
| Control, Medium | 22 | 20.8333 | 0.0653 |
| Control, High | 10 | 16.6667 | 2.6667 |
| Treatment, Low | 12 | 17.5000 | 1.7286 |
| Treatment, Medium | 28 | 29.1667 | 0.0467 |
| Treatment, High | 30 | 23.3333 | 1.9048 |

Summing the last column gives $\chi^2 = 8.8320$. The degrees of freedom are $(r-1)(c-1) = (2-1)(3-1) = 2$, and the p-value is 0.0121. The smaller of $r-1$ and $c-1$ is 1, so:

$$V = \sqrt{\frac{8.8320}{120 \times 1}} = 0.2713$$

The chi-square test is significant at the 0.05 level, and V of 0.2713 sits just above the 0.25 cutoff, so the rough social science guide used in this article would label it a strong association.

Here is the same calculation in Python:

```python
import numpy as np
from scipy.stats import chi2_contingency
table = np.array([[18,22,10],[12,28,30]])
chi2, p, dof, exp = chi2_contingency(table, correction=False)
n = table.sum()
V = np.sqrt(chi2 / (n * min(table.shape[0]-1, table.shape[1]-1)))
print(f"chi2={chi2:.4f}, p={p:.4f}, dof={dof}, V={V:.4f}")
```

Output:

```
chi2=8.8320, p=0.0121, dof=2, V=0.2713
```

## How to Interpret It

Cramer's V ranges from 0 to 1. A value of 0 means the two variables are independent, and a value of 1 means one variable is perfectly predicted by the other. There is no universal cutoff, because the meaning of "strong" depends on your field and your variables [4]. One widely used classification in social science research labels V below 0.10 as weak, 0.10 to 0.25 as moderate, and above 0.25 as strong [3].

Two cautions apply when you read a V value. First, always report it alongside the chi-square p-value, since V describes strength while the p-value describes whether the association is likely to be real [2]. Second, remember that a moderate V can still be meaningful in a large sample, and a large V in a tiny sample can be unstable. If you want a sense of the underlying variability, reviewing how a [mean and standard deviation](/blog/data-analysis/mean-and-standard-deviation) behave in your sample can help you judge how much noise to expect.

## When to Use It (and when not to)

Use Cramer's V when both variables are categorical and you have already run a chi-square test of independence. It handles nominal, ordinal, and binary variables, and it works for tables larger than 2x2, where Phi is no longer appropriate [1][5]. It is the most common strength measure reported after a significant chi-square result [2].

Avoid it in a few situations. For two binary variables, Phi is the standard choice, even though V returns the same value [1]. For two ordinal variables, measures designed for ordered categories are usually more informative [1]. If your table has a large imbalance between the number of rows and columns, V can overestimate the association [1]. And if your variables are continuous, use a correlation coefficient instead.

## Cramer's V vs Phi

Phi and Cramer's V are closely related. Phi is the square root of chi-square divided by the sample size, and it is designed for 2x2 tables [1]. Cramer's V generalizes that idea to larger tables by adding the $\min(r-1,\ c-1)$ correction. For a 2x2 table, the correction equals 1, so the two formulas produce identical values [1][5].

| Feature | Cramer's V | Phi |
|---|---|---|
| Table size | Any r x c table | 2x2 tables |
| Formula | $\sqrt{\chi^2 / (n \cdot \min(r-1, c-1))}$ | $\sqrt{\chi^2 / n}$ |
| Value on a 2x2 table | Same as Phi | Same as V |
| Typical use | Nominal or ordinal pairs, larger tables | Two binary variables |

## Common Mistakes

- Reporting V without the chi-square p-value. V measures strength, not significance. Report both so readers know whether the association is likely real [2].
- Using V for a 2x2 table when Phi is expected. The values match, but Phi is the conventional choice for two binary variables [1].
- Treating V as a measure of direction or causation. V is symmetric and non-directional, so it never tells you which variable drives the other [1].
- Comparing V across studies with very different sample sizes without caution. Small samples produce unstable estimates, and a large row-to-column imbalance can inflate V [1].
- Assuming a fixed cutoff for weak, moderate, and strong. Naming conventions vary by field, so state the benchmark you are using [4].
- Forgetting that V ignores category order. If your categories are ordered, V treats them as unordered labels, which can hide a monotonic trend.

## Limitations

Cramer's V cannot tell you the direction of an association or which categories drive it. It also cannot establish causation, since a third variable may explain the relationship. Because it depends on the chi-square statistic, it inherits the chi-square assumption that expected counts are large enough for the test to be valid, and sparse tables with many small expected counts can distort the result.

The value of V is also sensitive to table shape. When the number of rows and columns differ substantially, V tends to overestimate the association [1]. For ordinal variables, V ignores the ordering and can understate a real monotonic relationship. Treat V as one piece of evidence alongside the full contingency table, the p-value, and your knowledge of the data.

## Frequently Asked Questions

### What is a good Cramer's V value?

There is no single threshold. A common social science guide calls values below 0.10 weak, 0.10 to 0.25 moderate, and above 0.25 strong [3]. Other fields use different labels, so report the benchmark you rely on and interpret V in context [4].

### Can Cramer's V be negative?

No. Because it is a square root of a non-negative quantity, V ranges from 0 to 1. It measures strength only, so it never carries a sign. If you need direction, examine the individual cell counts or use a directional measure.

### What is the difference between Cramer's V and chi-square?

Chi-square tests whether an association exists and grows with sample size. Cramer's V rescales that statistic to a 0-to-1 range so you can judge strength independently of sample size [2]. They are usually reported together.

### When should I use Cramer's V instead of Phi?

Use Phi for 2x2 tables and Cramer's V for anything larger [1][5]. On a 2x2 table the two give the same number, but Phi is the conventional label for two binary variables.

### Does Cramer's V work for ordinal variables?

It can be computed, but it treats the categories as unordered, so it may miss a monotonic trend [1]. For ordinal pairs, a measure that respects category order is usually a better fit.

## References

1. [2.5: An In-Depth Look At Measures of Association - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Applied_Statistics/Social_Data_Analysis%3A_Qualitative_and_Quantitative_Approaches_(Arthur_and_Clark)/02%3A_Quantitative_Data_Analysis/2.05%3A_An_In-Depth_Look_At_Measures_of_Association)
2. [McHugh ML. (2013). The chi-square test of independence. Biochemia medica](https://pubmed.ncbi.nlm.nih.gov/23894860/)
3. [Two Basic Summary Statistics](https://people.uncw.edu/lowery/pls101/MicroCase/two_basic_summary_statistics.htm)
4. [Akoglu H. (2018). User's guide to correlation coefficients. Turkish journal of emergency medicine](https://pmc.ncbi.nlm.nih.gov/articles/PMC6107969/)
5. [lectur15](https://web.pdx.edu/~newsomj/pa551/lectur15.htm)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)

## Related Articles

- [Covariance Formula: Definition and Calculation Examples](/blog/data-analysis/covariance-formula-definition)
- [Bayes' Theorem: Definition, Formula and Examples](/blog/data-analysis/bayes-theorem-definition-formula-examples)
- [Mean and Standard Deviation: Definition, Formula and Examples](/blog/data-analysis/mean-and-standard-deviation)
- [Expected Value: Definition, Formula and Examples](/blog/data-analysis/expected-value-definition-formula-examples)
- [Sample Variance Equation: Formula, Steps and Examples](/blog/data-analysis/sample-variance-equation-formula-examples)