Cramer's V: Definition, Formula and Examples

By Dr. Zubair Khalid, DVM, MS, PhD ·

Cramer's V: Definition, Formula and Examples

Cramer's V is a measure of association between two categorical variables. It takes a chi-square statistic and converts it into a number between 0 and 1, where 0 means no association and 1 means a perfect association. It works for tables of any size, which makes it the default strength measure after a significant chi-square test [1][2].

Quick Answer

  • Cramer's V measures how strongly two categorical variables are related, on a scale from 0 to 1.
  • The formula is $V = \sqrt{\dfrac{\chi^2}{n \cdot \min(r-1,\ c-1)}}$, where $\chi^2$ is the chi-square statistic, $n$ is the sample size, and $r$ and $c$ are the table's row and column counts.
  • It is symmetric, so it does not assume one variable causes the other [1].
  • It applies to nominal, ordinal, or binary variables, though Phi is preferred for two binary variables [1].
  • A common rough guide treats values below 0.10 as weak, 0.10 to 0.25 as moderate, and above 0.25 as strong in social science research [3].

What Cramer's V Means

In plain terms, Cramer's V tells you how much knowing one variable's category helps you predict the other variable's category. If every group has the same mix of outcomes, V is near 0. If each category of one variable lines up almost perfectly with one category of the other, V is near 1.

The precise definition: Cramer's V is the square root of the chi-square statistic divided by the product of the sample size and the smaller of (rows minus one) or (columns minus one) [1]. It is a symmetric, non-directional measure, meaning it treats both variables the same way and does not point to a cause [1]. The size of the table and the number of categories in each variable do not change the scale, so you can compare V values across tables with different dimensions [1].

How It Works

The formula is:

$$V = \sqrt{\frac{\chi^2}{n \cdot \min(r-1,\ c-1)}}$$

Each symbol means:

  • $\chi^2$ is the chi-square statistic from the test of independence, computed as $\sum \frac{(O-E)^2}{E}$, where $O$ is an observed cell count and $E$ is the expected count under independence.
  • $n$ is the total number of observations in the table.
  • $r$ is the number of rows.
  • $c$ is the number of columns.
  • $\min(r-1,\ c-1)$ is the smaller of the two values, which corrects for the table's shape.

The chi-square statistic alone tells you whether an association exists, but its size grows with the sample size, so it cannot measure strength on its own [2]. Dividing by $n$ removes that dependence, and dividing by $\min(r-1,\ c-1)$ rescales the result so the maximum possible value is 1. The square root brings the value back to a linear scale.

Worked Example

A survey records responses (Low, Medium, High) for a Control group and a Treatment group. The contingency table is:

GroupLowMediumHighRow total
Control18221050
Treatment12283070
Column total305040120

The grand total is $n = 120$. Expected counts come from $E = \frac{\text{row total} \times \text{column total}}{n}$. For the Control/Low cell, $E = \frac{50 \times 30}{120} = 12.5000$, and the cell contribution is $\frac{(18-12.5000)^2}{12.5000} = 2.4200$.

Repeating this for every cell gives:

CellObservedExpected$(O-E)^2/E$
Control, Low1812.50002.4200
Control, Medium2220.83330.0653
Control, High1016.66672.6667
Treatment, Low1217.50001.7286
Treatment, Medium2829.16670.0467
Treatment, High3023.33331.9048

Summing the last column gives $\chi^2 = 8.8320$. The degrees of freedom are $(r-1)(c-1) = (2-1)(3-1) = 2$, and the p-value is 0.0121. The smaller of $r-1$ and $c-1$ is 1, so:

$$V = \sqrt{\frac{8.8320}{120 \times 1}} = 0.2713$$

The chi-square test is significant at the 0.05 level, and V of 0.2713 sits just above the 0.25 cutoff, so the rough social science guide used in this article would label it a strong association.

Here is the same calculation in Python:

import numpy as np
from scipy.stats import chi2_contingency
table = np.array([[18,22,10],[12,28,30]])
chi2, p, dof, exp = chi2_contingency(table, correction=False)
n = table.sum()
V = np.sqrt(chi2 / (n * min(table.shape[0]-1, table.shape[1]-1)))
print(f"chi2={chi2:.4f}, p={p:.4f}, dof={dof}, V={V:.4f}")

Output:

chi2=8.8320, p=0.0121, dof=2, V=0.2713

How to Interpret It

Cramer's V ranges from 0 to 1. A value of 0 means the two variables are independent, and a value of 1 means one variable is perfectly predicted by the other. There is no universal cutoff, because the meaning of "strong" depends on your field and your variables [4]. One widely used classification in social science research labels V below 0.10 as weak, 0.10 to 0.25 as moderate, and above 0.25 as strong [3].

Two cautions apply when you read a V value. First, always report it alongside the chi-square p-value, since V describes strength while the p-value describes whether the association is likely to be real [2]. Second, remember that a moderate V can still be meaningful in a large sample, and a large V in a tiny sample can be unstable. If you want a sense of the underlying variability, reviewing how a mean and standard deviation behave in your sample can help you judge how much noise to expect.

When to Use It (and when not to)

Use Cramer's V when both variables are categorical and you have already run a chi-square test of independence. It handles nominal, ordinal, and binary variables, and it works for tables larger than 2x2, where Phi is no longer appropriate [1][5]. It is the most common strength measure reported after a significant chi-square result [2].

Avoid it in a few situations. For two binary variables, Phi is the standard choice, even though V returns the same value [1]. For two ordinal variables, measures designed for ordered categories are usually more informative [1]. If your table has a large imbalance between the number of rows and columns, V can overestimate the association [1]. And if your variables are continuous, use a correlation coefficient instead.

Cramer's V vs Phi

Phi and Cramer's V are closely related. Phi is the square root of chi-square divided by the sample size, and it is designed for 2x2 tables [1]. Cramer's V generalizes that idea to larger tables by adding the $\min(r-1,\ c-1)$ correction. For a 2x2 table, the correction equals 1, so the two formulas produce identical values [1][5].

FeatureCramer's VPhi
Table sizeAny r x c table2x2 tables
Formula$\sqrt{\chi^2 / (n \cdot \min(r-1, c-1))}$$\sqrt{\chi^2 / n}$
Value on a 2x2 tableSame as PhiSame as V
Typical useNominal or ordinal pairs, larger tablesTwo binary variables

Common Mistakes

  • Reporting V without the chi-square p-value. V measures strength, not significance. Report both so readers know whether the association is likely real [2].
  • Using V for a 2x2 table when Phi is expected. The values match, but Phi is the conventional choice for two binary variables [1].
  • Treating V as a measure of direction or causation. V is symmetric and non-directional, so it never tells you which variable drives the other [1].
  • Comparing V across studies with very different sample sizes without caution. Small samples produce unstable estimates, and a large row-to-column imbalance can inflate V [1].
  • Assuming a fixed cutoff for weak, moderate, and strong. Naming conventions vary by field, so state the benchmark you are using [4].
  • Forgetting that V ignores category order. If your categories are ordered, V treats them as unordered labels, which can hide a monotonic trend.

Limitations

Cramer's V cannot tell you the direction of an association or which categories drive it. It also cannot establish causation, since a third variable may explain the relationship. Because it depends on the chi-square statistic, it inherits the chi-square assumption that expected counts are large enough for the test to be valid, and sparse tables with many small expected counts can distort the result.

The value of V is also sensitive to table shape. When the number of rows and columns differ substantially, V tends to overestimate the association [1]. For ordinal variables, V ignores the ordering and can understate a real monotonic relationship. Treat V as one piece of evidence alongside the full contingency table, the p-value, and your knowledge of the data.

Frequently Asked Questions

What is a good Cramer's V value?

There is no single threshold. A common social science guide calls values below 0.10 weak, 0.10 to 0.25 moderate, and above 0.25 strong [3]. Other fields use different labels, so report the benchmark you rely on and interpret V in context [4].

Can Cramer's V be negative?

No. Because it is a square root of a non-negative quantity, V ranges from 0 to 1. It measures strength only, so it never carries a sign. If you need direction, examine the individual cell counts or use a directional measure.

What is the difference between Cramer's V and chi-square?

Chi-square tests whether an association exists and grows with sample size. Cramer's V rescales that statistic to a 0-to-1 range so you can judge strength independently of sample size [2]. They are usually reported together.

When should I use Cramer's V instead of Phi?

Use Phi for 2x2 tables and Cramer's V for anything larger [1][5]. On a 2x2 table the two give the same number, but Phi is the conventional label for two binary variables.

Does Cramer's V work for ordinal variables?

It can be computed, but it treats the categories as unordered, so it may miss a monotonic trend [1]. For ordinal pairs, a measure that respects category order is usually a better fit.

References

  1. 2.5: An In-Depth Look At Measures of Association - Statistics LibreTexts/02%3A_Quantitative_Data_Analysis/2.05%3A_An_In-Depth_Look_At_Measures_of_Association)
  2. McHugh ML. (2013). The chi-square test of independence. Biochemia medica
  3. Two Basic Summary Statistics
  4. Akoglu H. (2018). User's guide to correlation coefficients. Turkish journal of emergency medicine
  5. lectur15

Further Reading

Related Articles