Cronbach's Alpha: Definition, Formula and Example

By Dr. Zubair Khalid, DVM, MS, PhD ·

Cronbach's Alpha: Definition, Formula and Example

Cronbach's alpha is a number, usually between 0 and 1, that tells you how consistently the items on a scale measure the same underlying thing. If you have a questionnaire with several items that are supposed to tap one concept, alpha summarizes how well those items agree with each other. This article gives you the definition, the formula, a full worked example, and clear guidance on interpretation.

Quick Answer

  • Cronbach's alpha estimates the internal consistency of a set of items, meaning the extent to which they are related to each other [1].
  • It is the mean of all possible split-half reliability coefficients for a test [2].
  • The formula is $\alpha = \frac{k}{k-1}\left(1 - \frac{\sum \sigma^2_i}{\sigma^2_t}\right)$, where $k$ is the number of items.
  • Values usually run from 0 to 1, although alpha can be negative when items are negatively related. Higher means the items move together more closely.
  • A common rule of thumb treats 0.70 and above as acceptable for research scales, but the right cutoff depends on your purpose.

What Cronbach's Alpha Means

In plain terms, Cronbach's alpha answers one question: do the items on my scale behave as if they are measuring the same thing? If people who score high on item 1 also tend to score high on items 2, 3, 4 and 5, the items are consistent and alpha will be high. If the items pull in different directions, alpha drops.

The precise statistical definition is that alpha is the mean of all split-half coefficients that result from different ways of splitting a test into two halves [2]. It is therefore an estimate of the correlation between two random samples of items drawn from a universe of items like the ones in your test [2]. That is why alpha is described as an index of equivalence among items.

Cronbach's alpha is widely used in biomedical and social research as a measure of the internal consistency of a questionnaire [1]. It is closely related to the Kuder-Richardson 20 coefficient, which is the special case of alpha for items scored right or wrong [2].

How It Works

The formula is:

$$\alpha = \frac{k}{k-1}\left(1 - \frac{\sum_{i=1}^{k} \sigma^2_i}{\sigma^2_t}\right)$$

Each symbol means the following:

  • $k$ is the number of items on the scale.
  • $\sigma^2_i$ is the variance of scores on item $i$.
  • $\sum \sigma^2_i$ is the sum of those item variances.
  • $\sigma^2_t$ is the variance of the total scores, where each respondent's total is the sum of their item scores.

The logic is straightforward. The term $\frac{k}{k-1}$ is a correction factor that grows as the scale gets shorter. The fraction $\frac{\sum \sigma^2_i}{\sigma^2_t}$ compares the variability within items to the variability of total scores. When items are consistent, people who score high on one item score high on the others, so total scores spread out more than any single item. That makes the ratio small, and alpha large.

Worked Example

The dataset is a 5-item Likert survey scale scored 1 to 5, completed by 8 respondents.

RespondentQ1Q2Q3Q4Q5
145345
234434
355455
423323
544544
633233
754554
845445

Here are the steps with the computed values.

StepValue
Number of items ($k$)5
Number of respondents ($n$)8
Variance of Q11.0714
Variance of Q20.6964
Variance of Q31.0714
Variance of Q41.0714
Variance of Q50.6964
Sum of item variances4.6071
Total scores per respondent21, 18, 24, 13, 21, 14, 23, 22
Variance of total scores16.8571
Step 1: $k/(k-1)$5/4 = 1.2500
Step 2: sum of item variances / total variance4.6071/16.8571 = 0.2733
Step 3: $1 -$ ratio1 - 0.2733 = 0.7267
Cronbach's alpha1.2500 × 0.7267 = 0.9084

The item variances were computed with the sample variance (VAR.S in Excel, which divides by $n-1$). The same convention applies to the total score variance. Mixing population and sample variance in the same calculation will give you a slightly wrong answer.

You can reproduce this in Python:

import numpy as np
data = np.array([
    [4, 5, 3, 4, 5],
    [3, 4, 4, 3, 4],
    [5, 5, 4, 5, 5],
    [2, 3, 3, 2, 3],
    [4, 4, 5, 4, 4],
    [3, 3, 2, 3, 3],
    [5, 4, 5, 5, 4],
    [4, 5, 4, 4, 5],
])
k = data.shape[1]
item_vars = data.var(axis=0, ddof=1)
total_var = data.sum(axis=1).var(ddof=1)
alpha = (k/(k-1)) * (1 - item_vars.sum()/total_var)

Output:

alpha = 0.9084

The same value comes out of the Excel formula engine, so you can check your work either way.

How to Interpret It

Alpha is a reliability coefficient, so it describes consistency, not validity. A high alpha means your items hang together. It does not prove they measure the concept you intended.

AlphaTypical reading
Below 0.60Weak internal consistency for most research purposes
0.60 to 0.70Acceptable for exploratory work
0.70 to 0.90Good for most research scales
Above 0.90Very high, but check for redundant items

The 0.70 threshold is a convention, not a law. A scale used for individual decisions needs higher reliability than one used to compare group averages, because group averages average out measurement error. Report alpha alongside the number of items, since alpha depends heavily on scale length.

When to Use It (and when not to)

Use Cronbach's alpha when you have a set of items that are all intended to measure one underlying dimension and you want a single summary of how well they cohere. It fits Likert scales, knowledge tests, and multi-item indices in survey research.

Do not use it when your items deliberately measure different dimensions. If a questionnaire has distinct subtests, alpha should be computed separately for each subtest, because tests divisible into distinct subtests should be divided before the formula is applied [2]. Do not use alpha for a single item, since the formula requires at least two. Do not use it as a substitute for validity evidence, and do not use it on items with different response formats without thinking carefully about whether the variances are comparable.

Cronbach's Alpha vs Split-Half Reliability

Split-half reliability splits your items into two halves, correlates the half scores, and corrects for the shortened length. Alpha generalizes this idea across every possible split.

FeatureCronbach's alphaSplit-half reliability
Splits usedAll possible splits, averagedOne specific split
StabilitySame value every timeDepends on how you split
Items requiredTwo or moreTwo or more
RelationshipMean of all split-half coefficients [2]A single special case

Because alpha averages over all splits, it avoids the arbitrariness of choosing which items go in which half. That is the main reason it became the standard internal consistency statistic.

Common Mistakes

  • Using population variance instead of sample variance. The formula assumes sample variances that divide by $n-1$. If you mix conventions, your alpha will be off. Pick VAR.S in Excel or ddof=1 in NumPy and stay consistent.
  • Reporting alpha without the item count. A 20-item scale will almost always show higher alpha than a 4-item scale built from equally good items. Always state $k$ next to the value.
  • Treating a high alpha as proof of validity. Alpha says nothing about whether you measured the right construct. A scale can be internally consistent and still measure the wrong thing.
  • Chasing the highest possible alpha by deleting items. Dropping items raises alpha only if the removed item was inconsistent, but it also narrows your construct coverage. Check what you lose conceptually.
  • Computing one alpha for a multidimensional scale. If your items form two or three factors, a single alpha blends them and can mislead. Compute alpha per dimension.
  • Ignoring reverse-worded items. If some items are negatively worded and you forget to reverse-score them, alpha will collapse toward zero. Recode first, then compute.

Limitations

Alpha assumes a unidimensional set of items. When a scale is multidimensional, alpha can be high simply because the dimensions correlate, which hides the fact that your items are not measuring one thing. Alpha also assumes essential tau-equivalence, meaning all items load equally on the underlying factor. When that assumption fails and item errors are uncorrelated, alpha underestimates reliability. Correlated errors between items can make it overestimate reliability.

Alpha is also sensitive to the number of items and to the sample you collected. Small samples give unstable estimates, and a short scale can show low alpha even when the items are genuinely good. Finally, alpha is a property of the scores in your sample, not a fixed property of the instrument. A scale with alpha 0.90 in one population can show 0.75 in another.

Frequently Asked Questions

What is a good Cronbach's alpha value?

For most research scales, 0.70 or above is treated as acceptable and 0.80 to 0.90 as good. Values below 0.60 usually signal that the items do not cohere well enough to be summed into one score. The right threshold depends on your purpose, so report the value and let readers judge it against your context.

Can Cronbach's alpha be negative?

Yes. A negative alpha means the items are, on average, negatively related to each other. This usually happens when reverse-worded items were not recoded, or when the items genuinely measure different constructs. Check your scoring before you conclude anything about the scale.

Does a higher Cronbach's alpha always mean a better scale?

No. Very high alpha, above about 0.95, often means your items are nearly redundant and you are asking the same question several times. That adds respondent burden without adding information. A moderate value with good construct coverage is usually more useful.

How many items do I need for Cronbach's alpha?

You need at least two items, but two is rarely enough for a stable estimate. In practice, scales of four to ten items are common. Alpha tends to rise as you add items, so a short scale needs stronger inter-item correlations to reach the same value.

Is Cronbach's alpha the same as reliability?

Alpha is one type of reliability, specifically internal consistency. Other types include test-retest reliability, which measures stability over time, and inter-rater reliability, which measures agreement between judges. A scale can have high alpha and poor test-retest reliability, so match the reliability type to your question.

If you are working with several variables at once, it helps to understand how they vary together, which is what the covariance formula describes. For categorical agreement between raters, a different statistic such as Cramer's V is the right tool.

References

  1. "An R Function for Cronbach’s Alpha Analysis: A Case-Based Approach" by Himani Kotian, Aiswarya Liz Varghese et al.
  2. Cronbach LJ (1951). Coefficient Alpha and the Internal Structure of Tests. Psychometrika

Further Reading

Related Articles