# What Is Communality in Factor Analysis? Definition and Examples

Communality is the proportion of an item's variance that the factors in a factor analysis explain. It is usually written as $h^2$ and ranges from 0 to 1. A high communality means the item is well explained by the shared factors, while a low value means most of its variance is unique to that item.

## Quick Answer

- Communality $h^2$ is the sum of squared loadings for an item across all retained factors.
- It ranges from 0 to 1 and is the share of that item's variance the factors account for.
- The leftover, $1 - h^2$, is called uniqueness or unique variance.
- In a one-factor solution, communality is simply the squared loading on that factor.
- The sum of all communalities equals the total variance explained by the retained factors.

## What Communality Means

In plain terms, communality answers one question for each variable: how much of this item's behavior is driven by the same underlying thing the other items share? If four survey questions all tap a common attitude, each question's answers move partly with that attitude and partly for their own reasons. Communality is the size of the shared part.

The precise statistical definition is this. In the common factor model, the variance of a standardized variable $x_i$ is split into two pieces:

$$1 = h_i^2 + u_i^2$$

Here $h_i^2$ is the communality, the variance shared with the factors, and $u_i^2$ is the uniqueness, the variance specific to the item plus measurement error. Because variables are standardized to have variance 1, the two parts always add to 1. Communality is computed as the sum of squared loadings:

$$h_i^2 = \sum_{j=1}^{m} \lambda_{ij}^2$$

where $\lambda_{ij}$ is the loading of item $i$ on factor $j$ and $m$ is the number of retained factors. This is the same quantity you see in the "communalities" table of most factor analysis software.

## How It Works

The mechanism is straightforward once you accept that loadings are correlations between items and factors. A loading of 0.80 means the item correlates 0.80 with that factor. Squaring a correlation gives the proportion of shared variance, so squaring the loading gives the proportion of the item's variance tied to that factor.

Each symbol means the following:

- $\lambda_{ij}$ is the loading of item $i$ on factor $j$, a value between -1 and 1.
- $\lambda_{ij}^2$ is the variance in item $i$ explained by factor $j$ alone.
- $m$ is the number of factors you keep.
- $h_i^2$ is the total explained variance for item $i$ across those factors.
- $u_i^2 = 1 - h_i^2$ is what remains unexplained.

When factors are uncorrelated, the squared loadings add up cleanly. When you rotate obliquely and factors correlate, the sum of squared loadings is no longer the exact communality, so software reports the communality from the fitted model instead. In a single-factor solution there is only one term, so $h^2$ is just the squared loading.

## Worked Example

The data are a 50-row Likert survey (1 to 5) with four items, Q1 to Q4, measuring one shared latent construct. Here is the raw dataset.

| Q1 | Q2 | Q3 | Q4 |
|----|----|----|----|
| 2 | 1 | 2 | 2 |
| 3 | 3 | 3 | 3 |
| 4 | 3 | 4 | 3 |
| 2 | 2 | 1 | 2 |
| 3 | 2 | 3 | 3 |
| 4 | 4 | 4 | 3 |
| 3 | 2 | 2 | 2 |
| 3 | 3 | 3 | 4 |
| 5 | 4 | 5 | 4 |
| 2 | 2 | 2 | 2 |
| 4 | 4 | 3 | 4 |
| 1 | 2 | 2 | 2 |
| 3 | 3 | 3 | 3 |
| 4 | 4 | 5 | 4 |
| 2 | 2 | 2 | 1 |
| 3 | 3 | 3 | 3 |
| 4 | 4 | 4 | 4 |
| 2 | 3 | 2 | 3 |
| 3 | 3 | 4 | 4 |
| 5 | 5 | 4 | 4 |
| 2 | 2 | 2 | 2 |
| 4 | 4 | 4 | 4 |
| 2 | 2 | 2 | 2 |
| 3 | 2 | 3 | 2 |
| 4 | 4 | 4 | 5 |
| 2 | 2 | 2 | 2 |
| 3 | 3 | 3 | 4 |
| 4 | 4 | 4 | 3 |
| 2 | 2 | 2 | 3 |
| 3 | 3 | 3 | 2 |
| 4 | 5 | 5 | 4 |
| 2 | 3 | 3 | 3 |
| 3 | 3 | 3 | 4 |
| 1 | 1 | 1 | 1 |
| 3 | 3 | 3 | 2 |
| 4 | 5 | 3 | 4 |
| 1 | 2 | 2 | 3 |
| 3 | 3 | 3 | 4 |
| 5 | 4 | 5 | 4 |
| 2 | 2 | 2 | 3 |
| 3 | 3 | 3 | 3 |
| 4 | 5 | 5 | 5 |
| 2 | 2 | 2 | 2 |
| 4 | 4 | 3 | 3 |
| 2 | 2 | 1 | 2 |
| 3 | 3 | 3 | 3 |
| 4 | 4 | 4 | 4 |
| 2 | 2 | 2 | 3 |
| 3 | 3 | 2 | 3 |
| 4 | 4 | 4 | 5 |

The sample size is 50 and there are 4 items. The correlation matrix is:

| | Q1 | Q2 | Q3 | Q4 |
|---|---|---|---|---|
| Q1 | 1.0000 | 0.8654 | 0.8634 | 0.7029 |
| Q2 | 0.8654 | 1.0000 | 0.8267 | 0.7832 |
| Q3 | 0.8634 | 0.8267 | 1.0000 | 0.7498 |
| Q4 | 0.7029 | 0.7832 | 0.7498 | 1.0000 |

The eigenvalues of this correlation matrix are 3.3985, 0.3226, 0.1714, and 0.1074. Only the first exceeds 1, so a one-factor solution is reasonable. The first eigenvalue, 3.3985, is the variance carried by that factor.

The loadings on Factor 1 are 0.9331 for Q1, 0.9441 for Q2, 0.9344 for Q3, and 0.8738 for Q4. Squaring each loading gives the communality:

| Item | Loading | Communality $h^2$ | Uniqueness $1 - h^2$ |
|------|---------|-------------------|----------------------|
| Q1 | 0.9331 | 0.8706 | 0.1294 |
| Q2 | 0.9441 | 0.8912 | 0.1088 |
| Q3 | 0.9344 | 0.8731 | 0.1269 |
| Q4 | 0.8738 | 0.7635 | 0.2365 |

The sum of the communalities is 3.3985, which matches the first eigenvalue exactly. That is the total variance explained by the single factor.

Here is the code that produces these values.

```python
import numpy as np, pandas as pd
R = np.corrcoef(df[['Q1','Q2','Q3','Q4']].values, rowvar=False)
vals, vecs = np.linalg.eigh(R)
i = np.argsort(vals)[::-1]
loadings = vecs[:, i[0]] * np.sqrt(vals[i[0]])
loadings = loadings * np.sign(loadings.sum())  # eigenvector sign is arbitrary
communalities = loadings**2
print("loadings =", np.round(loadings, 4).tolist())
print("communalities =", np.round(communalities, 4).tolist())
```

Output:

```text
loadings = [0.9331, 0.9441, 0.9344, 0.8738]
communalities = [0.8706, 0.8912, 0.8731, 0.7635]
```

## How to Interpret It

Read communality as a percentage of explained variance. Q2 has $h^2 = 0.8912$, so the factor explains about 89 percent of its variance. Q4 has $h^2 = 0.7635$, so about 76 percent is explained and roughly 24 percent is unique to that item.

Common rules of thumb treat $h^2$ below about 0.40 as weak, meaning the item is poorly explained by the factor solution. Values above 0.70 are usually considered strong. These cutoffs are conventions, not laws, and they depend on how many factors you retain. Adding factors always raises communalities because you are summing more squared loadings.

Communality also tells you how much an item contributes to the factor. An item with a high $h^2$ sits close to the factor and helps define it. An item with a low $h^2$ drifts away from the factor and may belong to a different construct or carry a lot of measurement error.

If you are comparing a measurement model against theory, [confirmatory factor analysis](/blog/data-analysis/confirmatory-factor-analysis) reports the same communalities but tests them against a structure you specify in advance.

## When to Use It (and when not to)

Use communality when you want to judge item quality in an exploratory factor analysis, decide which items to keep, or report how much variance a factor solution captures. It is a standard part of the output and a quick way to spot weak items before you finalize a scale.

Do not use it as a standalone measure of whether your factor model is correct. A high communality does not prove the factor structure is right, and a low one does not always mean the item is bad. Communality is also unstable in small samples, so with 50 rows and 4 items the estimates here are reasonable but would shift with a different sample.

Avoid comparing communalities across studies with different numbers of factors or different extraction methods. The numbers are not on the same footing. If your variables are highly correlated with each other, check for [multicollinearity](/blog/data-analysis/multicollinearity-definition-detection) before trusting the loadings, since it can distort the solution.

## Communality vs Uniqueness

Communality and uniqueness are two halves of the same split. Communality is the variance an item shares with the factors. Uniqueness is the variance left over, which includes item-specific variance and measurement error.

| Feature | Communality $h^2$ | Uniqueness $u^2$ |
|---------|-------------------|------------------|
| Definition | Variance explained by factors | Variance not explained |
| Formula | Sum of squared loadings | $1 - h^2$ |
| Range | 0 to 1 | 0 to 1 |
| High value means | Item fits the factor well | Item is mostly unique or noisy |
| Typical use | Judging item quality | Diagnosing error or specificity |
| Relationship | $h^2 + u^2 = 1$ | $u^2 = 1 - h^2$ |

In the worked example, Q2 has the highest communality (0.8912) and the lowest uniqueness (0.1088). Q4 has the lowest communality (0.7635) and the highest uniqueness (0.2365), so it is the least well explained item.

## Common Mistakes

- Treating communality as a correlation. It is a squared quantity, so a communality of 0.87 corresponds to a loading of about 0.93, not 0.87. Square the loading or take the square root of the communality to move between them.
- Comparing $h^2$ across solutions with different numbers of factors. More factors means more squared loadings in the sum, so communalities rise mechanically. Compare only within the same solution.
- Using a fixed cutoff as a pass or fail test. The 0.40 and 0.70 thresholds are conventions. Judge items in context and alongside the loadings and theory.
- Forgetting that communality includes error. A high $h^2$ still leaves room for measurement error, and a low one may reflect a genuinely distinct construct instead of a bad item.
- Reading communalities from an oblique rotation as a simple sum of squared loadings. When factors correlate, use the model-based communality your software reports.
- Ignoring sample size. Communalities from small samples are noisy. With 50 rows, small differences between items may not be meaningful.

## Limitations

Communality cannot tell you whether your factor model is correct. It only describes how much variance the retained factors explain for each item. A model with the wrong number of factors can still produce high communalities, especially if you keep many factors.

The value also depends on the extraction method, the rotation, and the number of factors, so it is not an absolute property of an item. It is a property of the item within a specific analysis. Treat it as one piece of evidence alongside loadings, eigenvalues, theory, and fit statistics.

## Frequently Asked Questions

### What is a good communality value?

Many textbooks treat 0.40 as a minimum and 0.70 as strong, but these are conventions. What matters more is whether the item loads clearly on the intended factor and whether the value is consistent with your theory. A slightly low communality on a theoretically sound item is often acceptable.

### Is communality the same as a factor loading?

No. A loading is a correlation between an item and a factor, so it can be negative and ranges from -1 to 1. Communality is the squared loading summed across factors, so it is always between 0 and 1. In a one-factor solution, $h^2$ is just the loading squared.

### Can communality be greater than 1?

No, not in a proper solution. Because variables are standardized to variance 1, the explained and unexplained parts must sum to 1. A value above 1 signals a problem such as a Heywood case, which usually points to too many factors, too few items, or an unstable sample.

### Why does communality change when I add factors?

Each retained factor adds another squared loading to the sum. Adding factors can only increase or hold steady the explained variance for an item, so communalities rise as you keep more factors. This is why you should compare communalities only within the same solution.

### What does a low communality mean for my scale?

A low $h^2$ means the factors explain little of that item's variance. The item may measure something else, carry a lot of error, or be poorly worded. Check its loadings and content before dropping it, and consider whether the factor solution itself needs revising.

## References

This article draws on the standard references listed under Further Reading.

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)
- [Wasserstein RL, Lazar NA (2016). The ASA Statement on p -Values: Context, Process, and Purpose. The American Statistician](https://doi.org/10.1080/00031305.2016.1154108)
- [Greenland S, Senn SJ, Rothman KJ et al. (2016). Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology](https://doi.org/10.1007/s10654-016-0149-3)

## Related Articles

- [Confirmatory Factor Analysis: Definition and Example](/blog/data-analysis/confirmatory-factor-analysis)
- [Multicollinearity: Definition, Detection and Examples](/blog/data-analysis/multicollinearity-definition-detection)
- [Multivariate Analysis: Definition, Methods and Examples](/blog/data-analysis/multivariate-analysis)
- [Bivariate Data: Definition, Examples and Analysis](/blog/data-analysis/bivariate-data-definition-examples)