# Confirmatory Factor Analysis: Definition and Example

Confirmatory factor analysis (CFA) is a statistical method for testing whether a set of measured variables follows a factor structure you specify in advance. You decide how many factors exist and which items load on each one, then check how well that model reproduces the correlations in your data. It is the standard tool for validating questionnaires, scales and other multi-item instruments.

## Quick Answer

- CFA tests a hypothesis. You fix the number of factors and the item-to-factor links before estimation, then evaluate fit [1].
- It differs from exploratory factor analysis (EFA), which lets the data determine the factor structure [2].
- The model is a measurement model. Each item is a linear function of one or more latent factors plus error [2].
- Fit is judged with several indices together, typically chi-square, CFI, TLI, RMSEA and SRMR [3].
- Poor fit means your hypothesized structure is inconsistent with the sample data, not that the items are useless [4].

## What Confirmatory Factor Analysis Means

In plain terms, confirmatory factor analysis is a way to check whether the pattern of relationships among your items matches the pattern your theory predicts. If you believe a six-item survey measures two constructs, you tell the software which three items belong to each construct and let it estimate how strongly each item reflects its factor.

The precise definition: CFA is a structural equation modeling technique in which the researcher imposes an a priori measurement model on the data and tests how well that model fits [3]. The model specifies (a) the number of factors, (b) whether factors are correlated or uncorrelated, and (c) which items load on which factor [3]. Estimation compares the model-implied variance-covariance matrix to the observed variance-covariance matrix [3].

This is the key contrast with EFA. EFA is an exploratory tool for understanding the psychometric properties of an unknown scale. CFA verifies the psychometric structure of a previously developed scale [2]. EFA has the longer history, dating to Spearman's work in 1904, while CFA became widely used after Jöreskog's estimation method in 1969 and the growth of computing power [2].

## How It Works

The CFA measurement model is a linear regression in which the predictor is latent, meaning it is not directly observed [2]. For a single item, the model is:

$$y_{i} = \tau_{i} + \lambda_{i}\eta + \epsilon_{i}$$

- $y_{i}$ is the observed score on item $i$.
- $\tau_{i}$ is the intercept for item $i$.
- $\lambda_{i}$ is the factor loading, the strength of the item's relationship to the factor.
- $\eta$ is the latent factor score.
- $\epsilon_{i}$ is the unique error for item $i$, which includes measurement error and item-specific variance.

Factor loadings range from -1.0 to 1.0 and are interpreted much like correlation coefficients. They show how strongly each item relates to each factor and how much variation in responses the factor accounts for [3]. A common rule of thumb treats loadings at or above |.4| as strongly related to the underlying factor [3].

Estimation proceeds by minimizing the difference between the observed correlation or covariance matrix $R$ and the model-implied matrix $\Sigma$. A widely used discrepancy function is:

$$F_{ML} = \ln|\Sigma| - \ln|R| + \mathrm{tr}(R\Sigma^{-1}) - k$$

where $k$ is the number of items, $|\cdot|$ is the determinant and $\mathrm{tr}(\cdot)$ is the trace. The chi-square statistic is then $\chi^2 = (n-1)F_{ML}$, with degrees of freedom equal to the number of unique variances and covariances minus the number of estimated parameters.

For ordinal items such as Likert scales, a theoretically appropriate approach fits the model to polychoric correlations using weighted least squares or robust weighted least squares. Simulation work shows polychoric estimation is robust to modest violations of underlying normality, and robust WLS performed well across all conditions studied [5].

## Worked Example

The dataset below contains 12 respondents rating 6 survey items on a 1-5 Likert scale. Two factors are hypothesized: Engagement (Q1 to Q3) and Satisfaction (Q4 to Q6).

| Q1 | Q2 | Q3 | Q4 | Q5 | Q6 |
|----|----|----|----|----|----|
| 5 | 4 | 5 | 4 | 5 | 4 |
| 4 | 5 | 4 | 5 | 4 | 5 |
| 5 | 5 | 4 | 4 | 4 | 5 |
| 3 | 4 | 3 | 3 | 4 | 3 |
| 4 | 4 | 5 | 5 | 5 | 4 |
| 5 | 3 | 4 | 4 | 3 | 4 |
| 2 | 3 | 2 | 2 | 3 | 2 |
| 4 | 5 | 5 | 5 | 4 | 5 |
| 5 | 4 | 4 | 4 | 5 | 4 |
| 3 | 3 | 4 | 3 | 3 | 3 |
| 4 | 5 | 3 | 5 | 4 | 4 |
| 5 | 5 | 5 | 4 | 5 | 5 |

Step 1. Sample size and items: $n = 12$ respondents, $k = 6$ items.

Step 2. Hypothesized factors: Engagement covers Q1, Q2, Q3. Satisfaction covers Q4, Q5, Q6.

Step 3. Average within-block correlation for Factor 1 is $\bar{r}_1 = 0.4772$. For Factor 2 it is $\bar{r}_2 = 0.5873$.

Step 4. As a simple hand approximation (CFA software instead estimates every loading separately by maximum likelihood), standardized loadings come from the square root of those averages. For Factor 1, $\lambda_1 = \sqrt{0.4772} = 0.6908$. For Factor 2, $\lambda_2 = \sqrt{0.5873} = 0.7664$.

Step 5. The factor correlation is the mean cross-block correlation divided by the product of the loadings: $\phi = 0.9900$.

Step 6. Evaluated at these approximate values, the ML discrepancy function is $F_{ML} = 2.8417$, giving $\chi^2 = (12-1) \times 2.8417 = 31.2584$. A full ML fit would minimize $F_{ML}$, so it would give a value no larger than this.

Step 7. Degrees of freedom: $df = k(k+1)/2 - (k + k + 1) = 8$.

Step 8. The independence model gives $\chi^2_{ind} = 65.8094$ with $df_{ind} = 15$.

Step 9. Fit indices follow from these values. CFI $= 1 - \max(\chi^2 - df, 0)/\max(\chi^2_{ind} - df_{ind}, 0) = 0.5422$. TLI $= (\chi^2_{ind}/df_{ind} - \chi^2/df)/(\chi^2_{ind}/df_{ind} - 1) = (4.3873 - 3.9073)/(4.3873 - 1) = 0.1417$. RMSEA $= \sqrt{\max(F_{ML}/df - 1/(n-1), 0)} = 0.5141$. SRMR $= \sqrt{\text{mean(residual}^2)} = 0.1491$.

```python
import numpy as np, pandas as pd
df = pd.DataFrame(rows, columns=['Q1','Q2','Q3','Q4','Q5','Q6'])
R = df.corr().values
lam1 = np.sqrt(R[0:3,0:3][np.triu_indices(3,1)].mean())
lam2 = np.sqrt(R[3:6,3:6][np.triu_indices(3,1)].mean())
```

Output: `lambda1 = 0.6908, lambda2 = 0.7664, phi = 0.9900, chi2 = 31.2584, df = 8, CFI = 0.5422, TLI = 0.1417, RMSEA = 0.5141, SRMR = 0.1491`

Standardized CFA loadings: Engagement items load at 0.6908, Satisfaction items at 0.7664, factor correlation 0.9900.

## How to Interpret It

Start with the loadings. Both blocks show loadings above 0.4, so each item relates meaningfully to its hypothesized factor [3]. The factor correlation of 0.9900 is very high, which suggests the two factors may not be empirically distinguishable in this sample.

Now the fit indices. The chi-square of 31.2584 with 8 degrees of freedom is large relative to its df, which signals misfit. CFI is 0.5422 and TLI is 0.1417, both far below the conventional 0.90 or 0.95 thresholds. RMSEA of 0.5141 is far above the usual 0.06 or 0.08 cutoffs, and SRMR of 0.1491 exceeds the common 0.08 benchmark. Every index points the same way: this model does not fit.

That result is expected here. With only 12 respondents, the sample is far too small for stable CFA estimation, and the near-perfect factor correlation suggests a one-factor solution might describe the data just as well. A poor fit means the constraints you imposed are inconsistent with the sample data and the model is rejected [4]. It does not automatically mean your theory is wrong. It may mean the sample, the item set or the model specification needs work.

## When to Use It (and when not to)

Use CFA when you have a clear theory or prior evidence about the factor structure and you want to test it. It is the standard method for evaluating the internal structural validity of measurement instruments [6]. Typical uses include validating a scale after piloting, checking measurement invariance across groups, and building the measurement portion of a larger structural equation model.

Do not use CFA to discover structure. If you have no prior hypothesis about how many factors exist or which items belong together, run exploratory factor analysis first [2]. Do not run CFA on samples that are too small for the number of parameters you are estimating. Do not use maximum likelihood estimation on ordinal items without considering the polychoric correlation approach, since ordinal scales are common in the applied social sciences and several popular commercial packages handle this case poorly [6].

If you are still deciding between related techniques, a broader look at multivariate analysis methods helps you place CFA among its alternatives.

## Confirmatory Factor Analysis vs Exploratory Factor Analysis

| Feature | Confirmatory Factor Analysis | Exploratory Factor Analysis |
|---------|------------------------------|-----------------------------|
| Factor structure | Specified by the researcher in advance [1] | Determined by the data [2] |
| Purpose | Test a hypothesis about structure [3] | Explore an unknown scale [2] |
| Number of factors | Fixed before estimation [3] | Chosen after extraction |
| Item-to-factor links | Constrained by theory [4] | Free to load on any factor |
| Typical output | Fit statistics and loadings [3] | Loadings and rotated solutions |
| Historical origin | Popularized after Jöreskog (1969) [2] | Dates to Spearman (1904) [2] |

The two methods share many concepts. CFA simply adds constraints based on your a priori hypotheses, forcing the model to be consistent with your theory [4]. Recent work has blurred the boundary between them, but they remain distinct in practice [2].

## Common Mistakes

- **Running CFA without a prior hypothesis.** CFA requires a structure to test. If you have none, start with EFA and use its output to build a testable model [2].
- **Reporting only chi-square.** Chi-square is sensitive to sample size and model complexity. Report CFI, TLI, RMSEA and SRMR together so readers see a consistent picture [3].
- **Treating a poor fit as proof the theory is wrong.** Poor fit can come from a small sample, misspecified error covariances or items that do not behave as expected. Diagnose before discarding the model [4].
- **Ignoring the measurement scale.** Likert items are ordinal. Fitting them with methods designed for continuous normal variables can distort results, so consider polychoric correlations with robust weighted least squares [5].
- **Keeping items with weak loadings.** Loadings below about |.4| suggest the item is not strongly related to the factor and may need to be dropped or reassigned [3].
- **Forgetting that factors can be correlated or uncorrelated.** This is a modeling decision you must make explicitly. Leaving it to software defaults can change your results [1].

## Limitations

CFA cannot tell you which of several plausible models is correct. It only tells you how well each model fits. Two different structures can produce similar fit statistics, so model comparison and theory both matter. A good fit is also not proof of validity, since fit indices can look acceptable for models that are substantively wrong.

Sample size is a hard constraint. Complex models with many items and factors need large samples for stable estimates, and ordinal data methods can struggle at smaller sizes even when robust estimators are used [5]. CFA also assumes the relationships between items and factors are linear and that the factor structure is the same for everyone in the sample, which may not hold across subgroups.

## Frequently Asked Questions

### What is the difference between confirmatory and exploratory factor analysis?

EFA lets the data reveal the factor structure, while CFA lets you pre-specify the structure and test whether it holds [1]. Use EFA when you are exploring a new scale and CFA when you are validating one that already has theoretical backing [2].

### What sample size do I need for confirmatory factor analysis?

There is no single cutoff, but larger is better and the requirement grows with model complexity. Ordinal data methods in particular can run into estimation problems at small sample sizes, while robust weighted least squares performs well across conditions [5]. Plan your sample size before collecting data using established sample size formulas and practical considerations.

### What fit indices should I report?

Report chi-square with its degrees of freedom, CFI, TLI, RMSEA and SRMR. These capture different aspects of fit, and software produces them as standard output [3]. Presenting several together gives a more complete picture than any single number.

### Can I use confirmatory factor analysis with Likert scale data?

Yes, but treat the items as ordinal. Fitting the model to polychoric correlations with weighted least squares or robust weighted least squares is the theoretically appropriate approach for Likert-type items [5]. This matters because ordinal scales are common in the social sciences and standard maximum likelihood methods assume continuous normal variables [6].

### What does a high factor correlation mean?

A very high correlation between two factors, such as 0.9900 in the example above, suggests the factors may not be empirically distinct. You may need to test a one-factor model and compare fit, or reconsider whether your items truly measure separate constructs. Factor correlations are part of the model specification you control [1].

If you want to understand how the individual item relationships feed into the model, review correlation versus covariance and communality in factor analysis before you specify your next model.

## References

1. [A Practical Introduction to Factor Analysis: Confirmatory Factor Analysis](https://stats.oarc.ucla.edu/spss/seminars/introduction-to-factor-analysis/a-practical-introduction-to-factor-analysis-confirmatory-factor-analysis/)
2. [Confirmatory Factor Analysis (CFA) in R with lavaan](https://stats.oarc.ucla.edu/r/seminars/rcfa/)
3. [An Overview of Confirmatory Factor Analysis and Item Response Analysis Applied to Instruments to Evaluate Primary Healthcare - PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC3399444/)
4. [Confirmatory factor analysis - Wikipedia](https://en.wikipedia.org/wiki/Confirmatory_factor_analysis)
5. [Flora DB, Curran PJ. (2004). An empirical evaluation of alternative methods of estimation for confirmatory factor analysis with ordinal data. Psychological methods](https://pmc.ncbi.nlm.nih.gov/articles/PMC3153362/)
6. [Rogers P. (2024). Best practices for your confirmatory factor analysis: A JASP and lavaan tutorial. Behavior research methods](https://pubmed.ncbi.nlm.nih.gov/38480677/)

## Related Articles

- [Multivariate Analysis: Definition, Methods and Examples](/blog/data-analysis/multivariate-analysis)
- [Bivariate Data: Definition, Examples and Analysis](/blog/data-analysis/bivariate-data-definition-examples)
- [Ranking Correlation Coefficient: Spearman and Kendall](/blog/data-analysis/ranking-correlation-coefficient)
- [What Is Communality in Factor Analysis? Definition and Examples](/blog/data-analysis/communality-factor-analysis)
- [Quadratic Regression Analysis: Equation and Example](/blog/data-analysis/quadratic-regression-analysis)
- [Qualitative Data Analysis: Coding, Theming, and Interpretation](/blog/guides/qualitative-data-analysis-coding-theming-and-interpretation)
- [Validity of a Study: Understanding Internal and External Validity](/blog/guides/validity-of-a-study-understanding-internal-and-external-validity)
- [What is Regression Analysis? A Practical Introduction](/blog/guides/what-is-regression-analysis-a-practical-introduction)