# Split-Half Reliability: Definition, Formula and Example

Split half reliability is a measure of internal consistency that correlates scores on two halves of a test or scale. Because splitting a test shortens it, the raw correlation between halves underestimates the reliability of the full test, so you apply the Spearman-Brown formula to correct it. This article explains the definition, the formula, and a full worked example with real numbers.

## Quick Answer

- Split half reliability splits a test into two comparable halves, correlates the half scores, and reports the result as an estimate of internal consistency [1].
- The raw half-to-half correlation $r$ describes only half-length tests, so it is corrected with the Spearman-Brown formula: $r_{sb} = \frac{2r}{1+r}$ [1].
- The correction assumes the two halves are parallel, meaning equal means, equal variances, and equal error variances.
- A common splitting rule is odd versus even items, which keeps item order and difficulty balanced across halves [1].
- Values closer to 1 indicate stronger internal consistency, and the corrected value is the one you report for the full test [1].

## What Split Half Reliability Means

In plain terms, split half reliability asks whether two parts of the same test give you consistent information about the same underlying construct. If a respondent scores high on one half, they should score high on the other half. When the two halves agree, the test is internally consistent.

The precise statistical definition is the correlation between scores on two equivalent halves of a test, adjusted for the fact that each half is only half as long as the full test. The raw correlation is a reliability estimate for a test of half the length. The Spearman-Brown correction rescales it to the length of the full instrument [1]. The correction was developed independently by Spearman (1910) and Brown (1910), which is why it carries both names [1].

This approach belongs to the family of single-session reliability metrics. It is widely used for self-report scales and for cognitive task data, where it often performs better than coefficient alpha [2].

## How It Works

The procedure has two stages. First, split the items into two halves and compute each respondent's total score on each half. Second, correlate the two sets of half scores with the Pearson correlation:

$$r = \frac{\text{cov}(X_{odd}, X_{even})}{s_{odd} \cdot s_{even}}$$

Here $X_{odd}$ and $X_{even}$ are the half scores, $\text{cov}$ is the covariance between them, and $s_{odd}$ and $s_{even}$ are the sample standard deviations of each half. If you need a refresher on those spread measures, see [measures of variability](/blog/data-analysis/measures-of-variability) and [mean and standard deviation](/blog/data-analysis/mean-and-standard-deviation).

That $r$ is the reliability of a half-length test, not the full test. The Spearman-Brown formula steps it back up:

$$r_{sb} = \frac{2r}{1+r}$$

Each symbol means the following.

- $r$ is the Pearson correlation between the two half scores.
- $2$ is the number of halves, since the full test is twice as long as each part.
- $r_{sb}$ is the estimated reliability of the full-length test.

The formula generalizes to any number of parts $k$ as $r_{sb} = \frac{kr}{1+(k-1)r}$, but the two-part version is what split half reliability uses.

## Worked Example

The dataset is a 10-item survey split into odd and even halves, with 20 respondents. Each respondent has a total score on the odd items and a total score on the even items.

| Respondent | Odd items | Even items |
|---|---|---|
| 1 | 18 | 17 |
| 2 | 22 | 21 |
| 3 | 15 | 16 |
| 4 | 25 | 24 |
| 5 | 19 | 20 |
| 6 | 21 | 19 |
| 7 | 14 | 15 |
| 8 | 23 | 22 |
| 9 | 17 | 18 |
| 10 | 20 | 21 |
| 11 | 16 | 14 |
| 12 | 24 | 25 |
| 13 | 18 | 19 |
| 14 | 22 | 23 |
| 15 | 13 | 14 |
| 16 | 26 | 25 |
| 17 | 19 | 18 |
| 18 | 21 | 22 |
| 19 | 15 | 17 |
| 20 | 23 | 21 |

The steps run as follows.

1. Number of respondents: $n = 20$.
2. Odd-item half mean: $19.5500$.
3. Even-item half mean: $19.5500$.
4. Odd-item half SD (n-1): $3.7763$.
5. Even-item half SD (n-1): $3.4255$.
6. Covariance of halves: $12.1553$.
7. Pearson correlation between halves: $r = 12.1553 / (3.7763 \times 3.4255) = 0.9397$.
8. Spearman-Brown correction: $r_{sb} = (2 \times 0.9397) / (1 + 0.9397) = 0.9689$.

The two halves have identical means, which is a good sign for the parallel-halves assumption. The corrected reliability of $0.9689$ is higher than the raw correlation of $0.9397$, exactly as the correction predicts.

Here is the same computation in Python.

```python
import numpy as np
odd = np.array([18,22,15,25,19,21,14,23,17,20,16,24,18,22,13,26,19,21,15,23])
even = np.array([17,21,16,24,20,19,15,22,18,21,14,25,19,23,14,25,18,22,17,21])
r = np.corrcoef(odd, even)[0,1]
r_sb = 2*r/(1+r)
print(f"r(half) = {r:.4f}")
print(f"Spearman-Brown reliability = {r_sb:.4f}")
```

Output:

```
r(half) = 0.9397
Spearman-Brown reliability = 0.9689
```

A scatterplot of the two half scores would show points clustered tightly along a rising line, consistent with $r = 0.9397$ and a corrected reliability of $0.9689$.

## How to Interpret It

The corrected coefficient runs from 0 to 1 in typical use. Higher values mean the two halves rank respondents in nearly the same order, which supports treating the items as measuring one consistent construct. Lower values mean the halves disagree, which weakens the case for combining all items into a single score.

There is no universal cutoff, but many researchers treat values of 0.70 and above as acceptable for research purposes and 0.90 and above as strong. The value of $0.9689$ in the example is very high, which suggests the odd and even halves are nearly interchangeable.

Interpret the corrected value, not the raw one. Reporting $0.9397$ as the reliability of a 10-item test understates it, because that number describes a 5-item test.

## When to Use It (and when not to)

Use split half reliability when you have a single administration of a test and want an internal consistency estimate without a second testing session. It suits unidimensional scales where all items are meant to measure the same thing. It is also a strong choice for reaction-time and cognitive task data, where split-half methods tend to be more accurate than coefficient alpha [2].

Avoid it when the test is deliberately multidimensional. If a scale contains distinct subtests, the halves may not be comparable, and the estimate will be misleading. Avoid it when the two halves are clearly not parallel, for example when the odd items are much harder than the even items. In that case the halves differ in difficulty and the correction rests on an assumption that fails.

You should also be careful with very short tests. With only a handful of items, each half contains very few items, and the estimate becomes unstable.

## Split Half Reliability vs Cronbach's Alpha

Both estimate internal consistency, but they get there differently. Cronbach's alpha is the mean of all possible split-half coefficients across every way of dividing the test [3]. Split half reliability uses one specific split.

| Feature | Split Half Reliability | Cronbach's Alpha |
|---|---|---|
| Basis | One chosen split of items | All possible splits averaged [3] |
| Correction | Spearman-Brown applied [1] | Built into the formula |
| Sensitivity to split | Depends on how you split | Averages over splits |
| Typical use | Scales and cognitive tasks [2] | Self-report scales |
| Behavior on RT data | Often more accurate [2] | Can be negatively biased [2] |

Because alpha averages over all splits, it avoids the arbitrariness of choosing one split. Because split half reliability uses a single split, it can be higher or lower depending on that choice, which is why the splitting method matters [4].

## Common Mistakes

- Reporting the raw half correlation as the test's reliability. Fix: always apply the Spearman-Brown correction and report $r_{sb}$ [1].
- Using a single arbitrary split and ignoring alternatives. Fix: try odd-even, first-second half, and random splits, and compare results [4].
- Assuming the halves are parallel without checking. Fix: compare the half means and standard deviations, as in the worked example where both means equal 19.5500.
- Splitting a multidimensional scale into halves that measure different constructs. Fix: divide distinct subtests before estimating reliability [3].
- Applying the correction when the halves differ in length. Fix: use the general $k$-part Spearman-Brown formula, or make the halves equal length.
- Treating a high value as proof of validity. Fix: remember that reliability is about consistency, not about measuring the right construct.

## Limitations

Split half reliability cannot tell you whether your test measures the construct you intend. A scale can be highly consistent and still measure the wrong thing. It also depends on the specific split you choose, so two analysts can get different answers from the same data [4]. The Spearman-Brown correction assumes parallel halves, and when that assumption fails, the corrected value can be biased.

The method also says nothing about stability over time. It uses a single administration, so it captures consistency within one session, not test-retest reliability. For cognitive task data, the choice of splitting method can meaningfully change the estimate, and Monte Carlo splitting is often recommended as the most resistant to confounding effects [4]. Permutation-based split-half methods, which average many random splits, tend to be more accurate than a single split [2].

## Frequently Asked Questions

### What is a good split half reliability value?

There is no fixed threshold, but values of 0.70 and above are commonly treated as acceptable for research, and values of 0.90 and above as strong. The example here produced 0.9689, which is very high. The right cutoff depends on your field and how the scores will be used.

### Why do I need the Spearman-Brown formula?

The raw correlation is computed on half-length tests, and shorter tests are less reliable. The Spearman-Brown formula corrects the estimate back to full-test length [1]. Without it, you understate the reliability of the instrument you actually plan to use.

### Should I use odd-even or first-second half splitting?

Odd-even splitting keeps item order and difficulty balanced across halves, which is why it is a common default [1]. First-second half splitting can be confounded by fatigue, practice, or warm-up effects that change across the test [4]. If you are unsure, compare both and check whether the estimates agree.

### Is split half reliability the same as Cronbach's alpha?

No. Alpha is the mean of all possible split-half coefficients, while split half reliability uses one specific split [3]. They usually agree closely on unidimensional scales, but they can diverge on reaction-time data, where alpha is often lower [2].

### Can I compute split half reliability in R?

Yes. The `item_split_half` function in the performance package computes split-half reliability for items in a data frame, including the Spearman-Brown adjustment, and splits by selecting odd versus even columns [1]. It returns both the raw split-half reliability and the corrected value.

## References

1. [R: Split-Half Reliability](https://search.r-project.org/CRAN/refmans/performance/html/item_split_half.html)
2. [Reaction-time task reliability is more accurately computed with permutation-based split-half correlations than with Cronbach’s alpha - PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC12000231/)
3. [Cronbach LJ (1951). Coefficient Alpha and the Internal Structure of Tests. Psychometrika](https://doi.org/10.1007/bf02310555)
4. [Pronk T, Molenaar D, Wiers RW, Murre J. (2022). Methods to split cognitive task data for estimating split-half reliability: A comprehensive review and systematic assessment. Psychonomic bulletin & review](https://pmc.ncbi.nlm.nih.gov/articles/PMC8858277/)

## Further Reading

- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)

## Related Articles

- [Measures of Variability: Range, Variance and Standard Deviation](/blog/data-analysis/measures-of-variability)
- [Negative Binomial Distribution: Formula and Examples](/blog/data-analysis/negative-binomial-distribution-formula-examples)
- [Student's t-Distribution: Definition, Formula and Examples](/blog/data-analysis/students-t-distribution-definition-formula)
- [Uniform Distribution: Definition, Formula and Examples](/blog/data-analysis/uniform-distribution)
- [Mean and Standard Deviation: Definition, Formula and Examples](/blog/data-analysis/mean-and-standard-deviation)
- [Median Absolute Deviation: Formula and Worked Example](/blog/research-skills/median-absolute-deviation-formula-and-worked-example)
- [Statistical Range: Definition, Calculation, and Applications](/blog/guides/statistical-range-definition-calculation-and-applications)
- [Replication Fork Short: Definition and Key Concepts](/knowledge/molecular-biology/replication-fork-short)