Cohen's Kappa: How to Calculate and Interpret Inter-Rater Agreement

By Dr. Zubair Khalid, DVM, MS, PhD ·

Cohen's Kappa: How to Calculate and Interpret Inter-Rater Agreement

Cohen's kappa is a statistic that measures how much two raters agree when they classify the same items into categories, after subtracting the agreement you would expect from chance alone. It was introduced by Jacob Cohen in 1960 in "A Coefficient of Agreement for Nominal Scales" [1], and it remains a widely used measure of inter-rater reliability for nominal data in clinical research, diagnostic studies and machine learning evaluation.

You will meet kappa whenever two people score the same set of things: two pathologists reading biopsy slides, two radiologists labeling images, two coders tagging interview transcripts, or a human annotator checked against a model's predictions. Reviewers and editors expect it in methods sections, and getting it wrong is easy because raw agreement and kappa can tell very different stories about the same data.

Quick Answer

  • Cohen's kappa measures agreement between exactly two raters on categorical items, corrected for the agreement expected by chance [2].
  • The formula is $\kappa = \frac{p_o - p_e}{1 - p_e}$, where $p_o$ is the observed proportion of agreement and $p_e$ is the proportion expected by chance [2].
  • Kappa runs from -1 to +1: 0 means chance-level agreement, 1 means perfect agreement, and negative values mean agreement worse than chance [2].
  • A rough reading guide: below 0.60 is generally inadequate for health research, 0.60 to 0.79 is moderate, 0.80 to 0.90 is strong, and above 0.90 is almost perfect [2].
  • Always report kappa with its 95% confidence interval, the number of items and raters, the category totals and the raw percent agreement.

What Is Cohen's Kappa?

Inter-rater reliability is the extent to which data collectors assign the same score to the same variable. It was traditionally measured as percent agreement, the number of agreements divided by the total number of scores [2]. Percent agreement is intuitive, but Cohen pointed out that it cannot account for agreement expected by chance when raters guess on uncertain items, so he developed kappa to correct for it [2].

The logic is straightforward. Suppose two raters classify 100 slides and agree on 85 of them. Some of that agreement happened because both raters used the same common category often, not because they truly saw the same thing. Kappa estimates how much of the agreement above chance the raters actually achieved, and expresses it as a proportion of the maximum possible agreement above chance.

The chance component comes from each rater's own label frequencies, called marginals. If rater A calls 90% of slides positive and rater B calls 90% positive, they will agree on many positives even if neither is reading the slides carefully. Kappa estimates the expected agreement when both annotators assign labels randomly from their own label frequencies [10].

How to Calculate Cohen's Kappa

Start with a contingency table of the two raters' labels. For two categories, that is a 2x2 table. The observed agreement $p_o$ is the sum of the diagonal cells (where the raters match) divided by the total number of items.

The chance agreement $p_e$ for a 2x2 table is:

$$p_e = \frac{\left(\frac{R_1 \times C_1}{n}\right) + \left(\frac{R_2 \times C_2}{n}\right)}{n}$$

Here $R_1$ and $R_2$ are the row totals, $C_1$ and $C_2$ are the column totals, and $n$ is the number of items rated, not the number of raters [2]. Each term inside the parentheses is the count of agreements expected by chance in that cell; dividing by $n$ converts the sum to a proportion.

Then plug both into the kappa formula:

$$\kappa = \frac{p_o - p_e}{1 - p_e}$$

The numerator is the agreement achieved beyond chance. The denominator is the agreement available beyond chance. The ratio tells you what fraction of the available headroom the raters captured.

For a confidence interval, McHugh gives the large-sample standard error:

$$SE = \sqrt{\frac{p_o(1 - p_o)}{n(1 - p_e)^2}}$$

and the 95% CI as $\kappa \pm 1.96 \times SE$ [2]. This is a simplified approximation; exact large-sample variance formulas differ slightly, and software such as statsmodels uses a different variance formula that produces nearly identical intervals in practice.

Worked Example

Two pathologists, A and B, classify the same 100 biopsy slides as positive or negative. The 2x2 table (rows are A, columns are B):

B positiveB negativeRow total
A positive401050
A negative54550
Column total4555100

Observed agreement: $p_o = (40 + 45)/100 = 0.85$.

Chance agreement: $p_e = (50 \times 45 + 50 \times 55)/100^2 = (2250 + 2750)/10000 = 0.50$.

Kappa: $\kappa = (0.85 - 0.50)/(1 - 0.50) = 0.70$.

Standard error: $SE = \sqrt{0.85 \times 0.15 / (100 \times 0.50^2)} = 0.0714$.

95% CI: $0.70 \pm 1.96 \times 0.0714 = 0.56$ to $0.84$.

Cross-checks: statsmodels gives kappa 0.700, SE 0.0711, CI 0.561 to 0.839; scikit-learn's cohen_kappa_score on the 100 paired labels gives 0.700. The Landis and Koch benchmarks label 0.70 as substantial (0.61 to 0.80) [4], while McHugh's stricter scale calls it moderate (0.60 to 0.79) [2]. Note that the lower confidence limit, 0.56, falls below McHugh's 0.60 adequacy line, so the study's agreement is not as solid as the point estimate suggests.

How to Interpret Cohen's Kappa

The number itself is only half the story. Kappa has no units and no natural scale, so interpretation depends on benchmarks and context.

The most widely reproduced benchmarks come from Landis and Koch (1977), who presented methods for observer agreement on categorical data [3]. As commonly reproduced: below 0 poor, 0.00 to 0.20 slight, 0.21 to 0.40 fair, 0.41 to 0.60 moderate, 0.61 to 0.80 substantial, 0.81 to 1.00 almost perfect [4]. Reproductions differ on whether 0.00 belongs to "poor" or "slight," and Landis and Koch are widely reported to have described their divisions as arbitrary.

McHugh argues these benchmarks are too lenient for health research because they allow kappa as low as 0.41 to be seen as acceptable and very little agreement to be called "substantial" [2]. Her stricter scale: 0 to 0.20 none, 0.21 to 0.39 minimal, 0.40 to 0.59 weak, 0.60 to 0.79 moderate, 0.80 to 0.90 strong, above 0.90 almost perfect [2]. She summarizes that any kappa below 0.60 indicates inadequate agreement and that little confidence should be placed in such study results [2].

KappaLandis and Koch [4]McHugh [2]
0.00 to 0.20slightnone (0 to 0.20)
0.21 to 0.40fairminimal (0.21 to 0.39)
0.41 to 0.60moderateweak (0.40 to 0.59)
0.61 to 0.80substantialmoderate (0.60 to 0.79)
0.81 to 1.00almost perfectstrong (0.80 to 0.90) to almost perfect (above 0.90)

Two cautions apply to any benchmark. First, confidence intervals matter more for kappa than for percent agreement, because kappa is an estimate, not a direct measurement [2]. Second, the acceptable threshold depends on the stakes: a kappa of 0.65 may be fine for a screening triage task and unacceptable for a diagnostic label that drives treatment.

Kappa vs Percent Agreement, Correlation and Other Statistics

Percent agreement and kappa answer different questions. Percent agreement tells you how often the raters matched. Kappa tells you how much of that matching exceeds chance. In the worked example, 85% raw agreement corresponds to kappa 0.70 once 50% chance agreement is removed. McHugh's own example shows the gap can be much larger: percent agreement 0.94, chance agreement 0.57, n = 222, kappa 0.85 [2]. The greater the expected chance agreement, the lower the kappa for a given observed agreement [2].

McHugh concludes that perhaps the best advice is to calculate both percent agreement and kappa; percent agreement is a direct measure, whereas kappa is an estimate, so confidence intervals matter more for kappa [2]. Many texts recommend 80% agreement as the minimum acceptable inter-rater agreement [2].

Correlation coefficients such as Pearson's r are a poor reflection of agreement between raters, because they measure association, not identical scoring [2]. Two raters can correlate almost perfectly while one consistently scores two points higher than the other; r will not notice, but kappa will.

For more than two raters, Cohen's kappa does not apply directly. Cohen specifically discussed two raters in his papers [2]. Fleiss (1971) generalized kappa to the case where each subject is rated by the same number of raters but the raters need not be the same people for every subject [8]. McHugh lists Fleiss kappa as the adaptation of Cohen's kappa for three or more raters, alongside alternatives such as the intraclass correlation coefficient and Krippendorff's alpha [2]. Fleiss kappa is not a direct multi-rater average of Cohen's kappa; with two raters it reduces to Scott's pi, not Cohen's kappa.

For ordered categories, weighted kappa gives partial credit for disagreements. Cohen introduced it in 1968 [7]. Linear weights penalize disagreements in proportion to their distance in categories; quadratic weights penalize by the squared distance, so near-misses cost little. scikit-learn's cohen_kappa_score supports weights=None, 'linear' and 'quadratic' [10]. For ordinal data, weighted kappa is usually higher than unweighted kappa when most disagreements are between adjacent categories.

Common Mistakes

  • Reporting kappa without percent agreement. Kappa is an estimate that depends on prevalence and bias; percent agreement is a direct measure. Report both [2].
  • Treating kappa as a correlation. Kappa measures agreement on identical labels, not association. A high Pearson's r can coexist with systematic disagreement [2].
  • Using Cohen's kappa with three or more raters. Cohen's kappa is for two raters [2]. Use Fleiss kappa or another multi-rater statistic [8].
  • Ignoring prevalence and bias. Kappa is affected by bias between observers and by the distribution of data across categories [6]. Report marginal totals so readers can see these effects.
  • Applying a single benchmark to every study. Landis and Koch's divisions are widely described as arbitrary, and McHugh argues they are too lenient for health research [2]. Choose a threshold that matches the stakes.
  • Skipping the confidence interval. A kappa of 0.70 with a CI of 0.56 to 0.84 is a different claim from a kappa of 0.70 with a CI of 0.65 to 0.75. The interval tells you how much to trust the estimate [2].
  • Running too few comparisons. As a general heuristic, a kappa study should include no fewer than 30 comparisons [2].

Limitations

Kappa has two well-documented paradoxes. Feinstein and Cicchetti (1990) described how a high observed agreement can be drastically lowered by strong imbalance in the marginal totals, and how kappa is higher with asymmetrical than with symmetrical imbalance [5]. Byrt, Bishop and Carlin (1993) showed kappa is affected by bias between observers and by the distribution of data across categories, and recommend reporting indicators of bias and prevalence alongside kappa [6].

A concrete illustration: two tables with 100 items can have $p_o = 0.85$ and kappa = 0.70, or $p_o = 0.90$ and kappa = 0.44, because skewed prevalence in the second table raises $p_e$ to 0.82. Higher raw agreement, lower kappa. This is why reporting kappa alone can mislead.

Sim and Wright (2005) list prevalence, bias and non-independent ratings as factors that influence the magnitude of kappa, and discuss confidence intervals and sample size for kappa studies [9]. Non-independent ratings occur, for example, when raters discuss cases before scoring.

The standard error formula given by McHugh is a simplified approximation; exact large-sample variance formulas differ slightly, as the statsmodels comparison shows. Software implementations may use different variance formulas, so check the current documentation for the tool you use.

Prevalence- and bias-adjusted kappa (PABAK = $2p_o - 1$) is associated with Byrt et al. 1993, but the formula has not been checked here against the original paper. If you use it, check the original source.

Frequently Asked Questions

What is a good Cohen's kappa?

There is no universal cutoff. McHugh argues that any kappa below 0.60 indicates inadequate agreement for health research [2], while the Landis and Koch benchmarks treat 0.61 to 0.80 as substantial [4]. The right threshold depends on the stakes of the classification task and should be set before data collection.

How do I calculate Cohen's kappa by hand?

Build a contingency table of the two raters' labels, compute observed agreement as the diagonal sum divided by the total, compute chance agreement from the marginal totals, then apply $\kappa = (p_o - p_e)/(1 - p_e)$ [2]. For the worked example above, that gives 0.70.

What is the difference between kappa and percent agreement?

Percent agreement is the raw proportion of items on which the raters matched. Kappa subtracts the agreement expected by chance and rescales the remainder [2]. Percent agreement is a direct measure; kappa is an estimate, so confidence intervals matter more for kappa [2].

When should I use weighted kappa instead of unweighted kappa?

Use weighted kappa when the categories are ordered, such as mild, moderate and severe. Weighted kappa gives partial credit for disagreements on adjacent categories, so a mild-versus-moderate disagreement costs less than a mild-versus-severe one [7]. Linear weights penalize by distance; quadratic weights penalize by squared distance [10].

What is the difference between Cohen's kappa and Fleiss kappa?

Cohen's kappa handles exactly two raters [2]. Fleiss kappa handles three or more raters, and the raters need not be the same people for every subject, as long as each subject is rated by the same number of raters [8]. Fleiss kappa is not a direct average of Cohen's kappa values.

References

  1. Cohen J. A Coefficient of Agreement for Nominal Scales. Educ Psychol Meas 1960;20:37-46
  2. McHugh ML. Interrater reliability: the kappa statistic. Biochem Med 2012;22:276-282
  3. Landis JR, Koch GG. The Measurement of Observer Agreement for Categorical Data. Biometrics 1977;33:159-174
  4. Piron S et al. Interrater reliability of 18F-PSMA-11 imaging (Landis and Koch categories, Table 2). EJNMMI Res 2020;10
  5. Feinstein AR, Cicchetti DV. High agreement but low kappa: I. J Clin Epidemiol 1990;43:543-54990158-L)
  6. Byrt T, Bishop J, Carlin JB. Bias, prevalence and kappa. J Clin Epidemiol 1993;46:423-42990018-V)
  7. Cohen J. Weighted kappa. Psychol Bull 1968;70:213-220
  8. Fleiss JL. Measuring nominal scale agreement among many raters. Psychol Bull 1971;76:378-382
  9. Sim J, Wright CC. The kappa statistic in reliability studies. Phys Ther 2005;85:257-268
  10. scikit-learn: sklearn.metrics.cohen_kappa_score

Related Articles