# ROC Curve and AUC Explained: How to Evaluate a Classifier

A receiver operating characteristic (ROC) curve plots the true positive rate against the false positive rate as you move a classifier's decision threshold from strict to lenient. The area under that curve (AUC) condenses the whole plot into one number between 0 and 1, where 0.5 means random guessing and 1.0 means perfect separation. This article shows how the roc and auc curve is built from real scores, what each point means, and when the metric misleads you.

## Quick Answer

- A classifier usually outputs a probability or score, not a label. You pick a threshold to convert scores into yes/no decisions.
- Each threshold produces one point on the ROC curve: false positive rate on the x-axis, true positive rate on the y-axis [1].
- Sweeping the threshold across all possible values traces the full curve. A curve hugging the top-left corner is better [2].
- AUC is the area under that curve. Random guessing gives 0.5, and a perfect classifier gives 1.0 [3].
- AUC is a ranking measure. It tells you how well the model separates classes, not how good any single threshold is.

## What ROC and AUC Mean

In plain terms, the ROC curve is a map of every trade-off your classifier can make between catching positives and raising false alarms. The AUC is the single score that summarizes that map.

The precise statistical definition: the ROC curve is the set of points $(FPR(t), TPR(t))$ for every threshold $t$, where $TPR(t)$ is the proportion of true positives correctly flagged and $FPR(t)$ is the proportion of true negatives incorrectly flagged. The AUC is the integral of that curve over the false positive rate from 0 to 1. Equivalently, AUC is the probability that the classifier ranks a randomly chosen positive case above a randomly chosen negative case [3].

The name comes from radar engineering, where operators had to separate real signals from noise. The same math now applies to spam filters, medical tests, and credit scoring.

## How It Works

Start with a confusion matrix at one threshold. Four counts come out of it: true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN). Two rates are built from those counts.

$$TPR = \frac{TP}{TP + FN} \qquad FPR = \frac{FP}{FP + TN}$$

- $TP$ is the number of positive cases the model flagged as positive.
- $FN$ is the number of positive cases the model missed.
- $FP$ is the number of negative cases the model wrongly flagged.
- $TN$ is the number of negative cases the model correctly left alone.
- $TPR$ is also called sensitivity or recall. It answers: of all real positives, how many did we catch?
- $FPR$ is the false alarm rate. It answers: of all real negatives, how many did we wrongly flag?

Now move the threshold. A high threshold flags only the most confident cases, so both TPR and FPR are low. Lower the threshold and you catch more positives, but you also flag more negatives. Each threshold gives one $(FPR, TPR)$ pair, and plotting all of them gives the curve [2].

The AUC is then computed with the trapezoidal rule, which sums the area of the strips under the curve:

$$AUC = \sum_{i} \frac{(FPR_{i+1} - FPR_i)(TPR_{i+1} + TPR_i)}{2}$$

In Python, `sklearn.metrics.roc_curve` returns the FPR and TPR arrays, and `sklearn.metrics.roc_auc_score` returns the AUC directly from scores [4][5]. If you are new to the tooling, the [Python for Machine Learning guide](/blog/data-analysis/python-for-machine-learning-guide) covers the setup.

## Worked Example

Take 10 test cases with a classifier score and a true label (1 = positive). Four cases are positive and six are negative, so $P = 4$ and $N = 6$.

| case_id | score | true_label |
|---------|-------|------------|
| 1 | 0.10 | 0 |
| 2 | 0.35 | 0 |
| 3 | 0.40 | 0 |
| 4 | 0.55 | 0 |
| 5 | 0.60 | 1 |
| 6 | 0.65 | 0 |
| 7 | 0.70 | 1 |
| 8 | 0.80 | 1 |
| 9 | 0.85 | 0 |
| 10 | 0.95 | 1 |

Sort the cases by score descending: 0.95, 0.85, 0.80, 0.70, 0.65, 0.60, 0.55, 0.40, 0.35, 0.10. Then walk down the list, treating each score as the threshold. Everything at or above the threshold is predicted positive.

| threshold | TP | FP | FN | TN | TPR | FPR |
|-----------|----|----|----|----|-----|-----|
| 0.95 | 1 | 0 | 3 | 6 | 0.2500 | 0.0000 |
| 0.85 | 1 | 1 | 3 | 5 | 0.2500 | 0.1667 |
| 0.80 | 2 | 1 | 2 | 5 | 0.5000 | 0.1667 |
| 0.70 | 3 | 1 | 1 | 5 | 0.7500 | 0.1667 |
| 0.65 | 3 | 2 | 1 | 4 | 0.7500 | 0.3333 |
| 0.60 | 4 | 2 | 0 | 4 | 1.0000 | 0.3333 |
| 0.55 | 4 | 3 | 0 | 3 | 1.0000 | 0.5000 |
| 0.40 | 4 | 4 | 0 | 2 | 1.0000 | 0.6667 |
| 0.35 | 4 | 5 | 0 | 1 | 1.0000 | 0.8333 |
| 0.10 | 4 | 6 | 0 | 0 | 1.0000 | 1.0000 |

Plotting those (FPR, TPR) pairs gives the ROC points: (0.00, 0.00), (0.00, 0.00), (0.00, 0.25), (0.17, 0.25), (0.17, 0.50), (0.17, 0.75), (0.33, 0.75), (0.33, 1.00), (0.50, 1.00), (0.67, 1.00), (0.83, 1.00), (1.00, 1.00). Applying the trapezoidal rule over those points gives **AUC = 0.8333**.

Here is the same computation in code.

```python
import numpy as np
y_true  = np.array([0,0,0,0,1,0,1,1,0,1])
y_score = np.array([0.10,0.35,0.40,0.55,0.60,0.65,0.70,0.80,0.85,0.95])
order = np.argsort(-y_score)
y_s, s_s = y_true[order], y_score[order]
P, N = int(y_true.sum()), len(y_true)-int(y_true.sum())
tp = fp = 0; fpr=[0.0]; tpr=[0.0]; prev=None
for i in range(len(s_s)):
    if s_s[i] != prev:
        tpr.append(tp/P); fpr.append(fp/N); prev=s_s[i]
    if y_s[i]==1: tp+=1
    else: fp+=1
tpr.append(tp/P); fpr.append(fp/N)
auc = np.trapezoid(tpr, fpr)  # NumPy 2.0+; older versions use np.trapz
print(f"{auc:.4f}")  # 0.8333
```

Output:

```
0.8333
```

The curve sits well above the diagonal, so the model separates the two classes better than chance.

## How to Interpret It

The AUC has a clean probability reading. An AUC of 0.8333 means that if you pick one positive case and one negative case at random, the model gives the positive case a higher score about 83% of the time [3]. That is the ranking interpretation, and it holds regardless of which threshold you eventually choose.

The curve shape matters too. A steep rise near the left edge means the model catches positives while keeping false alarms low, which is the ideal region [1]. A curve that only lifts off near the right side is weak. The points closest to the top-left corner mark the best-performing thresholds for that model [3].

AUC of 0.5 is the diagonal line, which is what random guessing or a coin flip produces [3]. An AUC below 0.5 means the model ranks in the wrong direction, which usually signals a labeling or sign error somewhere in the pipeline.

## When to Use It (and when not to)

Use AUC when you want to compare two models on the same task without committing to a threshold yet, and when the classes are roughly balanced [3]. It is also useful when you care about ranking quality, such as ordering cases for manual review.

Avoid relying on AUC alone when the classes are heavily imbalanced. With rare positives, a large negative class can inflate the false positive rate's stability while the metric hides poor performance on the minority class. Precision-recall curves and the area under them often give a clearer comparison in that setting [3]. Also skip AUC when the real cost of a false positive and a false negative is very different, because AUC treats all thresholds as equally interesting while your business only cares about one.

## ROC and AUC vs Precision-Recall Curves

The closest related idea is the precision-recall curve, which plots precision against recall across thresholds [3]. Both summarize threshold behavior, but they weight errors differently.

| Aspect | ROC / AUC | Precision-Recall |
|--------|-----------|------------------|
| Axes | TPR vs FPR | Precision vs Recall |
| Baseline | 0.5 (diagonal) | Depends on positive rate |
| Best for | Balanced classes, model ranking | Imbalanced classes |
| Sensitive to class balance | Less so | Yes, directly |
| Single score | AUC | Average precision or PR AUC |

If your positive class is rare, the precision-recall view usually tells the more honest story [3]. If you need a threshold-free ranking score on a balanced problem, AUC is the standard choice.

## Common Mistakes

- **Reporting AUC without ever picking a threshold.** AUC compares models, but production needs a decision rule. Fix: after comparing AUCs, choose a threshold from the curve based on the cost of each error type [3].
- **Treating 0.5 as a passing grade.** An AUC of 0.5 is random guessing, not a mediocre model [3]. Fix: set a target based on the problem, not on the midpoint of the scale.
- **Using AUC on badly imbalanced data and calling it done.** The metric can look healthy while minority-class recall is poor [3]. Fix: check a precision-recall curve alongside it.
- **Computing AUC from hard labels instead of scores.** Passing 0/1 predictions into an AUC function collapses the curve to a single point. Fix: pass probabilities or decision scores, as `roc_auc_score` expects [5].
- **Comparing AUCs across different datasets.** The score depends on the class mix and the problem. Fix: compare models on the same test set, ideally with cross-validated curves [1].
- **Ignoring the curve shape.** Two models can share an AUC while behaving very differently at the operating threshold you need [6]. Fix: plot both curves and inspect the region you will actually use.

## Limitations

AUC cannot tell you whether your chosen threshold is good for your use case. It averages performance over every possible threshold, including ones you would never deploy. A model with a strong AUC can still fail at the specific cutoff your workflow requires.

The metric also hides class-specific behavior. It does not report precision, and it does not tell you how many false positives you will actually see at a given operating point. For multiclass problems, you must choose an averaging strategy such as macro or micro, and the choice changes the number you get [7]. Finally, AUC says nothing about calibration. A model can rank cases perfectly and still output probabilities that are far from the true rates.

## Frequently Asked Questions

### What is a good AUC value?

There is no universal cutoff, but 0.5 is random guessing and 1.0 is perfect separation [3]. In many applied settings, values above 0.8 are considered strong, though the bar depends on the problem and the cost of errors. Always compare against a baseline model rather than an abstract number.

### Does a higher AUC always mean a better model?

Not always. AUC measures ranking across all thresholds, so a model with higher AUC can still perform worse at the one threshold you deploy [6]. Check the curve near your operating point and confirm with precision, recall, or cost-based metrics.

### Can AUC be below 0.5?

Yes. A value below 0.5 means the model tends to rank negatives above positives, which is worse than chance [3]. This often points to a flipped label encoding or a sign error in the score. Flipping the scores usually turns it into a value above 0.5.

### How do I compute ROC and AUC in Python?

Use `sklearn.metrics.roc_curve` to get the false positive and true positive rates, and `sklearn.metrics.roc_auc_score` to get the AUC from prediction scores [4][5]. Pass probabilities or decision function values, not hard class labels. For multiclass targets, choose an averaging strategy such as macro or micro [7].

### What is the difference between ROC AUC and accuracy?

Accuracy counts correct predictions at one fixed threshold and ignores how confident the model was. AUC uses the full range of scores and measures ranking quality across all thresholds [3]. Accuracy can look high on imbalanced data while AUC reveals weak separation, so the two answer different questions.

If you want to see how a probabilistic classifier produces the scores that feed this curve, the [Naive Bayes explainer](/blog/data-analysis/bayesian-classifiers-naive-bayes) walks through the mechanics, and the [introduction to statistical learning](/blog/guides/introduction-to-statistical-learning) covers the broader evaluation context.

## References

1. [Receiver Operating Characteristic (ROC) with cross validation, scikit-learn 1.9.1 documentation](https://scikit-learn.org/stable/auto_examples/model_selection/plot_roc_crossval.html)
2. [UDRC](https://udrc.ushe.edu/news/2021/research_skills/20210407AucRocCurve.html)
3. [Classification: ROC and AUC | Machine Learning | Google for Developers](https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc)
4. [3.4. Metrics and scoring: quantifying the quality of predictions, scikit-learn 1.9.1 documentation](https://scikit-learn.org/stable/modules/model_evaluation.html)
5. [roc_auc_score, scikit-learn 1.9.1 documentation](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_auc_score.html)
6. [ROC Curves and AUC for Models Used for Binary Classification | UVA Library](https://library.virginia.edu/data/articles/roc-curves-and-auc-for-models-used-for-binary-classification)
7. [Multiclass Receiver Operating Characteristic (ROC), scikit-learn 1.9.1 documentation](https://scikit-learn.org/stable/auto_examples/model_selection/plot_roc.html)

## Related Articles

- [Bayesian Classifiers: How Naive Bayes Works](/blog/data-analysis/bayesian-classifiers-naive-bayes)
- [Python for Machine Learning: A Beginner's Guide](/blog/data-analysis/python-for-machine-learning-guide)
- [Support Vector Machines (SVM): Definition and Examples](/blog/data-analysis/support-vector-machines-svm)
- [How to Interpret a Bradford Assay Standard Curve: Linearity, R², and Outliers](/knowledge/diagnostics/molecular/interpret-bradford-assay-standard-curve)
- [Introduction To Statistical Learning](/blog/guides/introduction-to-statistical-learning)