# F1 Score: Definition, Formula and When to Use It

The F1 score is a single number that balances precision and recall, so you can judge a classifier when both false positives and false negatives matter. It is the harmonic mean of the two, which means it stays low unless both are reasonably high. This article covers the definition, the formula, a worked example, and the situations where the F1 score is a better choice than accuracy.

## Quick Answer

- The F1 score is the harmonic mean of precision and recall, ranging from 0 (worst) to 1 (best).
- Formula: $F_1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}$.
- It rewards models that do well on both metrics at once, so a high precision with poor recall still gives a low F1.
- Use it when classes are imbalanced or when false positives and false negatives carry similar cost.
- Accuracy can look great on imbalanced data while the model misses the class you care about, which is exactly the gap F1 fills.

## What the F1 Score Means

In plain terms, the F1 score answers the question: "How well does this model find the positive class without also flagging too many wrong cases?" It is a single summary of two competing goals. Precision asks how many of the items you flagged as positive were actually positive. Recall asks how many of the truly positive items you managed to find.

The precise definition: the F1 score is the harmonic mean of precision and recall, reaching its optimal value at 1 and its worst value at 0 [1]. The harmonic mean is used instead of the arithmetic mean because it punishes imbalance. If precision is 1.0 and recall is 0.1, the arithmetic mean is 0.55, which sounds acceptable. The harmonic mean is about 0.18, which correctly signals that the model is failing at one of its two jobs.

The F1 score is a special case of the more general F-beta score, where the beta parameter sets the weight of precision in the combined score. A beta below 1 lends more weight to precision, while a beta above 1 favors recall [1]. F1 is simply the case where beta equals 1, so precision and recall are weighted equally.

## How It Works

Everything starts with the confusion matrix, which counts four outcomes for a binary classifier.

| Term | Meaning |
|---|---|
| TP (true positive) | Predicted positive, actually positive |
| FP (false positive) | Predicted positive, actually negative |
| FN (false negative) | Predicted negative, actually positive |
| TN (true negative) | Predicted negative, actually negative |

From those counts you get precision and recall.

$$\text{Precision} = \frac{TP}{TP + FP}$$

$$\text{Recall} = \frac{TP}{TP + FN}$$

Precision is the share of positive predictions that were correct. Recall is the share of actual positives that were caught. They pull against each other: flag more items as positive and recall rises while precision usually falls.

The F1 score combines them with the harmonic mean.

$$F_1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}$$

There is also a form that skips the intermediate steps and works directly from the counts.

$$F_1 = \frac{2 \cdot TP}{2 \cdot TP + FP + FN}$$

Both forms give the same answer. The count-based version is handy when you have the confusion matrix but have not computed precision and recall separately.

## Worked Example

Take a binary classifier evaluated on 240 cases. The confusion matrix is TP = 80, FP = 20, FN = 40, TN = 100.

| TP | FP | FN | TN |
|---|---|---|---|
| 80 | 20 | 40 | 100 |

Step 1, precision:

$$80 / (80 + 20) = 0.8000$$

Step 2, recall:

$$80 / (80 + 40) = 0.6667$$

Step 3, F1 from precision and recall:

$$2 \cdot (0.8000 \cdot 0.6667) / (0.8000 + 0.6667) = 0.7273$$

Step 4, F1 from the count-based form, as a cross-check:

$$2 \cdot 80 / (2 \cdot 80 + 20 + 40) = 0.7273$$

Both routes land on 0.7273. The same values come out of scikit-learn.

```python
from sklearn.metrics import precision_score, recall_score, f1_score
y_true = [1]*120 + [0]*120
y_pred = [1]*80 + [0]*40 + [1]*20 + [0]*100
p = precision_score(y_true, y_pred)  # 0.8000
r = recall_score(y_true, y_pred)     # 0.6667
f = f1_score(y_true, y_pred)         # 0.7273
print(f"precision={p:.4f}, recall={r:.4f}, f1={f:.4f}")
```

Output:

```
precision=0.8000, recall=0.6667, f1=0.7273
```

Notice that accuracy on this same data is 0.75, which is higher than the F1 score. The model gets three quarters of all predictions right, yet it misses 40 of the 120 positive cases. Accuracy hides that miss because the 100 true negatives inflate the total. F1 does not.

## How to Interpret It

The F1 score runs from 0 to 1, and higher is better. A value of 1 means perfect precision and perfect recall. A value of 0 means the model found no true positives at all.

The score is only meaningful next to the numbers it came from. An F1 of 0.73 could come from precision 0.80 and recall 0.67, or from precision 0.67 and recall 0.80, or from a dozen other pairs. Always report precision and recall alongside F1 so readers can see which side is weaker.

A useful habit is to compare F1 against a baseline. On imbalanced data, a naive model that always predicts the negative majority class can score high accuracy with an F1 of 0, so a real model's F1 of 0.73 would be a large improvement. The absolute number matters less than the gap between your model and the trivial baseline.

Because F1 is a harmonic mean, it is sensitive to the smaller of the two values. If either precision or recall drops toward zero, F1 drops toward zero with it. That property is the whole point: it refuses to let a strong score on one metric mask a weak score on the other.

## When to Use It (and when not to)

Use F1 when the positive class is rare and you care about finding it. Fraud detection, disease screening, and rare-event tagging are typical cases. With 1 percent positives, a model that predicts "negative" for everything can hit 99 percent accuracy while catching nothing. F1 exposes that failure.

Use F1 when false positives and false negatives cost roughly the same. If both mistakes are equally bad, equal weighting is the right call, and that is exactly what F1 does.

Use F1 when you need one number to rank models or tune a threshold. It is a compact way to compare candidates without staring at two separate metrics.

Do not use F1 when the two error types have very different costs. If a missed fraud case is far worse than a false alarm, you want recall weighted more heavily, which points to an F-beta score with beta above 1 [1]. If false alarms are the bigger problem, weight precision instead.

Do not use F1 when you need a probability-calibrated view of performance or when true negatives matter. F1 ignores TN entirely, so it says nothing about how well the model handles the negative class. For balanced problems where every class matters equally, accuracy or a per-class breakdown may serve you better.

## F1 Score vs Accuracy

Accuracy is the share of all predictions that are correct. F1 is the harmonic mean of precision and recall. They answer different questions, and they diverge sharply on imbalanced data.

| Property | F1 Score | Accuracy |
|---|---|---|
| What it measures | Balance of precision and recall | Share of all correct predictions |
| Uses true negatives | No | Yes |
| Behavior on imbalanced data | Reflects performance on the positive class | Can look high while the minority class is ignored |
| Range | 0 to 1 | 0 to 1 |
| Best for | Rare or costly positive class | Balanced classes with equal error costs |

In the worked example, accuracy was 0.75 and F1 was 0.7273. The gap is small because the classes are close to balanced. Push the positive class down to 5 percent of the data and the gap widens dramatically, with accuracy staying high while F1 collapses.

## Common Mistakes

- **Reporting F1 without precision and recall.** The single number hides which side is weak. Fix: always show all three so the trade-off is visible.
- **Using F1 on balanced data by default.** When classes are balanced and errors cost the same, accuracy is simpler and easier to explain. Fix: check the class balance first and pick the metric that matches the problem.
- **Assuming a high F1 means a good model in every sense.** F1 ignores true negatives, so a model can score well while misclassifying most of the negative class. Fix: inspect the full confusion matrix.
- **Comparing F1 scores across datasets with different class balances.** The same model can post different F1 values on different samples. Fix: compare models on the same data and the same split.
- **Tuning the threshold to maximize F1 when the costs are unequal.** F1 assumes equal weight on precision and recall. Fix: use an F-beta score with a beta that reflects your actual costs [1].
- **Forgetting how the averaging mode changes the number in multiclass problems.** Macro, micro, and weighted averages give different results, and macro averaging can produce an F-score that is not between precision and recall [1]. Fix: state which average you used.

## Limitations

F1 cannot tell you whether your model is well calibrated, and it cannot tell you how the model performs on the negative class, because true negatives never enter the formula. A model that flags almost everything as positive can still post a respectable F1 if it catches most true positives, even though it floods you with false alarms. The score also depends on the decision threshold, so a single F1 value describes one operating point, not the model as a whole.

F1 treats precision and recall as equally important, which is a modeling choice, not a fact about your problem. When the real-world costs are lopsided, F1 can point you toward the wrong model. It is also sensitive to how you define the positive class and, in multiclass settings, to the averaging method you choose. Treat F1 as one view of performance, not the final word.

## Frequently Asked Questions

### What is a good F1 score?

There is no universal cutoff. A good F1 depends on the problem, the class balance, and the baseline. Compare your score against a trivial model that always predicts the majority class, and against any existing system you are replacing. A model that beats the baseline by a clear margin is doing real work, even if the absolute number looks modest.

### What is the difference between F1 and F-beta?

F1 is the special case of the F-beta score where beta equals 1, so precision and recall carry equal weight. Setting beta below 1 lends more weight to precision, and setting beta above 1 favors recall [1]. Use F-beta when one type of error costs more than the other.

### Can the F1 score be higher than accuracy?

Yes. When the positive class is the majority, a model can have high precision and recall on it while getting most of the few negative cases wrong, which lowers accuracy below F1. For example, with 90 positives and 10 negatives, TP = 85, FN = 5, FP = 10 and TN = 0 give accuracy 0.85 and F1 about 0.92. The two metrics measure different things, so neither is always larger.

### Does F1 work for multiclass problems?

Yes, but you must choose an averaging method. Macro averaging computes the metric per class and takes the unweighted mean, micro averaging pools all classes together, and weighted averaging weights each class by its size. Each gives a different number, so report which one you used.

### Why use the harmonic mean instead of the arithmetic mean?

The harmonic mean drops sharply when either input is small, so it will not reward a model that is excellent at precision but terrible at recall. The arithmetic mean would let a strong value paper over a weak one. The harmonic mean keeps both metrics honest.

If you want to see how a single score can summarize a distribution, the same logic behind standardizing values appears in [what a z-score measures](/blog/data-analysis/what-is-a-z-score).

## References

1. [sklearn.metrics.fbeta_score, scikit-learn 0.16.1 documentation](http://scikit-learn.org/0.16/modules/generated/sklearn.metrics.fbeta_score.html)

## Further Reading

- [Lever J, Krzywinski M, Altman N (2016). Classification evaluation. Nature Methods](https://doi.org/10.1038/nmeth.3945)
- [scikit-learn User Guide](https://scikit-learn.org/stable/user_guide.html)
- [Wilson G, Bryan J, Cranston K et al. (2017). Good enough practices in scientific computing. PLOS Computational Biology](https://doi.org/10.1371/journal.pcbi.1005510)
- [Lever J, Krzywinski M, Altman N (2016). Model selection and overfitting. Nature Methods](https://doi.org/10.1038/nmeth.3968)
- [Saito T, Rehmsmeier M (2015). The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE](https://doi.org/10.1371/journal.pone.0118432)

## Related Articles

- [What Is a Z-Score? Definition, Formula and Examples](/blog/data-analysis/what-is-a-z-score)
- [What Is Boosting? Definition and Examples for Analysts](/blog/data-analysis/what-is-boosting-algorithms)
- [Eigenvalues and Eigenvectors: Definition and Examples](/blog/data-analysis/eigenvalues-eigenvectors-definition-examples)
- [Euclidean Distance: Definition, Formula and Examples](/blog/data-analysis/euclidean-distance-definition-formula)
- [Sigmoid Function: Definition, Formula and Examples](/blog/data-analysis/sigmoid-function-definition-formula)