Bayes' Theorem: Definition, Formula and Examples

By Dr. Zubair Khalid, DVM, MS, PhD ·

Bayes' Theorem: Definition, Formula and Examples

Bayes' theorem is a formula for reversing a conditional probability. If you know how likely a positive test is among sick people, the theorem tells you how likely sickness is among people who test positive. It combines a prior belief with new evidence to produce an updated probability, which is why the Bayesian theorem sits at the center of Bayesian statistics [1].

Quick Answer

  • Bayes' theorem computes $P(H \mid E)$, the probability of a hypothesis given evidence, from the reverse conditional probability $P(E \mid H)$.
  • The formula is $P(H \mid E) = \dfrac{P(E \mid H)\,P(H)}{P(E)}$.
  • $P(H)$ is the prior, $P(E \mid H)$ is the likelihood, $P(E)$ is the marginal probability of the evidence, and $P(H \mid E)$ is the posterior.
  • The posterior is a weighted compromise: strong evidence moves the prior a lot, weak evidence moves it a little [2].
  • In a disease test with 1% prevalence, 99% sensitivity and 95% specificity, a positive result raises the probability of disease from 0.01 to only 0.1667.

What Bayes' Theorem Means

In plain terms, Bayes' theorem is a rule for updating what you believe after you see data. You start with a prior probability, look at how well the hypothesis predicts the data you observed, and end with a posterior probability that reflects both.

The precise definition is a relationship between conditional probabilities. For a hypothesis $H$ and evidence $E$, the probability of $H$ conditional on $E$ is the ratio of the unconditional probability of the conjunction of the hypothesis with the data to the unconditional probability of the data alone, provided both terms of that ratio exist [1]. Written out:

$$P(H \mid E) = \frac{P(H \cap E)}{P(E)}$$

Because $P(H \cap E) = P(E \mid H)\,P(H)$, this rearranges into the familiar form of the theorem. The Stanford Encyclopedia of Philosophy notes that the theorem's central insight is that a hypothesis is confirmed by any body of data that its truth renders probable [1]. That single idea is the cornerstone of subjectivist and Bayesian approaches to evidence and learning [1].

How It Works

The formula has four parts. Each one has a name and a job.

$$P(H \mid E) = \frac{P(E \mid H)\,P(H)}{P(E)}$$

SymbolNameMeaning
$P(H)$PriorYour probability for the hypothesis before seeing the evidence
$P(E \mid H)$LikelihoodHow probable the evidence is if the hypothesis is true
$P(E)$Marginal likelihoodHow probable the evidence is overall, under all hypotheses
$P(H \mid E)$PosteriorYour updated probability for the hypothesis after the evidence

The denominator is usually the part people forget. When there are only two hypotheses, $H$ and its complement $\text{no } H$, you expand it with the law of total probability:

$$P(E) = P(E \mid H)\,P(H) + P(E \mid \text{no } H)\,P(\text{no } H)$$

That expansion is what makes the theorem usable with numbers you can actually estimate. It also explains the base rate effect: when $P(H)$ is tiny, the second term in the denominator can dominate, and the posterior stays small even after a positive result. This is the same logic behind conditional probability, where you restrict attention to a subset of outcomes.

Worked Example

The table below only illustrates the data layout, one row per patient with a disease status and a test result. The calculation itself uses the population rates given next, not these 10 rows.

patient_iddisease_statustest_result
1D+
2D+
3no D-
4no D-
5no D-
6no D-
7no D-
8no D-
9no D-
10no D-

The scenario uses a disease prevalence of 0.01, a test sensitivity of 0.99 and a specificity of 0.95. Sensitivity is $P(+ \mid D)$ and specificity is $P(- \mid \text{no } D)$, so the false positive rate is $1 - 0.95 = 0.05$.

Step 1. Set the priors. Before any test, $P(D) = 0.01$ and $P(\text{no } D) = 0.99$.

Step 2. Set the likelihoods. $P(+ \mid D) = 0.99$ and $P(+ \mid \text{no } D) = 0.05$.

Step 3. Compute the marginal probability of a positive test.

$$P(+) = 0.99 \times 0.01 + 0.05 \times 0.99 = 0.0594$$

Step 4. Apply the theorem.

$$P(D \mid +) = \frac{0.99 \times 0.01}{0.0594} = 0.1667$$

Step 5. Check the complement. $P(\text{no } D \mid +) = \dfrac{0.05 \times 0.99}{0.0594} = 0.8333$. The two posteriors sum to 1.

Here is the same calculation in Python:

prevalence = 0.01
sensitivity = 0.99
specificity = 0.95
p_pos = sensitivity*prevalence + (1-specificity)*(1-prevalence)
p_d_given_pos = sensitivity*prevalence / p_pos
print(f"P(D|+) = {p_d_given_pos:.4f}")  # 0.1667

Output:

P(D|+) = 0.1667
QuantityValue
Prior $P(D)$0.01
Prior $P(\text{no } D)$0.99
Likelihood $P(+ \mid D)$0.99
Likelihood $P(+ \mid \text{no } D)$0.05
Marginal $P(+)$0.0594
Posterior $P(D \mid +)$0.1667
Posterior $P(\text{no } D \mid +)$0.8333

How to Interpret It

The posterior of 0.1667 means that among people who test positive, about 17 in 100 actually have the disease. The other 83 are false positives. That result surprises most people, because the test looks accurate at 99% sensitivity.

The reason is the base rate. Only 1% of the population has the disease, so the small false positive rate applies to a very large healthy group. Roughly 5% of 99 healthy people produce about 5 false positives, while 99% of 1 sick person produces about 1 true positive. The positives are therefore mostly false.

This is why a Bayesian approach is useful in diagnostic testing. It lets you combine prior knowledge, such as prevalence, with the test result to produce a posterior probability rather than reading the raw accuracy as if it were the answer [3]. The same updating logic applies when you revise an estimate of a project's value after new information arrives [2].

When to Use It (and when not to)

Use Bayes' theorem when you have a prior probability and a likelihood, and you want the reverse conditional probability. Typical cases include diagnostic testing, spam filtering, and any setting where you update beliefs as data arrive [4].

Use it when the prior is meaningful and defensible. In clinical work, prevalence from a real population is a reasonable prior. In a Bayesian model, the prior is a full probability distribution over parameters, and the theorem updates that distribution into a posterior.

Do not use it when you have no basis for the prior and the result is sensitive to that choice. Do not use it as a substitute for good data collection. The theorem cannot fix a biased sample or a miscalibrated test. If your likelihoods are wrong, the posterior is wrong, and it will look just as confident as a correct one.

Bayes' Theorem vs Conditional Probability

Conditional probability is the general concept. Bayes' theorem is a specific tool for inverting it.

AspectConditional probabilityBayes' theorem
What it gives$P(A \mid B)$ directly from a defined sample space$P(H \mid E)$ from $P(E \mid H)$ and priors
DirectionForward, as definedReversed
Key inputsJoint and marginal probabilitiesPrior, likelihood, marginal likelihood
Typical useRestricting to a subgroupUpdating beliefs with evidence
Main riskMisreading the conditionA bad or unjustified prior

If you already know $P(A \mid B)$ and just need to state it, you do not need the theorem. You need it when the conditional you want runs opposite to the one you can estimate. The same reversal idea appears in maximum likelihood estimation, which asks which parameter value makes the observed data most probable.

Common Mistakes

  • Ignoring the base rate. Treating a 99% sensitive test as 99% accurate for a positive patient. Fix: always include $P(H)$ in the denominator.
  • Confusing $P(E \mid H)$ with $P(H \mid E)$. These are different numbers. Fix: write both out in symbols before plugging in values.
  • Forgetting the second term in the denominator. Using only $P(E \mid H)P(H)$ understates $P(E)$. Fix: expand $P(E)$ with the law of total probability.
  • Using a prior you cannot defend. A convenient prior can drive the posterior. Fix: state where the prior comes from and test sensitivity to it.
  • Dropping the complement check. Posteriors over all hypotheses must sum to 1. Fix: compute the complement and confirm it adds up.
  • Treating the posterior as final truth. It is a probability given the inputs, not a verdict. Fix: report it with the assumptions behind it.

Limitations

Bayes' theorem is exact given its inputs, but the inputs are estimates. The prior and the likelihoods come from data, expert judgment, or both, and each carries uncertainty that the formula does not represent. A posterior of 0.1667 is only as good as the 0.01 prevalence and the 0.99 sensitivity you fed it.

The theorem also scales poorly by hand. With two hypotheses it is a short calculation. With many parameters you need numerical methods such as Markov chain Monte Carlo to approximate the posterior, because the denominator becomes an integral over the whole parameter space. Finally, the theorem says nothing about causation. It updates probabilities, and a strong posterior is not evidence that one variable causes another.

Frequently Asked Questions

What is Bayes' theorem in simple terms?

It is a formula for updating a probability when new evidence arrives. You multiply your prior belief by how likely the evidence is under that belief, then divide by how likely the evidence is overall. The result is your updated belief, called the posterior.

What is the difference between prior and posterior?

The prior is your probability before seeing the data. The posterior is your probability after seeing the data. In the disease example, the prior for disease is 0.01 and the posterior after a positive test is 0.1667.

Why does a positive test not mean you have the disease?

Because the test also fires on healthy people. When the disease is rare, false positives from the large healthy group outnumber true positives from the small sick group. That is why the posterior stays low even with a highly sensitive test.

Can Bayes' theorem be used with more than two hypotheses?

Yes. Replace the two-term denominator with a sum over all hypotheses: $P(E) = \sum_i P(E \mid H_i)\,P(H_i)$. The posterior for each hypothesis is then its own likelihood times prior, divided by that sum. This is the form used in naive Bayes classifiers.

What does the denominator in Bayes' theorem represent?

It is the total probability of the evidence under all possible hypotheses. It acts as a normalizing constant, scaling the numerator so the posteriors sum to 1. You compute it with the law of total probability.

References

  1. Bayes’ Theorem (Stanford Encyclopedia of Philosophy/Spring 2021 Edition)
  2. Bayes Theorem Supplement
  3. Thi Mai H, He S, Alexandrou E, Alfred Frost S. (2022). An introduction to Bayes' theorem and examples of its application to a diagnostic test and a clinical trial. Nurse researcher
  4. Rindskopf D. (2020). Overview of Bayesian Statistics. Evaluation review

Further Reading

Related Articles