What Is Shannon Entropy? Definition, Formula and Examples

By Dr. Zubair Khalid, DVM, MS, PhD ·

What Is Shannon Entropy? Definition, Formula and Examples

Shannon entropy is a number that measures how uncertain you are about the outcome of a random process. It is usually written $H$ and measured in bits. A fair coin has Shannon entropy of 1 bit, while a heavily biased coin has less, because you can guess its outcome more often.

Quick Answer

  • Shannon entropy $H$ measures the average uncertainty in a probability distribution over outcomes.
  • The formula is $H = -\sum_i p_i \log_2 p_i$, where $p_i$ is the probability of outcome $i$.
  • It is measured in bits when you use base-2 logarithms.
  • Maximum entropy occurs when all outcomes are equally likely. Minimum entropy (zero) occurs when one outcome has probability 1.
  • It is also the lower limit on the average number of yes/no questions per outcome needed to identify outcomes, a limit approached when many outcomes are identified together [1].

What Shannon Entropy Means

In plain terms, Shannon entropy tells you how surprised you should expect to be. If every outcome is equally likely, you cannot predict what happens next, so uncertainty is high. If one outcome almost always occurs, you can predict it well, so uncertainty is low.

The precise statistical definition: for a discrete random variable $X$ with possible outcomes $x_1, x_2, \dots, x_n$ and probabilities $p_i = P(X = x_i)$, the Shannon entropy is the expected value of $-\log_2 p_i$. That is, it is the probability-weighted average of the information content of each outcome. Rare outcomes carry more information when they occur, and entropy averages that information across the whole distribution.

Shannon entropy is the foundation of information theory and appears throughout data analysis, from feature selection to measuring the impurity of decision tree splits. It is closely related to the Shannon Diversity Index used in ecology, which applies the same mathematics to species abundances.

How It Works

The formula for Shannon entropy is:

$$H = -\sum_{i=1}^{n} p_i \log_2 p_i$$

Each symbol means the following:

  • $H$ is the entropy, measured in bits.
  • $n$ is the number of possible outcomes.
  • $p_i$ is the probability of the $i$-th outcome, and all probabilities sum to 1.
  • $\log_2$ is the base-2 logarithm, which makes the unit bits. Using natural logarithms gives nats instead.

The negative sign in front makes the result positive, because $\log_2 p_i$ is negative for any probability between 0 and 1. By convention, a term with $p_i = 0$ contributes 0, since $0 \log_2 0$ is treated as 0.

Two properties follow directly from the formula. When all $n$ outcomes are equally likely, each has $p_i = 1/n$ and entropy reaches its maximum of $\log_2 n$ bits. When one outcome has probability 1, every other term is 0 and entropy is 0 bits, meaning no uncertainty at all [1].

Worked Example

The dataset below holds ten coin-flip probability rows: five fair coins with $p = 0.5$ and five biased coins with $p = 0.9$ for heads.

coinp_headsp_tails
fair0.50.5
fair0.50.5
fair0.50.5
fair0.50.5
fair0.50.5
biased0.90.1
biased0.90.1
biased0.90.1
biased0.90.1
biased0.90.1

For the fair coin, the two probability terms are:

  • Term 1: $-0.5 \times \log_2(0.5) = 0.5000$
  • Term 2: $-0.5 \times \log_2(0.5) = 0.5000$

Adding them gives $H = 1.0000$ bits. This is the maximum for two outcomes, since $\log_2 2 = 1$.

For the biased coin, the terms are:

  • Term 1: $-0.9 \times \log_2(0.9) = 0.1368$
  • Term 2: $-0.1 \times \log_2(0.1) = 0.3322$

Adding them gives $H = 0.4690$ bits. The entropy drop from fair to biased is $1.0000 - 0.4690 = 0.5310$ bits. The biased coin is more predictable, so it carries less uncertainty.

Here is the same calculation in Python:

import math
def shannon_entropy(ps):
    return -sum(p * math.log2(p) for p in ps if p > 0)
H_fair = shannon_entropy([0.5, 0.5])   # 1.0000
H_biased = shannon_entropy([0.9, 0.1]) # 0.4690
print(f"H_fair = {H_fair:.4f} bits, H_biased = {H_biased:.4f} bits")

Output:

H_fair = 1.0000 bits, H_biased = 0.4690 bits

How to Interpret It

Entropy is easiest to read as an average number of yes/no questions. A fair coin needs exactly one question ("is it heads?") to identify the outcome, so $H = 1$ bit [1]. A single biased flip still needs one question, but when you identify many biased flips together you can get by with fewer than one question per flip on average, approaching 0.4690.

Higher entropy means more uncertainty and more information gained when you observe an outcome. Lower entropy means the distribution is concentrated and outcomes are more predictable. Entropy is never negative, and for $n$ outcomes it never exceeds $\log_2 n$.

When you compare distributions, keep the number of outcomes in mind. A distribution over 8 equally likely outcomes has $H = 3$ bits, which is not directly comparable to a 2-outcome distribution with $H = 1$ bit unless you normalize by dividing by $\log_2 n$.

When to Use It (and when not to)

Use Shannon entropy when you need a single number for the uncertainty of a categorical or discrete distribution. Common uses include measuring class imbalance, ranking features by information gain, evaluating decision tree splits, and quantifying diversity in a population.

Do not use it when your data are continuous without first binning or estimating a density, since the discrete formula assumes countable outcomes. Do not use it to compare distributions with different numbers of categories unless you normalize. Do not treat entropy as a distance between distributions. If you need to compare two distributions directly, a divergence measure is the right tool.

If you are summarizing a numeric column instead, the mean and standard deviation describe center and spread, which entropy does not.

Shannon Entropy vs Gini Impurity

Both measure the impurity of a distribution, and both are used to choose splits in decision trees. They differ in how they weight outcomes.

PropertyShannon EntropyGini Impurity
Formula$-\sum p_i \log_2 p_i$$1 - \sum p_i^2$
UnitsBitsNone
Maximum (n outcomes)$\log_2 n$$1 - 1/n$
ComputationRequires logarithmsUses only squares
Typical useInformation gainCART splitting

Entropy tends to penalize mixed distributions slightly more strongly, while Gini impurity is cheaper to compute because it avoids logarithms.

Common Mistakes

  • Forgetting the negative sign. The logarithm of a probability is negative, so without the leading minus the result is negative. Fix: always apply $-\sum$ or take the absolute value of the sum.
  • Using the wrong logarithm base. Base 2 gives bits, base $e$ gives nats, base 10 gives hartleys. Fix: decide your unit first and stay consistent.
  • Treating $p = 0$ as an error. A zero-probability outcome contributes nothing. Fix: skip zero terms, as the code above does with if p > 0.
  • Comparing entropy across different numbers of categories. A 10-outcome distribution can reach higher entropy than a 2-outcome one. Fix: normalize by $\log_2 n$ before comparing.
  • Using probabilities that do not sum to 1. The formula assumes a valid distribution. Fix: check that your probabilities sum to 1 before computing.
  • Applying it to continuous data directly. The discrete formula does not apply to raw continuous values. Fix: bin the data or use a differential entropy estimator.

Limitations

Shannon entropy summarizes a distribution in one number, so it discards the identity of outcomes. Two very different distributions can share the same entropy. It also says nothing about the shape of the distribution beyond its spread across categories.

For continuous variables, the discrete formula does not apply without binning, and the result depends on your bin width. Entropy is also sensitive to how you define categories. Merging or splitting categories changes the value, which makes cross-study comparisons fragile unless the category scheme is identical.

Frequently Asked Questions

What is Shannon entropy in simple terms?

It is a measure of how uncertain you are about an outcome. High entropy means outcomes are hard to predict, and low entropy means one outcome is likely. A fair coin has 1 bit of entropy, and a coin that always lands heads has 0 bits.

What is the Shannon entropy formula?

The formula is $H = -\sum_i p_i \log_2 p_i$, where $p_i$ is the probability of the $i$-th outcome. The sum runs over all possible outcomes. Using base-2 logarithms gives the result in bits.

Why is Shannon entropy measured in bits?

Bits come from using base-2 logarithms, which match binary yes/no questions. One bit is the uncertainty of a single fair coin flip [1]. This makes entropy directly interpretable as the average number of binary questions needed to identify an outcome.

Can Shannon entropy be negative?

No. Probabilities lie between 0 and 1, so each $\log_2 p_i$ is zero or negative, and the leading minus sign makes every term non-negative. The minimum possible value is 0, which occurs when one outcome has probability 1.

What is the difference between Shannon entropy and thermodynamic entropy?

They share a mathematical form but describe different systems. Shannon entropy measures uncertainty in a probability distribution over outcomes [1]. Thermodynamic entropy measures the number of microstates compatible with a macrostate, and it is maximized when all microstates are equally likely [1].

References

  1. Entropy

Further Reading

Related Articles