Probability Distributions: Definition, Types and Examples
By Dr. Zubair Khalid, DVM, MS, PhD ·

A probability distribution is a complete list of the possible values of a random variable together with the probability of each value. It answers one question: how likely is each outcome? Once you have a distribution, you can compute averages, spread, and the chance of any range of results.
Quick Answer
- A probability distribution pairs every possible outcome of a random variable with its probability [1].
- Discrete distributions assign probability to countable values, and their probabilities must sum to one [1].
- Continuous distributions assign probability to intervals, and the area under the curve equals one [1].
- Every distribution is described by location, spread and shape [2].
- Common examples include the binomial, uniform, normal, exponential and Poisson distributions [3][4].
What a Probability Distribution Means
In plain terms, a probability distribution is the rulebook for a random variable. It tells you which values can occur and how often you should expect each one. If you flip a coin, the distribution is simple: heads and tails each get 0.5. If you measure the heights of adult men, the distribution is a smooth curve spread across a wide range.
The precise statistical definition is stricter. For a discrete random variable, a probability mass function assigns a probability to each value, and those probabilities must be between zero and one and sum to one [1]. For a continuous random variable, a probability density function $f(x)$ satisfies two conditions: the probability that $x$ falls between two points $a$ and $b$ is the area under the curve between them, and the total integral equals one [1].
$$P[a \le x \le b] = \int_{a}^{b} f(x)\,dx \qquad \int_{-\infty}^{\infty} f(x)\,dx = 1$$
A key consequence for continuous variables is that the probability at any single point is zero. Probabilities are measured over intervals, not single points [1]. This is why the height of a density curve can exceed one even though probabilities never can.
How It Works
Every distribution is characterized by three features: location, spread and shape [2]. Location is the expected value, the center around which outcomes cluster. Spread is the expected amount of variation. Shape describes whether the variation is symmetric about the mean or skewed [2].
For a discrete distribution, the mean and variance come from weighted sums:
$$\mu = \sum_{k} k \cdot P(X = k) \qquad \sigma^2 = \sum_{k} (k - \mu)^2 \cdot P(X = k)$$
Here $k$ runs over every possible value, $P(X = k)$ is the probability of that value, $\mu$ is the mean, and $\sigma^2$ is the variance. The standard deviation $\sigma$ is the square root of the variance.
The binomial distribution is a good illustration of a discrete probability function. It models the number of successes in $N$ independent trials where each trial has two outcomes and a fixed success probability $p$ [3]:
$$P(x;p,n) = \binom{n}{x} p^{x}(1-p)^{n-x} \quad \text{for } x = 0, 1, 2, \cdots, n$$
The term $\binom{n}{x}$ counts the number of ways to arrange $x$ successes among $n$ trials, $p^x$ is the probability of those successes, and $(1-p)^{n-x}$ is the probability of the remaining failures [3]. The binomial is probably the most commonly used discrete distribution [3].
Worked Example
The dataset below is 50 dice-roll outcomes, a deterministic sequence recorded to compare an empirical distribution against a theoretical one.
| Roll values |
|---|
| 3, 1, 4, 2, 6, 5, 3, 2, 1, 4, 5, 6, 2, 3, 4, 1, 5, 6, 3, 2, 4, 1, 6, 5, 3, 2, 4, 6, 1, 5, 2, 3, 4, 5, 6, 1, 2, 3, 4, 5, 6, 1, 2, 3, 4, 5, 6, 2, 3, 4 |
Step 1. Count each face. With $n = 50$ rolls, the counts are 1:7, 2:9, 3:9, 4:9, 5:8, 6:8.
Step 2. Convert counts to empirical probabilities. Divide each count by $n$:
$$P(1)=0.1400,\ P(2)=0.1800,\ P(3)=0.1800,\ P(4)=0.1800,\ P(5)=0.1600,\ P(6)=0.1600$$
Step 3. Compare with the theoretical uniform distribution. A fair die gives $P(X=k) = 1/6 = 0.1667$ for every face.
Step 4. Compute the empirical mean.
$$1(0.1400) + 2(0.1800) + 3(0.1800) + 4(0.1800) + 5(0.1600) + 6(0.1600) = 3.5200$$
The theoretical mean is $(1+2+3+4+5+6)/6 = 3.5000$.
Step 5. Compute the empirical variance.
$$(1-3.5200)^2(0.1400) + (2-3.5200)^2(0.1800) + (3-3.5200)^2(0.1800) + (4-3.5200)^2(0.1800) + (5-3.5200)^2(0.1600) + (6-3.5200)^2(0.1600) = 2.7296$$
The theoretical variance is $2.9167$, and the empirical standard deviation is $\sqrt{2.7296} = 1.6522$ against a theoretical $1.7078$.
import numpy as np
rolls = [3,1,4,2,6,5,3,2,1,4,5,6,2,3,4,1,5,6,3,2,
4,1,6,5,3,2,4,6,1,5,2,3,4,5,6,1,2,3,4,5,
6,1,2,3,4,5,6,2,3,4]
faces, counts = np.unique(rolls, return_counts=True)
p_emp = counts / len(rolls)
print(dict(zip(faces, p_emp.round(4))))
Output: Empirical P: 1:0.1400, 2:0.1800, 3:0.1800, 4:0.1800, 5:0.1600, 6:0.1600 | mean=3.5200, var=2.7296, sd=1.6522
The empirical probabilities sit close to 0.1667 but not exactly on it, which is what you expect from a finite sample. The empirical mean of 3.5200 is slightly above the theoretical 3.5000, and the empirical variance is slightly below the theoretical value.
How to Interpret It
Read a distribution in three passes. First, look at the center: where do the values pile up? Second, look at the spread: how far do values stray from that center? Third, look at the shape: is it symmetric, or does it lean to one side [2]?
For the dice example, the center is near 3.5, the spread is about 1.65, and the shape is roughly flat because a fair die has no preferred face. If you saw a distribution with a long right tail, you would know extreme high values are possible but rare. That kind of reading is the same whether you are looking at a bar chart of counts or a smooth density curve.
You can also read probabilities directly off the distribution. The chance of rolling a 6 is 0.1600 in this sample and 0.1667 in theory. The chance of rolling a 4 or higher is the sum of the individual probabilities for 4, 5 and 6.
When to Use It (and when not to)
Use a probability distribution when you need to quantify uncertainty about a random variable. It is the right tool for computing expected values, setting prediction intervals, running simulations, and choosing a statistical test. If your data are counts of successes in fixed trials, the binomial fits [3]. If your data are counts of events in a fixed interval, the Poisson fits. If your data are measurements on a continuous scale, a density function such as the normal or exponential applies.
Do not use a distribution when the underlying process is not random or when the assumptions are clearly violated. A binomial model assumes a fixed number of independent trials with a constant success probability [3]. If the probability changes from trial to trial, the binomial is the wrong choice. If you have no repeated structure at all, a distribution may add false precision.
Probability Distribution vs Frequency Distribution
These two ideas are related but not the same. A frequency distribution describes what you observed in a dataset. A probability distribution describes what you expect from a random process.
| Feature | Probability distribution | Frequency distribution |
|---|---|---|
| Describes | Theoretical chances | Observed counts |
| Values | Probabilities between 0 and 1 | Counts or relative frequencies |
| Sums to | 1 | Sample size $n$ |
| Depends on | A model or assumption | The collected data |
| Example | Fair die, $P = 1/6$ | 50 rolls, counts 7 to 9 |
In the worked example, the empirical probabilities form a frequency distribution, and the uniform $1/6$ rule is the probability distribution. Comparing the two is a standard way to check whether a model fits.
Common Mistakes
- Treating a continuous density value as a probability. The height of a density curve is not a probability, and it can exceed one. Fix: always integrate over an interval to get a probability [1].
- Forgetting that discrete probabilities must sum to one. A set of values that sums to 0.8 or 1.3 is not a valid probability function [1]. Fix: check the sum before using the distribution.
- Assuming the binomial applies when $p$ changes. The binomial requires a fixed success probability across trials [3]. Fix: use a model that allows varying probability, or split the data into homogeneous groups.
- Confusing the mean with the most likely value. The mean is a weighted average, not necessarily the peak. Fix: look at the shape as well as the center [2].
- Ignoring spread when comparing distributions. Two distributions can share a mean but differ widely in variance. Fix: report location, spread and shape together [2].
- Reading a single point probability for a continuous variable. That probability is zero. Fix: ask for the probability of a range instead [1].
Limitations
A probability distribution is only as good as its assumptions. If you pick the wrong family, every probability you compute inherits that error. The dice example works because the uniform model is known to be correct. In real data, you rarely know the true distribution, so you estimate it and carry uncertainty forward.
Distributions also say nothing about cause. They describe how often values occur, not why. A well-fitting distribution can hide a biased sampling process, and a poorly fitting one can still produce useful summaries. Treat the distribution as a description of variation, not an explanation of it.
Frequently Asked Questions
What is the difference between discrete and continuous probability distributions?
Discrete distributions assign probability to countable values such as 0, 1, 2, and their probabilities sum to one. Continuous distributions assign probability to intervals, and the area under the density curve equals one [1]. The probability at any single point of a continuous distribution is zero.
What is a probability mass function versus a probability density function?
Discrete probability functions are called probability mass functions, and continuous probability functions are called probability density functions [1]. The umbrella term "probability functions" covers both. In software, the density or mass function is often named with a leading d, the cumulative function with p, the quantile function with q, and random generation with r [4].
How do I calculate the mean of a probability distribution?
For a discrete distribution, multiply each value by its probability and add the results. In the dice example, that gives 3.5200 for the empirical distribution and 3.5000 for the theoretical one. For a continuous distribution, you integrate $x$ times the density over all values.
Which probability distribution should I use for count data?
It depends on the structure. Use the binomial when you count successes in a fixed number of independent trials with a constant probability [3]. Use the Poisson when you count events in a fixed interval of time or space. If counts are overdispersed, a negative binomial often fits better.
Can a probability density be greater than one?
Yes. Because continuous probabilities are areas under the curve, the height of the density can exceed one as long as the total area equals one [1]. Only the area, not the height, is a probability.
If you want to compute these values quickly, the Probability Calculator handles common distributions, and the Binomial Distribution Calculator is built for trial-based problems. For background on the building blocks, see Fundamental Statistics: Core Concepts Explained and Statistical Parameter: Definition, Types, and Estimation. To go deeper on specific families, read the Normal Distribution, the Uniform Distribution, the Probability Density Function, and the Poisson Distribution.
References
- 1.3.6.1. What is a Probability Distribution
- 3.1.3.1. Distribution (Location, Spread and Shape)
- 1.3.6.6.18. Binomial Distribution
- R: Distributions in the 'stats' Package
Further Reading
- 5.2: Binomial Probability Distribution - Statistics LibreTexts
- NIST/SEMATECH e-Handbook of Statistical Methods
Related Articles
- Probability Density Function: Definition, Formula and Examples
- Normal Distribution: Definition, Properties and Examples
- Uniform Distribution: Definition, Formula and Examples
- Exponential Distribution: Definition, Formula and Examples
- Frequency Distribution: Definition, Table and Examples
- Poisson Distribution: Formula and Examples
- Fundamental Statistics: Core Concepts Explained
- Statistical Parameter: Definition, Types, and Estimation