# Probability Density Function: Definition, Formula and Examples

The probability density function (PDF) describes how probability is spread over the values of a continuous random variable. You cannot read a probability off the height of the curve. You get probability from the area under it. The distribution and density function are two views of the same variable, and knowing which one to use prevents most beginner errors.

## Quick Answer

- A PDF, written $f(x)$, is a non-negative function whose total area under the curve equals 1 [1].
- Probability for a continuous variable comes from intervals, not single points, so $P(X = a) = 0$ for any exact value [2].
- The cumulative distribution function (CDF), written $F(x)$, gives the probability that $X$ is less than or equal to $x$ [3].
- The CDF is the integral of the PDF, so $F(x) = \int_{-\infty}^{x} f(t)\,dt$ [4].
- To find $P(a \le X \le b)$, subtract two CDF values: $F(b) - F(a)$.

## What a Probability Density Function Means

A probability density function is a mathematical rule that assigns a density to each value of a continuous random variable. Think of it like mass density along a wire. The density at a point is not a mass, but the density times a small length gives you the mass in that slice [5]. A PDF works the same way. The height tells you how concentrated probability is near that value, not the probability of landing exactly there.

The precise statistical definition: a function $f(x)$ is a probability density function for a continuous random variable $X$ if it satisfies two conditions. First, $f(x) \ge 0$ for all $x$. Second, the total area under the curve is 1, written as

$$\int_{-\infty}^{\infty} f(x)\,dx = 1.$$

That integral-equals-one property is the continuous version of the rule that all probabilities in a discrete distribution sum to 1 [1]. Discrete variables use a probability mass function, where each value carries its own probability. Continuous variables use a density, where probability lives in intervals [1].

The distribution function, also called the cumulative distribution function, is a different object. It is defined for every random variable, discrete or continuous, and it returns the probability that $X$ is less than or equal to a given value [3]. For a continuous variable, the CDF is the running integral of the PDF.

## How It Works

The core mechanism is integration. To get a probability from a PDF, you find the area under the curve between two bounds:

$$P(a \le X \le b) = \int_{a}^{b} f(x)\,dx.$$

Each symbol means the following.

- $X$ is the continuous random variable.
- $f(x)$ is the probability density at the value $x$.
- $a$ and $b$ are the interval bounds, with $a < b$.
- The integral sign sums the density across the interval, which produces an area.

The CDF packages that same idea into one function. By definition,

$$F(x) = P(X \le x) = \int_{-\infty}^{x} f(t)\,dt,$$

so the CDF at any point is the area to the left of that point [4]. Because the CDF accumulates area, it always runs from 0 to 1 and never decreases. The PDF is the derivative of the CDF, which is why the two are linked: one is the rate of accumulation, the other is the accumulated total [2].

For the normal distribution, the PDF is

$$f(x) = \frac{1}{\sigma\sqrt{2\pi}}\,e^{-\frac{(x-\mu)^2}{2\sigma^2}},$$

where $\mu$ is the mean and $\sigma$ is the standard deviation. There is no closed-form formula for the normal CDF, so you use tables or software. A standard normal table or a function like `stats.norm.cdf` handles it. If you want to try values yourself, the [Probability Calculator](/tools/probability-calculator) computes these areas directly.

## Worked Example

Suppose a class of 15 students takes an exam, and you want to model the scores with a normal distribution. The raw scores are below.

| student_id | score |
|---|---|
| 1 | 88 |
| 2 | 92 |
| 3 | 76 |
| 4 | 105 |
| 5 | 118 |
| 6 | 95 |
| 7 | 82 |
| 8 | 110 |
| 9 | 99 |
| 10 | 73 |
| 11 | 101 |
| 12 | 87 |
| 13 | 114 |
| 14 | 90 |
| 15 | 96 |

The sample size is 15. The sample mean is 95.0667 and the sample standard deviation (using $n-1$) is 13.1718. For the model, you choose a normal distribution with $\mu = 100.0000$ and $\sigma = 15.0000$, a common choice for test scores.

You want the probability that a score falls between 85 and 115. First convert the bounds to z-scores:

$$z_{lower} = \frac{85.0000 - 100.0000}{15.0000} = -1.0000,$$

$$z_{upper} = \frac{115.0000 - 100.0000}{15.0000} = 1.0000.$$

Now evaluate the PDF and CDF at both bounds.

| Quantity | Value |
|---|---|
| PDF at 85, $f(85.0000)$ | 0.0161 |
| PDF at 115, $f(115.0000)$ | 0.0161 |
| PDF at the mean, $f(100.0000)$ | 0.0266 |
| CDF at 85, $F(85.0000)$ | 0.1587 |
| CDF at 115, $F(115.0000)$ | 0.8413 |
| Probability between bounds | 0.8413 - 0.1587 = 0.6827 |

Notice that the PDF at 85 and at 115 are equal, both 0.0161, because the curve is symmetric around the mean. The PDF is highest at the mean, 0.0266, but that height is still not a probability. The probability comes from the CDF difference: 0.6827.

Here is the code that produced these values.

```python
from scipy import stats
mu, sigma = 100, 15
lo, hi = 85, 115
pdf_lo = stats.norm.pdf(lo, mu, sigma)   # 0.0161
pdf_hi = stats.norm.pdf(hi, mu, sigma)   # 0.0161
cdf_lo = stats.norm.cdf(lo, mu, sigma)   # 0.1587
cdf_hi = stats.norm.cdf(hi, mu, sigma)   # 0.8413
prob   = cdf_hi - cdf_lo                 # 0.6827
for name, v in [('pdf_lo', pdf_lo), ('pdf_hi', pdf_hi), ('cdf_lo', cdf_lo), ('cdf_hi', cdf_hi), ('prob  ', prob)]:
    print(f"{name} = {v:.4f}")
```

Output:

```text
pdf_lo = 0.0161
pdf_hi = 0.0161
cdf_lo = 0.1587
cdf_hi = 0.8413
prob   = 0.6827
```

The shaded region between 85 and 115 has area 0.6827, and the matching CDF shows $F(85) = 0.1587$ and $F(115) = 0.8413$.

## How to Interpret It

Read the PDF as a concentration map. Taller regions mean values are more likely to appear nearby. The height itself is a density, so it can exceed 1 for narrow distributions, and that is not an error [1]. Only the area carries probability.

Read the CDF as a running total. $F(85) = 0.1587$ means about 15.87 percent of the distribution sits at or below 85. $F(115) = 0.8413$ means about 84.13 percent sits at or below 115. The gap between them, 0.6827, is the share inside the interval. This matches the well-known rule that roughly 68 percent of a normal distribution lies within one standard deviation of the mean.

When you compare two distributions, the PDF shape tells you about spread and skew. A wider PDF means more variability. The CDF tells you about position, since it answers "what fraction is below this value" directly. For ranking or threshold questions, the CDF is usually the faster tool.

## When to Use It (and when not to)

Use a PDF when your variable is continuous and you care about intervals. Heights, weights, times, temperatures and test scores all fit. Use the CDF when you need cumulative probabilities, percentiles, or threshold comparisons. Both are standard in [probability distributions](/blog/data-analysis/probability-distributions-explained) work.

Do not use a PDF for a discrete variable. Counts like the number of heads in ten coin flips need a probability mass function, such as the [binomial distribution](/blog/data-analysis/binomial-distribution-formula-examples) or the [Poisson distribution](/blog/research-skills/poisson-distribution-formula-and-examples). Do not use a PDF to get the probability of an exact value, because that probability is zero for continuous variables [2]. Do not assume every continuous variable is normal. The [uniform distribution](/blog/data-analysis/uniform-distribution) and the [exponential distribution](/blog/data-analysis/exponential-distribution) are common alternatives, and picking the wrong model distorts every probability you compute [4].

## Distribution Function vs Density Function

The two terms are often confused because both describe the same variable. The table below separates them.

| Feature | Probability Density Function (PDF) | Cumulative Distribution Function (CDF) |
|---|---|---|
| Symbol | $f(x)$ | $F(x)$ |
| Returns | Density at a point | Probability up to a point |
| Range of output | 0 or positive, can exceed 1 | 0 to 1 |
| Probability from it | Area between two bounds | Difference of two values |
| Relationship | Derivative of the CDF | Integral of the PDF |
| Exact value $P(X=a)$ | 0 | $F(a) - F(a^-)$ |

The PDF answers "how concentrated is probability here." The CDF answers "how much probability has accumulated so far." For a continuous variable, the CDF is the one that returns a probability directly.

## Common Mistakes

- Treating the PDF height as a probability. The value $f(85) = 0.0161$ is not the chance of scoring exactly 85. Fix: always integrate over an interval or subtract two CDF values.
- Thinking a PDF cannot exceed 1. A narrow distribution can have a peak above 1. Fix: check that the total area equals 1, not the height.
- Using the PDF for a discrete variable. Counts need a mass function. Fix: match the tool to the variable type.
- Forgetting that $P(X = a) = 0$ for continuous variables. Fix: ask for a range, not a single point [2].
- Subtracting PDF values instead of CDF values. $f(b) - f(a)$ is meaningless. Fix: use $F(b) - F(a)$.
- Assuming normality without checking. Fix: plot the data and compare, as with the sample mean 95.0667 and SD 13.1718 above, which differ from the chosen $\mu = 100$ and $\sigma = 15$.

## Limitations

A PDF describes a model, not the data itself. Choosing $\mu = 100$ and $\sigma = 15$ is a modeling decision, and the sample values of 95.0667 and 13.1718 show the fit is close but not exact. If the true distribution is skewed or has heavy tails, normal-based probabilities will be wrong in the tails, which is often where the decisions matter most.

The CDF also has limits. For many distributions there is no closed-form formula, so you depend on tables or software, and those carry rounding. The normal CDF is the classic case. Numerical answers like 0.1587 and 0.8413 are accurate to the digits shown but are still approximations. For very small samples, the empirical distribution may be a better guide than any fitted curve.

## Frequently Asked Questions

### What is the difference between a distribution and density function?

The density function $f(x)$ gives the concentration of probability at a point, and probability comes from its area. The distribution function $F(x)$ gives the cumulative probability up to a point. The CDF is the integral of the PDF, and the PDF is the derivative of the CDF [4][2].

### Can a probability density function be greater than 1?

Yes. The height of a PDF is a density, not a probability, so it can exceed 1 for narrow distributions. The only hard rule is that the total area under the curve equals 1 [1].

### Why is the probability of an exact value zero for a continuous variable?

A single point has no width, so the area above it is zero. That is why $P(X = a) = 0$ and why you compute probabilities over intervals instead [2]. This does not mean the value is impossible, only that its exact probability is zero.

### How do I find the probability between two values?

Subtract the CDF at the lower bound from the CDF at the upper bound: $P(a \le X \le b) = F(b) - F(a)$. In the worked example, $F(115) - F(85) = 0.8413 - 0.1587 = 0.6827$.

### Do I always need calculus to use a PDF?

No. For common distributions you can use tables or software. The [Gaussian function](/blog/data-analysis/gaussian-function-formula-examples) page covers the normal case, and a calculator or a function like `stats.norm.cdf` handles the integration for you [4].

## References

1. [1.3.6.1. What is a Probability Distribution](https://www.itl.nist.gov/div898/handbook/eda/section3/eda361.htm)
2. [Probability 101](https://www.cs.cornell.edu/courses/cs664/1997sp/probability.htm)
3. [Distributions and Densities](http://www.stat.yale.edu/Courses/1997-98/101/distrib.htm)
4. [5.2: Continuous Probability Functions - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Introductory_Statistics/Introductory_Statistics_2e_(OpenStax)/05%3A_Continuous_Random_Variables/5.02%3A_Continuous_Probability_Functions)
5. [6.2: Continuous Probability Distributions and Probability Density - Physics LibreTexts](https://phys.libretexts.org/Courses/University_of_California_Davis/UCD%3A_Physics_9D__Modern_Physics/6%3A_The_Universe_is_Inherently_Probabilistic/6.2%3A_Continuous_Probability_Distributions_and_Probability_Density)

## Further Reading

- [7: Distribution and Density Functions - Statistics LibreTexts](https://stats.libretexts.org/Bookshelves/Probability_Theory/Applied_Probability_(Pfeiffer)/07%3A_Distribution_and_Density_Functions)

## Related Articles

- [Probability Distributions: Definition, Types and Examples](/blog/data-analysis/probability-distributions-explained)
- [Uniform Distribution: Definition, Formula and Examples](/blog/data-analysis/uniform-distribution)
- [Bernoulli Distribution: Definition, Formula and Examples](/blog/data-analysis/bernoulli-distribution)
- [Binomial Distribution: Formula, Mean and Examples](/blog/data-analysis/binomial-distribution-formula-examples)
- [Gaussian Function: Definition, Formula and Examples](/blog/data-analysis/gaussian-function-formula-examples)
- [Poisson Distribution: Formula and Examples](/blog/research-skills/poisson-distribution-formula-and-examples)
- [Fundamental Statistics: Core Concepts Explained](/blog/research-skills/fundamental-statistics-core-concepts-explained)
- [Statistical Range: Definition, Calculation, and Applications](/blog/guides/statistical-range-definition-calculation-and-applications)