# CDF vs PDF: Differences and When to Use Each

The probability density function (PDF) and the cumulative distribution function (CDF) describe the same random variable from two different angles. The PDF gives you the relative likelihood of values near a point, while the CDF gives you the probability that the variable falls at or below a value. If your question is "how concentrated is the data here," use the PDF. If your question is "what fraction is below this threshold," use the CDF. This article covers cdf versus pdf in plain terms, with a comparison table, a worked example, and guidance on picking the right one.

## Quick Answer

- The PDF is a density curve. Its height at a point is not a probability for continuous variables, and it can exceed 1 [1].
- The CDF is a running total. It returns $P(X \le x)$, always between 0 and 1, and it never decreases.
- Probabilities for continuous variables come from areas under the PDF between two points, which equals the difference between two CDF values [1].
- Use the PDF to compare shapes, peaks, and relative likelihood across regions.
- Use the CDF to answer threshold questions, find percentiles, or read off tail probabilities directly.

## Key Differences

| Feature | PDF | CDF |
|---|---|---|
| Definition | Density at a point, $f(x)$ | Cumulative probability, $F(x) = P(X \le x)$ |
| Output range | 0 or greater, can exceed 1 [1] | 0 to 1 |
| Shape | Can rise and fall, may have multiple peaks | Non-decreasing, flat or rising |
| Reading a single point | Height is a density, not a probability [1] | Height is a probability |
| Getting an interval probability | Integrate the curve between two points | Subtract $F(b) - F(a)$ |
| Typical question | Where is the data concentrated? | What share falls below a cutoff? |
| Discrete counterpart | Probability mass function (PMF) [1] | Same CDF idea, sums instead of integrates |

The relationship between the two is direct. The CDF is the integral of the PDF:

$$F(x) = \int_{-\infty}^{x} f(t)\,dt$$

And the PDF is the derivative of the CDF:

$$f(x) = \frac{d}{dx} F(x)$$

That pair of formulas is the whole story. Every fact in the table follows from it.

## PDF Explained

The probability density function describes how densely probability is packed around each value. For a continuous variable, the probability of any single exact point is zero. Probabilities are measured over intervals, not single points, so the area under the curve between two distinct points defines the probability for that interval [1]. That is why the height of the density can be greater than one without breaking any rule. The only constraint is that the total area under the curve equals one [1].

A practical consequence: you cannot read $f(115) = 0.0161$ as "a 1.61% chance of exactly 115." You can read it as a density, and you can compare it to $f(100)$ to say which value is more likely to appear in a narrow band around it.

The most common continuous PDF is the normal density:

$$f(x) = \frac{1}{\sigma\sqrt{2\pi}} \exp\left(-\frac{(x-\mu)^2}{2\sigma^2}\right)$$

If you want to see this formula applied step by step, the guide on the [normal PDF formula with examples](/blog/data-analysis/normal-pdf-formula-examples) walks through several cases.

Discrete variables work differently. They use a probability mass function, and each value carries an actual probability you can read directly [1]. The term probability density function is sometimes used generically to cover both cases, but the interpretation differs [1].

## CDF Explained

The cumulative distribution function answers a threshold question: what is the probability that the variable is at or below $x$? It accumulates everything to the left, so it starts near 0, rises, and approaches 1.

For an empirical dataset, the CDF is a step function that increases at each data point and stays flat in between. With a large number of data points, this step function is a very good approximation to the true CDF [2]. That makes the empirical CDF a useful descriptive tool even when you have no assumed distribution.

The CDF is also the natural home for percentiles. If you want the value below which 80% of observations fall, you read the CDF backward. One tool describes this as finding the x-axis value where the CDF reaches a given percentage, so a CDF of 80% corresponds to the value exceeded 20% of the time [2]. The same logic gives you medians, quartiles, and any other quantile.

For the normal distribution, the CDF has no closed form in elementary functions, so it is written as $\Phi(z)$ and evaluated numerically. The article on the [normal CDF with calculator examples](/blog/data-analysis/normal-cdf) shows how to get these values from tables and software.

## Worked Example

The dataset below is a set of exam scores from a class of 15 students. We use it to show the sample mean and standard deviation, then apply both functions to a normal model.

| student_id | score |
|---|---|
| 1 | 88 |
| 2 | 92 |
| 3 | 76 |
| 4 | 100 |
| 5 | 85 |
| 6 | 95 |
| 7 | 78 |
| 8 | 82 |
| 9 | 90 |
| 10 | 105 |
| 11 | 70 |
| 12 | 98 |
| 13 | 84 |
| 14 | 91 |
| 15 | 87 |

The class sample mean is 88.0667 and the sample standard deviation is 9.4148 (computed with ddof=1). Suppose a separate reference population is modeled as normal with $\mu = 100$ and $\sigma = 15$, and we want to evaluate that model at $x = 115$.

Step 1. Standardize the value.

$$z = \frac{115 - 100}{15} = 1.0000$$

Step 2. Evaluate the PDF at $x = 115$.

$$f(x) = \frac{1}{\sigma\sqrt{2\pi}} \exp\left(-\frac{z^2}{2}\right) = 0.0161$$

Step 3. Evaluate the CDF at $x = 115$.

$$F(x) = P(X \le 115) = \Phi(1.0000) = 0.8413$$

The two numbers answer different questions. The density 0.0161 tells you how tall the curve is at 115. The cumulative value 0.8413 tells you that about 84% of the distribution sits at or below 115.

Here is the code that produced both values.

```python
from scipy.stats import norm
mu, sigma, x = 100, 15, 115
pdf = norm.pdf(x, loc=mu, scale=sigma)  # 0.0161
cdf = norm.cdf(x, loc=mu, scale=sigma)  # 0.8413
print(f"pdf = {pdf:.4f}")
print(f"cdf = {cdf:.4f}")
```

Output:

```text
pdf = 0.0161
cdf = 0.8413
```

Notice how you would use the CDF to get an interval probability. The probability of a score between 100 and 115 is $F(115) - F(100) = 0.8413 - 0.5000 = 0.3413$. With the PDF alone you would have to integrate the curve over that range to get the same answer.

## Which One Should You Use?

Match the function to the question you are asking.

| Your question | Use |
|---|---|
| Where is the data most concentrated? | PDF |
| Which of two values is more likely in a narrow band? | PDF |
| What is the chance of a value at or below 115? | CDF |
| What value marks the 90th percentile? | CDF, read backward |
| What is the chance of a value between 100 and 115? | Either, but the CDF is faster |
| How do two groups' distributions compare in shape? | PDF |
| How do two groups compare in the tails? | CDF |

The PDF is the better visual for shape. Peaks, skew, and multimodality show up immediately. The CDF is the better visual for tails and thresholds, because tail probabilities are read directly off the vertical axis instead of estimated as slivers of area.

When you plot distributions for a report, the choice also affects how readable the chart is. The trade-offs between tables and charts are covered in [data table vs graph](/blog/data-analysis/data-table-vs-graph), and the same reasoning applies to picking between a density plot and a cumulative plot.

If you are comparing groups, remember that the mean and the median can tell different stories about the same data, which is covered in [mean vs median differences](/blog/data-analysis/mean-vs-median-differences). The CDF makes that difference visible, because the median is simply the point where the curve crosses 0.5.

## Common Mistakes

- Treating PDF height as a probability. For continuous variables, $f(x)$ is a density and can exceed 1 [1]. Fix: integrate over an interval, or use the CDF to get an actual probability.
- Assuming the PDF must stay below 1. It does not, because only the total area is constrained to equal 1 [1]. Fix: check the area, not the peak height.
- Reading a CDF value as "the probability of exactly this value." The CDF gives $P(X \le x)$, which includes everything to the left. Fix: subtract two CDF values to isolate an interval.
- Forgetting that discrete variables use a PMF. The PDF label does not apply to counts and categories [1]. Fix: use the PMF for discrete data and the PDF for continuous data.
- Using a smooth theoretical CDF on a small sample without checking the fit. The empirical CDF is a step function, and it only approximates the true CDF well when the sample is large [2]. Fix: plot the empirical version first and compare.
- Confusing the density peak with the most probable value in a practical sense. The peak marks the highest density, but interval probabilities depend on width. Fix: compare areas, not heights, when the intervals differ in width.

## Limitations

Neither function tells you whether your distributional assumption is correct. A normal PDF and a normal CDF will produce smooth, plausible numbers even when the underlying data are skewed or heavy-tailed. The functions describe the model you chose, not the data you have. Always plot the empirical distribution before trusting a fitted curve.

The CDF also hides local structure. Two very different densities can produce similar-looking cumulative curves, so a CDF comparison can miss a bimodal shape that a density plot would reveal instantly. And for continuous variables, the PDF cannot give you a probability at a single point at all, which trips up readers who expect one number per value. Use both views together when the shape matters.

## Frequently Asked Questions

### What is the difference between CDF and PDF in simple terms?

The PDF shows how densely probability is packed around each value, like a profile of the distribution. The CDF shows the running total of probability up to each value. One is a shape, the other is a cumulative share.

### Can the PDF be greater than 1?

Yes. For continuous variables, the height of the density is not a probability, and only the total area under the curve must equal 1 [1]. A narrow distribution can have a tall peak, well above 1, and still be valid.

### How do I get a probability from a PDF?

Integrate the density between the two endpoints of the interval you care about. Equivalently, subtract the CDF at the lower endpoint from the CDF at the upper endpoint. For the example above, $F(115) - F(100) = 0.3413$.

### When should I use the CDF instead of the PDF?

Use the CDF whenever your question involves a threshold, a percentile, or a tail probability. Questions like "what share falls below this cutoff" or "what value marks the top 10%" are CDF questions by nature.

### Does the CDF work for discrete data too?

Yes. Every random variable has a CDF, discrete or continuous. The difference is that a discrete CDF jumps at each possible value, while a continuous CDF rises smoothly. The discrete counterpart of the PDF is the probability mass function [1].

If you want more practice data to test these ideas on, the list of [sample datasets for practice](/blog/data-analysis/sample-datasets-for-practice) includes options for building both density and cumulative plots.

## References

1. [1.3.6.1. What is a Probability Distribution](https://www.itl.nist.gov/div898/handbook/eda/section3/eda361.htm)
2. [PDF / CDF](https://samrepo.nlr.gov/help/pdfcdf.html)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)

## Related Articles

- [Normal PDF: Formula, Definition and Examples](/blog/data-analysis/normal-pdf-formula-examples)
- [Data Table vs Graph: Differences and When to Use Each](/blog/data-analysis/data-table-vs-graph)
- [Mean vs Median: Differences and When to Use Each](/blog/data-analysis/mean-vs-median-differences)
- [Structured vs Unstructured Data: Differences and Examples](/blog/data-analysis/structured-vs-unstructured-data)
- [Normal CDF: Definition, Formula and Calculator Examples](/blog/data-analysis/normal-cdf)
- [Discrete vs Continuous Variables: Key Differences](/blog/research-skills/discrete-vs-continuous-variables-key-differences)
- [SnpEff vs. VEP: Which Variant Annotation Tool Should You Choose?](/knowledge/bioinformatics/snpeff-vs-vep-which-variant-annotation-tool-should-you-choose)
- [mzML vs. mzXML vs. RAW: Choosing the Right Data Format for Proteomics Data Deposition and Sharing](/knowledge/bioinformatics/mzml-vs-mzxml-vs-raw-choosing-the-right-data-format-for-proteomics-data-deposition-and-sharing)