# What Is a Marginal Distribution? Definition and Examples

To define marginal distribution, you take a joint distribution of two variables and sum it over one of them, leaving the distribution of the other variable alone. In a contingency table, those sums sit in the margins, which is where the name comes from. This article explains the definition, the formula, and a full worked example.

## Quick Answer

- A marginal distribution is the probability distribution of one variable when you ignore the other variable in the pair [1].
- You get it by summing the joint probabilities over every value of the variable you are discarding [1].
- In a contingency table, row totals give the marginal distribution of the row variable, and column totals give the marginal distribution of the column variable.
- The discarded variable is said to be "marginalized out" [1].
- Marginalizing keeps every probability that involves only the retained variable, but it throws away information about how the two variables depend on each other [2].

## What Marginal Distribution Means

Start with the plain version. Suppose you collect two pieces of information about each person in a survey, such as gender and product preference. The joint distribution tells you how often each combination occurs, like "male who prefers A." The marginal distribution ignores the pairing and answers a simpler question: how many males are there in total, regardless of preference?

The precise statistical definition is close to the plain one. Given a known joint distribution of two discrete random variables, X and Y, the marginal distribution of X is the probability distribution of X when the values of Y are not taken into consideration [1]. You compute it by summing the joint probability distribution over all values of Y, and you can do the same for Y by summing over all values of X [1].

The word "marginal" is not a comment on importance. It refers to the physical location of the numbers. Marginal variables are the variables you keep, and their distribution is found by summing values in a table along rows or columns and writing the sum in the margins of the table [1]. That is also why the process is called marginalizing: you focus on the sums in the margin and discard the rest [1].

## How It Works

For discrete variables, the formula is a sum. If $p(x_i, y_j)$ is the joint probability that X takes value $x_i$ and Y takes value $y_j$, then:

$$p_X(x_i) = \sum_{j} p(x_i, y_j)$$

$$p_Y(y_j) = \sum_{i} p(x_i, y_j)$$

Each symbol does one job:

- $x_i$ is one specific value of the variable X, such as "Male."
- $y_j$ is one specific value of the variable Y, such as "Prefer A."
- $p(x_i, y_j)$ is the joint probability of that exact pair.
- $\sum_{j}$ means add up over every value of Y while holding $x_i$ fixed.
- $p_X(x_i)$ is the marginal probability of that single value of X.

For continuous variables the same idea uses integration instead of summation. If the pair has a joint probability density function, integrating out the unwanted component gives the marginal density [2].

One property matters for checking your work. Because you are summing a probability distribution, the marginal probabilities must add to 1 for each variable you keep. If they do not, you have summed the wrong thing or dropped a category.

## Worked Example

The dataset is a survey of 200 respondents cross-classified by gender and product preference. Here are the joint counts:

| Gender | Prefer A | Prefer B |
|---|---|---|
| Male | 60 | 40 |
| Female | 30 | 70 |

**Step 1: Read the joint counts.** The interior of the table is the joint distribution in counts: $[[60, 40], [30, 70]]$. Each cell is one combination of gender and preference.

**Step 2: Sum across columns to get row totals.** Male: 60 + 40 = 100. Female: 30 + 70 = 100. These are the row marginals in counts.

**Step 3: Sum down rows to get column totals.** Prefer A: 60 + 30 = 90. Prefer B: 40 + 70 = 110. These are the column marginals in counts.

**Step 4: Find the grand total.** 100 + 100 = 200.

**Step 5: Convert the row totals to the marginal distribution of gender.** Male: 100/200 = 0.5000. Female: 100/200 = 0.5000.

**Step 6: Convert the column totals to the marginal distribution of preference.** Prefer A: 90/200 = 0.4500. Prefer B: 110/200 = 0.5500.

**Step 7: Check that the marginals sum to 1.** 1.0000 and 1.0000.

Here is the same calculation in Python:

```python
import pandas as pd
df = pd.DataFrame(
    [[60, 40],
     [30, 70]],
    index=['Male', 'Female'], columns=['Prefer A', 'Prefer B'])
row_marg = df.sum(axis=1) / df.values.sum()
col_marg = df.sum(axis=0) / df.values.sum()
```

Output:

```
row_marg: Male=0.5000, Female=0.5000; col_marg: Prefer A=0.4500, Prefer B=0.5500
```

For reference, the joint proportions are 0.3, 0.2, 0.15 and 0.35. Notice that the marginal for gender (0.5 and 0.5) says nothing about preference. The sample is split evenly by gender, but that fact alone tells you nothing about which product people like.

## How to Interpret It

A marginal distribution answers a question about one variable in isolation. In the example, the marginal distribution of preference says that 45% of respondents prefer A and 55% prefer B. That is a complete, valid summary of preference on its own.

What it does not tell you is whether preference differs by gender. To see that, you need the joint distribution or the conditional distributions. In this dataset the joint counts do differ by gender, so the two variables are related. The marginals hide that relationship entirely.

This is the tradeoff described in the formal definition. Marginalization preserves all probabilities involving only the retained components, but it generally discards information about their dependence on the components that were summed out [2]. You lose the pairing on purpose, and you should know what you gave up.

A useful habit is to compare a marginal proportion with the corresponding conditional proportions. If they match, the variables look independent in your sample. If they differ, there is an association worth investigating. The same logic underlies [marginal means](/blog/data-analysis/marginal-means-definition-examples), which average a response over the levels of another factor.

## When to Use It (and when not to)

Use a marginal distribution when you want a clean summary of one variable and the other variable is not part of your question. Common cases include reporting overall response rates, describing the composition of a sample, and building the inputs for further probability work.

Use it when you need the denominator for conditional probabilities. The marginal probability of the conditioning event is exactly what you divide by in the definition of a conditional probability.

Do not use it when your question is about a relationship. If you want to know whether preference depends on gender, the marginal distribution of preference cannot answer that. You need the joint or conditional view.

Do not use it as a substitute for the full table when the categories interact. Collapsing a table to its margins can reverse or hide patterns, and no amount of care in the arithmetic will bring that information back.

If you are summarizing a single variable and want simple descriptive statistics instead of a full distribution, a [mean, median and mode calculator](/tools/mean-median-mode-calculator) is often the faster route.

## Marginal Distribution vs Joint Distribution

The joint distribution covers every combination of the two variables. The marginal distribution covers one variable with the other summed out. They are related by addition, but they answer different questions.

| Feature | Joint distribution | Marginal distribution |
|---|---|---|
| Describes | Both variables together | One variable alone |
| Values in the example | 0.3, 0.2, 0.15, 0.35 | 0.5, 0.5 and 0.45, 0.55 |
| Number of entries | One per cell | One per row or column |
| Sums to | 1 across all cells | 1 for each variable |
| Shows dependence | Yes | No |
| How to get it | Collect or model the pairs | Sum the joint over one variable |

The joint distribution is the richer object. The marginal is what you get when you deliberately give up that richness. For a broader tour of these objects, see [probability distributions explained](/blog/data-analysis/probability-distributions-explained).

## Common Mistakes

- **Dividing by the wrong total.** Marginal proportions use the grand total, not the row or column total. In the example, 100/200 = 0.5000 is the marginal for Male, while 60/100 = 0.6 would be a conditional proportion. Fix: decide first whether your denominator is the grand total or a subgroup total.
- **Reading a marginal as a conditional.** A marginal of 0.45 for Prefer A does not mean 45% of males prefer A. Fix: if the question mentions a subgroup, compute the conditional distribution for that subgroup.
- **Forgetting to sum to 1.** Each marginal distribution must total 1. Fix: add the marginals as a check before you report them.
- **Dropping a category.** If a table has a row for "Other" or "No response," leaving it out changes every marginal. Fix: include all categories, or state clearly which ones you excluded.
- **Assuming independence from marginals alone.** Two variables can have identical marginals and still be strongly associated. Fix: inspect the joint or conditional table before claiming independence.
- **Mixing counts and proportions.** Row totals are counts, and the marginal distribution is those counts divided by the grand total. Fix: label each number as a count or a proportion so the two never get confused.

## Limitations

A marginal distribution cannot show dependence. Once you sum over a variable, the information about how the two variables move together is gone, and it cannot be recovered from the margins alone [2]. Many different joint distributions share the same margins, so the marginal is not a unique summary of the pair.

Marginals can also mislead when groups have very different sizes. A marginal proportion is a weighted average of the group proportions, so a large group can dominate the number and make a pattern in a small group invisible. If group sizes matter for your question, report the conditional distributions alongside the marginal.

## Frequently Asked Questions

### What does marginal mean in statistics?

It refers to the margins of a table. Marginal variables are the variables you keep, and their distribution is found by summing values along rows or columns and writing the sums in the margins [1]. The term describes where the numbers sit, not how important they are.

### How do you define marginal distribution for a continuous variable?

You integrate instead of summing. If the pair has a joint probability density function, integrating out the unwanted component gives the marginal density [2]. The interpretation is the same: the distribution of one variable with the other averaged out.

### Is a marginal distribution the same as a conditional distribution?

No. A marginal distribution ignores the other variable. A conditional distribution fixes the other variable at a specific value. In the worked example, the marginal for Prefer A is 0.4500, while the conditional proportion of Prefer A among males is 60/100 = 0.6.

### Why do marginal probabilities sum to 1?

Because you are summing a probability distribution over all its possible values. Every joint probability is counted exactly once in the row marginals and exactly once in the column marginals. In the example, both marginals total 1.0000.

### Can two different joint distributions have the same marginals?

Yes. The margins do not determine the joint distribution. This is why marginalizing discards information about dependence, and why you should not infer a relationship from marginals alone [2]. If dependence matters, work with the joint or conditional distributions.

If you want to see the same summing logic applied to a continuous family, the [uniform distribution](/blog/data-analysis/uniform-distribution) and the [normal distribution](/blog/data-analysis/normal-distribution-definition-examples) both show how a density behaves when you look at one variable at a time.

## References

1. [Marginal distribution - Wikipedia](https://en.wikipedia.org/wiki/Marginal_distribution)
2. [Marginal Distribution -- from Wolfram MathWorld](https://mathworld.wolfram.com/MarginalDistribution.html)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)

## Related Articles

- [What Are Marginal Means? Definition and Examples](/blog/data-analysis/marginal-means-definition-examples)
- [Uniform Distribution: Definition, Formula and Examples](/blog/data-analysis/uniform-distribution)
- [Probability Distributions: Definition, Types and Examples](/blog/data-analysis/probability-distributions-explained)
- [Normal Distribution: Definition, Properties and Examples](/blog/data-analysis/normal-distribution-definition-examples)
- [Exponential Distribution: Definition, Formula and Examples](/blog/data-analysis/exponential-distribution)