# Beta Distribution: Definition, Parameters and Examples

The beta distribution is a continuous probability distribution defined on the interval from 0 to 1, which makes it the natural model for proportions, rates and probabilities. Its beta density is controlled by two positive shape parameters, usually written $\alpha$ and $\beta$, that determine whether the curve is flat, U-shaped, or skewed toward one end. This article explains the definition, the parameters, and how to compute and interpret values with a real example.

## Quick Answer

- The beta distribution describes a continuous variable $x$ that can only take values between 0 and 1, such as a proportion of defectives or a conversion rate.
- It has two shape parameters, $\alpha$ and $\beta$, both greater than 0. They control the shape, center and spread of the beta density.
- The mean is $\alpha/(\alpha+\beta)$ and the variance is $\alpha\beta / [(\alpha+\beta)^2(\alpha+\beta+1)]$.
- When $\alpha = \beta$, the curve is symmetric around 0.5. When $\alpha < \beta$ it leans left, and when $\alpha > \beta$ it leans right.
- It is the conjugate prior for the binomial proportion, which is why it appears so often in Bayesian statistics.

## What the Beta Distribution Means

In plain terms, the beta distribution is the go-to model when your data are proportions. If you measure the fraction of defective items in a batch, the share of visitors who click a button, or the percentage of a budget spent, the value lives between 0 and 1 and the beta distribution can describe it.

The precise statistical definition is this. A random variable $X$ follows a beta distribution with shape parameters $\alpha > 0$ and $\beta > 0$ if its probability density function on the interval $[0, 1]$ is

$$f(x; \alpha, \beta) = \frac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha, \beta)}$$

where $B(\alpha, \beta)$ is the beta function, a normalizing constant that makes the total area under the curve equal 1. The beta function is defined as

$$B(\alpha, \beta) = \int_{0}^{1} t^{\alpha-1}(1-t)^{\beta-1}\,dt = \frac{\Gamma(\alpha)\Gamma(\beta)}{\Gamma(\alpha+\beta)}$$

and $\Gamma$ is the gamma function [1]. The standard beta distribution uses the bounds 0 and 1. A general form shifts and stretches the curve to any interval $[a, b]$ using lower and upper bounds, which act like location and scale parameters [1]. The cumulative distribution function is the incomplete beta function ratio $I_x(\alpha, \beta)$, and it does not have a simple closed form, so it is computed numerically [1]. The beta density is a flexible family. By changing $\alpha$ and $\beta$ you can produce a wide range of shapes, which is why it fits so many real proportions. If you want the broader picture of how continuous distributions work, see [probability distributions explained](/blog/data-analysis/probability-distributions-explained).

## How It Works

The formula has three moving parts, and each symbol has a clear job.

| Symbol | Meaning |
|---|---|
| $x$ | The value of interest, a proportion between 0 and 1 |
| $\alpha$ | First shape parameter, greater than 0 |
| $\beta$ | Second shape parameter, greater than 0 |
| $B(\alpha, \beta)$ | Beta function, the normalizing constant |
| $\Gamma(\cdot)$ | Gamma function, a continuous extension of the factorial |

The numerator $x^{\alpha-1}(1-x)^{\beta-1}$ sets the shape. When $\alpha = 1$ and $\beta = 1$, the numerator is constant and the density is flat, which is the uniform distribution on $[0, 1]$. When both parameters are below 1, the curve dips in the middle and rises at the edges, producing a U shape. When both are above 1, the curve has a single peak. The denominator $B(\alpha, \beta)$ rescales the curve so the area equals 1.

Two summary formulas matter most. The mean is

$$\mu = \frac{\alpha}{\alpha+\beta}$$

and the variance is

$$\sigma^2 = \frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}$$

The mean tells you where the bulk of the probability sits. The variance tells you how concentrated it is. Larger values of $\alpha + \beta$ shrink the variance, so the distribution becomes tighter around the mean. This behavior is exactly what you want when modeling a proportion estimated from many observations.

## Worked Example

The dataset below shows 12 production batches with an observed proportion of defectives, $p\_hat$, and the beta shape parameters used to model each batch.

| batch | p_hat | alpha | beta |
|---|---|---|---|
| 1 | 0.08 | 2 | 5 |
| 2 | 0.15 | 2 | 5 |
| 3 | 0.22 | 2 | 5 |
| 4 | 0.31 | 2 | 5 |
| 5 | 0.44 | 5 | 2 |
| 6 | 0.52 | 5 | 2 |
| 7 | 0.61 | 5 | 2 |
| 8 | 0.70 | 5 | 2 |
| 9 | 0.35 | 2 | 2 |
| 10 | 0.48 | 2 | 2 |
| 11 | 0.55 | 2 | 2 |
| 12 | 0.66 | 2 | 2 |

We evaluate the beta density at $x = 0.4$ for three parameter sets. The formula is

$$f(x; \alpha, \beta) = \frac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha, \beta)}$$

For $\alpha = 2$ and $\beta = 5$, the numerator is $0.4^{1}(0.6)^{4} = 0.05184$ and dividing by $B(2,5)$ gives a density of 1.5552. For $\alpha = 5$ and $\beta = 2$, the numerator is $0.4^{4}(0.6)^{1} = 0.01536$, giving a density of 0.4608. For $\alpha = 2$ and $\beta = 2$, the numerator is $0.4^{1}(0.6)^{1} = 0.24$, giving a density of 1.4400.

The mean and variance for the first set follow directly. The mean is $2/(2+5) = 0.2857$, and the variance is $10 / (7^2 \times 8) = 0.0255$.

```python
from scipy.stats import beta
beta.pdf(0.4, 2, 5)  # 1.5552
beta.pdf(0.4, 5, 2)  # 0.4608
beta.pdf(0.4, 2, 2)  # 1.4400
```

Output:

```
Beta(2,5) pdf at 0.4 = 1.5552; Beta(5,2) pdf at 0.4 = 0.4608; Beta(2,2) pdf at 0.4 = 1.4400
```

Notice how the same value $x = 0.4$ gets very different densities depending on the parameters. With $\alpha = 2, \beta = 5$ the curve leans left, so 0.4 sits in the tail and the density is lower than at the peak near 0.2. With $\alpha = 5, \beta = 2$ the curve leans right, so 0.4 is on the low side and the density is smaller still. With $\alpha = 2, \beta = 2$ the curve is symmetric around 0.5, and 0.4 is close to the center, giving a high density. This is the core intuition behind the beta density: the parameters move the mass, and the density at any point reflects how much probability sits there.

## How to Interpret It

The height of the beta density at a point is not a probability. It is a density, so you read probabilities from areas under the curve, not from single heights. To get the probability that a proportion falls below some value, you use the cumulative distribution function, which is the incomplete beta function ratio [1]. For example, the probability that $X \le 0.4$ is $I_{0.4}(\alpha, \beta)$.

The shape tells a story. A left-leaning curve with $\alpha < \beta$ means small proportions are more likely. A right-leaning curve with $\alpha > \beta$ means large proportions are more likely. A symmetric curve with $\alpha = \beta$ means the proportion is centered near 0.5. The spread narrows as $\alpha + \beta$ grows, so a tight curve signals confidence in the estimate and a wide curve signals uncertainty. If you are new to densities, the [probability density function guide](/blog/data-analysis/probability-density-function-definition-formula) walks through how to read them.

## When to Use It (and when not to)

Use the beta distribution when your outcome is a proportion or probability bounded between 0 and 1, when you want a flexible shape that can be skewed or symmetric, and when you are modeling a rate estimated from a limited number of trials. It is also the standard choice as a prior for a binomial proportion in Bayesian analysis, because the beta and binomial pair cleanly.

Do not use it when your data can fall outside $[0, 1]$, when you need a discrete count, or when a simple normal approximation already fits well. For counts of successes in a fixed number of trials, the [binomial distribution](/blog/data-analysis/binomial-distribution-formula-examples) is the right tool. For counts of events over time or space, use the [Poisson distribution](/blog/research-skills/poisson-distribution-formula-and-examples). The beta distribution is for the proportion itself, not the count behind it.

## Beta Distribution vs the Binomial Distribution

These two are often confused because they both involve proportions, but they answer different questions.

| Feature | Beta distribution | Binomial distribution |
|---|---|---|
| Variable type | Continuous proportion on $[0, 1]$ | Discrete count of successes |
| Parameters | $\alpha, \beta$ (shape) | $n$ (trials), $p$ (probability) |
| Typical use | Modeling an unknown proportion | Counting successes in $n$ trials |
| Mean | $\alpha/(\alpha+\beta)$ | $np$ |
| Relationship | Conjugate prior for the binomial | Likelihood for the beta prior |

The binomial counts how many successes occur. The beta describes what the underlying success probability might be. In Bayesian work they are used together, with the beta as the prior and the binomial as the data. If you want to see how the binomial behaves on its own, the [standard deviation of a binomial distribution](/blog/data-analysis/standard-deviation-binomial-distribution) is a good next step.

## Common Mistakes

- Treating the density height as a probability. The fix is to integrate over an interval or use the cumulative distribution function to get a probability.
- Assuming $\alpha$ and $\beta$ must be integers. They can be any positive real numbers, and fractional values are common in practice.
- Forgetting that the standard form is bounded by 0 and 1. If your data range over a different interval, you must apply the location and scale transformation [2].
- Confusing the beta function $B(\alpha, \beta)$ with the beta distribution. The beta function is just the normalizing constant inside the density.
- Reading a symmetric curve as always correct. Symmetry only happens when $\alpha = \beta$, which is a special case, not the default.
- Ignoring the variance formula when comparing fits. Two curves can share a mean but differ sharply in spread, and the variance captures that.

## Limitations

The beta distribution cannot model values outside its support. If your proportion can exceed 1 or drop below 0, or if your data are counts rather than proportions, the beta density will mislead you. It also assumes the underlying quantity is continuous, so it is an approximation when the true proportion comes from a small number of discrete trials.

Fitting the parameters is another constraint. There is no simple closed form for the cumulative distribution function, so probabilities and quantiles are computed numerically [1]. With small samples the estimated parameters can be unstable, and the fitted curve may look reasonable while hiding large uncertainty. Always check the fit against the data before trusting the shape.

## Frequently Asked Questions

### What does the beta distribution model?

It models a continuous proportion or probability that lies between 0 and 1. Common examples include defect rates, conversion rates, and the share of a total allocated to one category. The two shape parameters let the curve take many forms, from flat to sharply peaked.

### What do the parameters alpha and beta control?

They control the shape, center and spread of the beta density. The mean is $\alpha/(\alpha+\beta)$, so the ratio of the two parameters sets the center. Their sum sets the spread, with larger sums producing a tighter curve. When $\alpha = \beta$ the curve is symmetric.

### Is the beta distribution the same as the beta function?

No. The beta function $B(\alpha, \beta)$ is a mathematical constant used to normalize the density. The beta distribution is the full probability model built around that constant. The function appears in the denominator of the density formula.

### Can alpha and beta be less than 1?

Yes. When both are below 1, the density becomes U-shaped, with more probability near 0 and 1 and less in the middle. This is useful when extreme proportions are more likely than moderate ones.

### How is the beta distribution related to the binomial distribution?

The beta distribution is the conjugate prior for the binomial proportion. If you start with a beta prior and observe binomial data, the posterior is also a beta distribution with updated parameters. This makes Bayesian updating for proportions straightforward.

## References

1. [1.3.6.6.17. Beta Distribution](https://www.itl.nist.gov/div898/handbook/eda/section3/eda366h.htm)
2. [scipy.stats.beta, SciPy v1.18.0 Manual](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.beta.html)

## Further Reading

- [R: The Beta Distribution](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/Beta.html)
- [R: Density function of Truncated Beta Distribution](https://search.r-project.org/CRAN/refmans/cascsim/html/tbeta.html)
- [scipy.stats.betaprime, SciPy v1.18.0 Manual](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.betaprime.html)
- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)

## Related Articles

- [Probability Density Function: Definition, Formula and Examples](/blog/data-analysis/probability-density-function-definition-formula)
- [Probability Distributions: Definition, Types and Examples](/blog/data-analysis/probability-distributions-explained)
- [Negative Binomial Distribution: Formula and Examples](/blog/data-analysis/negative-binomial-distribution-formula-examples)
- [Student's t-Distribution: Definition, Formula and Examples](/blog/data-analysis/students-t-distribution-definition-formula)
- [Chi-Square Distribution: Definition, Formula and Examples](/blog/data-analysis/chi-square-distribution)
- [Poisson Distribution: Formula and Examples](/blog/research-skills/poisson-distribution-formula-and-examples)
- [Fundamental Statistics: Core Concepts Explained](/blog/research-skills/fundamental-statistics-core-concepts-explained)
- [Statistical Parameter: Definition, Types, and Estimation](/blog/guides/statistical-parameter-definition-types-and-estimation)