# What Is a Dichotomous Variable? Definition and Examples

A dichotomous variable is a variable that can take exactly two possible values or categories. Common examples include yes/no, alive/dead, and improved/not improved. The dichotomous definition matters in statistics because two-category variables behave differently from variables with three or more categories, and they open the door to specific tools such as proportions, odds ratios, and logistic regression.

## Quick Answer

- A dichotomous variable has exactly two mutually exclusive categories, and every case must fall into one of them [1].
- The two categories are often coded 0 and 1, but the codes are labels, not measurements.
- Binary, indicator, and dummy variable are near-synonyms for a dichotomous variable.
- A dichotomous variable is a special case of a nominal variable, since the categories have no natural order.
- You summarize it with counts and proportions, not means and standard deviations.

## What a Dichotomous Variable Means

In plain terms, a dichotomous variable splits your data into two groups. Each observation belongs to one group or the other, and nothing in between. A patient improved or did not improve. A survey respondent agreed or did not agree. A coin landed heads or tails.

The precise statistical definition is stricter. A dichotomous variable is a categorical variable whose value set contains exactly two elements, and those two elements form a partition of the whole. That partition must be jointly exhaustive, meaning every case belongs to one of the two parts, and mutually exclusive, meaning no case belongs to both [1]. If a category is missing from your coding scheme, or if a case can sit in both categories at once, you do not have a clean dichotomous variable.

The word itself comes from the Greek roots for "cut in two." The same idea appears outside statistics. In astronomy, dichotomy describes the moment when the Moon or an inferior planet is exactly half-lit as seen from Earth [1]. In botany, dichotomous branching happens when a terminal bud divides into two equal branches [1]. The statistical use keeps the core meaning: one whole, two parts.

## How It Works

The mechanism behind a dichotomous variable is simple counting. Once you code the two categories, usually as 0 and 1, you can describe the variable with a single proportion.

$$p = \frac{k}{n}$$

- $p$ is the proportion of cases in the category you care about, often the category coded 1.
- $k$ is the number of cases in that category.
- $n$ is the total number of cases.

Because there are only two categories, the proportion in the other category is $1 - p$. That single number carries all the information about the distribution. This is why dichotomous variables are so compact to report and so easy to compare across groups.

When you have two dichotomous variables, you can cross them in a contingency table. A 2x2 table shows how the two variables relate, and from it you can compute row proportions, column proportions, risk differences, and odds ratios. The coding choice of 0 and 1 is arbitrary in the sense that swapping the labels changes nothing about the underlying data, but it does change how you read a regression coefficient. The category coded 1 is the one the coefficient refers to.

## Worked Example

A small survey recorded 30 patients, noting their treatment group (0 = placebo, 1 = drug) and their outcome (0 = not improved, 1 = improved). Both variables are dichotomous.

| Patient group | Not improved (0) | Improved (1) | Total |
|---|---|---|---|
| Placebo (0) | 9 | 6 | 15 |
| Drug (1) | 4 | 11 | 15 |
| Total | 13 | 17 | 30 |

Walking through the steps:

1. Total patients surveyed: $n = 30$.
2. Treatment group counts: placebo = 15, drug = 15.
3. Outcome counts: not improved = 13, improved = 17.
4. The 2x2 contingency table is [[9, 6], [4, 11]].
5. Proportion improved in the placebo group: $6 / (9 + 6) = 6/15 = 0.4000$.
6. Proportion improved in the drug group: $11 / (4 + 11) = 11/15 = 0.7333$.
7. Overall proportion improved: $17 / 30 = 0.5667$.
8. Proportion assigned to placebo: $15 / 30 = 0.5000$.
9. Proportion assigned to drug: $15 / 30 = 0.5000$.

Here is the code that produces the table and the row proportions.

```python
import pandas as pd
df = pd.DataFrame({'treatment': [0]*15 + [1]*15,
                   'outcome':   [0]*9 + [1]*6 + [0]*4 + [1]*11})
ct = pd.crosstab(df['treatment'], df['outcome'])
print(ct)
print(ct.div(ct.sum(axis=1), axis=0).round(4))
```

Output:

```text
outcome    0   1
treatment       
0          9   6
1          4  11
outcome         0       1
treatment                
0          0.6000  0.4000
1          0.2667  0.7333
```

The treatment variable is balanced at 50.0% in each arm, which is what you would expect from a simple equal allocation. The outcome variable is not balanced: 56.7% of patients improved overall. The drug group improved at 73.3% versus 40.0% for placebo, a difference of 33.3 percentage points.

## How to Interpret It

Read a dichotomous variable through its proportions, not its codes. Saying "the mean of outcome is 0.5667" is technically correct when the variable is coded 0 and 1, because the mean of a 0/1 variable equals the proportion of ones. But the sentence "56.7% of patients improved" is what a reader actually needs.

When you compare two groups, report both proportions and the difference between them. In the example, the risk difference is $0.7333 - 0.4000 = 0.3333$. You can also express the comparison as a ratio: $0.7333 / 0.4000 = 1.833$, meaning the drug group improved at about 1.8 times the rate of the placebo group. Each of these summaries answers a slightly different question, so pick the one that matches your research question.

Watch the base rates. A proportion of 0.7333 built on 15 people is far less stable than the same proportion built on 1,500 people. Small cells in a 2x2 table, especially the 4 patients in the drug group who did not improve, make every downstream estimate noisy.

## When to Use It (and when not to)

Use a dichotomous variable when the underlying question is genuinely two-sided. Did the patient respond to treatment or not? Did the machine fail or not? Did the customer churn or not? A well-constructed dichotomous outcome can have clinical sense and specificity, and under realistic conditions its statistical power can come close to that of a continuous outcome measure [2].

Do not dichotomize a continuous variable just to simplify it. If you measure a pain score from 0 to 100 and then split it at 50, you throw away the information in every value between the extremes. Simulation studies comparing dichotomous and continuous outcomes in rheumatoid arthritis and ankylosing spondylitis found that continuous outcomes are typically more powerful than dichotomous ones, though there are situations where the gap narrows [2]. Splitting a continuous variable also creates a sharp boundary where none exists in the data, so two patients scoring 49 and 51 land in different categories while two patients scoring 10 and 49 land in the same one.

Use a dichotomous variable when the two categories are meaningful on their own terms, and keep the continuous version when the gradations carry information you care about.

## Dichotomous vs Nominal

A dichotomous variable is a nominal variable with exactly two categories. Nominal variables can have any number of unordered categories, such as blood type or country of birth. Dichotomous variables are the two-category special case.

| Feature | Dichotomous variable | Nominal variable |
|---|---|---|
| Number of categories | Exactly 2 | 2 or more |
| Ordering | None | None |
| Typical coding | 0 and 1 | Arbitrary labels or codes |
| Summary statistic | Proportion | Counts and proportions per category |
| Regression handling | Single 0/1 predictor | Dummy coding with k-1 indicators |
| Example | Improved / not improved | Blood type (A, B, AB, O) |

The practical difference shows up in modeling. A dichotomous predictor needs one coefficient. A nominal predictor with four categories needs three dummy variables, with one category left out as the reference. If you want to see how multi-category unordered variables are handled, the guide to [nominal variables](/blog/data-analysis/nominal-variable-definition-examples) walks through the coding.

## Common Mistakes

- Treating the 0/1 codes as numbers with real magnitude. The fix: remember the codes are labels. A mean of 0.5667 is a proportion, not a measurement on a scale.
- Splitting a continuous variable at its median to "make analysis easier." The fix: keep the continuous variable and use a method designed for it, since dichotomizing discards information and usually reduces power [2].
- Forgetting that the two categories must be mutually exclusive and jointly exhaustive. The fix: check your coding scheme for cases that fit both categories or neither before you analyze [1].
- Reporting only one proportion when comparing groups. The fix: report both group proportions and the difference, so readers can judge the size of the gap.
- Ignoring small cell counts in a 2x2 table. The fix: report the raw counts alongside every percentage, and treat estimates from cells under about five cases with caution.
- Confusing a dichotomous variable with an ordinal one. The fix: ask whether the categories have a natural order. "Mild, moderate, severe" is ordinal, while "improved, not improved" is dichotomous. The [ordinal variable guide](/blog/data-analysis/ordinal-variable-definition-examples) covers the ordered case.

## Limitations

A dichotomous variable cannot represent gradation. Once you collapse a measurement into two categories, every difference within a category disappears, and the choice of cut point drives the results. Move the threshold and the proportions change, sometimes substantially. This makes dichotomous outcomes sensitive to decisions that have nothing to do with the underlying phenomenon.

Dichotomous variables also limit which statistical methods apply. The standard deviation of a 0/1 variable is fully determined by the proportion, $\sqrt{p(1-p)}$, so it adds nothing beyond $p$, and ordinary linear regression on a 0/1 outcome can produce predicted values outside the 0 to 1 range. Methods built for binary outcomes, such as logistic regression, avoid that problem but require their own assumptions and their own interpretation of coefficients as log odds. Finally, a dichotomous variable tells you nothing about why a case fell into one category. It records the outcome, not the process behind it.

## Frequently Asked Questions

### What is a simple dichotomous definition?

A dichotomous variable is a variable with exactly two possible categories, where every case belongs to one category or the other. Examples include yes/no, pass/fail, and alive/dead. The two categories must be mutually exclusive and jointly exhaustive [1].

### Is a dichotomous variable the same as a binary variable?

Yes, in practice the terms are interchangeable. Binary variable, indicator variable, and dummy variable all describe a variable with two categories. Some writers reserve "dummy variable" for a 0/1 variable used as a predictor in a regression model.

### Can a dichotomous variable be used in regression?

Yes. A 0/1 variable works as a predictor in linear regression, where its coefficient represents the average difference between the two groups. When a dichotomous variable is the outcome, logistic regression is the usual choice because it keeps predicted probabilities between 0 and 1.

### What is the difference between dichotomous and ordinal?

A dichotomous variable has two categories with no order, while an ordinal variable has categories that can be ranked. "Improved or not improved" is dichotomous. "No improvement, some improvement, full recovery" is ordinal because the categories have a natural sequence.

### How do I choose which category gets coded 1?

Pick the category that matches your research question, usually the event you are studying, such as improvement, failure, or a positive test result. The choice does not change the underlying data, but it does change how you read coefficients and odds ratios, so state your coding clearly in any write-up.

## References

1. [Dichotomy - Wikipedia](https://en.wikipedia.org/wiki/Dichotomy)
2. [Anderson JJ. (2007). Mean changes versus dichotomous definitions of improvement. Statistical methods in medical research](https://pubmed.ncbi.nlm.nih.gov/17338291/)

## Further Reading

- [NIST/SEMATECH e-Handbook of Statistical Methods](https://www.itl.nist.gov/div898/handbook/index.htm)
- [OpenStax. Introductory Statistics 2e](https://openstax.org/details/books/introductory-statistics-2e)
- [Krzywinski M, Altman N (2013). Importance of being uncertain. Nature Methods](https://doi.org/10.1038/nmeth.2613)
- [Krzywinski M, Altman N (2013). Significance, P values and t-tests. Nature Methods](https://doi.org/10.1038/nmeth.2698)

## Related Articles

- [What Is an Ordinal Variable? Definition and Examples](/blog/data-analysis/ordinal-variable-definition-examples)
- [What Is an Independent Variable? Definition and Examples](/blog/data-analysis/what-is-an-independent-variable)
- [Dependent Variable Examples: Definition and Study Design](/blog/data-analysis/dependent-variable-examples)
- [Disjoint Events: Definition and Probability Examples](/blog/data-analysis/disjoint-events-definition)
- [What Is a Nominal Variable? Definition and Examples](/blog/data-analysis/nominal-variable-definition-examples)