# What Is a Decision Tree? Definition, Splits and Examples

A decision tree is a predictive model that splits data into smaller and smaller groups by asking a sequence of yes-or-no questions about the input features. Each question sends a record down one branch, and the record ends up in a leaf that gives a prediction. You can read the whole model as a flowchart, which is why a decision tree is one of the easiest machine learning methods to explain to a non-technical audience.

## Quick Answer

- A decision tree predicts an outcome by routing each record through a series of if-then splits on feature values.
- Each internal node tests one feature against a threshold, and each branch represents one outcome of that test.
- The tree picks the split that makes the resulting groups as pure as possible, meaning each group contains mostly one class or a tight range of values.
- Classification trees predict a category, and regression trees predict a number [1].
- Trees are easy to read and need little data preparation, but a single tree can overfit and is often weaker than an ensemble of trees [1].

## What Decision Tree Means

In plain terms, a decision tree is a flowchart that turns a chain of questions into a prediction. You start at the top, answer the first question, follow the matching branch, and keep going until you reach a final answer at the bottom.

The precise definition is a recursive partitioning model. The algorithm repeatedly divides the training data into subsets based on a feature and a threshold, choosing the division that best separates the target values. In scikit-learn, `DecisionTreeClassifier` handles multi-class classification and takes an array `X` of shape `(n_samples, n_features)` plus an array `y` of class labels [1]. A companion class, `DecisionTreeRegressor`, does the same job when the target is a floating-point number [1].

## How It Works

A tree is built top-down. At each node, the algorithm searches over every feature and every candidate threshold and scores how well that split separates the target. The split with the best score becomes the node, and the process repeats on each resulting subset until a stopping rule is met.

For classification, the usual score is impurity. Gini impurity for a node with class proportions $p_k$ is:

$$G = 1 - \sum_{k=1}^{K} p_k^2$$

Here $K$ is the number of classes, and $p_k$ is the fraction of samples in the node that belong to class $k$. A node where every sample has the same label has $G = 0$, which is perfectly pure. A node split evenly between two classes has $G = 0.5$.

The algorithm picks the split that reduces impurity the most, measured as the weighted average of the child impurities subtracted from the parent impurity. For regression, the same logic applies, but the score is usually variance or mean squared error instead of Gini.

Each symbol in the tree means one thing:

- **Node**: a test on a single feature, such as "is petal length ≤ 2.45?"
- **Branch**: the answer to that test, true or false.
- **Leaf**: a terminal node with no further test, holding the predicted class or value.
- **Depth**: the number of splits from the root to a leaf. Limiting depth is a common way to control overfitting [1].

## Worked Example

Imagine a tiny dataset of six flowers with one feature, petal length, and two classes, setosa and versicolor. Suppose setosa flowers have petal lengths of 1.0, 1.4 and 1.8, and versicolor flowers have lengths of 3.5, 4.0 and 4.5.

The tree tries a threshold at 2.0. The left group holds the three setosa flowers, and the right group holds the three versicolor flowers. Both groups are perfectly pure, so Gini impurity is 0 on each side. The tree stops and creates two leaves.

Now classify a new flower with a petal length of 1.6. It is less than 2.0, so it goes left and is predicted as setosa. A flower with a length of 4.2 goes right and is predicted as versicolor. That is the entire model: one question, two answers.

Real datasets need more than one split, and the tree keeps adding nodes until the groups are pure enough or a stopping rule fires. You can print the structure as text with `export_text`, which does not require external libraries [1].

## How to Interpret It

Read a tree from the root down. Each node shows the feature and threshold being tested. Follow the branch that matches your record's value. The leaf you land on gives the prediction.

Two numbers help you judge a node. The first is the impurity, which tells you how mixed the node is. The second is the number of samples that reach it. A leaf with very few samples gives a prediction based on thin evidence, so treat it with caution.

For classification, the leaf usually reports the majority class. For regression, the leaf reports the mean target value of the training samples that landed there. The path from root to leaf is also a readable rule, such as "if petal length ≤ 2.0, predict setosa." That readability is the main reason people reach for a decision tree.

## When to Use It (and when not to)

Use a decision tree when you need a model you can explain, when your features are a mix of numeric and categorical values, or when you want a quick baseline before trying heavier methods. Trees handle missing values and outliers better than many linear models, and they capture interactions between features without you specifying them.

Avoid a single tree when accuracy is the top priority and you have plenty of data. A lone tree often overfits, and the greedy split search cannot guarantee a globally optimal tree [1]. In those cases, train many trees in an ensemble, where features and samples are randomly drawn with replacement, which reduces the overfitting problem [1].

Also avoid trees when the underlying pattern is hard to express as axis-aligned splits. Concepts like XOR, parity and multiplexer problems are difficult for a single tree to learn [1].

## Decision Tree vs Logistic Regression

Both models can classify records, but they work in very different ways. A decision tree carves the feature space into rectangles with if-then rules. Logistic regression fits a single smooth equation that outputs a probability.

| Aspect | Decision Tree | Logistic Regression |
|---|---|---|
| Shape of boundary | Axis-aligned rectangles | A single smooth line or curve |
| Interpretability | Readable as a flowchart | Readable as coefficients |
| Handles interactions | Yes, automatically | Only if you add them |
| Output | Class or value | Probability |
| Overfitting risk | High for a single deep tree | Lower, but needs regularization |
| Best for | Mixed features, explainable rules | Linear trends, probability estimates |

If you need a probability and a simple linear story, logistic regression is often the better fit. If you need rules a stakeholder can follow, a decision tree wins.

## Common Mistakes

- **Letting the tree grow without limits.** A deep tree memorizes the training data. Fix it by setting a maximum depth or a minimum number of samples per leaf [1].
- **Ignoring class imbalance.** Decision tree learners create biased trees when some classes dominate, so balance the dataset before fitting [1].
- **Reading a leaf with two samples as a strong rule.** Check the sample count at each leaf before trusting its prediction.
- **Assuming the tree found the best possible splits.** The greedy search cannot guarantee a globally optimal tree [1].
- **Forgetting to prune or validate.** Always check performance on held-out data, not just the training set.
- **Treating one tree as the final model.** If accuracy matters, compare it against an ensemble.

## Limitations

A single decision tree cannot express every pattern. It struggles with problems where the correct answer depends on combinations of features in a way that axis-aligned splits cannot capture, such as XOR or parity [1]. Its greedy construction means the tree you get is good but not guaranteed to be the best possible [1].

Trees also mislead when classes are imbalanced, because the majority class pulls the splits toward itself [1]. A tree can look accurate while quietly ignoring a minority class you care about. Finally, small changes in the data can produce a very different tree, which makes a single tree less stable than an ensemble.

## Frequently Asked Questions

### What is a decision tree in simple terms?

It is a flowchart of yes-or-no questions about your data. Each answer sends a record down a branch until it reaches a final prediction. You can trace any prediction by hand, which makes the model easy to explain.

### How does a decision tree decide where to split?

It tests every feature and threshold and scores how well each split separates the target. For classification it usually minimizes Gini impurity, and for regression it minimizes variance or mean squared error. The best-scoring split becomes the node.

### Can a decision tree be used for regression?

Yes. `DecisionTreeRegressor` applies the same splitting logic when the target is a floating-point number instead of a class label [1]. The leaf then predicts the mean target value of the samples that reach it.

### Why does my decision tree overfit?

A tree that grows until every leaf is pure will fit noise in the training data. Limit the depth, require a minimum number of samples per leaf, or train an ensemble of trees instead of one [1].

### Is a decision tree better than a random forest?

A random forest usually predicts more accurately because it averages many trees trained on random samples and features [1]. A single decision tree is easier to read and faster to explain, so the choice depends on whether you value accuracy or transparency more.

## References

1. [1.10. Decision Trees, scikit-learn 1.9.1 documentation](https://scikit-learn.org/stable/modules/tree.html)

## Further Reading

- [Krzywinski M, Altman N (2017). Classification and regression trees. Nature Methods](https://doi.org/10.1038/nmeth.4370)
- [Lever J, Krzywinski M, Altman N (2016). Classification evaluation. Nature Methods](https://doi.org/10.1038/nmeth.3945)
- [scikit-learn User Guide](https://scikit-learn.org/stable/user_guide.html)
- [Wilson G, Bryan J, Cranston K et al. (2017). Good enough practices in scientific computing. PLOS Computational Biology](https://doi.org/10.1371/journal.pcbi.1005510)
- [Lever J, Krzywinski M, Altman N (2016). Model selection and overfitting. Nature Methods](https://doi.org/10.1038/nmeth.3968)

## Related Articles

- [What Is Boosting? Definition and Examples for Analysts](/blog/data-analysis/what-is-boosting-algorithms)
- [Eigenvalues and Eigenvectors: Definition and Examples](/blog/data-analysis/eigenvalues-eigenvectors-definition-examples)
- [Euclidean Distance: Definition, Formula and Examples](/blog/data-analysis/euclidean-distance-definition-formula)
- [Sigmoid Function: Definition, Formula and Examples](/blog/data-analysis/sigmoid-function-definition-formula)
- [Apriori Algorithm: Definition, Steps and Examples](/blog/data-analysis/apriori-algorithm-definition-steps-examples)