# What Are SHAP Values? Definition and Examples

SHAP (SHapley Additive exPlanations) is a method for explaining why a machine learning model made one specific prediction. It assigns each input feature a number, called a SHAP value, that shows how much that feature pushed the prediction up or down compared with the model's average output. The values for all features sum exactly to the prediction, which makes them easy to read and audit.

## Quick Answer

- A SHAP value is the contribution of one feature to one prediction, measured in the units of the model output.
- Positive values push the prediction higher, negative values push it lower.
- The base value (the model's average prediction) plus all SHAP values equals the individual prediction.
- SHAP comes from cooperative game theory, where the "players" are features and the "payout" is the prediction.
- You can compute SHAP values for any model, but exact values are cheapest for linear models and tree models.

## What SHAP Means

In plain terms, a SHAP value answers the question: "How much did this feature change the prediction for this one row?" If a house is predicted to cost $157,080 and the average house in the training data costs $242,870, SHAP tells you which features dragged the price down and by how much.

The precise definition comes from Shapley values in cooperative game theory. For a model $f$, a feature set $N$, and a row of inputs, the SHAP value of feature $i$ is its average marginal contribution across all possible orderings of features:

$$\phi_i = \sum_{S \subseteq N \setminus \{i\}} \frac{|S|!\,(|N|-|S|-1)!}{|N|!}\left[f(S \cup \{i\}) - f(S)\right]$$

Here $S$ is a subset of features that does not include $i$, $f(S)$ is the model's prediction using only the features in $S$, and the fraction is a weight that counts how many orderings produce that subset. The formula averages the change in prediction caused by adding feature $i$ to every possible subset of the other features.

This construction gives SHAP values three properties that make them trustworthy. They are locally accurate (they sum to the prediction), they treat missing features fairly (a feature that never changes the output gets a value of zero), and they are consistent (if a feature becomes more important in the model, its SHAP value does not decrease).

SHAP is widely used in explainable AI research. One study applies it to explain not just model outputs but also model uncertainty, showing which input features drive how confident a model is [1]. Another uses SHAP-derived directions to make explanations more stable over time in financial models [2].

## How It Works

The mechanism rests on a simple additive identity. For any single prediction:

$$f(x) = \phi_0 + \sum_{i=1}^{M} \phi_i$$

Each symbol means the following.

- $f(x)$ is the model's prediction for the row you are explaining.
- $\phi_0$ is the base value, usually the mean prediction over the background dataset.
- $\phi_i$ is the SHAP value for feature $i$.
- $M$ is the number of features.

For linear models, the SHAP value of feature $i$ has a closed form:

$$\phi_i = \beta_i (x_i - \mathbb{E}[x_i])$$

where $\beta_i$ is the model coefficient, $x_i$ is the feature value for this row, and $\mathbb{E}[x_i]$ is the feature's mean over the background data. This is why linear SHAP is fast and exact: you multiply the coefficient by how far the row sits from the average.

For tree models, the TreeSHAP algorithm computes exact values in polynomial time instead of enumerating all feature subsets. For arbitrary models, KernelSHAP estimates the values by sampling, which is slower and approximate. Newer architectures can avoid sampling entirely. SHAPformer, a Transformer-based forecasting model, produces exact explanations in under one second, with speedups of 50 to 1000 times compared with PermutationSHAP [3].

## Worked Example

Take a dataset of 30 houses with size in square feet, number of rooms, age in years, and price in thousands of dollars. The first house has 1,200 sqft, 2 rooms, and is 10 years old.

| size | rooms | age | price |
|---|---|---|---|
| 1200 | 2 | 10 | 159 |
| 1500 | 3 | 8 | 197 |
| 1800 | 3 | 6 | 228 |
| 2100 | 4 | 4 | 274 |
| 2400 | 4 | 2 | 297 |
| 1350 | 2 | 12 | 166 |
| 1650 | 3 | 9 | 209 |
| 1950 | 3 | 7 | 240 |
| 2250 | 4 | 5 | 278 |
| 2550 | 5 | 3 | 324 |
| 1280 | 2 | 11 | 162.4 |
| 1580 | 3 | 8 | 204.4 |
| 1880 | 3 | 6 | 236.4 |
| 2180 | 4 | 4 | 276.4 |
| 2480 | 4 | 2 | 308.4 |
| 1420 | 3 | 10 | 189.6 |
| 1720 | 3 | 8 | 214.6 |
| 2020 | 4 | 6 | 261.6 |
| 2320 | 4 | 4 | 288.6 |
| 2620 | 5 | 2 | 329.6 |
| 1250 | 2 | 9 | 162 |
| 1550 | 3 | 7 | 207 |
| 1850 | 3 | 5 | 232 |
| 2150 | 4 | 3 | 279 |
| 2450 | 4 | 1 | 305 |
| 1380 | 3 | 11 | 181.4 |
| 1680 | 3 | 9 | 212.4 |
| 1980 | 4 | 7 | 254.4 |
| 2280 | 4 | 5 | 284.4 |
| 2580 | 5 | 3 | 324.4 |

**Step 1. Fit an ordinary least squares model.** The fitted coefficients are an intercept of 52.8108, a size coefficient of 0.0826, a rooms coefficient of 13.1661, and an age coefficient of -2.1199.

**Step 2. Compute the base value.** The mean price across all 30 houses is 242.8667. This is the base value.

**Step 3. Compute the SHAP value for size.** The mean size is 1,913.3333 sqft. The first house is 1,200 sqft, which is 713.3333 below average. Multiply by the coefficient: $0.0826 \times (1200 - 1913.3333) = -58.9307$.

**Step 4. Compute the SHAP value for rooms.** The mean is 3.4333 rooms. The first house has 2 rooms: $13.1661 \times (2 - 3.4333) = -18.8713$.

**Step 5. Compute the SHAP value for age.** The mean age is 6.2333 years. The first house is 10 years old: $-2.1199 \times (10 - 6.2333) = -7.9850$.

**Step 6. Check additivity.** Add the base value and the three contributions: $242.8667 + (-58.9307) + (-18.8713) + (-7.9850) = 157.0796$. The model's prediction for this house is 157.0796, so the identity holds exactly.

The figure below shows these values as a horizontal bar chart: size -58.93, rooms -18.87, age -7.98, with the base of 242.87 plus the contributions equaling the prediction of 157.08.

Here is the code that produces these numbers.

```python
import shap, pandas as pd
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(df[['size','rooms','age']], df['price'])
explainer = shap.LinearExplainer(model, df[['size','rooms','age']])
sv = explainer.shap_values(df[['size','rooms','age']].iloc[[0]])
```

Output:

```
base=242.8667; phi_size=-58.9307; phi_rooms=-18.8713; phi_age=-7.9850; prediction=157.0796
```

## How to Interpret It

Read a SHAP value as a signed push in the model's output units. For this house, the size of 1,200 sqft is far below the average of 1,913 sqft, so it pushes the predicted price down by 58.93 thousand dollars. The two rooms push it down by 18.87 thousand, and the age of 10 years pushes it down by 7.99 thousand. The house starts from the average prediction of 242.87 and ends at 157.08.

The sign tells you direction, and the magnitude tells you strength. A feature with a SHAP value near zero had little effect on this particular prediction, even if it matters a lot on average across the dataset.

When you aggregate SHAP values across many rows, you get global importance. The mean absolute SHAP value per feature ranks features by how much they move predictions overall. This is often more informative than impurity-based importance because it reflects the actual direction and size of effects.

## When to Use It (and when not to)

Use SHAP when you need to explain a single prediction to a person, such as a loan decision, a price quote, or a forecast. Use it when you want global feature importance that respects the model's actual behavior. Use it when you need to compare explanations across rows to spot patterns, such as detecting that a model behaves differently in one operating regime than another [3].

Avoid SHAP when you need a causal claim. SHAP values describe what the model did, not what the world does. A large SHAP value for a feature does not mean changing that feature would change the real outcome.

Avoid it when computation is tight and the model is not linear or tree-based. KernelSHAP on a large neural network can be slow. If you only need a rough ranking of features, a simpler method may be enough.

## SHAP vs Feature Importance

Feature importance and SHAP both rank features, but they answer different questions. Feature importance is a single global number per feature. SHAP gives a per-row breakdown that also aggregates to a global view.

| Aspect | SHAP | Feature importance |
|---|---|---|
| Granularity | Per prediction, per feature | One number per feature |
| Direction | Signed, shows up or down | Usually unsigned magnitude |
| Additivity | Sums to the prediction | No such guarantee |
| Cost | Higher, especially for non-tree models | Low |
| Interpretation | Local and global | Global only |

If you need to explain one row, use SHAP. If you only need a quick global ranking and the model is already interpretable, feature importance is cheaper.

## Common Mistakes

- **Treating SHAP values as causal effects.** They describe the model, not the world. Fix: state clearly that explanations are about model behavior, and run a separate causal analysis if you need one.
- **Comparing SHAP values across different models.** A value of 0.5 in one model is not the same as 0.5 in another because the output scale differs. Fix: compare within a single model, or normalize before comparing.
- **Ignoring the base value.** A SHAP value of -58.93 means nothing without knowing the starting point of 242.87. Fix: always report the base value alongside the contributions.
- **Using the wrong background dataset.** The base value and the reference point for each feature depend on the background data. Fix: choose a background that represents the population you care about.
- **Reading a near-zero value as "no effect."** A feature can have a small SHAP value on one row and a large one on another. Fix: look at the distribution across rows before drawing conclusions.
- **Assuming KernelSHAP is exact.** It estimates values by sampling and can vary between runs. Fix: increase the sample budget or use an exact method when precision matters.

## Limitations

SHAP explains the model you give it, so if the model is wrong, the explanations are wrong in the same way. It cannot tell you whether a feature is a genuine cause of the outcome, and it cannot detect problems the model never learned about. Correlated features also cause trouble: when two features carry the same information, SHAP splits their contribution between them, which can make each look less important than it is in practice.

Computational cost is a real constraint. Exact values are fast for linear and tree models, but sampling-based methods for other model types can be slow and noisy. The values are also sensitive to the background dataset you choose, so two analysts using different backgrounds can get different numbers for the same row. For time-series models, standard SHAP requires sampling from background data, which newer sampling-free approaches avoid [3]. SHAP has also been extended beyond outputs to explain model uncertainty, which shows how flexible the framework is but also how much depends on what you choose to explain [1].

## Frequently Asked Questions

### What is a SHAP value in simple terms?

A SHAP value is a number that says how much one feature moved one prediction away from the model's average output. Positive means it pushed the prediction up, negative means it pushed it down. All the values for a row add up to the prediction itself.

### Do SHAP values always add up to the prediction?

Yes, for the standard formulation. The base value plus every feature's SHAP value equals the model output for that row. In the worked example, 242.8667 minus 58.9307 minus 18.8713 minus 7.9850 equals 157.0796, which matches the prediction exactly.

### What is the difference between SHAP and Shapley values?

They are the same mathematical object. Shapley values come from cooperative game theory and describe how to fairly divide a payout among players. SHAP applies that idea to machine learning by treating features as players and the prediction as the payout.

### How do I compute SHAP values for my model?

Use the shap library. For linear models, use LinearExplainer. For tree models, use TreeExplainer, which is exact and fast. For other models, use KernelSHAP or a model-specific explainer. The code in the worked example shows the linear case.

### Can SHAP values be negative?

Yes. A negative SHAP value means the feature lowered the prediction relative to the base value. In the example, all three features are negative because this house is smaller, has fewer rooms, and is older than average, so each one pulls the price down.

### Is SHAP only for tabular data?

No. SHAP has been applied to images, text, and time-series forecasting. For time-series Transformers, sampling-free variants produce exact explanations quickly, with reported speedups of 50 to 1000 times over PermutationSHAP [3]. The interpretation stays the same: each input contributes a signed amount to the output.

## References

1. [Uncertainty explanation of artificial intelligence models by SHAP - Experts@Minnesota](https://experts.umn.edu/en/publications/uncertainty-explanation-of-artificial-intelligence-models-by-shap/)
2. [SHAP-Guided Adversarial Training for Robust Interpretability in Asset Pricing | UChicago Knowledge](https://knowledge.uchicago.edu/records/sp6gm-4pp45)
3. [Explainable time-series forecasting with sampling-free SHAP fo...](https://publikationen.bibliothek.kit.edu/1000193702)

## Further Reading

- [Lever J, Krzywinski M, Altman N (2016). Classification evaluation. Nature Methods](https://doi.org/10.1038/nmeth.3945)
- [scikit-learn User Guide](https://scikit-learn.org/stable/user_guide.html)
- [Wilson G, Bryan J, Cranston K et al. (2017). Good enough practices in scientific computing. PLOS Computational Biology](https://doi.org/10.1371/journal.pcbi.1005510)

## Related Articles

- [What Is Boosting? Definition and Examples for Analysts](/blog/data-analysis/what-is-boosting-algorithms)
- [Eigenvalues and Eigenvectors: Definition and Examples](/blog/data-analysis/eigenvalues-eigenvectors-definition-examples)
- [Euclidean Distance: Definition, Formula and Examples](/blog/data-analysis/euclidean-distance-definition-formula)
- [Sigmoid Function: Definition, Formula and Examples](/blog/data-analysis/sigmoid-function-definition-formula)
- [Apriori Algorithm: Definition, Steps and Examples](/blog/data-analysis/apriori-algorithm-definition-steps-examples)