# What Is ARIMA? Time Series Model Explained

ARIMA is a forecasting model that describes a time series using its own past values, its past forecast errors, and one or more rounds of differencing. The name is an acronym for AutoRegressive Integrated Moving Average, and the three letters map directly onto the three orders you specify: p, d and q. This article explains each order, shows a full worked fit on 36 months of sales data, and covers how to interpret the coefficients and forecasts.

## Quick Answer

- ARIMA models a time series as a combination of an autoregressive (AR) part, an integrated (differencing) part, and a moving average (MA) part [1].
- The order is written ARIMA(p, d, q): p is the number of AR lags, d is the number of differences taken, and q is the number of MA (error) lags [2].
- Differencing removes trend so the series becomes stationary, meaning its statistical properties no longer drift over time [1].
- You fit the model by maximum likelihood, then compare candidate orders using AIC or BIC, where lower is better [2].
- Forecasts come with widening confidence intervals as you project further ahead, because uncertainty compounds.

## What ARIMA Means

In plain terms, ARIMA says: the next value of a series depends on some number of recent values (the AR part), on some number of recent surprises or errors (the MA part), and on how many times you had to difference the series to make it stable (the I part).

The precise definition: an ARIMA(p, d, q) model is an ARMA(p, q) model fitted to the d-th difference of the series. If $y_t$ is the original series and $d = 1$, you model $y_t - y_{t-1}$. The general form is:

$$y'_t = c + \phi_1 y'_{t-1} + \dots + \phi_p y'_{t-p} + \theta_1 \varepsilon_{t-1} + \dots + \theta_q \varepsilon_{t-q} + \varepsilon_t$$

where $y'_t$ is the differenced series, $c$ is a constant, the $\phi$ terms are autoregressive coefficients, the $\theta$ terms are moving average coefficients, and $\varepsilon_t$ is white noise [1][2]. ARIMA models are, in theory, the most general class of models for forecasting a series that can be made stationary by differencing [1].

## How It Works

Each order does one job.

**p, the AR order.** The model regresses the current value on its own previous values. An ARIMA(1,1,0) is a first-order autoregressive model on the differenced series [1]. Positive autocorrelation at low lags usually points to adding AR terms [1].

**d, the order of differencing.** Differencing subtracts each value from the previous one, which removes a linear trend. If the series has positive autocorrelations out to many lags, it probably needs a higher order of differencing [3]. If the lag-1 autocorrelation is zero or negative, or the autocorrelations are small and patternless, the series does not need more differencing [3]. A lag-1 autocorrelation of -0.5 or more negative suggests the series may be overdifferenced [3].

**q, the MA order.** The model uses past forecast errors. Negative autocorrelation is usually best treated by adding an MA term, and differencing itself reduces positive autocorrelation and can even flip its sign [1]. That is why the ARIMA(0,1,1) model, where differencing is paired with an MA term, is used more often than an ARIMA(1,1,0) [1].

You choose the orders by inspecting the autocorrelation function (ACF) and partial autocorrelation function (PACF) of the differenced series, then fitting candidates and comparing information criteria [3][2].

## Worked Example

The dataset is 36 months of monthly sales, from January 2021 to December 2023.

| Month | Sales | Month | Sales | Month | Sales |
|---|---|---|---|---|---|
| 2021-01 | 112 | 2022-01 | 163 | 2023-01 | 207 |
| 2021-02 | 118 | 2022-02 | 159 | 2023-02 | 203 |
| 2021-03 | 124 | 2022-03 | 168 | 2023-03 | 212 |
| 2021-04 | 131 | 2022-04 | 174 | 2023-04 | 218 |
| 2021-05 | 127 | 2022-05 | 170 | 2023-05 | 214 |
| 2021-06 | 135 | 2022-06 | 179 | 2023-06 | 223 |
| 2021-07 | 142 | 2022-07 | 185 | 2023-07 | 229 |
| 2021-08 | 138 | 2022-08 | 181 | 2023-08 | 225 |
| 2021-09 | 146 | 2022-09 | 190 | 2023-09 | 234 |
| 2021-10 | 152 | 2022-10 | 196 | 2023-10 | 240 |
| 2021-11 | 149 | 2022-11 | 192 | 2023-11 | 236 |
| 2021-12 | 157 | 2022-12 | 201 | 2023-12 | 245 |

The series length is n = 36. The first difference is $y_t - y_{t-1}$, so the first differenced value is 118 - 112 = 6.0. The series trends upward, so d = 1 is a reasonable starting point.

Fitting an ARIMA(1,1,1) gives these estimates:

| Quantity | Value |
|---|---|
| AR(1) coefficient, $\phi_1$ | -0.0354 |
| MA(1) coefficient, $\theta_1$ | 0.0115 |
| Innovation variance, $\sigma^2$ | 42.9194 |
| Log-likelihood | -115.4511 |
| AIC | 236.9021 |
| BIC | 241.5682 |

The AIC is computed as $-2(-115.4511) + 2 \times 3 = 236.9021$, counting three estimated parameters. The BIC is $-2(-115.4511) + 3 \times \ln(36) = 241.5682$.

Here is the code:

```python
import pandas as pd
from statsmodels.tsa.arima.model import ARIMA
y = pd.Series([112,118,124,131,127,135,142,138,146,152,149,157,
               163,159,168,174,170,179,185,181,190,196,192,201,
               207,203,212,218,214,223,229,225,234,240,236,245])
model = ARIMA(y, order=(1,1,1)).fit()
print(model.summary())
print(model.get_forecast(steps=3).summary_frame())
```

Key values from the printed summary and forecast table, rounded to four decimals:

```
ARIMA(1,1,1) fit: ar.L1=-0.0354, ma.L1=0.0115, sigma2=42.9194, AIC=236.9021, BIC=241.5682
Forecast: 244.7839, 244.7915, 244.7912
95% CI step 1: [231.9436, 257.6242]
```

The three-step forecast is 244.7839, then 244.7915, then 244.7912. The 95% interval for step 1 runs from 231.9436 to 257.6242, and for step 3 it widens to 222.8978 to 266.6847.

## How to Interpret It

Start with the coefficients. The AR(1) estimate of -0.0354 is small and close to zero, which means the previous differenced value carries little information about the next one. The MA(1) estimate of 0.0115 is also near zero. Together they say the differenced series behaves close to white noise, so the model is close to a random walk. Because statsmodels adds no drift term by default when d = 1, the forecast settles quickly to a nearly flat level around 244.8 and ignores the upward trend of about 3.8 units per month.

Read the information criteria as relative scores, not absolute ones. AIC of 236.9021 and BIC of 241.5682 only mean something when compared against other orders fitted to the same data. Lower values indicate a better trade-off between fit and complexity [2]. Because BIC penalizes parameters more heavily, it tends to pick smaller models.

Read the forecast interval as the honest part of the output. The point forecast barely moves across three steps, but the interval grows from roughly 26 units wide to roughly 44 units wide. That widening is the model telling you that uncertainty compounds with horizon.

## When to Use It (and when not to)

Use ARIMA when you have a single series with autocorrelation, a trend you can remove by differencing, and enough observations to estimate the parameters. It is a standard tool for interrupted time series analysis of large-scale interventions, where segmented regression is not adequate because of seasonality and autocorrelation [4]. Seasonal ARIMA extends the same idea to repeating patterns, and the most commonly used seasonal form is the (0,1,1)x(0,1,1) model, an MA(1) with a seasonal MA(1) and both a seasonal and non-seasonal difference [5]. You generally want four or five seasons of data to fit a seasonal ARIMA model [5].

Do not reach for ARIMA when the series is driven by external variables you can measure. A regression with ARIMA errors or a transfer function handles that better [4]. Do not use it when the series has structural breaks, regime changes, or variance that explodes, since differencing does not fix those. And do not expect it to capture long-run nonlinear behavior, because linear time series models such as ARIMA and exponential smoothing are limited in that respect [3].

## ARIMA vs ARMA

ARMA is the same model without the I. ARMA assumes the series is already stationary, so it has only two orders. ARIMA adds the differencing step, which is what lets it handle trending data.

| Feature | ARMA(p, q) | ARIMA(p, d, q) |
|---|---|---|
| Orders | Two: p, q | Three: p, d, q |
| Differencing | None | d differences applied first |
| Assumes stationarity | Yes | Achieved by differencing |
| Typical use | Stationary series | Trending or drifting series |
| Example | ARMA(1,1) | ARIMA(1,1,1) |

An ARIMA(p, 0, q) model is exactly an ARMA(p, q) model [2].

## Common Mistakes

- **Differencing when the series is already stationary.** If the lag-1 autocorrelation is zero or negative, or the autocorrelations are small and patternless, you do not need more differencing [3]. Overdifferencing shows up as a lag-1 autocorrelation of -0.5 or lower [3]. Fix: check the ACF before choosing d.
- **Reading AIC as an absolute quality score.** AIC and BIC are only comparable across models fitted to the same data with the same differencing. Fix: compare candidates on one consistent series and pick the lowest.
- **Ignoring the confidence interval.** A point forecast of 244.8 looks precise, but the 95% band at step 3 spans 222.8978 to 266.6847. Fix: always report the interval alongside the point forecast.
- **Adding too many seasonal parameters.** Using more than one or two seasonal parameters in the same model is likely to lead to overfitting [3]. Fix: keep the seasonal part small and check out-of-sample error.
- **Treating near-zero coefficients as proof the model is useless.** Small AR and MA estimates can still be correct, and they tell you the series is close to a random walk. Fix: judge the model on forecast error, not on coefficient size alone.
- **Forgetting that differencing changes the meaning of the constant.** The constant in a differenced model represents an average trend, not an average level [1]. Fix: interpret the constant in the units of the differenced series.

## Limitations

ARIMA cannot explain why a series moves. It only extrapolates patterns in the series itself, so it will not tell you that a price change or a policy caused a shift unless you add external regressors or a transfer function [4]. It also assumes the relationship between past and future stays stable, which fails around structural breaks.

Long-horizon forecasts from linear time series models are unreliable in general [3]. The intervals widen fast, and the point forecast often collapses to a flat line, as in the worked example where steps 1 through 3 barely differ. Treat anything beyond a few steps as a rough guide, and validate with a held-out period before trusting the model in production.

## Frequently Asked Questions

### What do p, d and q stand for in ARIMA?

The three integer components (p, d, q) are the AR order, the degree of differencing, and the MA order [2]. p counts how many past values enter the equation, d counts how many differences you take first, and q counts how many past errors enter. Together they define the shape of the model.

### How do I choose the values of p, d and q?

Inspect the ACF and PACF of the series, using the rules for identifying the order of differencing and the numbers of AR or MA terms [3]. Then fit several candidate orders and compare AIC or BIC, where lower is better [2]. Positive autocorrelation at low lags suggests AR terms, and negative autocorrelation suggests MA terms [1].

### What is the difference between ARIMA and ARMA?

ARMA has only p and q, and it assumes the series is already stationary. ARIMA adds the d term, so it can handle a trending series by differencing it first [1]. An ARIMA model with d = 0 is an ARMA model [2].

### How do I fit an ARIMA model in R or Python?

In R, the `arima` function fits a univariate time series and takes an `order` argument for the non-seasonal part plus a `seasonal` list for the seasonal part [2]. In Python, `statsmodels.tsa.arima.model.ARIMA` accepts the same order tuple, as shown in the worked example. Both estimate parameters by maximum likelihood by default.

### Can ARIMA handle seasonality?

Yes, through seasonal orders written as (P, D, Q) with a period. The seasonal part has the same structure as the non-seasonal part, with its own AR factor, MA factor and order of differencing [5]. The (0,1,1)x(0,1,1) model is the most commonly used seasonal form [5].

## References

1. [Introduction to ARIMA models](https://people.duke.edu/~rnau/411arim.htm)
2. [R: ARIMA Modelling of Time Series](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/arima.html)
3. [Rules for identifying ARIMA models](https://people.duke.edu/~rnau/arimrule.htm)
4. [Schaffer AL, Dobbins TA, Pearson SA. (2021). Interrupted time series analysis using autoregressive integrated moving average (ARIMA) models: a guide for evaluating large-scale health interventions. BMC medical research methodology](https://pmc.ncbi.nlm.nih.gov/articles/PMC7986567/)
5. [General seasonal ARIMA models -- (0,1,1)x(0,1,1) etc.](https://people.duke.edu/~rnau/seasarim.htm)

## Further Reading

- [R: ARIMA Modelling of Time Series - Preliminary Version](https://web.mit.edu/R/current/lib/R/library/stats/html/arima0.html)

## Related Articles

- [No Correlation: Definition, Graphs and Examples](/blog/data-analysis/no-correlation-definition-graphs-examples)
- [Extrapolation in Regression: Definition and Dangers](/blog/data-analysis/extrapolation-in-regression-definition)
- [What Are Residuals in Statistics? Definition and Formula](/blog/data-analysis/what-are-residuals-in-statistics)
- [What Is Slope? Definition, Formula and Regression Examples](/blog/data-analysis/what-is-slope-regression)
- [OLS Regression: What Ordinary Least Squares Means](/blog/data-analysis/ols-regression-ordinary-least-squares)