What Does Truncating Data Mean? Definition and Examples
By Dr. Zubair Khalid, DVM, MS, PhD ·

Truncating meaning in data work is simple: you cut off part of a value at a fixed point instead of rounding it to the nearest value. Truncating 2.83 to whole seconds gives 2, while rounding gives 3. The same word also describes cutting observations out of a dataset entirely, and both uses change what your analysis can tell you.
Quick Answer
- Truncation cuts a value or a dataset at a fixed boundary. Digits are dropped, or observations outside a range are removed.
- Truncating a number always moves it toward zero (for positive values, downward). Rounding moves it to the nearest value, so it can go up or down.
- Truncating a dataset removes observations whose values fall outside a boundary, so those rows never enter the analysis [1].
- Censoring is different. With censored data you keep every observation but do not know the true value of some of them [1].
- Truncating a dataset usually shrinks variance, and both kinds of truncation can shift the mean. In the worked example below, truncating reaction times to whole seconds moved the mean from 1.9900 to 1.5000, a shift of -0.4900 seconds.
What Truncating Means
In plain language, to truncate is to shorten or cut off at the end [2]. When you truncate a number, you keep the leading digits and discard the rest without adjusting what remains. When you truncate a dataset, you keep the observations inside a boundary and discard the rest.
The precise statistical definition is narrower. Truncation means some observations are not included in the analysis because of the value of the outcome variable [1]. If you study reaction times but only collect data from trials that finished within 3 seconds, every slower trial is missing from your file. The data you have are real values, but the sample is no longer a random draw from the full population.
That distinction matters because truncation is a sampling mechanism, not a measurement error. Nothing was recorded incorrectly. The values that remain are accurate. What changed is which values exist in your dataset at all.
How It Works
For a single number, truncation to $d$ decimal places follows this rule:
$$T(x) = \frac{\lfloor x \cdot 10^{d} \rfloor}{10^{d}}$$
- $x$ is the original value.
- $d$ is the number of decimal places you keep.
- $\lfloor \cdot \rfloor$ is the floor function, which returns the largest integer less than or equal to its input.
- $T(x)$ is the truncated value.
For positive numbers, this always moves the value down or leaves it unchanged. For negative numbers, the floor function moves the value toward negative infinity, so truncation toward zero needs a different rule. Most spreadsheet and programming tools handle the sign for you.
For a dataset, truncation is a filter on the outcome variable. If you truncate from below at $a$ and from above at $b$, you keep only observations where:
$$a \le y_i \le b$$
- $y_i$ is the outcome value for observation $i$.
- $a$ is the lower truncation point.
- $b$ is the upper truncation point.
Every observation outside that interval is dropped before analysis. The remaining sample is smaller, and its distribution is a slice of the original one.
Worked Example
The dataset is 10 reaction times in seconds from a simple lab task. Suppose you decide to record only whole seconds by truncating each value.
| Trial | Reaction time (s) | Truncated (floor to whole seconds) |
|---|---|---|
| 1 | 0.82 | 0 |
| 2 | 1.34 | 1 |
| 3 | 1.91 | 1 |
| 4 | 2.47 | 2 |
| 5 | 2.05 | 2 |
| 6 | 3.12 | 3 |
| 7 | 1.58 | 1 |
| 8 | 2.83 | 2 |
| 9 | 1.12 | 1 |
| 10 | 2.66 | 2 |
The steps, with the computed values:
- Original reaction times (s): 0.82, 1.34, 1.91, 2.47, 2.05, 3.12, 1.58, 2.83, 1.12, 2.66
- Truncated values: 0, 1, 1, 2, 2, 3, 1, 2, 1, 2
- Sum of original values: 19.90
- Mean before truncation: 19.90 / 10 = 1.9900
- Sum of truncated values: 15
- Mean after truncation: 15 / 10 = 1.5000
- Mean shift: 1.5000 - 1.9900 = -0.4900
Here is the same operation in Python:
import math, statistics
times = [0.82, 1.34, 1.91, 2.47, 2.05, 3.12, 1.58, 2.83, 1.12, 2.66]
trunc = [math.floor(t) for t in times]
print(f"{statistics.mean(times):.4f} {statistics.mean(trunc):.4f}")
Output:
1.9900 1.5000
Truncating cost you 0.4900 seconds of average reaction time. Every value moved down by up to one second, and the losses accumulated.
How to Interpret It
The mean shift is the headline. Truncation toward zero on positive data pulls the mean down, and the size of the pull depends on how much fractional value you discard. If most values sit just above a whole number, the loss is small. If they sit just below the next whole number, the loss approaches a full unit.
Variance behaves differently for the two kinds of truncation. When you truncate a dataset, the variance of the outcome variable is reduced compared to the distribution that is not truncated [1]. Truncating digits does not reliably shrink spread. In the worked example the sample variance rises from 0.5987 to 0.7222.
Direction matters when you truncate a dataset instead of a number. If the lower part of the distribution is truncated, the mean of the truncated variable is greater than the mean from the untruncated variable. If truncation is from above, the mean of the truncated variable is less than the untruncated variable [1]. Dropping slow trials makes your sample look faster. Dropping fast trials makes it look slower.
When to Use It (and when not to)
Use numeric truncation when the extra precision is meaningless or when a system requires it. Timestamps stored to the second, IDs cut to a fixed length, and integer bins built from continuous measurements are all common cases. Use dataset truncation when the study design genuinely excludes a range, such as a task that cannot record responses faster than 200 milliseconds.
Do not truncate when rounding would serve you better. Rounding keeps the value close to the truth on average, while truncation introduces a one-directional bias. Do not truncate a dataset to remove outliers you dislike. That is a modeling decision, and it belongs in the methods section with a stated boundary, not in a silent filter step.
If your data are truncated, standard regression on the remaining rows gives biased estimates. Truncated regression models exist for exactly this situation, and they are related to Heckman selection models used to correct for sampling selection bias [1]. If you only lost the true values of some observations but kept the rows, you have censored data, and a censored regression model is the right tool [1].
Truncation vs Rounding and Censoring
| Feature | Truncation (number) | Rounding | Censoring |
|---|---|---|---|
| Direction of change | Always toward zero | To nearest value | Value replaced by a boundary |
| Bias on positive data | Downward | Roughly balanced | Depends on the limit |
| Observations kept | All | All | All |
| True value known | Yes, after cutting | Yes, approximately | No, for censored rows |
| Typical fix | Accept the loss or keep more digits | None needed | Censored regression |
Truncation and censoring are the pair people mix up most. With censored data we have all of the observations, but we do not know the true values of some of them. With truncation, some observations are not included in the analysis because of the value of the outcome variable [1]. One keeps rows with unknown values. The other deletes rows with known values.
Common Mistakes
- Confusing truncation with rounding. Rounding 2.83 gives 3, truncation gives 2. Fix: check which function your tool applies, since
ROUNDandTRUNCbehave differently on the same input. - Treating truncated data as censored data. Censored rows stay in the file with a boundary value, truncated rows are gone. Fix: count your rows before and after filtering to see which case you have.
- Truncating too early in a text search. Cutting a term to a short root retrieves many irrelevant documents [2]. Fix: keep the root long enough to stay specific, and check the database help for the symbol it uses [3].
- Assuming truncation is harmless because the values are still real. The remaining values are accurate, but the sample is biased. Fix: report the truncation boundary and re-estimate with a truncated model.
- Forgetting that truncation symbols vary by database. Common symbols include *, !, ?, or # [3]. Fix: confirm the symbol in the database help before running the search.
- Truncating a search term in PubMed and expecting synonym expansion. Truncating a search term in PubMed disables automatic term mapping, so synonyms and MeSH terms are not added [4]. Fix: add the synonyms yourself when you truncate.
Limitations
Truncation cannot recover information you discarded. Once digits are gone or rows are removed, the original values are unavailable unless you kept a copy. For numeric truncation, the bias is systematic and one-directional, so it does not cancel out across a large sample the way random measurement error tends to.
For dataset truncation, ordinary summary statistics and regressions are misleading because the sample no longer represents the population you want to describe. Truncation also interacts badly with small samples, where losing a few rows can change the mean substantially. In text searching, truncation broadens a query and can flood your results with irrelevant records, which is why some words are poor candidates for it [5].
Frequently Asked Questions
What is the truncating meaning in simple terms?
Truncating means cutting something off at a fixed point. For a number, you drop the digits past that point. For a dataset, you drop the observations outside a boundary. In both cases the cut is deliberate and happens at a position you choose.
Is truncating the same as rounding?
No. Rounding looks at the digit after your cutoff and moves the value to the nearest representable number, so it can go up or down. Truncation ignores that digit and always moves positive values down. Truncating 1.91 gives 1, rounding gives 2.
What is the difference between truncation and censoring?
Truncation removes observations from the analysis because of their outcome value [1]. Censoring keeps every observation but leaves the true value unknown for some of them [1]. A trial that was never recorded is truncated. A trial recorded as "3 seconds or longer" is censored.
Does truncation change the mean?
Yes. Truncating positive values downward lowers the mean, and the size of the drop depends on how much fractional value you discard. In the worked example, the mean fell from 1.9900 to 1.5000 seconds, a shift of -0.4900.
How do I handle truncated data in a regression?
Do not run ordinary least squares on the remaining rows, because the estimates will be biased. Use a truncated regression model with the truncation point specified, such as the ll() option for left truncation or the ul() option for right truncation [1]. If your rows are censored instead, use a censored regression model [1].
References
- Truncated Regression | Stata Data Analysis Examples
- Using Truncation and Wildcards - Systematic Reviews - Research Guides at UC Davis
- Truncation - Database Search Tips - LibGuides at MIT Libraries
- Truncation - Getting the Most out of PubMed Medline - Research Guides at University of Hawaii at Manoa
- Home - Search Tips: Truncation and Boolean Searching - Library Research Guides at Wellesley College