Data Cleaning in Biostatistics
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Systematic Data Integrity is Paramount: Biological data quality directly dictates the validity of biostatistical conclusions. Prioritize a reproducible workflow using versioned scripts (R, Python, SPSS) to meticulously document and trace every data transformation, ensuring that missing values, outliers, and inconsistencies are systematically identified and corrected.
- Preserve Raw Data and Document Every Change: Never overwrite original datasets. Maintain a separate, untouched raw data repository and create a detailed cleaning log that records the variable, original value, new value, reason for change, and method used for every modification, enabling full auditability and reproducibility.
- Domain Knowledge is Indispensable for Biological Data: Distinguish genuine biological variation from technical artifacts by leveraging specific biological context. For instance, a blood glucose of 500 mg/dL might be plausible in a diabetic patient but an error in a healthy control, requiring expert judgment beyond statistical thresholds.
- Structured Decision-Making Prevents Ad Hoc Cleaning: Implement a decision matrix to classify data problems (e.g., implausible values, outliers, inconsistencies, duplicates) and define evidence requirements, decision options, and documentation standards. This framework ensures consistent and defensible cleaning choices, moving beyond subjective judgment calls.
- Rigorous Validation and Transparency are Non-Negotiable: After cleaning, validate the dataset by rerunning summary statistics, performing cross-variable checks (e.g., ensuring male participants do not have pregnancy test results), and comparing a sample of cleaned data against raw records. Transparently report all cleaning decisions and limitations, as data cleaning cannot rectify fundamental study design flaws or sampling errors.
Quick Answer
- Clean biological data systematically before analysis by documenting every change, checking missing values and outliers, and validating consistency across variables.
- Use a reproducible workflow with versioned scripts in R, Python, or SPSS so every transformation is traceable and repeatable.
- Data cleaning cannot recover fundamentally flawed study designs or sampling errors, so document limitations and report cleaning decisions transparently.
Understanding Data Cleaning in Biostatistics
Data cleaning is the process of detecting and correcting errors, inconsistencies, and missing values in biological datasets before statistical analysis. In biostatistics, the quality of your conclusions depends directly on the quality of your data. A dataset with duplicate records, misaligned variables, or unrecorded missing values can produce statistically significant results that are biologically meaningless.
Biological data arrives in many forms. Laboratory instruments generate continuous measurements such as gene expression levels, protein concentrations, or cell counts. Clinical records contain categorical variables like disease status, treatment group, or sex. Field observations may include spatial coordinates, environmental readings, or specimen identifiers. Each data type presents distinct cleaning challenges.
The goal of data cleaning is not to make the data look perfect. The goal is to make the data accurately represent the biological system you studied. This means you must distinguish between genuine biological variation and technical artifacts. A high outlier may be a real biological signal or a pipetting error. A missing value may indicate a failed assay or a censored observation. Your cleaning decisions shape the statistical results, so you must document every choice.
This article provides a software-agnostic workflow for data cleaning in biostatistics. It covers the core principles, practical steps, common failure patterns, and reporting expectations. The workflow applies to R, Python, and SPSS, and the concepts transfer across platforms.
At a Glance
| Cleaning Stage | Primary Question | Common Tools | Key Output |
|---|---|---|---|
| Data inventory | What variables and observations exist? | Spreadsheet review, metadata files | Variable dictionary, data dictionary |
| Missing value assessment | Why are values absent? | Summary tables, pattern plots | Missingness report |
| Outlier screening | Which values are biologically implausible? | Boxplots, z-scores, domain knowledge | Outlier log with decisions |
| Consistency checks | Do values follow defined rules? | Validation scripts, range checks | Error report with corrections |
| Duplicate detection | Are any records repeated? | Key matching, fuzzy matching | Duplicate resolution log |
| Final validation | Is the cleaned dataset ready for analysis? | Re-run summaries, cross-checks | Cleaned dataset with version control |
Core Principles of Data Cleaning
Preserve the Raw Data
The first principle of data cleaning is to preserve the original raw data. Never overwrite your raw files. Create a separate cleaned dataset and keep the raw data untouched. This allows you to trace any cleaning decision back to the original value and verify that your transformations are correct.
In practice, this means you should store raw data files in a read-only folder and write cleaned data to a separate location. If you use R or Python, your cleaning script should read the raw file and write a new cleaned file. If you use SPSS, you can save the cleaned dataset under a new name. The raw file remains unchanged.
Document Every Change
Every change you make to the data must be documented. This includes corrections, deletions, imputations, and transformations. Your documentation should record the variable name, the original value, the new value, and the reason for the change.
A cleaning log is a practical tool for this purpose. You can maintain it as a spreadsheet or as comments in your analysis script. The log should be detailed enough that another researcher can understand why each change was made. This documentation is essential for reproducibility and for reporting in your final publication.
Separate Cleaning from Analysis
Data cleaning and statistical analysis are distinct stages. You should complete the cleaning process before you begin the analysis. This separation prevents you from making ad hoc cleaning decisions during analysis, which can introduce bias. If you discover a data problem during analysis, you should return to the cleaning stage, update the cleaned dataset, and then rerun the analysis.
Use Domain Knowledge
Biological knowledge is essential for data cleaning. You cannot clean biological data without understanding the biological system. For example, a blood glucose measurement of 500 mg/dL may be plausible in a diabetic patient but implausible in a healthy control. A gene expression value of zero may be a true biological result or a technical failure. You must use your knowledge of the assay, the organism, and the experimental design to make these judgments.
Practical Workflow for Data Cleaning
Step 1: Inventory Your Data
Before you clean anything, you must know what you have. Create a data inventory that lists every variable, its type, its units, and its expected range. This inventory becomes your variable dictionary.
For each variable, record the following:
- Variable name as it appears in the dataset
- Variable label or description
- Data type (numeric, categorical, date, text)
- Units of measurement
- Expected range or allowed values
- Number of missing values
- Source of the data
This inventory helps you identify problems early. For example, you may find that a variable labeled "age" contains text values, or that a categorical variable has more categories than expected. The inventory also serves as a reference for your cleaning decisions.
Step 2: Assess Missing Values
Missing values are a common feature of biological data. They can arise from failed assays, lost samples, equipment errors, or participant dropout. The first step is to determine the pattern of missingness.
Missing data can be classified into three types:
- Missing completely at random (MCAR): The missingness is unrelated to any variable. For example, a sample tube breaks during storage.
- Missing at random (MAR): The missingness is related to other observed variables. For example, older participants are more likely to skip a follow-up visit.
- Missing not at random (MNAR): The missingness is related to the missing value itself. For example, a very high blood glucose measurement is more likely to be missing because the assay failed.
The pattern of missingness affects your cleaning decisions. If the missingness is MCAR or MAR, you may be able to use imputation methods. If the missingness is MNAR, imputation may introduce bias, and you should consider sensitivity analyses.
You should also check whether missing values are coded consistently. Some datasets use blank cells, others use "NA", "999", or "-999". You must standardize the missing value code across the dataset.
Step 3: Identify and Handle Outliers
Outliers are values that are far from the rest of the distribution. They can be genuine biological signals or technical errors. You must not automatically remove outliers. Instead, you should investigate each one.
The first step is to identify potential outliers. You can use graphical methods such as boxplots, histograms, or scatterplots. You can also use statistical methods such as z-scores or the interquartile range (IQR) rule. However, these methods are only screening tools. They do not tell you whether a value is an error.
For each potential outlier, you should ask:
- Is the value biologically plausible?
- Does the value fall within the expected range for this variable?
- Is the value consistent with other variables in the same observation?
- Could the value be a measurement error or a data entry error?
If the value is biologically plausible and consistent with the rest of the data, you should keep it. If the value is implausible, you should investigate the source. If you can confirm an error, you may correct or remove the value. If you cannot confirm an error, you should keep the value and note it in your cleaning log.
Step 4: Check Consistency and Range
Consistency checks verify that the values in your dataset follow the rules you defined. These rules may be based on the variable type, the expected range, or the relationships between variables.
Range checks verify that each value falls within the expected range. For example, a percentage should be between 0 and 100, a pH should be between 0 and 14, and a body temperature should be within a plausible biological range.
Cross-variable checks verify that values are consistent across variables. For example, if a participant is recorded as male and also has a pregnancy test result, there is an inconsistency. If a participant has a date of death that is earlier than the date of enrollment, there is an error.
You can implement these checks in R, Python, or SPSS. In R, you can use the dplyr package to filter and summarize. In Python, you can use pandas. In SPSS, you can use the VALIDATE command or write syntax for range checks.
Step 5: Detect and Remove Duplicates
Duplicate records can arise from data entry errors, merging multiple files, or repeated measurements. You must identify and handle duplicates before analysis.
The first step is to define what constitutes a duplicate. In a clinical study, a duplicate may be a participant with the same ID. In a laboratory experiment, a duplicate may be a sample with the same sample ID. In some cases, you may need to use multiple variables to identify duplicates, such as participant ID and visit date.
Once you identify duplicates, you must decide how to handle them. If the duplicates are exact copies, you can remove one. If the duplicates differ, you must investigate the source. You may need to check the original records to determine which record is correct.
Step 6: Standardize and Transform Variables
Biological data often require standardization before analysis. This includes:
- Converting units (e.g., from mg/dL to mmol/L)
- Recoding categorical variables (e.g., from "M" and "F" to "male" and "female")
- Creating derived variables (e.g., body mass index from height and weight)
- Transforming variables (e.g., log transformation for skewed data)
Each transformation must be documented. You should also verify that the transformation was applied correctly by checking a few values manually.
Step 7: Validate the Cleaned Dataset
After cleaning, you must validate the cleaned dataset. This involves rerunning the summary statistics and checking that the data meets your expectations. For example, you should check that the number of rows matches the expected number of observations, that the missing value counts are as expected, and that the range of each variable is correct.
You should also cross-check the cleaned dataset against the raw data for a random sample of records. This helps you catch errors in your cleaning script.
Software-Specific Implementation
Data Cleaning in R
R is a powerful tool for data cleaning. The tidyverse package provides a set of functions for data manipulation. The dplyr package includes functions for filtering, selecting, mutating, and summarizing data. The tidyr package provides functions for reshaping data.
A typical R cleaning script might look like this:
library(tidyverse)
## Read raw data
raw_data <- read_csv("raw_data.csv")
## Check missing values
summary(raw_data)
## Filter out implausible values
cleaned_data <- raw_data %>%
filter(age >= 0 & age <= 120) %>%
filter(weight > 0)
## Recode categorical variables
cleaned_data <- cleaned_data %>%
mutate(sex = recode(sex, "M" = "male", "F" = "female"))
## Write cleaned data
write_csv(cleaned_data, "cleaned_data.csv")
The key advantage of R is reproducibility. Your cleaning script is a record of every change you made. You can rerun the script on the raw data to produce the same cleaned dataset.
Data Cleaning in Python
Python with the pandas library is another common choice. The pandas library provides functions for reading, filtering, and transforming data.
A common Python cleaning script might look like this:
import pandas as pd
## Read raw data
raw_data = pd.read_csv("raw_data.csv")
## Check missing values
print(raw_data.isnull().sum())
## Filter out implausible values
cleaned_data = raw_data[
(raw_data["age"] >= 0) & (raw_data["age"] <= 120)
]
## Recode categorical variables
cleaned_data["sex"] = cleaned_data["sex"].map({"M": "male", "F": "female"})
## Write cleaned data
cleaned_data.to_csv("cleaned_data.csv", index=False)
Python is particularly useful when you need to integrate data cleaning with other data processing steps, such as merging multiple files or working with large datasets.
Data Cleaning in SPSS
SPSS provides a graphical interface for data cleaning. You can use the "Data" menu to sort cases, identify duplicates, and recode variables. You can also use the "Transform" menu to compute new variables.
For reproducibility, you should use SPSS syntax instead of the graphical interface. The syntax is a text file that records your cleaning steps. You can rerun the syntax on the raw data to produce the same cleaned dataset.
A simple SPSS syntax for data cleaning might look like this:
GET FILE='raw_data.sav'.
FREQUENCIES VARIABLES=age sex weight.
SELECT IF (age >= 0 AND age <= 120).
RECODE sex ('M'='male') ('F'='female').
SAVE OUTFILE='cleaned_data.sav'.
Records and Measurements
The Cleaning Log
The cleaning log is the central record of your data cleaning process. It should contain the following information for each change:
- Date of the change
- Variable name
- Original value
- New value
- Reason for the change
- Method used (e.g., manual correction, script, imputation)
The cleaning log should be stored with the cleaned dataset. It is the evidence that your data cleaning was systematic and transparent.
The Data Dictionary
The data dictionary is a separate document that describes each variable in the dataset. It should include the variable name, description, type, units, and allowed values. The data dictionary is essential for anyone who uses the dataset after you.
Version Control
You should use version control for your data files and cleaning scripts. This allows you to track changes over time and to revert to earlier versions if needed. Tools like Git are commonly used for this purpose.
Common Failure Patterns
Overcleaning
Overcleaning occurs when you remove too many values or make too many corrections. This can happen when you rely on statistical methods without domain knowledge. For example, you might remove all outliers based on a z-score threshold, even though some outliers are genuine biological signals. Overcleaning reduces the sample size and can introduce bias.
Undercleaning
Undercleaning occurs when you fail to identify and correct errors. This can happen when you skip the inventory step or when you do not check for duplicates. Undercleaning can produce results that are not reproducible.
Inconsistent Missing Value Codes
Biological datasets often use different codes for missing values. Some use blank cells, others use "NA", "999", or "-999". If you do not standardize these codes, your analysis will treat missing values as real values. This can produce incorrect results.
Ignoring the Pattern of Missingness
The pattern of missingness matters. If you ignore the pattern, you may use an imputation method that is not appropriate for the data. For example, if the missingness is MNAR, imputation will introduce bias.
Not Documenting Changes
If you do not document your cleaning decisions, you cannot reproduce your results. This is a serious problem for scientific integrity. Reviewers and other researchers need to know what you did to the data.
Reporting and Transparency
Reporting Guidelines
Transparent reporting of your data cleaning process is essential. The EQUATOR Network provides reporting guidelines for health research. These guidelines help you report your methods clearly and completely. You should consult the relevant guideline for your study type and follow its recommendations for reporting data cleaning.
Publication Ethics
The Committee on Publication Ethics (COPE) provides core practices for ethical research and publication. These practices include the responsible handling of data. You must report your data cleaning methods honestly and completely. You must not hide or misrepresent the changes you made to the data.
Data Management and Sharing
The National Institutes of Health (NIH) has a Data Management and Sharing Policy. This policy requires that you plan for the management and sharing of your data. Your data management plan should describe how you will clean, document, and share your data. The policy is available at the NIH Data Management and Sharing Policy.
Limitations of Data Cleaning
Data cleaning cannot fix a fundamentally flawed study design. If your sample is biased, your measurements are invalid, or your variables are poorly defined, no amount of cleaning will produce valid results. Data cleaning can only correct errors in the data you have collected.
Data cleaning also cannot recover data that is missing. If a sample was lost or an assay failed, you cannot create the missing value. You can only decide how to handle the missingness in your analysis.
Finally, data cleaning decisions are subjective. Different researchers may make different decisions about how to handle outliers or missing values. This is why documentation is so important. You must report your decisions so that others can evaluate them.
Professional Escalation Criteria
You should escalate data cleaning issues to a supervisor or a statistician when:
- You cannot determine whether a value is a genuine biological signal or an error.
- The pattern of missingness is complex and you are unsure how to handle it.
- You find a large number of errors in the data, which may indicate a problem with the data collection process.
- You are unsure whether a transformation is appropriate for your data.
- You are planning to use imputation and you are unsure which method is appropriate.
Building a Cleaning Decision Log and Audit Trail for Biological Datasets
A recurring failure in biological data cleaning is not the absence of effort but the absence of a structured decision framework that forces consistency and accountability. Researchers often clean data interactively, making judgment calls on outliers, missing values, and inconsistencies without a formal record of why each decision was made. When the analysis is later questioned, the researcher cannot reconstruct the reasoning behind a particular exclusion or correction. This section provides a practical decision framework, a record system, and a troubleshooting method that you can implement in R, Python, or SPSS, or even with a spreadsheet and paper forms.
The Core Problem: Decisions Without a Decision Log
The existing workflow describes what to do at each cleaning stage, but it does not provide a mechanism for capturing the reasoning behind each action. A cleaning log that records the variable name, original value, new value, and reason is a start, but it lacks structure. It does not force you to classify the type of decision, the evidence used, or the confidence level. Without this structure, the log becomes a narrative that is difficult to audit, difficult to compare across datasets, and difficult to defend in peer review.
The decision framework presented here addresses this gap. It is a formal, repeatable process that forces you to classify every data problem into a defined category, apply a pre-specified decision rule, and record the outcome in a structured format. This approach reduces the risk of ad hoc decisions that are inconsistent with the rest of the dataset.
The Decision Matrix: A Structured Approach to Data Problems
The decision matrix is a table that maps the type of data problem to a set of possible actions and the conditions under which each action is appropriate. The matrix is not a substitute for domain knowledge. It is a tool that forces you to articulate the reasoning behind each decision and to apply the same reasoning consistently across the dataset.
The matrix has four columns: Problem Type, Evidence Required, Decision Options, and Documentation Requirement. The Problem Type column lists the categories of problems you will encounter. The Evidence Required column specifies what information you need to gather before making a decision. The Decision Options column lists the possible actions. The Documentation Requirement column specifies what must be recorded in the decision log.
Problem Type 1: Implausible Value
An implausible value is a value that falls outside the biologically possible range for the variable. For example, a human body temperature of 45 degrees Celsius is implausible. A blood pH of 8.5 is implausible. A negative cell count is implausible.
Evidence Required: You must verify the value against the raw data source. Check the original instrument output, the data entry form, or the laboratory record. Confirm that the value was not a transcription error.
Decision Options: Correct the value if you can identify the correct value from the source. Remove the observation if the correct value cannot be determined and the value is clearly an error. Keep the value if you cannot confirm an error and the value is within a plausible range for the biological system.
Action: Record the decision in the log with the problem type, the evidence used, and the action taken.
Problem Type 2: Outlier
An outlier is a value that is far from the rest of the distribution but may be biologically plausible. For example, a very high blood glucose measurement in a diabetic patient is an outlier but is biologically plausible. A very high gene expression value in a sample from a tumor may be a real signal.
Evidence Required: Check the value against the expected biological range. Check the value against other variables in the same observation. For example, if a participant has a very high weight, check the height and body mass index. If the value is consistent with other variables, it is more likely to be a real signal.
Decision Options: Keep the value if it is biologically plausible and consistent with other variables. Investigate the value if it is implausible or inconsistent. Remove the value only if you can confirm an error from the original data.
Action: Record the decision in the decision log with the problem type, the value, the evidence used, and the decision taken.
Problem Type 3: Inconsistent
An inconsistent value is a value that contradicts another value in the same observation. For example, a participant recorded as male with a pregnancy test result. A date of death that is earlier than the date of enrollment. A weight that is recorded in kilograms but is clearly in pounds.
Decision: Check the original data source. Determine which value is correct. If the inconsistency is due to a data entry error, correct the value. If the inconsistency is due to a coding error, recode the value. If the inconsistency cannot be resolved, flag the observation for review.
Action: Record the decision in the decision log with the problem type, the values, the resolution, and the action taken.
Problem Type 4: Duplicate
A duplicate is a record that appears more than once in the dataset. Duplicates can be exact copies or they can differ in some variables.
Decision: Define the key variables that identify a unique observation. For a clinical study, this is the participant ID and visit date. For a laboratory experiment, this is the sample ID and the assay date. Check the duplicate records to determine if they are exact copies or if they differ.
Decision Options: If the duplicates are exact copies, remove one. If the duplicates differ, investigate the source. If the differences are due to data entry errors, correct the records. If the differences are due to a repeated measurement, keep both records and note the reason.
Action: Record the decision in the decision log with the type, the duplicate records, the resolution, and the decision.
The Decision Log Structure
The decision log is the structured record of every decision you make during data cleaning. It is a table with the following columns:
- Decision ID: A unique identifier for each decision.
- Date: The date the decision was made.
- Variable: The variable name.
- Observation ID: The unique identifier for the observation (e.g., participant ID, sample ID).
- Problem Type: The category of the problem (implausible, outlier, inconsistent, duplicate).
- Original Value: The value as it appeared in the raw data.
- New Value: The value after the decision was applied.
- Evidence Used: A description of the evidence you used to make the decision.
- Decision: The action taken (keep, correct, remove, impute).
- Reason: A brief explanation of the reasoning behind the decision.
- Confidence Level: A rating of your confidence in the decision (high, medium, low).
The decision log is a living document. You add a new row for every decision you make. At the end of the cleaning process, the decision log is a complete record of every change you made to the data.
Implementing the Decision Framework in R, Python, and SPSS
The decision framework is software-agnostic. You can implement it in any tool you use for data cleaning. The key is to create a structured process that forces you to record each decision.
Implementation in R
In R, you can create a decision log as a data frame. Each row of the data frame is a decision. You can add rows to the log as you make decisions.
## Create an empty decision log
decision_log <- data.frame(
decision_id = integer(),
date = character(),
variable = character(),
observation_id = character(),
problem_type = character(),
original_value = character(),
new_value = character(),
evidence = character(),
decision = character(),
reason = character(),
confidence = character(),
stringsAsFactors = FALSE
)
## Example: Add a decision to the log
decision_log <- rbind(decision_log, data.frame(
decision_id = 1,
date = "2025-01-15",
variable = "body_temp",
observation_id = "P001",
problem_type = "Implausible",
original_value = "45.2",
new_value = "NA",
evidence = "Checked raw data entry form, value was 36.2, transcription error",
decision = "Correct",
reason = "Transcription error confirmed",
confidence = "High"
))
The decision log is a data frame that you can save as a CSV file at the end of the cleaning process. This provides a permanent record of every decision.
Implementation in Python
In Python, you can use the pandas library to create a decision log as a DataFrame.
import pandas as pd
## Create an empty decision log
decision_log = pd.DataFrame(columns=[
"decision_id", "date", "variable", "observation_id",
"problem_type", "original_value", "new_value",
"evidence", "decision", "reason", "confidence"
])
## Add a decision to the log
new_decision = pd.DataFrame([{
"decision_id": 1,
"date": "2025-01-15",
"variable": "body_temp",
"observation_id": "P001",
"problem_type": "Implausible",
"original_value": "45.2",
"new_value": "NA",
"evidence": "Checked raw data entry form. Value was 45.2, transcription error",
"decision": "Correct",
"reason": "Transcription error from original",
"confidence": "High"
}])
decision_log = pd.concat([decision_log, new_decision], ignore_index=True)
The decision log is a DataFrame that you can save to a CSV file at the end of the cleaning process.
Implementation in SPSS
In SPSS, you can create a decision log as a separate data file. You can use the DATA LIST command to define the structure of the log, and then use ADD FILES to add new records.
DATA LIST FREE / decision_id (F4) date (A10) variable (A20) observation_id (A10) problem_type (A20) original_value (A10) new_value (A10) evidence (A100) decision (A20) reason (A100) confidence (A10).
BEGIN DATA
1 2025-01-15 body_temp P001 Implausible 45.2 NA "Checked raw data entry form" Correct "Transcription error" High
END DATA.
SAVE OUTFILE='decision_log.sav'.
The decision log is a separate data file that you can review and share with your team.
The Troubleshooting Method: The Five-Pass Review
The decision framework is most effective when combined with a structured troubleshooting method. The five-pass review is a systematic approach to reviewing the cleaned dataset for errors that may have been missed during the initial cleaning.
Pass 1: Variable-Level Review
The first pass is a variable-level review. For each variable, you check the summary statistics. You look at the minimum, maximum, mean, median, and standard deviation. You check the number of missing values. You compare these statistics to the expected values from your data dictionary.
If a variable has a minimum or maximum that is outside the expected range, you investigate. If a variable has a high number of missing values, you investigate the pattern of missingness.
Pass 2: Observation-Level Review
The second pass is an observation-level review. You select a random sample of observations and check the values across all variables. You look for inconsistencies. For example, if a participant has a body mass index of 35, you check that the height and weight values are consistent with that BMI.
You also check for duplicate records. You sort the data by the key variables and look for adjacent rows that are identical or nearly identical.
Pass 3: Cross-Variable Review
The third pass is a cross-variable review. You look for relationships between variables that should hold. For example, if a participant is recorded as male, the pregnancy test variable should be missing or negative. If a participant has a date of death, the date of death should be after the date of enrollment.
You can implement these checks in R, Python, or SPSS. In R, you can use the dplyr package to filter and summarize. In Python, you can use pandas. In SPSS, you can use the SELECT IF command.
Pass 4: Source Verification
The fourth pass is a source verification. You select a random sample of observations and compare the cleaned data to the raw data. You check that the values in the cleaned data match the values in the raw data, except for the changes you documented in the decision log.
This pass is critical for catching errors in your cleaning script. If you made a mistake in a transformation, you will find it here.
Pass 5: Final Review
The fifth pass is a final review. You review the decision log to ensure that every decision is documented. You check that the decision log is complete and that the confidence levels are recorded. You review the data dictionary to ensure that it is up to date.
The final review is the last step before you begin the analysis. It is your opportunity to catch any remaining problems.
Common Failure Patterns in Decision Making
The decision framework is designed to prevent common failure patterns. However, it is important to be aware of these patterns so that you can recognize them in your own work.
Pattern 1: The Automatic Removal
The automatic removal occurs when you remove an outlier or an implausible value without investigating it. This is a common failure pattern because it is easy to apply a statistical rule, such as a z-score threshold, without thinking about the biological context. The decision framework forces you to record the evidence you used to make the decision. If you cannot provide evidence, you should not remove the value.
Pattern 2: The Inconsistent Decision
The inconsistent decision occurs when you make different decisions for the same type of problem. For example, you might remove an outlier in one variable but keep an outlier in another variable. The decision framework forces you to apply the same decision rule to the same problem type. If you have a rule that says "remove an outlier if it is outside the 3 standard deviation range," you must apply that rule to all variables.
Pattern 3: The Unrecorded Decision
The unrecorded decision occurs when you make a change to the data but do not record it in the decision log. This is a serious problem because it makes the cleaning process unreproducible. The decision framework requires that you record every decision. If you cannot record a decision, you should not make the change.
Pattern 4: The Overconfident Decision
The overconfident decision occurs when you make a decision with high confidence but you have not gathered enough evidence. For example, you might remove a value because it is outside the expected range, but you have not checked the original data source. The decision framework requires that you record your confidence level. If you are not confident, you should gather more evidence or escalate the decision.
Professional Escalation Criteria
The decision framework includes a clear escalation path. You should escalate a decision to a supervisor or a statistician when:
- You cannot determine the correct action for a problem type.
- The evidence is insufficient to make a confident decision.
- The decision has a high impact on the analysis results.
- You find a pattern of errors that suggests a problem with the data collection process.
The escalation criteria are not a sign of failure. They are a sign that you are following a rigorous process. A supervisor or statistician can provide the domain knowledge or statistical expertise that you need to make a sound decision.
The Role of the Decision Framework in Reporting
The decision framework is beyond a tool for the cleaning process. It is also a tool for reporting. When you write your methods section, you can describe the decision framework and the decision log. You can report the number of decisions made, the types of problems encountered, and the actions taken. This transparency is essential for reproducibility and for peer review.
The EQUATOR Network provides reporting guidelines for health research. These guidelines help you report your methods clearly and completely. You should consult the relevant guideline for your study type and follow its recommendations for reporting data cleaning.
The Committee on Publication Ethics provides core practices for ethical research and publication. These practices include the responsible handling of data. You must report your data cleaning methods honestly and completely. You must not hide or misrepresent the changes you made to the data.
The NIH Data Management and Sharing Policy requires that you plan for the management and sharing of your data. Your data management plan should describe how you will clean, document, and share your data. The decision log is a key part of this documentation.
Practical Implementation Steps
To implement the decision framework in your own work, follow these steps:
- Create a decision log template. Use the structure described above or adapt it to your needs.
- Define your problem types. Use the four types described above or add your own.
- Define your decision rules. For each problem type, specify the evidence required and the possible actions.
- Apply the framework to your data. For each problem you encounter, classify the problem type, gather evidence, make a decision, and record the decision in the log.
- Perform the five-pass review. After the initial cleaning, review the data using the five-pass method.
- Review the decision log. Ensure that all decisions are recorded and that the log is complete.
- Report the decision framework in your methods. Describe the framework and the decision log in your publication.
The decision framework is a practical tool that you can implement today. It does not require special software or advanced statistical knowledge. It requires a commitment to transparency and a willingness to document your decisions. The result is a cleaner dataset, a reproducible process, and a stronger publication.
Frequently Asked Questions
What is the difference between data cleaning and data transformation?
Data cleaning is the process of correcting errors and handling missing values. Data transformation is the process of changing the scale or distribution of a variable, such as log transformation. Cleaning is done before transformation.
How do I decide whether to remove an outlier?
You should investigate each outlier individually. If the value is biologically plausible and consistent with the rest of the data, keep it. If the value is implausible and you can confirm an error, you may remove it. If you cannot confirm an error, keep the value and note it in your cleaning log.
What is the best way to handle missing values?
The best way depends on the pattern of missingness. If the missingness is MCAR or MAR, you may use imputation. If the missingness is MNAR, imputation may introduce bias. You should consult a statistician for complex missingness patterns.
Should I use a statistical test to identify outliers?
Statistical tests can screen for potential outliers, but they cannot tell you whether a value is an error. You must use domain knowledge to make the final decision. A z-score threshold is a screening tool, not a decision rule.
How do I document my data cleaning process?
Maintain a cleaning log that records every change you make. The log should include the variable name, original value, new value, and reason for the change. Store the log with the cleaned dataset.
What should I do if I find a large number of errors in my data?
If you find a large number of errors, you should investigate the source. The errors may indicate a problem with the data collection process. You should escalate the issue to your supervisor or the data collection team.
How does data cleaning affect my statistical results?
Data cleaning can change the results of your analysis. Removing outliers can change the mean and standard deviation. Imputing missing values can change the sample size and the standard errors. You must report your cleaning decisions so that readers can understand the impact.
What are the reporting guidelines for data cleaning?
The EQUATOR Network provides reporting guidelines for health research. These guidelines help you report your methods clearly. You should consult the relevant guideline for your study type and follow its recommendations.
Using the Evidence
| Source | Best use in this topic | Important limitation |
|---|---|---|
| Research Methods Resources | official guidance | Check the linked page for current local requirements |
| EQUATOR Network | official guidance | Check the linked page for current local requirements |
| Core Practices | official guidance | Check the linked page for current local requirements |
Related Bioinformatics Guides
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
- Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight
- Metagenomics Data Analysis: From Raw Reads to Biological Insights
- Proteomics Data Analysis Workflow: From Raw Spectra to Biological Insights
- Spatial Omics Data Analysis: From Image Processing to Biological Interpretation
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Research Methods Resources. National Library of Medicine.
- EQUATOR Network. EQUATOR Network.
- Core Practices. Committee on Publication Ethics.
- NIH Grants and Funding. National Institutes of Health.
- ORCID for Researchers. ORCID.
- Data Management and Sharing Policy. National Institutes of Health.
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Cleaning at Home and at Work in Relation to Lung Function Decline and Airway Obstruction.. American journal of respiratory and critical care medicine, 2018.
- Chlorhexidine-alcohol compared with povidone-iodine-alcohol skin antisepsis protocols in major cardiac surgery: a randomized clinical trial.. Intensive care medicine, 2024.
- A Primer of Data Cleaning in Quantitative Research: Handling Missing Values and Outliers.. Journal of advanced nursing, 2026.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.