# What Is Data? Definition, Meaning and Examples in Science

Data is any recorded observation, measurement or fact that describes something about the world. The formal definition of data is a set of values, characters or symbols that represent properties of objects or events and that can be stored, processed and analyzed. In science, data is the raw material that evidence is built from, and it only becomes information after someone interprets it in context.

## Quick Answer

- Data is recorded observation. A temperature reading, a species name, a yes/no answer and a date are all data.
- Data becomes information when it is processed, organized or interpreted so it answers a question.
- In science, data is the evidence produced by measurement, observation or experiment, and it must be reproducible to count.
- Data comes in two broad types: quantitative (numbers) and qualitative (labels or categories).
- The measurement scale matters. Nominal, ordinal, interval and ratio scales allow different operations, so they support different analyses.

## What Data Means

In everyday speech, people use "data" for anything from a spreadsheet to a phone bill. The plain definition is simpler: data is a recorded value that stands for something you observed. If you write down that a patient's temperature is 23.4 C, that number is a datum. If you record that a sample is "blue," that word is also a datum.

The precise statistical definition adds two conditions. First, data must be recorded in a form that can be stored and processed, whether that is a number, a category label or a timestamp. Second, data must be tied to a defined unit of observation, so each value describes a specific thing. A response variable and a predictor variable in a study are both data, and they are only meaningful together with the case they describe [1].

This is why database theory treats data as a representation that can change without changing what it means. Codd's relational model was built on the idea that users should be protected from how data is organized internally, because the representation can change while the underlying facts stay the same [2]. The definition of data therefore covers both the values and the structure that holds them.

A useful test: if you cannot say what was measured, on whom or what, and in what unit, you have a number but not usable data.

## How It Works

Data is classified along two independent axes. The first is type, which splits into quantitative and qualitative. The second is measurement scale, which determines what math you can do.

Quantitative data is numeric and answers "how much" or "how many." It splits further:

- Discrete data takes countable, separate values, such as the number of students in a room.
- Continuous data can take any value in a range, limited only by measurement precision, such as a mass or a temperature.

Qualitative data describes qualities or categories. It splits by scale:

- Nominal data is unordered labels, such as color or sex.
- Ordinal data has a meaningful order but no fixed spacing, such as "agree" or a rank.

The measurement scales follow a hierarchy of what operations are valid:

- Nominal: equality only. You can count frequencies.
- Ordinal: equality and order. You can rank and take medians.
- Interval: equal spacing between values, but no true zero. You can add and subtract.
- Ratio: equal spacing plus a true zero. You can multiply, divide and take ratios.

The normal distribution is the most frequently used distribution in both statistical theory and application, and it is defined by two parameters, the mean $\mu$ and the standard deviation $\sigma$ [3]:

$$ f(x) = \frac{1}{\sigma \sqrt{2 \pi}} e^{ -\frac{1}{2} \left(\frac{x - \mu}{\sigma} \right)^2 } $$

Here $x$ is a value of the variable, $\mu$ is the population mean, and $\sigma$ is the population standard deviation. The notation $X \sim N(\mu, \sigma^2)$ states that $X$ is normally distributed with those parameters [3]. This matters for data because the scale you assign decides whether a normal model is even appropriate.

## Worked Example

Suppose you have 12 short data entries from a mixed study and you want to classify each one by type, discreteness and scale. The table below shows the entries and their classifications.

| Value | Data Type | Discreteness | Scale |
|---|---|---|---|
| 23.4 C | Quantitative | Continuous | Interval |
| blue | Qualitative | Nominal | Nominal |
| yes | Qualitative | Nominal | Nominal |
| 2024-03-01 | Quantitative | Discrete | Interval |
| 7 | Quantitative | Discrete | Ratio |
| 3.14 | Quantitative | Continuous | Ratio |
| male | Qualitative | Nominal | Nominal |
| 150 kg | Quantitative | Continuous | Ratio |
| rank 2 | Qualitative | Ordinal | Ordinal |
| 0.85 | Quantitative | Continuous | Ratio |
| 12 students | Quantitative | Discrete | Ratio |
| agree | Qualitative | Ordinal | Ordinal |

Walking through the counts:

- Total entries classified: n = 12.
- Quantitative count: 7 of 12 = 58.3333%.
- Qualitative count: 5 of 12 = 41.6667%.
- Discrete count: 3.
- Continuous count: 4.
- Nominal scale count: 3.
- Ordinal scale count: 2.
- Interval scale count: 2.
- Ratio scale count: 5.
- Check: quant + qual = 7 + 5 = 12 = n.

The code below reproduces the type counts.

```python
import pandas as pd
df = pd.DataFrame([
    ['23.4 C', 'Quantitative', 'Continuous', 'Interval'],
    ['blue', 'Qualitative', 'Nominal', 'Nominal'],
    ['yes', 'Qualitative', 'Nominal', 'Nominal'],
    ['2024-03-01', 'Quantitative', 'Discrete', 'Interval'],
    ['7', 'Quantitative', 'Discrete', 'Ratio'],
    ['3.14', 'Quantitative', 'Continuous', 'Ratio'],
    ['male', 'Qualitative', 'Nominal', 'Nominal'],
    ['150 kg', 'Quantitative', 'Continuous', 'Ratio'],
    ['rank 2', 'Qualitative', 'Ordinal', 'Ordinal'],
    ['0.85', 'Quantitative', 'Continuous', 'Ratio'],
    ['12 students', 'Quantitative', 'Discrete', 'Ratio'],
    ['agree', 'Qualitative', 'Ordinal', 'Ordinal'],
], columns=['Value','Data Type','Discreteness','Scale'])
print(df['Data Type'].value_counts())
```

Output:

```text
Data Type
Quantitative    7
Qualitative     5
Name: count, dtype: int64
```

The split of 7 quantitative and 5 qualitative is a feature of this small set, not a rule. What matters is that every entry has a defensible type and scale. The temperature in Celsius is interval because 0 C is not an absolute zero of temperature, while 150 kg is ratio because 0 kg means no mass. That distinction decides whether you can say one value is "twice" another.

## How to Interpret It

Read the scale before you read the number. A value of 7 in a ratio scale supports statements like "twice as many," while a rank of 2 in an ordinal scale does not support "twice as good."

Check the discreteness next. Continuous data is usually rounded when recorded, so a recorded 3.14 is a rounded value, not an exact one. Discrete data has no such rounding problem because the values are countable.

Then check the unit and the case. A number without a unit and a case is ambiguous. "12" could be 12 students, 12 seconds or 12 kilograms, and each implies a different analysis.

Finally, ask whether the data is raw or processed. A mean is a summary statistic, not raw data. Keeping the raw values lets you recompute summaries and check them, which is why presentation of numerical data deserves its own care [4].

## When to Use It (and when not to)

Use this classification whenever you start an analysis. It tells you which summary statistics are valid, which charts make sense, and which tests apply. If you are working with a full collection of values, it helps to think of it as a [dataset with defined structure](/blog/data-analysis/what-is-a-dataset) before you classify individual entries.

Use it when you plan a study, because the scale you choose at collection time limits what you can do later. Recording age as a category instead of a number throws away information you cannot recover.

Do not use the classification to decide causation. Knowing a variable is ratio scale tells you nothing about whether it causes another variable. Correlation and regression on repeated data need care because repeated measurements on the same unit are not independent [5].

Do not force every variable into a numeric scale. Converting "agree" to 1 and "disagree" to 0 is common, but the resulting numbers inherit only the properties you assume, and the assumption should be stated.

## Data vs Information

The closest related idea is information. Data is the raw recorded value. Information is data that has been processed, organized or interpreted so it answers a question. The same data can produce different information depending on the question you ask.

| Aspect | Data | Information |
|---|---|---|
| Form | Raw values, labels, timestamps | Processed, organized, interpreted |
| Meaning | Needs context to interpret | Carries meaning for a question |
| Example | 23.4 C, 150 kg, agree | "Average temperature rose 2 C this week" |
| Depends on | What was measured and how | The question and the analysis |
| Changes when | New observations are recorded | The question or method changes |

A single temperature reading is data. The statement that the average temperature this week was higher than last week is information. The reading did not change, but the interpretation did.

## Common Mistakes

- Treating all numbers as the same type. A zip code is numeric but nominal, so averaging it is meaningless. Fix: check whether arithmetic on the values makes sense before you compute anything.
- Averaging ordinal data. The mean of "agree" and "disagree" codes depends on the codes you chose. Fix: use frequencies or medians for ordinal data.
- Ignoring the unit. A value of 150 is not data until you know it is kilograms. Fix: record units with every value.
- Confusing discrete and continuous. Counting how many people are in a room gives discrete data even if the average is fractional. Fix: ask whether the underlying quantity is countable or measurable.
- Dropping the case identifier. Without knowing which subject or sample a value belongs to, you cannot link variables. Fix: keep an ID column.
- Treating transformed data as raw. Log transforms and standardization change the scale. Fix: keep the original values and document the transformation [6].

## Limitations

This classification describes the data you have, not the data you wish you had. It cannot tell you whether your sample represents the population, whether your measurements are accurate, or whether your study design is sound. A perfectly classified dataset can still be biased.

The scale labels also depend on judgment. Whether a temperature scale is interval or ratio depends on which zero you use, and whether a rating scale is ordinal or interval is a modeling choice. Multidimensional data adds another layer, because relationships between variables are not visible in any single column [7]. Treat the classification as a starting point for analysis, not a verdict on it.

## Frequently Asked Questions

### What is the simplest definition of data?

Data is any recorded observation or measurement that can be stored and processed. It includes numbers, category labels, text and timestamps. The key requirement is that each value describes a defined thing and can be linked back to what was observed.

### What is the difference between data and information?

Data is the raw recorded value, and information is data that has been processed or interpreted to answer a question. The same data can yield different information depending on the question. Processing does not change the data, it changes what you can conclude from it.

### What counts as data in science?

In science, data is the evidence produced by measurement, observation or experiment. It must be recorded in a reproducible way so another researcher can check it. A response variable and a predictor variable in a study are both data, and they are meaningful together with the case they describe [1].

### Is a date considered data?

Yes. A calendar date is interval-scale data. Dates have a meaningful order and equal spacing, so you can sort them and subtract them to get a duration, but the zero point is arbitrary, so ratios such as "twice as late" are meaningless.

### Can qualitative data be analyzed statistically?

Yes. Qualitative data supports frequency counts, proportions and, for ordinal data, medians and rank-based tests. What it does not support is arithmetic that assumes equal spacing, such as means of nominal categories. The valid analysis depends on the measurement scale, not on whether the values are numbers.

If you want to go deeper into how recorded values turn into conclusions, the next step is understanding [what data analysis actually involves](/blog/data-analysis/what-is-data-analysis-definition) and how [structured data is organized](/blog/data-analysis/what-is-structured-data) before classification begins.

## References

1. [4.6.3.1. Background and Data](https://www.itl.nist.gov/div898/handbook/pmd/section6/pmd631.htm)
2. [Codd EF (1970). A relational model of data for large shared data banks. Communications of the ACM](https://doi.org/10.1145/362384.362685)
3. [6.5.1. What do we mean by "Normal" data?](https://www.itl.nist.gov/div898/handbook/pmc/section5/pmc51.htm)
4. [Altman DG, Bland JM (1996). Statistics Notes: Presentation of numerical data. BMJ](https://doi.org/10.1136/bmj.312.7030.572)
5. [Bland JM, Altman DG (1994). Statistics Notes: Correlation, regression, and repeated data. BMJ](https://doi.org/10.1136/bmj.308.6933.896)
6. [Bland JM, Altman DG (1996). Statistics Notes: Transforming data. BMJ](https://doi.org/10.1136/bmj.312.7033.770)
7. [Krzywinski M, Savig E (2013). Multidimensional data. Nature Methods](https://doi.org/10.1038/nmeth.2531)

## Related Articles

- [What Is Discrete Data? Definition and Examples](/blog/data-analysis/what-is-discrete-data)
- [Quantitative Data Examples: Definition and Types](/blog/data-analysis/quantitative-data-examples)
- [What Is Data Analysis? Definition, Steps and Examples](/blog/data-analysis/what-is-data-analysis-definition)
- [What Is Continuous Data? Definition and Examples](/blog/data-analysis/what-is-continuous-data)
- [What Is a Dataset? Definition, Types and Examples](/blog/data-analysis/what-is-a-dataset)
- [Hypothesis vs Theory vs Law in Science: Definitions and Examples](/blog/research-skills/hypothesis-vs-theory-vs-law)