# Data Visualization Basics: How to Choose the Right Chart

Data visualization is the practice of representing data as images so patterns, outliers, and trends become easier to see. Choosing the right chart starts with two questions: what kind of data do you have, and what do you want the reader to learn? This guide walks through the principles, a worked example, and a decision table you can reuse.

## Quick Answer

- Match the chart to the question first. Trends over time need a line chart, comparisons across categories need a bar chart, and relationships between two numeric variables need a scatter plot [1].
- Match the chart to the data type second. Categorical data suits bars and dots, continuous data suits lines, histograms, and scatter plots, and geographic data suits maps [2].
- Start by thinking about your data and the story you will be telling, and confirm you have a clear purpose and reliable data before you draw anything [1].
- Visual features such as position, size, shape, hue, and saturation encode the values, so pick the feature that people read most accurately for your main message [2].
- Keep one idea per chart. If you need to show a trend, a total, and a correlation, use three panels instead of one crowded figure.

## What Data Visualization Means

In plain terms, data visualization is the graphical representation of data, used to enhance understanding and communication of data with a wide audience [1]. Common examples include charts, graphs, timelines, maps, and even tables [2].

The precise definition is broader. Data visualization is an umbrella term referring to the representation of raw data content in a graphic or pictorial manner, using visual features of encoded data such as position, size, shape, hue and saturation, and motion [2]. Related concepts include information visualization and scientific visualization [2].

Two purposes drive almost every chart you will make. Exploration, sometimes called sense-making or data analysis, uses visualization to examine and discover trends, associations, and outliers that may not be obvious in the raw data. Explanation, also called storytelling or communication, creates visualizations that make what the data is saying clearer and more obvious [2]. A chart built for exploration can look messy on purpose. A chart built for explanation should be clean.

## How It Works

Chart choice is a mapping problem. You have variables, and you have visual channels. The job is to assign each variable to a channel that readers decode accurately.

For a single continuous variable measured over time, the mechanism is a line chart, and a least-squares trend line through the points estimates the average rate of change:

$$
\text{slope} = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{\sum_{i=1}^{n}(x_i - \bar{x})^2}
$$

Here $x_i$ is the time index, $y_i$ is the measured value, $\bar{x}$ and $\bar{y}$ are the means, and $n$ is the number of periods. The slope tells you the average change per period.

For spread, the sample standard deviation summarizes how far values sit from their mean:

$$
s = \sqrt{\frac{\sum_{i=1}^{n}(x_i - \bar{x})^2}{n-1}}
$$

Here $s$ is the sample standard deviation, $x_i$ is each observation, $\bar{x}$ is the mean, and $n-1$ is the degrees of freedom used for a sample.

For a relationship between two numeric variables, the Pearson correlation coefficient summarizes the strength of a linear association:

$$
r = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_{i=1}^{n}(x_i - \bar{x})^2}\sqrt{\sum_{i=1}^{n}(y_i - \bar{y})^2}}
$$

Here $r$ ranges from -1 to 1, with values near 0 meaning little linear association. A scatter plot is the visual counterpart of $r$, and you should always plot the points before trusting the number.

## Worked Example

The dataset is monthly revenue in thousands of dollars for three regions over 12 months.

| month | region | revenue |
|---|---|---|
| Jan | North | 120 |
| Feb | North | 132 |
| Mar | North | 145 |
| Apr | North | 158 |
| May | North | 171 |
| Jun | North | 190 |
| Jul | North | 205 |
| Aug | North | 220 |
| Sep | North | 238 |
| Oct | North | 255 |
| Nov | North | 270 |
| Dec | North | 290 |
| Jan | South | 90 |
| Feb | South | 95 |
| Mar | South | 102 |
| Apr | South | 110 |
| May | South | 118 |
| Jun | South | 125 |
| Jul | South | 133 |
| Aug | South | 140 |
| Sep | South | 148 |
| Oct | South | 155 |
| Nov | South | 163 |
| Dec | South | 172 |
| Jan | West | 60 |
| Feb | West | 72 |
| Mar | West | 68 |
| Apr | West | 85 |
| May | West | 95 |
| Jun | West | 110 |
| Jul | West | 105 |
| Aug | West | 125 |
| Sep | West | 140 |
| Oct | West | 150 |
| Nov | West | 165 |
| Dec | West | 180 |

Step 1, the North total: 120+132+...+290 = 2394.

Step 2, the North mean: 2394 / 12 = 199.5000.

Step 3, the North sample standard deviation: sqrt(sum((x-199.5000)^2)/11) = 56.1208.

Step 4, the North linear trend slope: a least-squares fit over months 1..12 gives 15.5385 per month.

Step 5, the North versus West correlation: r = 0.9918.

Step 6, region totals: North=2394, South=1551, West=1355.

```python
import pandas as pd, numpy as np
df = pd.DataFrame({'month':[...],'region':[...],'revenue':[...]})
df.groupby('region')['revenue'].sum()
np.polyfit(range(1,13), df[df.region=='North'].revenue, 1)
```

Results (the snippet is a sketch: fill the lists from the table above and print each value): North total=2394, mean=199.5000, SD=56.1208, slope=15.5385, r(N,W)=0.9918.

Three panels from the same 12-month sales data: a line chart shows North rising from 120 to 290 ($k, slope 15.54/mo), a bar chart compares region totals (North 2394, South 1551, West 1355), and a scatter plot shows North versus West with r = 0.99.

Notice that one dataset produced three different correct charts. The line chart answers "how is North changing." The bar chart answers "which region earned the most." The scatter plot answers "do the two regions move together." The chart follows the question, not the table.

## How to Interpret It

Read the axis labels and units before you read the shape. A line that looks steep can be a small change on a narrow scale, and a line that looks flat can hide a large change on a wide one. The creator decides on axes, labels and scales that can exaggerate or minimize differences [3].

Check whether the chart shows a total, an average, or individual points. A bar of 2394 and a mean of 199.5 describe the same region but answer different questions. If you are comparing groups, a bar chart of totals works. If you are comparing typical performance, a bar chart of means works better, and you should mention the spread (SD = 56.1208 for North) so readers know how much variation sits behind the average.

For scatter plots, read the correlation as a description, not a cause. An r of 0.9918 between North and West means the two series move together almost linearly across these 12 months. It does not mean one region's revenue caused the other's.

## When to Use It (and when not to)

Use a line chart when the x-axis is time or another ordered sequence and you want to show change. Use a bar chart when you compare a numeric value across a small number of categories. Use a scatter plot when you have two numeric variables per observation and want to show association. Use a histogram when you want the distribution of one continuous variable. Use a map when location is the point of the analysis [2].

Do not use a pie chart for many categories or for values that are close in size, because small angle differences are hard to judge. Do not use a line chart for unordered categories, since the connecting line implies a sequence that does not exist. Do not use a 3D chart for a two-variable comparison, because the extra dimension adds no information and distorts the values. If you are unsure whether your figure is a chart or a graph in the strict sense, the distinction is mostly about axes and coordinate systems, which is covered in [Chart vs Graph: Differences and When to Use Each](/blog/data-analysis/chart-vs-graph-differences).

## Chart Choice vs Chart Type

These two ideas get mixed up. Chart choice is the decision process. Chart type is the resulting form. The table below separates them.

| Aspect | Chart choice | Chart type |
|---|---|---|
| What it is | A decision about which visual fits the question and data | The named form you end up drawing |
| Driven by | The question, the data type, the audience | The decision you made |
| Example | "I need to show change over 12 months" | Line chart |
| Failure mode | Picking a form before stating the question | Using a valid type for the wrong question |
| Fix | Write the question in one sentence first | Re-check the type against the question |

A related decision is whether a table would serve better than a figure. Tables preserve exact values, and figures reveal shape. The trade-offs are compared in [Data Table vs Graph: Differences and When to Use Each](/blog/data-analysis/data-table-vs-graph).

## Common Mistakes

- Choosing the chart before stating the question. Fix: write "what do I want my audience to learn" in one sentence, then pick the form [1].
- Skipping data cleaning. Fix: check for errors, duplicates, and missing values before plotting, because clean data ensures accuracy and prevents misleading visuals [1].
- Using a truncated y-axis on a bar chart. Fix: start bar charts at zero, since bar length encodes the value.
- Overloading one figure. Fix: split a trend, a total, and a relationship into separate panels, as in the worked example.
- Ignoring the audience. Fix: build exploratory charts for yourself and simplified explanatory charts for others [2].
- Leaving out the source and collection method. Fix: be transparent about where the data comes from and who collected it [4].

## Limitations

No chart type is neutral. Data is not neutral, and creators have a responsibility not to create misleading charts and graphs [4]. The same numbers can be framed to emphasize growth or to hide a weak month, and the choice of baseline, bin width, or color scale can shift what a reader concludes.

Charts also compress information. A single line through 12 points hides the individual values, and a bar of a total hides the distribution behind it. When the exact numbers matter, pair the figure with a table or report the summary statistics alongside it. Visualization is a means to an end, not an end in itself [2].

## Frequently Asked Questions

### What is the best chart for time series data?

A line chart is the default for time series because the connected line shows direction and rate of change. Use points on the line when the number of periods is small, and use a bar chart instead when the periods are few and discrete, such as quarterly totals. If you have multiple series, keep them on one chart only when the scales are comparable.

### How many categories can a bar chart show?

Bar charts stay readable with roughly five to ten categories. Beyond that, readers struggle to compare lengths and labels start to collide. For long category lists, sort by value, consider a horizontal bar layout, or group small categories into an "other" bucket. If the categories are parts of a whole, check [Pie Chart Examples: When They Work and When They Mislead](/blog/data-analysis/pie-chart-examples-when-to-use) before defaulting to a pie.

### Should I use a bar chart or a line chart for monthly revenue?

Use a line chart when the message is the trend, as with North rising from 120 to 290 with a slope of 15.5385 per month. Use a bar chart when the message is the comparison, as with region totals of North 2394, South 1551, and West 1355. The data is the same. The question decides.

### What does a correlation of 0.99 look like on a scatter plot?

It looks like points sitting close to a straight line with a positive tilt. In the example, North versus West gives r = 0.9918, so the points cluster tightly along an upward line. Always plot the points, because a high r can still hide outliers or a curved relationship that the single number does not capture.

### How do I choose a chart for survey data?

Start by identifying each variable's type. Counts of categories suit bar charts, and Likert-style responses suit stacked or segmented bars. If you are comparing two categorical variables at once, a segmented bar chart shows the composition of each group, which is explained in [Segmented Bar Chart: Definition and Examples](/blog/data-analysis/segmented-bar-chart). For mixed variable types, review [Structured vs Unstructured Data: Differences and Examples](/blog/data-analysis/structured-vs-unstructured-data) first, then map each variable to a channel.

For a broader decision framework across research settings, see [How to Choose the Right Chart Type for Your Research Data](/blog/research-skills/how-to-choose-the-right-chart-type-for-your-research-data-a-decision-tree-for-qpcr-rna-seq-and-beyon), and for honesty checks before publishing, see [How to Design Clear and Honest Data Visualizations for Research Slides](/blog/research-skills/how-to-design-clear-and-honest-data-visualizations-for-research-slides).

## References

1. [Data Visualization - BIOL 205: Biostatistics - Research Guides at Mount St. Mary's University](https://libguides.msmary.edu/biol205/DataViz)
2. [Data Vis Foundations - Data Visualization Foundations - Library Guides at University of South Carolina](https://guides.library.sc.edu/datavizfoundations)
3. [Reading data visualizations - Data Literacy - LibGuides at Chapman University](https://libguides.chapman.edu/data_literacy/reading_viz)
4. [Data Visualization - Digital Scholarship - Research Guides at University of North Dakota](https://libguides.und.edu/digitalscholarship/data-visualization)

## Further Reading

- [Data Visualization - LAS 384 Proseminar- Digital Humanities - LibGuides at University of Texas at Austin](https://guides.lib.utexas.edu/las-384-Proseminar/data-visualization)
- [Cleveland WS, McGill R (1984). Graphical Perception: Theory, Experimentation, and Application to the Development of Graphical Methods. Journal of the American Statistical Association](https://doi.org/10.1080/01621459.1984.10478080)
- [Altman DG, Bland JM (1996). Statistics Notes: Presentation of numerical data. BMJ](https://doi.org/10.1136/bmj.312.7030.572)

## Related Articles

- [Chart vs Graph: Differences and When to Use Each](/blog/data-analysis/chart-vs-graph-differences)
- [Data Table vs Graph: Differences and When to Use Each](/blog/data-analysis/data-table-vs-graph)
- [Segmented Bar Chart: Definition and Examples](/blog/data-analysis/segmented-bar-chart)
- [Data Analytics Methods: Types and When to Use Each](/blog/data-analysis/data-analytics-methods)
- [Pie Chart Examples: When They Work and When They Mislead](/blog/data-analysis/pie-chart-examples-when-to-use)
- [How to Design Clear and Honest Data Visualizations for Research Slides](/blog/research-skills/how-to-design-clear-and-honest-data-visualizations-for-research-slides)
- [Data Visualization Best Practices for Computational Biologists](/blog/careers/data-visualization-best-practices-for-computational-biologists-choosing-the-right-plot-for-your-data)
- [How to Choose the Right Chart Type for Your Research Data](/blog/research-skills/how-to-choose-the-right-chart-type-for-your-research-data-a-decision-tree-for-qpcr-rna-seq-and-beyon)