# ggplot2 Tutorial for Beginners: Publication-Ready Plots for Lab Data

ggplot2 is the plotting package that many R users reach for when a figure has to survive peer review. It implements the grammar of graphics: you supply a data frame, map columns to visual properties such as x, y and color, choose geometric objects such as points or boxes, and the package handles the drawing [1]. That structure is why the same three or four lines of code can produce a scatter plot, a grouped boxplot or a faceted panel figure with only small edits.

By the end of this tutorial you will be able to install and load ggplot2, build scatter plots and boxplots from lab data, add error bars that match your summary statistics, apply colorblind-safe palettes and clean themes, split a figure into panels, and export a 300 dpi image at an exact size. The worked example uses `ToothGrowth`, a dataset built into R, so you can run every line without downloading anything.

## Quick Answer

- Install once with `install.packages("tidyverse")` for the full tidyverse or `install.packages("ggplot2")` for ggplot2 alone, then load it each session with `library(ggplot2)` [1].
- Every plot starts with `ggplot(data, aes(x = ..., y = ...))` and gains layers with `+`, for example `ggplot(mpg, aes(displ, hwy, colour = class)) + geom_point()` [1].
- Use `geom_point()` for scatter plots, `geom_boxplot()` for boxplots, and `geom_jitter()` to show individual points on top of boxes [3].
- Add error bars with `stat_summary(fun.data = mean_se, geom = "errorbar")`, or compute means and SDs first and pass `ymin`/`ymax` to `geom_errorbar()` [4].
- Set titles and axis labels with `labs()`, and switch the look with `theme_classic()` or `theme_bw()` [2][5].
- Save with `ggsave("figure.png", p, width = 6, height = 4, units = "in", dpi = 300)`, which gives a 1800 x 1200 pixel image [7].

## Step 1: Install and Load ggplot2

Install the package once per machine. The ggplot2 site documents two options: the whole tidyverse, or ggplot2 alone [1]. The tidyverse also brings dplyr, tidyr and readr, which you will likely want anyway.

```r
install.packages("tidyverse")   # whole tidyverse
install.packages("ggplot2")     # ggplot2 only
```

The development version installs from GitHub with `pak::pak("tidyverse/ggplot2")` [1]. The ggplot2 site showed version 4.0.3 as of October 2026, so check the site for the current release before you rely on a specific argument or message [1].

Loading is a per-session step. Loaded packages do not persist between R sessions, so `library()` belongs at the top of every script [1][2].

```r
library(ggplot2)
```

On Windows and macOS the install step may prompt you to choose a CRAN mirror or to compile from source. If a source install fails, the usual cause is a missing compiler toolchain, not a problem with ggplot2 itself.

## Step 2: Understand the Grammar Before You Type

A ggplot2 call has three parts. The `data` argument names the data frame. The `mapping` argument, always wrapped in `aes()`, says which columns become which visual properties. The `+` operator adds layers and settings [2].

Mappings set inside `ggplot()` are global: they apply to every layer. Mappings set inside a geom are local to that layer only. This distinction matters. In the R for Data Science example, a local color mapping on `geom_point()` produces a single overall trend line from `geom_smooth()`, because the smooth layer never sees the color grouping [2].

```r
library(palmerpenguins)   # install.packages("palmerpenguins") first; R4DS uses this copy of the data
ggplot(data = penguins, mapping = aes(x = flipper_length_mm, y = body_mass_g)) +
  geom_point()
```

That is the whole pattern. Everything else in this tutorial is a variation on it.

## Step 3: Build a Scatter Plot

Scatter plots show the relationship between two continuous variables and every individual observation. For small lab samples they are usually a better default than a bar chart. Weissgerber and colleagues reviewed 703 articles in top physiology journals and found that 85.6% included at least one bar graph, while only 13.4% included univariate scatterplots [8]. With median sample sizes of about 3 to 6 per group, they note that many different data distributions can produce the same bar graph, and recommend univariate scatterplots instead [8].

```r
ggplot(ToothGrowth, aes(x = dose, y = len, colour = supp)) +
  geom_point() +
  geom_smooth(method = "lm") +
  theme_classic()
```

`geom_smooth(method = "lm")` adds a linear best-fit line [2]. Be careful with this on `ToothGrowth`: dose takes the values 0.5, 1 and 2, so the relationship is not strictly linear on that scale, and log2(dose) is often used instead. The line is a summary of the fitted model, not evidence that the underlying biology is linear.

If you want a quick look at your own numbers before writing R code, the site's [Box Plot Maker](/tools/box-plot-maker) will render a summary figure from pasted data.

## Step 4: Build a Boxplot with Individual Points

`geom_boxplot()` draws the median as a line, the lower and upper hinges as the first and third quartiles (25th and 75th percentiles), and whiskers that extend to the most extreme value no further than 1.5 times the IQR from the hinge. Points beyond the whiskers are plotted individually as outliers [3].

For n = 10 per group, showing the raw points alongside the box is more informative than the box alone. The documented approach suppresses the boxplot's own outliers so they are not drawn twice [3].

```r
tg <- ToothGrowth
tg$dose <- factor(tg$dose)

ggplot(tg, aes(x = dose, y = len, colour = supp)) +
  geom_boxplot(outlier.shape = NA) +
  geom_jitter(width = 0.2) +
  theme_classic()
```

Converting `dose` to a factor matters. If dose stays numeric, ggplot2 treats it as a continuous axis and will not draw a separate box for each dose. As a factor, each dose becomes a category.

Notched boxes are an option when you want a visual cue about medians. Notches extend 1.58 times the IQR divided by the square root of n, and if the notches of two boxes do not overlap, that suggests the medians differ [3]. It is a suggestion, not a hypothesis test.

## Step 5: Add Error Bars

There are two clean routes. The first uses `stat_summary()`, which defaults to `mean_se()`, meaning mean plus or minus one standard error, and draws a pointrange by default [4].

```r
ggplot(tg, aes(x = dose, y = len, colour = supp)) +
  geom_jitter(width = 0.15, alpha = 0.5) +
  stat_summary(fun.data = mean_se, geom = "errorbar", width = 0.2) +
  stat_summary(fun = mean, geom = "point", size = 3) +
  theme_bw()
```

`mean_se()` is built into ggplot2. The alternatives `mean_cl_boot()`, `mean_cl_normal()`, `mean_sdl()` and `median_hilow()` are wrappers around functions from the Hmisc package, which you must install first [4].

The second route computes the summary yourself and passes explicit limits to `geom_errorbar()`. This is the one to use when you need standard deviation, or when you want the summary table for a report anyway.

```r
summ <- aggregate(len ~ supp + dose, data = tg,
                  FUN = function(x) c(mean = mean(x), sd = sd(x)))
summ <- do.call(data.frame, summ)

ggplot(summ, aes(x = dose, y = len.mean, colour = supp)) +
  geom_point(size = 3) +
  geom_errorbar(aes(ymin = len.mean - len.sd, ymax = len.mean + len.sd), width = 0.2) +
  theme_classic()
```

One warning from the documentation: do not use `ylim` to zoom into a summary plot, because that throws the data away. Use `coord_cartesian()` to zoom instead [4]. The difference is that `ylim` drops rows before summarizing, while `coord_cartesian()` only changes the visible window.

## Step 6: Colors, Themes and Facets

The default ggplot2 theme is `theme_grey()`, with a grey background and white gridlines. `theme_bw()` is the classic dark-on-light option, `theme_classic()` has axis lines and no gridlines, and `theme_minimal()` drops background annotations entirely [5]. All complete themes accept `base_size` (default 11 pt) and `base_family`, so `theme_classic(base_size = 14)` enlarges every text element at once [5].

For color, the viridis scales are perceptually uniform in color and in black and white, and are designed to be perceived by viewers with common forms of color blindness [6]. Use `scale_colour_viridis_d()` for discrete variables and `scale_fill_viridis_d()` for fills. The `begin` and `end` arguments default to 0 and 1, and `direction = -1` reverses the order [6]. Viridis is already the default color scale for ordered factors [6].

R for Data Science makes a related point: representing information using only colors is generally a bad idea because of color blindness, so map a variable to both color and shape where you can [2]. The `ggthemes` package provides `scale_color_colorblind()` as another safe palette [2].

Facets split one plot into panels by a categorical variable. `facet_wrap(~island)` is the formula form, and `facet_wrap(vars(class))` is equivalent. The `scales` argument controls whether axes are shared: `"fixed"` is the default, and `"free"`, `"free_x"` and `"free_y"` let panels differ. Free scales make patterns visible but make visual comparison across panels harder, so use them deliberately.

```r
p1 <- ggplot(tg, aes(x = dose, y = len, colour = supp)) +
  geom_boxplot(outlier.shape = NA) +
  geom_jitter(width = 0.2) +
  facet_wrap(~supp) +
  scale_colour_viridis_d(end = 0.8) +
  labs(x = "Vitamin C dose (mg/day)", y = "Tooth length", colour = "Supplement") +
  theme_classic(base_size = 12)
```

`labs()` sets the title, subtitle, axis labels and legend titles, including `color = "Species"` and `shape = "Species"` style legend names [2]. Labeling the legend with a readable name instead of the raw column name is one of the cheapest improvements you can make to a figure.

## Step 7: Save at the Right Size

`ggsave()` picks the output device from the file extension, supporting eps, ps, tex, pdf, jpeg, tiff, png, bmp, svg and wmf (Windows only) [7]. It saves the last displayed plot by default, or the plot you pass explicitly [7].

```r
ggsave("toothgrowth_boxplot.png", p1, width = 6, height = 4, units = "in", dpi = 300)
```

If you omit `width` and `height`, ggsave uses the size of the current graphics device, which means the saved file can differ between computers and between RStudio sessions [7][2]. Always set them. The `dpi` argument defaults to 300 and also accepts `"retina"` (320), `"print"` (300) or `"screen"` (72); it matters for raster formats such as png and tiff [7]. A 6 x 4 inch figure at 300 dpi is 1800 x 1200 pixels. `limitsize = TRUE` blocks images larger than 50 x 50 inches [7].

For vector output, save as pdf or svg and dpi becomes irrelevant. Journals differ on required format, width in millimeters and minimum dpi, so check the author guidelines for your target journal before you finalize.

## Worked Example

This example uses `ToothGrowth`, which is built into R. It has 60 rows and columns `len`, `supp` (OJ or VC) and `dose` (0.5, 1, 2), with 10 observations per supplement by dose group.

```r
library(ggplot2)
tg <- ToothGrowth
tg$dose <- factor(tg$dose)   # treat dose as categories for boxplots
```

**1) Boxplot + individual points (recommended for n = 10 per group)**

```r
p1 <- ggplot(tg, aes(x = dose, y = len, colour = supp)) +
  geom_boxplot(outlier.shape = NA) +
  geom_jitter(width = 0.2) +
  facet_wrap(~supp) +
  scale_colour_viridis_d(end = 0.8) +
  labs(x = "Vitamin C dose (mg/day)", y = "Tooth length", colour = "Supplement") +
  theme_classic(base_size = 12)
ggsave("toothgrowth_boxplot.png", p1, width = 6, height = 4, units = "in", dpi = 300)   # 1800 x 1200 px
```

**2) Mean +/- SE drawn by stat_summary (mean_se is built in)**

```r
p2 <- ggplot(tg, aes(x = dose, y = len, colour = supp)) +
  geom_jitter(width = 0.15, alpha = 0.5) +
  stat_summary(fun.data = mean_se, geom = "errorbar", width = 0.2) +
  stat_summary(fun = mean, geom = "point", size = 3) +
  facet_wrap(~supp) + theme_bw()
```

**3) Mean +/- SD computed first, then geom_errorbar**

```r
summ <- aggregate(len ~ supp + dose, data = tg, FUN = function(x) c(mean = mean(x), sd = sd(x)))
summ <- do.call(data.frame, summ)   # columns len.mean, len.sd
p3 <- ggplot(summ, aes(x = dose, y = len.mean, colour = supp)) +
  geom_point(size = 3) +
  geom_errorbar(aes(ymin = len.mean - len.sd, ymax = len.mean + len.sd), width = 0.2) +
  facet_wrap(~supp) + theme_classic()
```

**4) Scatter + linear fit using numeric dose**

```r
p4 <- ggplot(ToothGrowth, aes(x = dose, y = len, colour = supp)) +
  geom_point() + geom_smooth(method = "lm") + theme_classic()
```

The numbers behind those plots: group means (SD, SE) are OJ 0.5: 13.23 (4.46, 1.41); OJ 1: 22.70 (3.91, 1.24); OJ 2: 26.06 (2.66, 0.84); VC 0.5: 7.98 (2.75, 0.87); VC 1: 16.77 (2.52, 0.80); VC 2: 26.14 (4.80, 1.52). Medians are OJ 12.25, 23.45, 25.95 and VC 7.15, 16.50, 25.95. Mean plus or minus SD limits run from 8.77 to 17.69 for OJ 0.5 and 21.34 to 30.94 for VC 2, while mean plus or minus SE limits are much tighter, 11.82 to 14.64 and 24.62 to 27.66 for the same groups. The overall length range is 4.2 to 33.9.

The linear fits are OJ: len = 11.55 + 7.81 * dose, and VC: len = 3.30 + 11.72 * dose. Read those as descriptions of the fitted lines, not as proof that the dose response is linear. The gap between the two supplements is clear at 0.5 and 1 mg/day and nearly gone at 2 mg/day, which is a pattern the boxplots show directly and a single overall trend line would hide.

## Common Mistakes and How to Fix Them

- **"object 'len' not found" or a blank plot.** The column name in `aes()` does not match the data frame. Check spelling and case, and confirm the data frame you named is the one you loaded.
- **You get one box per colour instead of one per dose.** `dose` or another grouping column is numeric. Convert it with `factor()` before plotting.
- **Outliers appear twice.** You added `geom_jitter()` without `outlier.shape = NA` in `geom_boxplot()`. Suppress the boxplot outliers so only the jittered points show [3].
- **The trend line ignores your color groups.** The color mapping is local to `geom_point()`. Move it into `ggplot()` so it applies globally [2].
- **Error bars vanish or look wrong after zooming.** `ylim` drops rows before summarizing. Use `coord_cartesian()` instead [4].
- **`mean_sdl()` or `mean_cl_boot()` fails.** Those wrappers need the Hmisc package installed. `mean_se()` does not [4].
- **The saved file is a different size on a colleague's machine.** `width` and `height` were omitted, so ggsave used the local graphics device [7][2].
- **Colors are hard to tell apart in print.** Switch to a viridis scale or map the variable to shape as well as color [2][6].

## Limitations

ggplot2 draws figures. It does not run statistics, correct for multiple comparisons or tell you whether a difference is real. A non-overlapping notch or a non-overlapping error bar is a visual cue, not a test result.

Version 4.0.3 was current as of October 2026, and function arguments and deprecation messages can change between releases [1]. The documented spelling is `scale_colour_viridis_d()` [6]; ggplot2 generally accepts the American `scale_color_*` spelling, but confirm against the current reference if a call fails.

The R code in the worked example was not executed in R for this article; the numbers were computed independently from the Rdatasets CSV copy of `ToothGrowth`, which matches the 60-row built-in dataset. Run the code yourself and compare against the values above before you reuse the figure.

`ggsave()` output depends on the graphics device when width and height are omitted, and `dpi` only affects raster formats [7]. Journal requirements for dpi, physical width and file format vary, so check the author guidelines for your target journal.

## Frequently Asked Questions

### How do I use ggplot2 in R for the first time?

Install with `install.packages("ggplot2")`, load with `library(ggplot2)`, then build a plot from `ggplot(data, aes(x = ..., y = ...))` plus one geom [1]. Start with `geom_point()` on a small dataset so you can see what each layer adds.

### What is the difference between a ggplot2 boxplot and a ggplot2 scatter plot?

A scatter plot draws one point per observation and shows the full distribution and any relationship between two continuous variables. A boxplot summarizes a distribution as median, quartiles and whiskers [3], which hides sample size and shape. For small lab samples, overlay jittered points on the box.

### How do I add ggplot2 error bars?

Use `stat_summary(fun.data = mean_se, geom = "errorbar")` for mean plus or minus standard error, since `mean_se()` is built in [4]. For standard deviation or a custom interval, compute the summary first and pass `ymin` and `ymax` to `geom_errorbar()`.

### How do I choose ggplot2 colors that are safe for colorblind readers?

Use `scale_colour_viridis_d()` for discrete variables and `scale_fill_viridis_d()` for fills; the viridis scales are perceptually uniform and designed for viewers with common forms of color blindness [6]. Mapping a variable to shape as well as color adds a second channel of information [2].

### Which ggplot2 theme should I use for a journal figure?

`theme_classic()` gives axis lines and no gridlines, and `theme_bw()` gives a dark-on-light look; both accept `base_size` to scale all text at once [5]. Pick one, apply it consistently across every figure in the paper, and set the size in `ggsave()` to match the journal's column width.

## References

1. [ggplot2: Create Elegant Data Visualisations Using the Grammar of Graphics (tidyverse)](https://ggplot2.tidyverse.org/)
2. [Wickham, Cetinkaya-Rundel, Grolemund. R for Data Science (2e): Data visualization](https://r4ds.hadley.nz/data-visualize)
3. [ggplot2 reference: geom_boxplot()](https://ggplot2.tidyverse.org/reference/geom_boxplot.html)
4. [ggplot2 reference: stat_summary()](https://ggplot2.tidyverse.org/reference/stat_summary.html)
5. [ggplot2 reference: Complete themes (theme_bw, theme_classic)](https://ggplot2.tidyverse.org/reference/ggtheme.html)
6. [ggplot2 reference: Viridis colour scales](https://ggplot2.tidyverse.org/reference/scale_viridis.html)
7. [ggplot2 reference: ggsave()](https://ggplot2.tidyverse.org/reference/ggsave.html)
8. [Weissgerber TL et al. Beyond Bar and Line Graphs: Time for a New Data Presentation Paradigm. PLOS Biol. 2015;13(4):e1002128](https://doi.org/10.1371/journal.pbio.1002128)
9. [Rougier NP, Droettboom M, Bourne PE. Ten Simple Rules for Better Figures. PLoS Comput Biol. 2014;10(9):e1003833](https://doi.org/10.1371/journal.pcbi.1003833)

## Related Articles

- [How to Use ggplot2 for RNA-seq Visualization: A Practical Guide for Custom Publication-Ready Plots](/knowledge/bioinformatics/how-to-use-ggplot2-for-rna-seq-visualization-a-practical-guide-for-custom-publication-ready-plots)
- [How to Make and Read a Box Plot: Quartiles, Whiskers and Outliers](/blog/research-skills/how-to-make-a-box-plot)
- [How to Make a Scatter Plot](/blog/research-skills/how-to-make-a-scatter-plot)
- [How to Present Error Bars and Uncertainty in Lab Report Figures](/blog/research-skills/how-to-present-error-bars-and-uncertainty-in-lab-report-figures-a-guide-for-students)
- [RStudio for Beginners: Projects, Working Directory and Importing CSV and Excel Files](/blog/research-skills/how-to-start-an-rstudio-project-and-import-data)
- [How to Install R Packages from CRAN, Bioconductor and GitHub (and Fix Install Errors)](/blog/research-skills/how-to-install-r-packages-cran-bioconductor-github)