RStudio for Beginners: Projects, Working Directory and Importing CSV and Excel Files

By Dr. Zubair Khalid, DVM, MS, PhD ·

RStudio for Beginners: Projects, Working Directory and Importing CSV and Excel Files

RStudio is a desktop environment many scientists use to write and run R code. It gives you a script editor, a console, a file browser and a plot viewer in one window, which matters when your analysis has to be rerun, checked or handed to someone else. The tool itself is free, and the same interface runs on Windows, macOS and Linux.

By the end of this article you will be able to create an RStudio project, understand what the working directory is and why it breaks file paths, and import CSV and Excel files with base R, readr and readxl. You will also run a complete plate-reader example from raw file to summary table, and know how to recognize the errors that come from a mismatched working directory.

Quick Answer

  • Create a project first: File > New Project, then New Directory, and name it after the analysis [2].
  • Put raw data in a data/ folder inside the project and never overwrite the raw file [10].
  • Write code in a script in the Source pane, not in the Console, so the analysis is saved [3].
  • Read CSV files with read_csv("data/plate.csv") from readr, or read.csv("data/plate.csv") in base R [4][9].
  • Read Excel files with read_excel() from readxl after install.packages("readxl"); readxl is not loaded by library(tidyverse) [6].
  • Use relative paths inside the project. Absolute paths like C:/Users/... will not work on another computer [3].

Step 1: Get Oriented in the RStudio Interface

RStudio opens with four main panes, plus an optional sidebar [1]. The Source pane sits top-left and holds your scripts. The Console sits bottom-left and runs code the moment you press Enter. The Environment pane sits top-right, with tabs for Environment, History, Connections, Build, VCS and Tutorial. The Output pane sits bottom-right, with tabs for Files, Plots, Packages, Help, Viewer and Presentation [1].

You can rearrange these under Tools > Global Options > Pane Layout, which also has an Add Source Column option [1]. Menu items and tab names shift between releases, so if a path looks different on your machine, check the current Posit RStudio User Guide. As of October 2026 the current release is RStudio 2026.09.0; check the download page for the latest version.

A few shortcuts are worth memorizing early. Ctrl+1 moves focus to the Source pane and Ctrl+2 moves it to the Console. Ctrl+L clears the Console, and Esc interrupts R when a command hangs [1].

The Console and the Source pane do different jobs. Code typed in the Console runs immediately but is not saved anywhere. A script in the Source pane is the saved record of the analysis, and R for Data Science recommends treating the script, not the saved workspace, as that record [3]. Cmd/Ctrl + Enter runs the current expression from a script, and Cmd/Ctrl + Shift + S runs the whole script [3]. Posit's code execution page also lists Ctrl+Shift+Enter for sourcing the whole document, so expect small differences between references.

One setting to change on day one: in Tools > Global Options, uncheck "Restore .RData into workspace at startup" and set "Save workspace to .RData on exit" to "Never" [3]. This forces you to keep everything that matters in the script, which is what makes an analysis reproducible later.

Step 2: Create an RStudio Project

An RStudio Project divides your work into multiple contexts, each with its own working directory, workspace, history and source documents [2]. That single feature removes most file-path problems before they start.

Go to File > New Project, or use the New Project toolbar button. You then choose New Directory, Existing Directory, or Version Control to clone a Git or Subversion repository [2]. For a new analysis, pick New Directory and name the directory after the project, for example plate-assay.

The project creates a file with the .Rproj extension. That file stores project options, and double-clicking it opens the project [2]. When a project opens, RStudio starts a new R session, sources .Rprofile if present, loads .RData and .Rhistory if present, sets the current working directory to the project directory, and restores any open source documents [2]. With the R4DS settings above, no .RData is saved, so nothing is restored from a previous session.

To switch between projects, use File > Open Project, the Projects menu at the top right of the RStudio window, or double-click the .Rproj file [2].

Inside the project, create a data/ folder for raw data and metadata, and an output/ folder for anything the script generates [10]. Keeping raw data separate from cleaned data is standard practice: save the raw file, do not overwrite it with a cleaned version, and consider making it read-only and backing it up in more than one location [10]. Scripted imports inside a project also satisfy the reproducibility rule that every result should be traceable to the code that produced it [11].

Step 3: Understand the Working Directory

The working directory is the folder R uses to resolve relative file names. getwd() shows the current working directory and setwd() changes it [3]. When you pass a relative name such as "data/plate.csv" to read.csv(), R resolves it against getwd(). That is why "cannot open file ... No such file or directory" usually means the working directory is not where you think it is [9].

R for Data Science recommends projects over setwd("/path/to/my/CoolProject"), because the project sets the working directory for you [3]. If you do use setwd(), keep it out of scripts you plan to share: the path is specific to your machine.

Inside a project, use relative paths such as "data/diamonds.csv". Absolute paths in scripts will not work on someone else's computer, and you should use forward slashes in paths even on Windows [3].

The here package builds paths from the project root, which it finds by markers such as a .Rproj file, a .here file or a .git directory [8]. Install it once and call it from any script:

install.packages("here")
here::here("data", "plate.csv")

The here::i_am() function declares a file's path within the project, which helps when scripts live in subfolders [8]. For most beginner projects, a plain relative path is enough, and here is a good upgrade once your folder structure grows.

Step 4: Import CSV Files

Base R reads comma-separated files with read.csv(). Its defaults are read.csv(file, header = TRUE, sep = ",", quote = "\"", dec = ".", fill = TRUE, comment.char = "") [9]. Two useful details: read.csv() no longer converts text columns to factors by default (stringsAsFactors = FALSE), and it has a fileEncoding argument so you can declare the file's encoding, for example fileEncoding = "latin1" [9]. read.csv2() is the variant that uses sep = ";" and dec = "," [9].

The readr package, loaded with the tidyverse, gives you read_csv(). It reads comma-separated files and prints the number of rows and columns plus the column types it guessed [4]. Defaults are col_names = TRUE, na = c("", "NA") and skip = 0, and locale controls encoding and regional settings [5]. read_csv() can also read .gz, .bz2, .xz and .zip files, and files at a URL [5]. read_csv2() uses ; as the separator and , as the decimal point, a format common in some European countries [5].

Real instrument exports rarely start with clean headers. skip = n drops the first n lines, which handles instrument metadata above the table. comment = "#" drops lines starting with #, col_names = FALSE treats the first row as data, and na = c("N/A", "") declares extra missing-value codes [4]. If a column imports with the wrong type, col_types overrides the guess with helpers such as col_double() and col_character(), and problems() shows the parsing failures [4].

For a quick check on summary statistics after import, the site's Mean, Median and Mode Calculator is a fast way to sanity-check what R reports.

Step 5: Import Excel Files

Excel files need the readxl package. Install it with install.packages("readxl") or install.packages("tidyverse"), then load it with library(readxl). readxl is not a core tidyverse package and is not loaded by library(tidyverse), so the explicit library(readxl) line is required [6].

readxl has no external dependency on Java or Perl and works on Windows, Mac and Linux, and it re-encodes non-ASCII text to UTF-8 [6]. Use excel_sheets(path) to list sheet names before reading, and readxl_example() to see the example files bundled with the package [6].

The main function is read_excel(path, sheet = NULL, range = NULL, col_names = TRUE, col_types = NULL, na = "", trim_ws = TRUE, skip = 0, ...). It detects .xls versus .xlsx automatically, while read_xls() and read_xlsx() force a format [7]. The sheet argument takes a sheet name as a string or a position as an integer, and defaults to the first sheet. The range argument takes cell ranges such as "B3:D87" or "Budget!B2:G14" [7].

library(readxl)
excel_sheets("data/plate.xlsx")
plate <- read_excel("data/plate.xlsx", sheet = "Run1")

One caution about saving results: write_csv() saves a data frame to CSV, but column type information is lost. RDS or Parquet keep it [4]. For a summary table that a collaborator will open in Excel, CSV is the right choice. For intermediate objects you plan to reload in R, RDS is safer.

Worked Example

This example runs a small plate-reader assay end to end. Start with File > New Project > New Directory and name it plate-assay. Create folders data/ and output/.

Save the following as data/plate.csv:

well,group,od600
A1,control,0.512
A2,control,0.498
A3,control,0.530
A4,control,0.505
B1,treated,0.341
B2,treated,0.362
B3,treated,0.329
B4,treated,0.355

Open a new script in the Source pane and run:

library(tidyverse)
plate <- read_csv("data/plate.csv")
plate_means <- plate |>
  group_by(group) |>
  summarise(n = n(), mean_od600 = mean(od600), sd_od600 = sd(od600))
plate_means
write_csv(plate_means, "output/plate_means.csv")

The base R alternative avoids the tidyverse:

plate <- read.csv("data/plate.csv")
aggregate(od600 ~ group, data = plate, FUN = mean)

Expected results: control has n = 4, mean 0.51125, sd 0.01374; treated has n = 4, mean 0.34675, sd 0.01471. A tibble prints these to about three significant digits, so you will see 0.511 and 0.347. The overall mean is mean(plate$od600) = 0.429, and the medians are 0.5085 and 0.348. These numbers were computed independently, so they are a reliable check on your import.

With the here package, the read line becomes read_csv(here::here("data", "plate.csv")) [8]. For the Excel version of the same file, use excel_sheets("data/plate.xlsx") to find the sheet name, then read_excel("data/plate.xlsx", sheet = "Run1") with whatever sheet name your file actually has. A European decimal-comma version of the same file, with ; separators and 0,512 style values, is read with read_csv2("data/plate_eu.csv") or read.csv2().

What the result means: the treated wells have a lower mean OD600 than the controls in this plate. What it does not mean: that the treatment caused the difference. Four wells per group on one plate cannot separate treatment effect from edge effects, pipetting order or plate-to-plate variation. The import step is correct; the experimental inference needs replication.

Common Mistakes and How to Fix Them

  • "cannot open file 'data/plate.csv': No such file or directory." The working directory is not the project folder, so the relative path resolves to the wrong place [9]. Fix: open the project through its .Rproj file and confirm getwd() points at the project directory.
  • The file opens on your laptop but not your labmate's. The script contains an absolute path such as C:/Users/you/.... Fix: convert to a relative path inside the project, and use forward slashes even on Windows [3].
  • could not find function "read_excel". readxl is not loaded by library(tidyverse) [6]. Fix: add library(readxl) after installing the package.
  • Numbers import as text, or a column of NA. Type guessing failed, often because of a stray header row or a thousands separator. Fix: use skip to drop metadata rows, col_types to set the type explicitly, and problems() to see what failed [4].
  • Accented characters turn into garbage. The file is not UTF-8. Fix: declare the encoding, for example fileEncoding = "latin1" in base R [9], or set locale = locale(encoding = "latin1") in readr. Check the current readr documentation for the exact argument form.
  • The Console has a wall of old output. Fix: press Ctrl+L to clear it [1].
  • A command seems to hang. Press Esc to interrupt R [1].
  • The analysis cannot be reproduced next month. Everything was typed in the Console and never saved. Fix: move the code into a script in the Source pane and rerun it from the top [3].

Limitations

RStudio is an interface, not an analysis engine. It cannot fix a malformed CSV, and it will not warn you that a column of numbers arrived as text unless you look at the column types that read_csv() prints [4].

Menu paths, pane tabs and shortcut assignments change between releases. The layout described here was checked against the Posit RStudio User Guide in October 2026, with RStudio 2026.09.0 current at that time. Verify anything version-specific against the current documentation. Prefer projects, or an explicit setwd() call, to menu-based ways of changing the working directory.

Projects solve path problems, not scientific ones. A project does not validate your data, does not record instrument settings, and does not tell you whether a plate was read correctly. It also does not preserve column types across a CSV round trip; use RDS or Parquet when types matter [4].

Encoding remains a common source of quiet errors. readr reads UTF-8 by default through locale(), and other encodings need an explicit setting. If your file came from an older Windows instrument, assume nothing about its encoding until you check the imported values.

Frequently Asked Questions

What is the fastest way to learn how to use RStudio?

Create a project, write one script, and run it from top to bottom with Cmd/Ctrl + Shift + S. That loop (project, script, rerun) teaches the parts of RStudio that matter for research. Pane layout and shortcuts come later.

Do I need RStudio projects if I already use setwd()?

Projects are the better default. R for Data Science recommends them over setwd("/path/to/my/CoolProject") because the project sets the working directory for you [3]. setwd() still works, but a hard-coded path in a shared script will fail on another machine.

How do I change the working directory in R?

Use setwd() to change it and getwd() to check it [3]. In practice, open the project instead and let RStudio set the directory when the project loads [2]. If you must call setwd(), keep it out of scripts you plan to share.

How do I import Excel data into RStudio?

Install readxl, load it with library(readxl), list sheets with excel_sheets(path), then read with read_excel(path, sheet = "Run1") [6][7]. readxl handles both .xls and .xlsx and needs no Java or Perl [6].

Which is better for reading CSV in R, read.csv or read_csv?

read.csv() is base R and needs no packages, while read_csv() from readr prints the column specification, and reads compressed files and URLs [4][5][9]. For a one-off file either works. For a script you will rerun, readr's printed column types make problems easier to spot.

References

  1. Posit: RStudio User Guide, Pane layout
  2. Posit: RStudio User Guide, Projects
  3. R for Data Science (2e), Workflow: scripts and projects
  4. R for Data Science (2e), Data import
  5. readr: read_delim / read_csv reference
  6. readxl package home
  7. readxl: read_excel reference
  8. here package documentation
  9. R documentation: read.table / read.csv {utils}
  10. Wilson G, et al. Good enough practices in scientific computing. PLoS Comput Biol. 2017;13(6):e1005510
  11. Sandve GK, et al. Ten Simple Rules for Reproducible Computational Research. PLoS Comput Biol. 2013;9(10):e1003285

Related Articles