How to Read CSV in R with read.csv (Step by Step)

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Read CSV in R with read.csv (Step by Step)

To read a CSV in R, call read.csv() with the file path as the first argument. The function returns a data frame, one row per line of the file and one column per field. This guide walks through the whole process, from finding your file to verifying that the data frame looks right.

Quick Answer

  • read.csv("path/to/file.csv") imports the file and returns a data frame [1].
  • header = TRUE is the default for read.csv, so the first row is treated as column names [2].
  • The path is relative to your working directory unless you give an absolute path. Check it with getwd() [1].
  • Use stringsAsFactors = FALSE if you want text columns kept as character instead of factors [1].
  • Always inspect the result with str(), head(), and dim() before you analyze anything.

Before You Start

You need two things: a CSV file and a known working directory.

A CSV file stores a table as plain text. Each line is a row, and commas separate the fields. Excel, Google Sheets, and most database exports can produce one. In Excel you save a worksheet with File > Save As and pick the CSV option, which gives the file a .csv extension instead of .xlsx [3].

R looks for files relative to its working directory. Run getwd() to see what that directory is right now [1]. If your file is not there, either give the full path or change the directory with setwd().

Path separators matter. Use a forward slash (/) or a double backslash (\\) in the path string. A single backslash is an escape character in R and will cause an error [3]. On a Mac you can start the path with a tilde (~) to stand for your home folder, so "~/Downloads/data.csv" points at your Downloads folder [3].

Step by Step

  1. Find the file path. Locate your CSV and note its full location. If it sits in your working directory, the file name alone is enough.
  1. Check the working directory. Run getwd() in the console. If the file is elsewhere, either pass the full path or call setwd("path/to/folder") first.
  1. Read the file. Assign the result to a name so you can reuse it.
dat <- read.csv("data.csv")  # basic import, header row assumed
  1. Confirm the dimensions. dim(dat) returns the number of rows and columns. Compare that against what you expect from the file.
dim(dat)  # rows and columns
  1. Inspect the structure. str(dat) shows each column's name and type. This is where you catch a numeric column that arrived as text.
str(dat)  # column names and types
  1. Look at the first rows. head(dat) prints the top six rows so you can eyeball the values.
head(dat)  # first six rows
  1. Check for missing values. is.na(dat) flags missing entries. colSums(is.na(dat)) counts them per column.
colSums(is.na(dat))  # missing count per column
  1. Adjust arguments if needed. If the header is missing, set header = FALSE. If the separator is a semicolon, use read.csv2() or set sep = ";" [2].

Worked Example

Suppose you have a small file called scores.csv with a header row and three students. The file holds four columns: name, quiz1, quiz2, and total.

scores <- read.csv("scores.csv")
dim(scores)
str(scores)
head(scores)

dim(scores) reports 3 rows and 4 columns, matching the three students and four fields. str(scores) shows name as a character column and the three score columns as integers. head(scores) prints all three rows, so you can confirm the header landed in the column names instead of the first data row.

Now check a derived value. If the first student scored 2 on quiz1 and 3 on quiz2, the total column should read 5, since $2 + 3 = 5$. If it reads something else, the file and your expectations disagree, and you should look at the raw CSV before trusting the import.

Finally, count missing values:

colSums(is.na(scores))

If every column returns 0, the file imported cleanly. A nonzero count means at least one field was blank or held a value listed in na.strings, which defaults to "NA" [1].

Other Ways to Do It

read.csv() is a wrapper around read.table() with defaults tuned for comma-separated files: header = TRUE, sep = ",", fill = TRUE, and comment.char = "" [2]. You can call read.table() directly when you need finer control.

FunctionSeparatorDecimalTypical use
read.csv()commaperiodStandard CSV files
read.csv2()semicoloncommaEuropean-style CSV files
read.delim()tabperiodTab-separated files
read.table()you set ityou set itAny delimited text file

For large files, read.table() and its wrappers can use a surprising amount of memory [2]. Two arguments help. Setting colClasses to one of the six atomic vector classes reduces memory, especially for numeric columns with many distinct values, because storing each value as a character string can take up to 14 times more space [2]. Setting nrows to a mild over-estimate of the row count also helps [2].

If your file has row names in the first column, read it with row.names = 1 so those names become the data frame's row names instead of a regular column [2].

Once the data is in, you may want to reshape it. See How to Collapse Data in R (Step by Step) for turning many rows into summary rows.

Troubleshooting

"cannot open the connection" or "No such file or directory." The path is wrong or the file is not in the working directory. Run getwd() and compare it against the file's real location [1].

"more columns than column names." Some rows have extra commas, often from unquoted text fields. Check the raw file and quote any text fields that contain commas. fill = TRUE is already the read.csv() default, so it does not fix this error [2].

Everything imported as one column. The separator is probably not a comma. Try read.csv2() for semicolons or read.delim() for tabs [2].

Numbers arrived as text. A stray character, a thousands separator, or a currency symbol in the column will force the whole column to character. Clean the source file or set colClasses explicitly.

Unexpected NA values. Blank fields count as missing in logical, integer, numeric, and complex columns [1]. If your file uses a different missing marker, pass it through na.strings.

Column names look mangled. check.names = TRUE is the default and makes names syntactically valid, which can add dots or prefixes. Set check.names = FALSE to keep the original names [2].

Common Mistakes

  • Forgetting the working directory. A bare file name only works if the file is in getwd(). Fix: use an absolute path or call setwd() first [1].
  • Using a single backslash in a Windows path. R reads it as an escape. Fix: use / or \\ [3].
  • Assuming text stays text. In older R versions, stringsAsFactors defaulted to TRUE and turned character columns into factors. Fix: set stringsAsFactors = FALSE and check with str() [1].
  • Skipping the inspection step. Reading without checking types or dimensions hides silent problems. Fix: always run str(), dim(), and head() after import.
  • Ignoring the header setting. If your file has no header row, the first data row becomes the column names. Fix: set header = FALSE [2].
  • Trusting the import on a huge file. Memory use can balloon. Fix: set colClasses and nrows to keep it in check [2].

Limitations

read.csv() is built for data frames, where columns can have different types. It is not the right tool for reading large matrices, especially ones with many columns [2]. For very large files, a package designed for fast or chunked reading will serve you better.

The function also makes assumptions you may not notice. It sets each column's type from all of its values, so a single non-numeric entry anywhere in a column makes the whole column import as character. It treats blank fields as missing in numeric columns without asking [1]. And it strips white space before testing na.strings, so a marker with its own surrounding spaces may not match unless you strip it in advance [1]. Always verify the result against the raw file.

Frequently Asked Questions

What is the difference between read.csv and read.table in R?

read.csv() is a convenience wrapper around read.table() with defaults set for comma-separated files: header = TRUE, sep = ",", fill = TRUE, and comment.char = "" [2]. Use read.table() when you need to control those settings yourself, such as a custom separator or comment character.

How do I read a CSV in R without a header row?

Set header = FALSE. The columns then get default names like V1, V2, and so on [2]. You can rename them afterward with colnames() or supply your own with the col.names argument.

Why did my numeric column import as a factor or character?

A non-numeric character somewhere in the column forces the whole column to text. Common culprits are thousands separators, currency symbols, and stray spaces. Inspect the raw file, clean the column, and re-import. Setting stringsAsFactors = FALSE keeps text as character instead of factor [1].

How do I read a CSV from a URL in R?

Pass the URL as the file argument. read.csv("https://example.com/data.csv") works because the function accepts a connection description, not only a local path [1]. The same argument rules apply.

How can I speed up reading a large CSV?

Set colClasses to the atomic vector classes your columns actually use, and set nrows to a slight over-estimate of the row count. Both reduce memory and time [2]. Setting comment.char = "" is also appreciably faster than leaving a comment character in place [2].

References

  1. read.csv function - RDocumentation
  2. R: Data Input
  3. 9.2 Directly Reading CSV Files | Analytics Using R

Further Reading

Related Articles