# R transform Function: Syntax and Examples

The **r transform function** evaluates tagged expressions inside a data frame and returns the modified data frame. You write `transform(data, new_column = expression)`, and each expression can refer to existing columns by name without repeating the data frame. It is a convenience function for interactive work, and it both adds new columns and replaces existing ones [1].

## Quick Answer

- `transform()` is a generic function in base R. The data frame method is `transform.data.frame`, and `transform.default` converts its first argument to a data frame when possible [1].
- Arguments after the data are tagged vector expressions. Tags that match existing column names replace those columns. Tags that do not match are appended as new columns [1].
- Expressions are evaluated in the data frame itself, so `mass_g / volume_mL` works without `data$` prefixes.
- It returns a new data frame. The original object is unchanged unless you assign the result back to it.
- For programming inside functions or loops, base R documentation recommends standard subsetting arithmetic instead, because the non-standard evaluation of `transform()` can have unanticipated consequences [1].

## Syntax

```r
transform(data, ...)
```

| Argument | Required? | Meaning |
|---|---|---|
| `data` | Yes | The object to transform. For `transform.default` it is converted to a data frame if possible [1]. |
| `...` | Yes, in practice | Tagged vector expressions such as `density = mass / volume`. Each tag is matched against `names(data)`. Matching tags replace the column, non-matching tags are appended [1]. |

The return value is a data frame. Untagged arguments are allowed syntactically but are not useful, because there is no name to match or append.

## How It Works

`transform()` is a generic function. When you pass a data frame, R dispatches to `transform.data.frame`. When you pass something else, `transform.default` tries to convert the first argument to a data frame and then calls the data frame method [1].

The important part is evaluation. Each expression in `...` is evaluated in the data frame as its environment. That means column names resolve directly to vectors of the right length. The result of each expression is then placed into the output: matched names replace the existing column, and unmatched names become new columns at the end [1].

Because all expressions are evaluated against the original data frame, one new column cannot refer to another new column created in the same call. If you need that, chain two `transform()` calls or compute in stages.

## Worked Example

The dataset below holds 5 lab samples with mass in grams, volume in milliliters, and temperature in Celsius.

| sample | mass_g | volume_mL | temp_C |
|---|---|---|---|
| S1 | 12.5 | 10.0 | 20.0 |
| S2 | 25.0 | 20.0 | 25.0 |
| S3 | 8.4 | 7.0 | 18.0 |
| S4 | 40.2 | 32.0 | 30.0 |
| S5 | 15.7 | 12.5 | 22.0 |

You want two derived columns: density in g/mL and temperature in kelvin. Density is mass divided by volume, and kelvin is Celsius plus 273.15.

```r
samples <- data.frame(
  sample = c("S1","S2","S3","S4","S5"),
  mass_g = c(12.5, 25.0, 8.4, 40.2, 15.7),
  volume_mL = c(10.0, 20.0, 7.0, 32.0, 12.5),
  temp_C = c(20.0, 25.0, 18.0, 30.0, 22.0)
)

result <- transform(samples,
  density_g_mL = mass_g / volume_mL,
  temp_K = temp_C + 273.15
)

print(result)
```

Output:

```text
  sample mass_g volume_mL temp_C density_g_mL temp_K
1     S1   12.5      10.0     20      1.25000 293.15
2     S2   25.0      20.0     25      1.25000 298.15
3     S3    8.4       7.0     18      1.20000 291.15
4     S4   40.2      32.0     30      1.25625 303.15
5     S5   15.7      12.5     22      1.25600 295.15
```

Walking through the arithmetic:

- Density for S1: $12.5 / 10.0 = 1.2500$ g/mL
- Density for S2: $25.0 / 20.0 = 1.2500$ g/mL
- Density for S3: $8.4 / 7.0 = 1.2000$ g/mL
- Density for S4: $40.2 / 32.0 = 1.2563$ g/mL
- Density for S5: $15.7 / 12.5 = 1.2560$ g/mL
- Temperature in kelvin for S1: $20.0 + 273.15 = 293.15$ K
- Temperature in kelvin for S2: $25.0 + 273.15 = 298.15$ K
- Temperature in kelvin for S3: $18.0 + 273.15 = 291.15$ K
- Temperature in kelvin for S4: $30.0 + 273.15 = 303.15$ K
- Temperature in kelvin for S5: $22.0 + 273.15 = 295.15$ K

The mean density across the five samples is 1.2425 g/mL and the mean temperature is 296.15 K. Both new columns appear at the right edge of the data frame, which is the standard behavior for appended columns [1].

## More Examples

**Replace an existing column.** If the tag matches a column name, the values overwrite that column in place instead of appending [1].

```r
samples2 <- transform(samples, temp_C = temp_C + 273.15)
```

Here `temp_C` keeps its position but now holds kelvin values. This is easy to misread later, so rename the column when the unit changes.

**Use a transformation on a single vector.** The documentation shows that `transform()` also works on a bare vector, converting it to a data frame first [1].

```r
attach(airquality)
transform(Ozone, logOzone = log(Ozone))
detach(airquality)
```

This returns a data frame with the original values and the log values side by side.

**Chain two calls when one column depends on another.** Because all expressions see the original data, you need two steps for dependent columns.

```r
step1 <- transform(samples, density_g_mL = mass_g / volume_mL)
step2 <- transform(step1, inverse_density = 1 / density_g_mL)
```

**Filter and transform together.** Combine `transform()` with logical indexing when you only want a subset.

```r
warm <- transform(samples[samples$temp_C > 20, ], temp_K = temp_C + 273.15)
```

If you need conditional logic inside a column, the vectorized approach in [ifelse in R](/blog/data-analysis/ifelse-in-r-vectorized-examples) pairs well with `transform()`.

## Errors and How to Fix Them

**"object 'x' not found".** You referenced a name that is not a column and not in the calling environment. Check spelling and confirm the column exists with `names(samples)`.

**Arguments imply differing number of rows.** The expression returned a vector whose length does not fit the data frame, for example `range(mass_g)`, which returns 2 values for 5 rows. A single value such as `mean(mass_g)` does not error, it is silently recycled down every row. Return one value per row, or compute summaries outside the call.

**Columns silently dropped.** If an expression returns `NULL`, the column is removed from the result. Guard against this when a helper function can return `NULL`.

**Unexpected results inside a function.** Non-standard evaluation means `transform()` looks for column names first. If your function has a local variable with the same name as a column, the column wins. This is the main reason the documentation recommends standard subsetting for programming [1].

## Common Mistakes

- **Assuming new columns can reference each other.** In `transform(d, a = x + 1, b = a * 2)`, `b` cannot see the new `a`. Fix: split into two `transform()` calls.
- **Forgetting to assign the result.** `transform(samples, temp_K = temp_C + 273.15)` prints the result but leaves `samples` untouched. Fix: assign with `<-`.
- **Overwriting a column without renaming it.** Replacing `temp_C` with kelvin values keeps the old name and misleads anyone reading the data later. Fix: use a new tag such as `temp_K`.
- **Using `transform()` inside a function or loop.** Non-standard evaluation can pick up the wrong variable [1]. Fix: use `d$new <- d$x + 1` or `within()` for interactive work only.
- **Passing an untagged expression.** `transform(samples, mass_g / volume_mL)` computes the values but gives them no name, so nothing useful is added. Fix: always tag the expression.
- **Expecting row order to change.** `transform()` preserves row order and row names. If you need sorting, do it separately.

## Limitations

`transform()` is built for interactive convenience, not for programmatic use. The base R documentation states this directly and recommends standard subsetting arithmetic functions for programming, because the non-standard evaluation of the `transform` argument can have unanticipated consequences [1]. If you are writing a package function or a script that runs unattended, plain assignment or a dedicated data manipulation package is safer.

It also does not group, summarize, or reshape. Every expression must return a vector of the same length as the data frame, or a value that recycles cleanly. For grouped operations you need a different tool. And because all expressions are evaluated against the input data frame, multi-step derivations require multiple calls, which can make long pipelines harder to read than a single assignment block.

## Frequently Asked Questions

### What does the transform function do in R?

It evaluates tagged expressions inside a data frame and returns the modified data frame. Tags that match existing column names replace those columns, and tags that do not match are appended as new columns [1]. It is a base R convenience function for interactive data preparation.

### Does transform() modify the original data frame?

No. It returns a new data frame and leaves the input unchanged. You must assign the result, as in `samples <- transform(samples, temp_K = temp_C + 273.15)`, for the change to persist in your object.

### Can I create two columns in one transform() call?

Yes, as long as neither new column depends on the other. Both expressions are evaluated against the original data frame, so `density_g_mL = mass_g / volume_mL` and `temp_K = temp_C + 273.15` work in a single call. If one depends on the other, chain two calls.

### Is transform() the same as mutate()?

They solve a similar problem but differ in evaluation rules and in how they handle sequential dependencies. `transform()` is base R and evaluates every expression against the input data frame. If you are working in a tidyverse pipeline, `mutate()` is the more common choice. For base R scripts, `transform()` avoids extra dependencies.

### Why does transform() fail inside my own function?

Non-standard evaluation makes `transform()` resolve names against the data frame first, so a column can shadow a local variable with the same name. The base R documentation flags this and recommends standard subsetting arithmetic for programming [1]. Rewrite the call as direct column assignment inside functions.

### Does transform() work on objects that are not data frames?

Yes, in many cases. `transform.default` converts its first argument to a data frame if possible and then calls the data frame method [1]. The documentation shows this with a bare vector, where `transform(Ozone, logOzone = log(Ozone))` returns a data frame with both the original and transformed values.

## References

1. [R: Transform an Object, for Example a Data Frame](https://stat.ethz.ch/R-manual/R-devel/library/base/html/transform.html)

## Further Reading

- [Wickham H (2014). Tidy Data. Journal of Statistical Software](https://doi.org/10.18637/jss.v059.i10)
- [Wickham H, Averick M, Bryan J et al. (2019). Welcome to the Tidyverse. Journal of Open Source Software](https://doi.org/10.21105/joss.01686)
- [An Introduction to R (R Core Team)](https://cran.r-project.org/doc/manuals/r-release/R-intro.html)
- [Wickham H, Cetinkaya-Rundel M, Grolemund G. R for Data Science (2e)](https://r4ds.hadley.nz/)
- [ggplot2 Reference](https://ggplot2.tidyverse.org/reference/index.html)

## Related Articles

- [Fisher r to z Transformation: Formula and Examples](/blog/data-analysis/fisher-r-to-z-transformation)
- [R and R-Squared: What They Mean and How to Interpret Them](/blog/data-analysis/r-and-r-squared-interpretation)
- [R letters Function: Generate Lowercase Letter Sequences](/blog/data-analysis/r-letters-function)
- [The %in% Operator in R: Syntax and Examples](/blog/data-analysis/in-operator-in-r-syntax-examples)
- [Excel FILTER Function: Syntax and Examples](/blog/data-analysis/excel-filter-function-syntax-examples)
- [How to Calculate Transformation Efficiency: Formula, Examples, and Common Pitfalls](/knowledge/diagnostics/molecular/calculate-transformation-efficiency-formula-examples)
- [Reproducible Research in R: A Practical Guide for Life Scientists](/blog/guides/reproducible-research-in-r-a-practical-guide-for-life-scientists)
- [Differential Gene Expression Analysis in R: A Practical Guide](/knowledge/molecular-biology/differential-gene-expression-analysis-in-r)