Apache Parquet File Format: What It Is and How to Use It

By Dr. Zubair Khalid, DVM, MS, PhD ·

Apache Parquet File Format: What It Is and How to Use It

Apache Parquet is an open source, column-oriented file format designed for analytical workloads. Instead of storing a table row by row, it stores each column together, which lets tools read only the columns a query needs and compress similar values efficiently. In Python, the pandas library reads and writes Parquet through read_parquet and to_parquet.

Quick Answer

  • Apache Parquet is a columnar file format for tabular data, common in analytics pipelines and data lakes.
  • It stores values column by column, so a query that needs two columns does not read the other columns.
  • Columns of similar values compress well, which usually shrinks files compared with plain CSV.
  • It keeps a schema, so data types like integers, floats and dates survive a round trip.
  • In Python, pandas.read_parquet() and DataFrame.to_parquet() handle reading and writing with one line each.

What Apache Parquet Means

In plain terms, Apache Parquet is a file format for tables. You save a table to a .parquet file, and any tool that understands the format can read it back with the same columns and types.

The precise definition: Parquet is an open source, column-oriented data file format that stores data in a binary, compressed, typed form, with metadata describing the schema and the location of each column chunk. The format was designed for efficient storage and retrieval of large analytical datasets, and it is readable from many programming languages through the Apache Arrow project, which provides in-memory columnar analytics across languages [1].

Two properties follow from that definition. First, the file is columnar, so the physical layout groups values from the same column. Second, the file is self-describing, so a reader knows the column names and types without a separate schema file.

How It Works

The core mechanism is the layout of values on disk. A row-oriented format writes all fields of record 1, then all fields of record 2, and so on. A columnar format writes all values of column A, then all values of column B.

For a table with columns $c_1, \dots, c_k$ and $n$ rows, the row layout is:

$$ \text{row layout} = (c_1^1, c_2^1, \dots, c_k^1),\ (c_1^2, c_2^2, \dots, c_k^2),\ \dots $$

The column layout is:

$$ \text{column layout} = (c_1^1, c_1^2, \dots, c_1^n),\ (c_2^1, c_2^2, \dots, c_2^n),\ \dots $$

Each symbol means the following. $c_j^i$ is the value in column $j$ of row $i$. $n$ is the number of rows. $k$ is the number of columns. The superscript indexes the row and the subscript indexes the column.

This layout produces three practical effects.

  1. Column pruning. A query that touches only some columns reads only those column chunks. The rest of the file is skipped.
  2. Better compression. Values in one column share a type and often repeat, so dictionary encoding and general compression work well on them.
  3. Cheaper scans. Analytical queries usually aggregate a few columns over many rows, which matches the column layout.

Parquet also stores metadata per column chunk, including statistics such as minimum and maximum values. A reader can use those statistics to skip chunks that cannot match a filter.

Worked Example

The dataset is a small sales table with 5 order records, each with an order id, date, region, product, quantity, unit price and revenue.

order_idorder_dateregionproductquantityunit_pricerevenue
10012024-01-05NorthWidget129.99119.88
10022024-01-06SouthGadget324.5073.50
10032024-01-07EastWidget209.99199.80
10042024-01-08WestDoohickey714.75103.25
10052024-01-09NorthGadget524.50122.50

The steps below write the table to both CSV and Parquet, then read two columns back from the Parquet file.

import pandas as pd
df = pd.DataFrame({
    'order_id': [1001,1002,1003,1004,1005],
    'region':   ['North','South','East','West','North'],
    'product':  ['Widget','Gadget','Widget','Doohickey','Gadget'],
    'quantity': [12,3,20,7,5],
    'unit_price':[9.99,24.50,9.99,14.75,24.50],
})
df['revenue'] = (df['quantity']*df['unit_price']).round(2)
df.to_parquet('sales.parquet', index=False)   # columnar, compressed
df.to_csv('sales.csv', index=False)           # row-based, plain text
back = pd.read_parquet('sales.parquet', columns=['order_id','revenue'])
import os
print('CSV size =', os.path.getsize('sales.csv'), 'bytes')
print('Parquet size =', os.path.getsize('sales.parquet'), 'bytes')
print('Total revenue =', round(back['revenue'].sum(), 2))

Output:

CSV size = 212 bytes
Parquet size = 4079 bytes
Total revenue = 618.93

Walking through the numbers:

  • Rows in dataset: 5.
  • CSV bytes: 212 bytes for the six columns in the code, including the header.
  • Parquet bytes: 4,079 bytes with pyarrow 25 (the exact size varies a little by library version), because the file carries a schema, per-column metadata and a footer.
  • Size ratio: $212 / 4079 \approx 0.05$, so the Parquet file is roughly 19 times larger here.
  • Read time: both files load in well under a millisecond, so timing differences on five rows are noise and say nothing about the format.
  • Total revenue: 618.93, the sum of the revenue column.

This result is the honest one for a five-row table. Parquet carries per-column metadata, schema information and compression headers, and those fixed costs dominate when the data is tiny. The format pays off as row counts grow into the thousands and beyond, where column pruning and compression have enough data to work on.

How to Interpret It

Read the size ratio and the read speedup together, and always check the row count first. On a handful of rows, a Parquet file can be larger than the CSV and slower to open, because the format's overhead is fixed and the payload is trivial. That is not a failure of the format, it is a scale effect.

On real analytical tables, the pattern reverses. A file with millions of rows and dozens of columns compresses well because each column holds values of one type, and a query that selects three columns reads roughly three columns' worth of data instead of the whole table. The gain grows with the number of columns you skip and the repetition inside each column.

The revenue total of 618.93 is a useful sanity check. Whatever format you use, the aggregated value should match, so a mismatch after a round trip points to a type or parsing problem, not a format problem.

When to Use It (and when not to)

Use Apache Parquet when you store analytical tables, when downstream queries touch a subset of columns, when files are large enough that compression matters, and when you want types preserved across tools. It is a common choice for data lake storage and for intermediate files in a pipeline. One dataset provider recommends Parquet over CSV for downloads because of its smaller size and its support across many languages through Apache Arrow [1].

Do not use it when you need a human-readable file you can open in a text editor, when you need to append single rows frequently, or when the data is a few dozen rows and a CSV is simpler. Parquet files are binary, and editing them by hand is not practical.

Apache Parquet vs CSV

PropertyApache ParquetCSV
LayoutColumnarRow-based
FormatBinaryPlain text
SchemaStored in the fileNone, types are inferred
CompressionBuilt in, per columnNone by default
Column pruningYesNo, the whole file is read
Human readableNoYes
Best forLarge analytical tablesSmall files, exchange, inspection

The comparison table above is the practical summary. CSV wins on simplicity and readability. Parquet wins on size and scan efficiency once the data is large.

Common Mistakes

  • Judging Parquet on tiny files. A five-row Parquet file can be larger than the CSV, as the worked example shows. Test on a realistic row count before drawing conclusions.
  • Assuming Parquet is always smaller. Compression depends on the data. High-entropy columns of unique strings may compress poorly. Measure your own files.
  • Losing the index on write. to_parquet with the default index=None stores a non-default pandas index as a column (a plain RangeIndex is kept only as metadata). Pass index=False when the index carries no information.
  • Reading every column when you need a few. Pass the columns argument to read_parquet so the reader can skip the rest. Reading all columns throws away the main advantage.
  • Treating Parquet as append-friendly. Rewriting a whole file is the normal pattern. Frequent single-row appends belong in a database, not a Parquet file.
  • Ignoring the engine. pandas needs a Parquet engine such as pyarrow or fastparquet installed. A missing engine raises an import error at read or write time.

Limitations

Parquet is not editable in place. Changing one value means reading the file, changing the value in memory and writing a new file. That makes it a poor fit for transactional workloads with frequent small updates.

The format is also binary and schema-bound. You cannot open it in a text editor, and a schema mismatch between writer and reader causes errors or unexpected nulls. Compression ratios vary widely with the data, so a Parquet file is not guaranteed to be smaller than a CSV for every table. Finally, small files carry fixed overhead, so many tiny Parquet files in a directory can be slower to process than one moderately sized file.

Frequently Asked Questions

Is Apache Parquet better than CSV?

For large analytical tables, usually yes. Parquet stores a schema, compresses each column and lets readers skip columns a query does not need. For small files, quick inspection or manual editing, CSV is simpler and often faster to open.

Can I open a Parquet file in Excel?

Not directly in a plain spreadsheet workflow. Parquet is a binary columnar format, so you need a tool that understands it, such as Python with pandas and pyarrow, or another analytics tool with Parquet support. Convert to CSV first if you need a spreadsheet.

Does Parquet keep my data types?

Yes. The schema is stored in the file, so integers, floats, strings and dates come back with their types intact. CSV stores text only, so a reader has to guess types, which can turn an id column into numbers or drop leading zeros.

Why was my Parquet file bigger than the CSV?

Small tables carry fixed overhead for schema, metadata and compression headers. With only a few rows, that overhead can exceed the raw text. The advantage appears as row counts and column counts grow.

How do I read only some columns from a Parquet file?

Pass the columns argument to read_parquet, for example pd.read_parquet('sales.parquet', columns=['order_id','revenue']). The reader then touches only those column chunks, which is the main performance benefit of the format.

References

  1. README

Further Reading

Related Articles