Back to Blog
Python

Read and Write CSV and Parquet Files with Polars DataFrames in Python

Read and write Polars DataFrames to CSV and Parquet files. Includes schema handling, compression, null values, and lazy scans for large files.

PolarsDataFrameCSVParquetPython
Illustration of a Polars DataFrame being read from a CSV file and written to a Parquet file, showing data flow between formats.

When you need to persist a Python Polars DataFrame, CSV and Parquet are the two most common formats. Both have different trade-offs, and Polars provides dedicated APIs for each. This article covers the syntax for reading and writing these formats, the options that affect behavior, and the situations where one format is preferable.

Reading a CSV into a Polars DataFrame

Polars reads CSV files with pl.read_csv. The function infers the schema by sampling the data, but you can override it when needed. The simplest call is:

import polars as pl df = pl.read_csv('data.csv')

This returns a DataFrame with columns inferred from the CSV header row. If your file has no header, pass has_header=False. You can also specify a custom separator with separator=';' for semicolon-delimited files.

For files with inconsistent null representations, use null_values to map them to null:

df = pl.read_csv( 'data.csv', null_values=['NA', 'N/A', 'NULL'], )

When the schema is known in advance, provide it with the schema parameter. This avoids inference mistakes and speeds up parsing:

schema = {'id': pl.Int64, 'name': pl.Utf8, 'score': pl.Float64} df = pl.read_csv('data.csv', schema=schema)

The try_parse_dates option attempts to convert date-like strings to date or datetime columns. Use it when your CSV contains ISO dates and you want to avoid a separate conversion step.

Writing a Polars DataFrame to CSV

To write a DataFrame to CSV, call write_csv on the DataFrame:

df.write_csv('output.csv')

By default, the header row is included. Set include_header=False to omit it. The separator option controls the delimiter, and quote sets the quoting character. For large DataFrames, batch_size controls how many rows are written at once, which affects memory usage during the write.

CSV has no native schema, so type information is lost unless you re-specify it on read. If you need to preserve the schema for later reads, consider writing a separate schema file or using Parquet instead.

Reading a Parquet File into a Polars DataFrame

Parquet is a columnar format that preserves schema and supports compression. Reading is done with pl.read_parquet:

df = pl.read_parquet('data.parquet')

You can read only a subset of columns with the columns parameter:

df = pl.read_parquet('data.parquet', columns=['id', 'name'])

The parallel option sets the parallel read strategy. For most workloads the default is fine, but on memory-constrained systems you can pass parallel='none' to reduce peak memory use.

Polars also supports reading partitioned Parquet datasets. Point read_parquet to a directory that follows Hive-style partitioning, and it will read all files and add the partition columns to the result.

Writing a Polars DataFrame to Parquet

The write_parquet method writes a DataFrame to a Parquet file:

df.write_parquet('output.parquet')

You can control compression with the compression parameter:

df.write_parquet('output.parquet', compression='gzip')

Polars supports several compression codecs, including snappy, gzip, lz4, and zstd. The choice affects file size and read/write speed. Gzip often produces smaller files but can take longer to write and read; Snappy is generally faster but less compact.

The row_group_size parameter controls how many rows are grouped in each row group. Larger row groups reduce metadata overhead but increase memory during reads. The default is usually a good balance, but you may adjust it for very wide or very tall DataFrames.

Choosing Between CSV and Parquet

ConsiderationCSVParquet
SchemaLost on write, inferred on readPreserved exactly
CompressionNot native; requires external toolsBuilt-in codecs
Read speedSlower due to parsing and type inferenceFaster for columnar access
File sizeLarger, text-basedSmaller, binary and compressed
Human readableYesNo
Use caseInterchange, debugging, small dataAnalytics, large datasets, pipelines

Use CSV when you need a portable, human-readable file that other tools can open without special libraries. Use Parquet when you are building a data pipeline, need to preserve types, or work with large datasets where read performance matters.

Performance and Memory Considerations

Polars is designed for efficient memory use and parallel execution. When reading CSV, the main cost is parsing text and inferring types. Providing an explicit schema avoids the inference pass and can reduce read time. For Parquet, the columnar layout allows Polars to read only the columns you need, which is a major advantage when working with wide tables.

Writing to Parquet is generally faster than writing to CSV because the columnar write path avoids text formatting and can compress data in parallel. CSV writing is bound by string formatting and I/O.

The eager read_csv and read_parquet functions build the whole DataFrame in memory. For files that are too large to fit comfortably, use pl.scan_csv or pl.scan_parquet to get a LazyFrame. You can filter or select before calling collect, and if your Polars version supports it, pass streaming=True to process the data in a streaming fashion:

lazy_df = pl.scan_csv('large.csv') filtered = lazy_df.filter(pl.col('value') > 100).collect(streaming=True)

The same pattern works with scan_parquet. Using lazy evaluation lets Polars push down predicates and projections, reducing the amount of data actually loaded.

Handling Edge Cases and Common Errors

CSV files often contain messy data. Missing values, inconsistent quoting, and unexpected types are common. Polars provides null_values and schema to handle these. If a row or value cannot be parsed, Polars raises an error by default. You can set ignore_errors=True to keep reading, but be careful: this can hide real problems.

Parquet files carry a fixed schema. On read, Polars uses that schema; if you request a column that is not present in the file, an error is raised. On write, the DataFrame schema becomes the Parquet file schema.

Another common type question is strings. CSV inference stores string columns as Utf8 (String in newer Polars versions). Parquet also has a string logical type, so Polars normally reads string columns back as Utf8/String. If you want categorical behavior for low-cardinality columns, cast explicitly with .cast(pl.Categorical).

For large files, avoid reading the entire file into memory if you only need a subset. Use scan_csv or scan_parquet with a filter and select before collect. This is the most effective way to keep memory usage under control.

Polars DataFrames: Read and Write CSV and Parquet in Python | RYUSLOG DEV