Data engineering workbench

DevToolBox
Data formats

How to Open and Inspect a Parquet File Without Python or Spark

You can open a Parquet file without Python or Spark by loading it in a browser-based Parquet viewer, which reads the file locally, or with the DuckDB command-line tool. Either way you can see the schema, page through rows and export to CSV or JSON. This guide uses the DataToolbox Parquet Viewer, which runs in your browser and does not upload the file, and shows the DuckDB equivalents.

What is inside a Parquet file?

A Parquet file is a columnar file: values are stored column by column, not row by row. The file starts and ends with the four bytes PAR1. The footer at the end holds the metadata: the schema, the number of rows, and where each row group and column chunk starts. A row group is a horizontal slice of rows; inside it, each column is a separately compressed column chunk.

This layout is why a viewer can open a large Parquet file quickly. It reads the footer first, then only the row groups and columns it needs for the page you are looking at.

Open a Parquet file in the browser

  1. Go to the Parquet Viewer.
  2. Click Open .parquet file or drop the file on the tool. To see how it works first, click Try sample file (500 rows of synthetic order data).
  3. The file line shows the row count, number of columns and number of row groups. The Rows tab lists the first 100 rows; use Next or a larger page size to keep going.
  4. Use Columns to hide what you don't need and Add filter to keep matching rows. Filters run before rows are listed, and the total number of matches is shown.

The file is read with the browser's File.slice() and decoded in a Web Worker on your device. It is not uploaded; the tool's browser tests check that no network request carries file content.

Read the schema and metadata

Open the Parquet Schema Viewer, or the Schema & metadata tab. Three things are worth checking in any Parquet file:

  • Logical vs physical types. A column stored as INT64 may be a DECIMAL(18,4) or a TIMESTAMP(MICROS, UTC). The logical type tells readers how to interpret the stored integers.
  • Timestamps and time zones. A Parquet timestamp is either adjusted to UTC or local (no time zone). Older writers such as Spark and Impala may use the legacy INT96 type, which carries no time zone flag. The viewer labels each one and only adds a Z to UTC-adjusted values; see Parquet timestamps explained.
  • Repetition. REQUIRED columns cannot be null, OPTIONAL ones can, and REPEATED fields form lists. Nested lists, structs and maps are shown indented under their parent.

The column chunk table lists each column's compression codecs, encodings, compressed and uncompressed sizes, and the min, max and null counts the writer stored in the footer. Those statistics are read from the file, not computed by scanning it, and a writer may leave them out.

Check values that are easy to get wrong

Two kinds of values often come out wrong in quick inspection tools, because JavaScript and JSON numbers are 64-bit floats:

  • 64-bit integers above 9,007,199,254,740,991 lose their last digits as floats. The viewer keeps them exact.
  • Decimals such as DECIMAL(38,10) are stored as integers with a scale. The viewer formats them from the stored integer, so 1234567890123456789012345678.0123456789 is shown digit for digit.

Also check how missing data is represented. The viewer shows null, an empty string "", and a missing struct field differently, so you can tell whether a blank cell is really empty or really NULL.

Export to CSV or JSON

To convert the file, use Parquet to CSV or Parquet to JSON, or the Export tab:

  • Rows to export: the current page, the rows matching your filters, or all rows. Each option shows its row count first.
  • CSV options: delimiter, how NULL is written (empty, NULL or \N), and whether nested columns are written as JSON text or left out. Empty strings are always quoted so they stay different from NULLs. Details and examples: Parquet to CSV without losing precision, NULLs or nested data.
  • JSON: a JSON array or JSON Lines. Decimals, and integers beyond the safe range, are written as strings so no digits are lost.

The same steps with DuckDB

If you prefer a command line, DuckDB reads Parquet directly, without Python:

-- first rows
SELECT * FROM 'orders.parquet' LIMIT 10;
-- column names and types
DESCRIBE SELECT * FROM 'orders.parquet';
-- schema and row group metadata from the footer
SELECT * FROM parquet_schema('orders.parquet');
SELECT * FROM parquet_metadata('orders.parquet');
-- export
COPY (SELECT * FROM 'orders.parquet') TO 'orders.csv' (HEADER);

DuckDB is the better choice for files beyond the browser limits below, or when you need SQL across several files.

When a Parquet file won't open

  • "Not a Parquet file": the file does not start and end with PAR1. It may be CSV, JSON or a compressed archive with a .parquet name.
  • "Incomplete Parquet file": it starts with PAR1 but has no footer — usually an interrupted download or copy.
  • "Encrypted Parquet file": files with an encrypted footer start or end with PARE and need the decryption keys.
  • "Too much data for one read": a single row group decodes to more than the tool's tested limit. Select fewer columns, or use DuckDB.

How large a file can the browser handle?

Memory, not file size, is the limit. In our measurements on a 16 GB Apple M4 in Chromium, a 5,000,000-row file split into 1,000,000-row groups opened in about 0.1 seconds and exported to a 338 MB CSV in about 9 seconds. A single 400 MB row group still opened, but exporting 400 MB of decoded data crashed the tab, so the tool refuses exports above 250 MB of decoded data and reads above 400 MB. These numbers are from that one machine and are not a guarantee: a device with less memory or a different browser can fail on smaller files. The full table is on the Parquet Viewer page.

Tools used in this guide