Compare Parquet schemas across a set of files and report the differences that actually break readers -- with the specific files and columns, ranked by severity, and CI-friendly exit codes.
$ parquet-schema-check data/*.parquet
Findings (2 error, 1 warning, 1 info):
[ERROR] id: physical type INT32 in data/a.parquet vs INT64 in data/b.parquet -- readers
coerce silently when values happen to fit and raise when they don't
files: data/a.parquet, data/b.parquet
[ERROR] event_time: logical type 'TIMESTAMP(unit=ms, local)' in data/a.parquet vs
'TIMESTAMP(unit=us, local)' in data/b.parquet -- same physical storage,
different meaning
files: data/a.parquet, data/b.parquet
A dataset is usually many Parquet files written over time by different
code: a Spark job, a pandas.to_parquet() call from a notebook, a Rust
service. Parquet's schema lives in each file's own footer, independently,
and nothing checks that footer against the others. When a column's type
drifts -- int32 becomes int64, a timestamp's unit or timezone changes,
a field becomes nullable -- some readers coerce the mismatch silently,
some fail the moment they touch it, and which one happens depends on the
data in the specific file, not the schema. The failure surfaces at query
time, in whichever file happens to trip it, far from whatever job wrote
that file weeks earlier.
Catching this needs the schema from every file's footer, compared against every other file's -- which means either a full Parquet engine or reading the footer directly.
A Parquet file ends with an 8-byte trailer: a 4-byte footer length and the
4-byte magic PAR1. Seek to file_size - 8, read the length, seek back
that far, and the bytes in between are the entire file's metadata --
schema, row groups, column chunks -- Thrift-compact-encoded as a
FileMetaData struct. None of the actual column data has to be read.
Thrift's compact protocol is a small, fully documented binary format --
varints, zigzag signed integers, a field-id-delta header byte, and
container headers for lists/sets/maps. Hand-rolling just enough of it to
walk FileMetaData turned out to be tractable in well under 500 lines
(see thrift_compact.py and
footer.py), so that's what this
does. There is no Parquet or Thrift dependency at runtime -- this
parses the footer itself, against the struct layout in Parquet's published
parquet.thrift. It was cross-checked field-by-field against pyarrow's
own schema output across every fixture in this repo's test suite and
matched exactly, including timestamp unit/timezone and per-column-chunk
encodings.
What it does not attempt: decoding row data, encrypted footers (Parquet's
encrypted-footer magic, PARE, is detected and rejected with a clear
error rather than misparsed), column indexes, or bloom filters. None of
those are needed to compare schemas.
Every claim below was checked by writing real Parquet files with
pyarrow and then forcing a mismatched schema through pyarrow's own
reader. This is the real output of
examples/reader_breakage_demo.py
(also exercised as regression tests in
tests/test_reader_breakage_evidence.py):
=== 1. int32 vs int64: silent when it fits, loud when it doesn't ===
in-range int64 data read as int32: OK -- [1, 2]
out-of-range int64 data read as int32: FAILED -- Integer value 5000000000 not in range: -2147483648 to 2147483647
=== 2. timestamp unit change: rescales or raises depending on the values ===
ns timestamps read as ms: FAILED -- Casting from timestamp[ns] to timestamp[ms] would lose data: 1
=== 3. timezone drop: silently reinterprets the instant, no error at all ===
UTC-tagged timestamp read with a naive schema: OK -- tz label dropped, value kept: [datetime.datetime(2024, 1, 1, 12, 0)]
=== 4. nullability: REQUIRED-declared schema over a file with real nulls ===
OK -- no error, 1 null(s) present despite the REQUIRED label
=== 5. missing column: filled with nulls by a schema-unifying reader ===
OK -- column 'b': ['x', 'y', None, None]
=== 6. dictionary vs plain encoding: fully transparent ===
same Arrow type either way: string == string
That maps onto the ranking this tool uses:
| Difference | What actually happened |
|---|---|
int32 vs int64, same column |
In-range values: silently coerced, no error. Out-of-range values (5_000_000_000 into an int32 slot): ArrowInvalid: Integer value ... not in range. Same schema difference, two outcomes, decided by which file has which rows. Ranked ERROR. |
Timestamp unit change (ms vs us vs ns) |
Widening a unit (ns forced through an ms schema) raised ArrowInvalid: ... would lose data. Narrowing with round values did not raise -- it silently rescaled. Ranked ERROR -- the difference is the same regardless of which failure mode a given file triggers. |
| Timestamp timezone change (naive vs UTC-adjusted) | No error at all. The wall-clock number is kept and the timezone label is dropped, silently reinterpreting what instant the value means. Worse than a raise, because nothing signals it happened. Ranked ERROR. |
Nullability change (REQUIRED vs OPTIONAL) |
A schema declaring REQUIRED read a file that actually contained nulls without complaint. The read itself never broke in testing. Ranked WARNING -- worth a look if a downstream consumer assumes non-null, but not a read failure. |
| Column present in some files, absent in others | A schema-unifying reader (pyarrow.dataset) filled the gap with nulls without complaint. Ranked WARNING -- fine for a dataset-style reader, not necessarily for one expecting an exact per-file match. |
Encoding (PLAIN vs dictionary/RLE_DICTIONARY) |
Decoded to the identical Arrow type and values either way -- encoding is a physical-page detail, not part of the logical schema. Ranked INFO. |
| Compression codec | Transparent to any reader that has the codec available. Ranked INFO. |
Repetition change to/from REPEATED (scalar vs list) |
Not exercised against a live reader here, but structurally these are different node shapes in the schema tree, not the same field with a relaxed constraint. Ranked ERROR on the same basis as a physical-type change. |
The takeaway that shaped the ranking: the data-dependent failures are the dangerous ones. A difference that always raises is annoying but safe -- CI catches it immediately. A difference that raises only for some rows is the one that reaches production.
pip install parquet-schema-checkZero runtime dependencies -- footer parsing is pure standard library, as described above.
parquet-schema-check data/*.parquet
parquet-schema-check --json data/*.parquet > report.json
parquet-schema-check --fail-on warning data/*.parquet # stricter CI gatefrom parquet_schema_check import read_schema, compare_schemas
schemas = [read_schema(p) for p in ("a.parquet", "b.parquet")]
result = compare_schemas(schemas)
for finding in result.errors:
print(finding.dotted_path, finding.message, finding.files)
for column in result.union_columns: # the schema a reader would need for every file
print(column.dotted_path, column.physical_types, column.missing_from)| Flag | Default | What it does |
|---|---|---|
--json |
off | Emit a machine-readable report instead of text: files, union_schema, findings, summary. |
--fail-on {error,warning,none} |
error |
Minimum severity that produces a non-zero exit code. |
Exit codes: 0 clean (or below the --fail-on threshold), 1 warnings
at or above the threshold, 2 errors present, 3 usage problem (fewer
than two readable files given).
- It does not read row data or validate values. This is a footer/schema comparison, not a data-quality checker -- it can't tell you a numeric column suddenly holds different values, only that its declared type changed.
- It does not read encrypted footers. Parquet Modular Encryption
encrypts the footer itself (magic
PAREinstead ofPAR1); this is detected and reported as a clear error, not silently misparsed. - It compares schemas, not partitioning or file layout. Hive-style partition columns encoded in directory names aren't part of any file's footer and aren't considered here.
- Statistics, column indexes, and bloom filters are skipped, not parsed -- they don't affect whether a reader can open the file.
python3 -m venv .venv && .venv/bin/pip install -e '.[dev]'
.venv/bin/python -m pytest -q
.venv/bin/python -m mypy src --strictTests generate real Parquet files with pyarrow (a dev-only dependency --
the package itself never imports it) into tmp_path, then read them back
with this package's own zero-dependency footer parser.
MIT