Skip to content

Fix reading gzipped files downloaded over HTTP - #1051

Merged
handecelikkanat merged 1 commit into
mainfrom
hande/fix-read-gzip-detection
Sep 30, 2026
Merged

handecelikkanat merged 1 commit into
mainfrom
hande/fix-read-gzip-detection

Conversation

@handecelikkanat

Copy link
Copy Markdown
Contributor

Refs #635

Problem

Gzipped CSV, TSV, JSON and JSON Lines files cannot be read when they are downloaded over HTTP.

  • mlcroissant saves HTTP downloads in its cache as croissant-<sha256 of URL>, without a file extension.
  • But Read in read.py uses filepath.suffix == ".gz" to detect gzipped files. (Introduced in git lfs download fileObject and read gzipped files #636.)
  • So downloaded files cannot be detected as gzipped, and sent directly to pandas:
    • Which fails with: UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8b in position 1: invalid start byte
  • This also affects the latest release (1.1.0).
  • Gzipped files read with unArchive (for example in manifest files for Common Crawl) are not affected. They go through extract.py, which already detects gzip by content.

We found this when we added a gzipped CSV file over HTTP to the Common Crawl croissants.

Fix

Use is_gzip from extract.py to detect gzipped files by magic bytes instead, which doesnt depend on the exact file name. (Introduced in #1001).

What it affects

  • This patch covers formats read through the gzip wrapper: {CSV, TSV, JSON and JSONL}.
  • As for Parquet files:
    • No change for regular Parquet files
    • Gzipped Parquet files (ie., .parquet.gz) did not work, they still dont work after this change. (The wrapper opens files in text mode.)
  • Gzipped files named .gz work as before.
  • Gzipped files without .gz now work, eg. HTTP downloads, local files without an extension, etc.
  • A file named .gz that is not gzipped is now correctly detected as a plain file, so it will not fail.
  • Plain files are not affected.

Tests

  • New test added that reads two gzipped CSVs, named file.csv.gz and file.
    • First test passes before and after this patch.
    • Second test passes after patch.
  • Also tested with a local HTTP server, loading a file.csv.gz URL, a URL without an extension (file), and a plain CSV (file.csv) all worked.

Note

This patch fixes the loader only. #635 introduces a broader spec question regarding how to best declare compression in the metadata, which is still an open question.

HTTP downloads are cached without a file extension, so the .gz
suffix check in read.py cannot catch those. Instead detect gzipped
files by magic number using is_gzip() defined in extract.py.

Refs #635
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@handecelikkanat
handecelikkanat marked this pull request as ready for review September 23, 2026 18:09
@handecelikkanat
handecelikkanat requested a review from a team as a code owner September 23, 2026 18:09
@handecelikkanat

Copy link
Copy Markdown
Contributor Author

@benjelloun If you have time to look at this in the next two days, it would be very helpful, we want to publish gzipped metadata files in October crawl, to be cited from Croissants in November. Ty!

@benjelloun benjelloun left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the fix!

@handecelikkanat
handecelikkanat merged commit 6896ea2 into main Sep 30, 2026
13 checks passed
@handecelikkanat
handecelikkanat deleted the hande/fix-read-gzip-detection branch September 30, 2026 16:05
handecelikkanat added a commit that referenced this pull request Oct 2, 2026
Bumps the mlcroissant version from 1.1.0 to 1.1.1 so that the gzip
detection fix (#1051) for #635 can be released to PyPI.

Before fix #1051, gzipped files over HTTP referenced by a croissant
could not be loaded, because the compression was taken from the file
name rather than from the content, and downloaded files wouldnt have
the .gz extension in the cache name.

Common Crawl publishes gzipped CSV files that depend on this fix, so a
PyPI release containing it would be really good for our downstream
users.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants