Fix reading gzipped files downloaded over HTTP - #1051
Merged
Merged
Conversation
HTTP downloads are cached without a file extension, so the .gz suffix check in read.py cannot catch those. Instead detect gzipped files by magic number using is_gzip() defined in extract.py. Refs #635
|
MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅ |
handecelikkanat
marked this pull request as ready for review
September 23, 2026 18:09
Contributor
Author
|
@benjelloun If you have time to look at this in the next two days, it would be very helpful, we want to publish gzipped metadata files in October crawl, to be cited from Croissants in November. Ty! |
handecelikkanat
added a commit
that referenced
this pull request
Oct 2, 2026
Bumps the mlcroissant version from 1.1.0 to 1.1.1 so that the gzip detection fix (#1051) for #635 can be released to PyPI. Before fix #1051, gzipped files over HTTP referenced by a croissant could not be loaded, because the compression was taken from the file name rather than from the content, and downloaded files wouldnt have the .gz extension in the cache name. Common Crawl publishes gzipped CSV files that depend on this fix, so a PyPI release containing it would be really good for our downstream users.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #635
Problem
Gzipped CSV, TSV, JSON and JSON Lines files cannot be read when they are downloaded over HTTP.
croissant-<sha256 of URL>, without a file extension.Readinread.pyusesfilepath.suffix == ".gz"to detect gzipped files. (Introduced in git lfs download fileObject and read gzipped files #636.)UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8b in position 1: invalid start byteunArchive(for example in manifest files for Common Crawl) are not affected. They go throughextract.py, which already detects gzip by content.We found this when we added a gzipped CSV file over HTTP to the Common Crawl croissants.
Fix
Use
is_gzipfromextract.pyto detect gzipped files by magic bytes instead, which doesnt depend on the exact file name. (Introduced in #1001).What it affects
.parquet.gz) did not work, they still dont work after this change. (The wrapper opens files in text mode.).gzwork as before..gznow work, eg. HTTP downloads, local files without an extension, etc..gzthat is not gzipped is now correctly detected as a plain file, so it will not fail.Tests
file.csv.gzandfile.file.csv.gzURL, a URL without an extension (file), and a plain CSV (file.csv) all worked.Note
This patch fixes the loader only. #635 introduces a broader spec question regarding how to best declare compression in the metadata, which is still an open question.