fix: prevent pandas_kwargs mutation during text file compression resolution - #3428
Open
hsusul wants to merge 1 commit into
Open
fix: prevent pandas_kwargs mutation during text file compression resolution#3428hsusul wants to merge 1 commit into
hsusul wants to merge 1 commit into
Conversation
…ession resolution
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This PR fixes a bug where
_get_read_detailsand_get_write_detailsmutated the passedpandas_kwargsdictionary in place when resolving file compression (e.g. settingpandas_kwargs["compression"] = infer_compression(...)).Symptom & Root Cause
When reading a list of S3 paths or a dataset containing files with different compression extensions (such as
["s3://bucket/file1.csv.gz", "s3://bucket/file2.csv"]or["s3://bucket/file1.json.gz", "s3://bucket/file2.json"]),_get_read_detailsmutated the sharedpandas_kwargsdictionary during the processing of the first file (e.g., setting"compression": "gzip").Because
pandas_kwargswas mutated in place:pandas_kwargs.get("compression", "infer")returned"gzip"for subsequent files rather than"infer".paths.gzip.BadGzipFile: Not a gzipped file,UnicodeDecodeError, or produced corrupted data.Solution
_get_read_detailsand_get_write_detailsto operate on a copy ofpandas_kwargsper-file, preservingpandas_kwargsacross multiple files so each file infers its own compression setting independently.tests/unit/test_moto.pycoveringread_csvandread_jsonwith mixed compression paths and chunked reads.Validation Results
pytest tests/unit/test_moto.py: All 47 tests passed.git diff --check: Clean (0 errors).ruff check: All checks passed.ruff format --check: All files formatted.mypy: Success (0 errors).