fix(serde): resolve persisted parquet paths relative to the configured directory - #945
Open
elijahbenizzy wants to merge 1 commit into
Open
elijahbenizzy wants to merge 1 commit into
elijahbenizzy wants to merge 1 commit into
Conversation
…d directory The pandas DataFrame serde now records only the parquet file name in persisted state and resolves it against pandas_kwargs["path"] on load. The resolved path must land inside that directory; URL-style values and paths resolving outside it raise ValueError. Absolute paths recorded by earlier versions still load when they resolve inside the configured directory. Deserialization now requires the same `path` entry serialization already required.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The pandas serde stored the full joined path of the parquet file in persisted state and passed whatever came back out of state straight to
pd.read_parquet. This change records only the file name and, on load, resolves it against the configuredpandas_kwargs["path"]directory, rejecting values that resolve outside that directory or that look like URLs.Behavior notes:
pathentry inpandas_kwargsthat serialization already required. All built-in persisters pass identicalserde_kwargsin both directions, so existing configurations are unaffected.Tests cover round-trip, nested relative directories, legacy absolute paths, and the rejected shapes.