Skip to content

[oss] Support refreshing the REST data token for open OSS streams - #10178

Open
sundapeng wants to merge 6 commits into
apache:masterfrom
sundapeng:oss-credentials-provider
Open

sundapeng wants to merge 6 commits into
apache:masterfrom
sundapeng:oss-credentials-provider

Conversation

@sundapeng

@sundapeng sundapeng commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Purpose

With a REST catalog that vends data tokens, RESTTokenFileIO builds the OSS FileIO with the token it holds at that moment. A refresh only reaches FileIOs created afterwards, so a stream that is already open, such as a long read or a multipart upload, keeps signing with the old token and fails once it expires.

OSSFileIO now signs each request through RESTTokenCredentialsProvider, which reloads the table's data token from the catalog before it expires, as Iceberg's VendedCredentialsProvider does for S3. A token is reloaded an hour before it expires, or halfway through when it lives shorter; while it is still valid, one caller reloads it and the others keep using it, and a failed reload is retried after 10 seconds. Since a delegate now reloads the token as a table and a catalog user, it is shared only by callers with the same ones. fs.oss.credentials.provider can now also replace the access keys.

Tests

RESTTokenRefresherTest: 6 new cases. RESTTokenCredentialsProviderTest: 3 new cases. RESTTokenFileIOOnOSSTest: 2 new cases. RESTTokenFileIOTest: 2 new cases, 12 tests passed. OSSLoaderTest: 1 new case.

…REST tokens

RESTTokenFileIO bakes the vended token into the delegate FileIO's options, so a
stream that is already open keeps signing with the token it was opened with and
fails once that token expires.

OSSFileIO now resolves credentials per request through a provider that reads the
current token from RESTTokenFileIO via CredentialsSupplierRegistry. OSSLoader
also accepts fs.oss.credentials.provider in place of the access keys.
@JingsongLi

Copy link
Copy Markdown
Contributor

Reviewed 93f7fc7. The REST-token/open-stream problem has clear end-to-end value, and the per-request OSS signing path is plausible. I found one production blocker in the supplier lifecycle:

[P1] Keep the supplier alive for the lifetime of an open OSS stream. CachedFileIO.close() unregisters its supplier when FILE_IO_CACHE evicts a delegate (10-hour access expiry or 1,000-entry size limit). The delegate returned by FileIO.get(oss://...) is OSSLoader.OSSPluginFileIO; PluginFileIO does not override FileIO.close(), whose default is a no-op. Thus eviction removes the supplier while the underlying OSS client and any open streams may still be live. RegisteredCredentialsProvider then silently falls back to its last credentials, and the next token expiry recreates the same mid-read failure this PR targets. The underlying uncached OSS filesystem is also not closed. Please make wrapper close reach the delegate and tie supplier removal to the actual lifetime of streams/requests, with a regression test that evicts the cache while a stream remains open and crosses a token refresh.

Validation: RegisteredCredentialsProviderTest (4), OSSLoaderTest (1), and RESTTokenFileIOTest (11) passed locally; git diff --check passed. CI run 36106203387 failed only in JDK 8/11 Core jobs because the unrelated S3FileIOTest could not pull the pinned MinIO image; its other selected jobs passed. The first local REST test attempt hit macOS sandbox denial of Mockito agent attachment; rerunning it with attachment allowed passed all 11 tests.

…uses it

Evicting a delegate from RESTTokenFileIO's cache unregistered its supplier even
though the OSS client and its open streams were still live, so they fell back to
their last credentials and failed at the next expiry.

The registry now holds suppliers weakly. OSSFileIO and RegisteredCredentialsProvider
keep a strong reference once they resolve one, so a supplier lives exactly as long
as a client or stream that can still sign with it.
@sundapeng sundapeng closed this Sep 28, 2026
@sundapeng sundapeng reopened this Sep 28, 2026
Passing a live supplier from RESTTokenFileIO to the OSS client through a static
registry tied the supplier to objects the client does not own, and got the
lifetime wrong on cache eviction.

Like Iceberg's VendedCredentialsProvider and Hadoop's credential providers, the
OSS provider is now built from configuration alone: RESTTokenFileIO names the
table and the token expiry in the delegate options, OSSFileIO hands the catalog
options to RESTTokenCredentialsProvider, and the provider loads the table's token
through RESTTokenRefresher when it is about to expire. The registry is removed.
… clock

The test waited on the wall clock for the first token to enter the refresh
window, so on a slow CI runner the refresh happened before the first read.
It now moves the refreshers' clock forward instead.
@sundapeng sundapeng closed this Sep 28, 2026
@sundapeng sundapeng reopened this Sep 28, 2026
RESTTokenRefresher held its lock across the catalog request, whose client may
retry for minutes, so every OSS request on the client waited while the current
token was still valid. Now one caller reloads and the others keep the valid
token; with an expired token, callers fail fast within the retry interval.

A reloaded token is layered over the catalog options as RESTTokenFileIO does,
OSSFileIO reads only a constant from RESTTokenRefresher so an older
paimon-common still loads it, and OSSLoader names its option keys.
@sundapeng sundapeng changed the title [oss] Support credentials provider so open streams pick up refreshed REST tokens [oss] Fix open OSS streams failing once the REST data token expires Sep 28, 2026
@sundapeng sundapeng changed the title [oss] Fix open OSS streams failing once the REST data token expires [oss] Support refreshing the REST data token for open OSS streams Sep 28, 2026
@sundapeng

Copy link
Copy Markdown
Member Author

Thanks, good catch. Instead of fixing the supplier's lifetime I removed the registry: RESTTokenCredentialsProvider is now built from the delegate's options and reloads the table's data token from the catalog on its own, like Iceberg's VendedCredentialsProvider for S3. Nothing ties it to RESTTokenFileIO or its cache, so a stream that outlives its cache entry keeps getting fresh tokens. RESTTokenFileIOOnOSSTest evicts the cache while a stream is open, crosses a reload and checks the next read is signed with the new token. CI is green on 002bdcc.

The evicted OSS FileSystem not being closed predates this PR: RESTTokenFileIO sets file-io.allow-cache=false and PluginFileIO.close() is a no-op. Forwarding close() would close the FileSystem under open streams on eviction, for every plugin FileIO, and closing only after the last stream would need stream tracking in RESTTokenFileIO that hides VectoredReadable and misses callers of fileIO(). I'd like to handle it in a separate issue. Does that work for you?

RESTTokenFileIO cached delegates by token alone. Now that a delegate reloads the
token as a specific table and catalog user, another table or user holding the
same token reused it and reloaded as the wrong one. The cache key now also
covers the token's table and the catalog options.

RESTTokenRefresher reloaded whenever less than an hour was left, so a token that
lives shorter was reloaded on every OSS request. Each token now carries its own
reload time: an hour before expiry, or halfway through when it lives shorter,
and at least ten seconds after it arrived. An expired token is never returned.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants