Skip to content

[fs] Optimize OSS retry strategy for QpsLimitExceeded and 5xx errors - #10299

Merged
JingsongLi merged 3 commits into
apache:masterfrom
sundapeng:oss-throttle-retry
Sep 28, 2026
Merged

JingsongLi merged 3 commits into
apache:masterfrom
sundapeng:oss-throttle-retry

Conversation

@sundapeng

@sundapeng sundapeng commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

Purpose

The Aliyun OSS SDK retries GET and PUT, but never retries a POST other than InitiateMultipartUpload. Completing a multipart upload and deleting objects in batch are both POSTs, so under 503 QpsLimitExceeded closing a data file larger than fs.oss.multipart.upload.size, or deleting a directory, fails on the first throttled response and takes the Flink or Spark job down with it. The SDK's backoff also has no cap and no jitter, so parallel writers retry in lockstep and a late attempt can sleep for minutes.

OSSFileIO now sets its own retry strategy on each OSS client, covering every request it sends:

  • Retry 429, 500, 502, 503, 504 and network errors, including those POSTs, which are safe to repeat.
  • Back off exponentially, capped at 10s and jittered so that parallel writers spread their retries.

fs.oss.attempts.maximum still bounds the number of attempts, and fs.oss.enhanced-retry.enabled=false falls back to the SDK default retry.

Tests

OSSRetryStrategyTest: 7 new cases, 4 of them against a fake OSS endpoint; 24 tests in paimon-oss-impl passed.

The Aliyun OSS SDK sends every POST except InitiateMultipartUpload with
NoRetryStrategy, so a single 503 QpsLimitExceeded on
CompleteMultipartUpload or DeleteObjects fails the write or delete at
once, while GET and PUT on the same client ride it out. A retry strategy
set on the ClientConfiguration takes precedence over that override.

OSSFileIO now installs OSSRetryStrategy on each hadoop-aliyun client:
- Retry 429/500/502/503/504 and network errors on every request. The
  POSTs Paimon sends are safe to repeat: a completed upload answers
  NoSuchUpload, and DeleteObjects is idempotent.
- Back off 300ms * 2^n, capped at 10s and jittered, so throttled writers
  do not retry in lockstep. The SDK default grows to a 5 minute sleep.

fs.oss.attempts.maximum still bounds the number of attempts.
@sundapeng sundapeng changed the title [fs] Fix OSS file writes failing on the first QpsLimitExceeded [fs] Fix large OSS file writes failing on the first QpsLimitExceeded Sep 28, 2026
@sundapeng sundapeng changed the title [fs] Fix large OSS file writes failing on the first QpsLimitExceeded [fs] Optimize OSS retry strategy for QpsLimitExceeded and 5xx errors Sep 28, 2026
@JingsongLi

Copy link
Copy Markdown
Contributor

+1

@JingsongLi
JingsongLi merged commit f43c0d3 into apache:master Sep 28, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants