Skip to content

fix(remote-cache): wait out the Registry's request limits and honour its 100 MiB artifact cap - #950

Open
tolgaergin wants to merge 9 commits into
mainfrom
remote-cache/request-limits
Open

tolgaergin wants to merge 9 commits into
mainfrom
remote-cache/request-limits

Conversation

@tolgaergin

@tolgaergin tolgaergin commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

Companion to tolgaergin/a-package-manager#233. The LPM.dev Registry now limits remote cache requests per account, answers a request past a limit with 429 and Retry-After, and stores artifacts up to 100 MiB.

What changes

crates/lpm-cli/src/commands/remote_cache.rs:

  • A 429 pauses only its kind of request, and only until Retry-After. The wait comes from the Retry-After header (seconds or an HTTP date), else the body's retryAfter, else a minute, and is kept between a second and a day. The lookup or upload is treated as a miss or skipped, the task builds as it would without a remote cache, and the CLI uses the cache again once the wait ends.
  • A lookup limit pauses uploads too. A task that skipped its lookup would likely upload an artifact the cache already holds, which costs the account an upload and spends the upload limits for nothing.
  • Pauses are kept per cache server and team for the whole run. Each task builds its own client, so the state is process-wide, but a 429 from one cache or team doesn't pause another.
  • Each kind of limit warns once, on its own, apart from the other remote cache warnings, for example Remote cache limit reached: 300 uploads a minute; skipping remote cache uploads for 42 seconds. Before this, any earlier warning hid the limit's.
  • An upload refused for a moment goes once more. A 429 with a short Retry-After (5 s at most), as the Registry answers when the account has as many uploads in flight as it may or another upload holds its cache, makes the CLI wait and send the upload again instead of dropping it and pausing other uploads. The wait adds up to half again, from the artifact's key hashed with a seed drawn afresh each run, so uploads refused together don't retry together, even in parallel CI jobs refused on the same key, and the upload doesn't go again if uploads or lookups were paused meanwhile. A longer wait is a spent limit and pauses uploads.
  • A lookup refused for a moment goes once more too, after the wait and its spread, unless lookups were paused meanwhile: a cache hit is worth a second's wait more than a rebuild.
  • Requests skipped while the cache is busy get their own notice, once a run for each kind, as a skipped row in LPM's own words: remote cache is busy; skipping some uploads for a moment (or lookups). An upload skipped because a lookup's short pause pauses uploads too gets the uploads notice, since the lookups notice alone wouldn't say so. It isn't given for a request a spent limit already pauses, whose warning covers it, nor, for uploads, by a read-only client. The CLI records which pauses are spent limits (a wait over 5 seconds) when it takes them, rather than guessing from the time left, so a spent limit's last seconds don't read as a busy server.
  • Requests the cache server refuses stop for the run: a 403 on an upload stops uploads, since the token can't write there and every later upload would only spend the server's limit for refused requests; a 403 on a lookup stops lookups and uploads. The warning says so once.
  • An upload answered 409 ends quietly. Another upload of the same artifact is storing it, as with parallel CI jobs, and serves the next lookup.
  • Artifact size: the 100 MiB cap applies only to the LPM.dev Registry, where the CLI skips uploading a larger archive and treats a larger download as a miss. Other cache servers keep the 500 MiB cap. Messages state both in MiB, as the Registry and the docs do.
  • lpm cache status reports a 429 as the limit and when to try again.

httpdate becomes a direct dependency; it was already in the lockfile through hyper.

The docs follow in lpm-dev/rust-client-docs#389.

Testing

Workflow tests against a mock cache server:

  • a lookup limit skips the run's later lookups and uploads (1 GET, 0 PUTs);
  • an upload limit skips later uploads while lookups continue;
  • the cache is used again once Retry-After has passed;
  • a limit warns even after another remote cache warning;
  • an artifact over 100 MiB from the Registry is a miss, while another cache server's is read up to 500 MiB (a raw TCP server advertises the size, so no test moves that much data);
  • lpm cache status reports the limit and the wait.

Unit tests cover Retry-After parsing, the wait's bounds and fallbacks against a real HTTP response, pause ordering, per-kind warnings, per-server-and-team state, the upload cap, a download streamed past the cap with no Content-Length, an upload refused for a moment that goes once more and pauses nothing, and uploads pausing after a refused retry or a long wait.

Mutation checks: 47 mutations across eight rounds, all caught. Round 3 adds tests for the 5-second boundary, the pause re-check, the spread and the warning rule; round 4 for the spread at the call site, the boundary for both kinds, the re-check seeing a lookup pause, the busy notice and its own flag, and the quiet 409; round 5 for the lookup retry and its re-check, the notice's text, and no notice during a spent limit; round 6 for spent limits told from busy pauses (uploads covered by a lookup limit, a long limit marked spent, a busy pause not), nothing paused while a lookup waits, the lookup's 5-second boundary and spread, and the busy notice when a lookup's retry is cut short; round 7 for the notice on uploads and lookups skipped while paused, none for a read-only client, and jitter that holds per key for the run; round 8 for uploads and lookups stopped after a 403, and a refused lookup stopping uploads too.

The whole CI gate passes locally on d14cfda, with every workspace member rebuilt from this branch:

  • cargo clippy --workspace --all-targets --locked -- -D warnings, cargo fmt --check, cargo build --workspace --locked
  • 7,148 workspace tests
  • cargo test -p lpm-cli --bin lpm-rs -- --test-threads=1: 5,435 passed
  • 119 binary-surface tests
  • the run and cache workflow tests with the hermetic CLI: 179 passed

🤖 Generated with Claude Code

tolgaergin and others added 2 commits October 11, 2026 01:41
…MiB artifact cap

The LPM.dev Registry now limits remote cache uploads and cache hits per
account and answers a request past a limit with 429. After a 429 the CLI
skips that kind of request, lookups or uploads, for the rest of the run
instead of sending more the limit would refuse, and warns once with the
Registry's message. Tasks build as they would without a remote cache.

The Registry stores artifacts up to 100 MiB, so the CLI skips uploading,
and treats as a miss, anything larger.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…he run

A 429 now pauses only its kind of request, lookups or uploads, until
the time Retry-After gives (the header in seconds or as a date, else
the body's retryAfter, else a minute; between a second and a day). The
CLI then uses the remote cache again. Pauses are kept per cache server
and team for the run, since each task builds its own client.

A lookup limit pauses uploads too: a task that skipped its lookup would
likely upload an artifact the cache already holds.

Each kind of limit warns once on its own, so an earlier remote cache
warning no longer hides it, and the warning says how long the CLI
skips that kind of request. `lpm cache status` reports a 429 as the
limit and when to try again.

The 100 MiB artifact cap applies only to the LPM.dev Registry; other
cache servers keep 500 MiB.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@tolgaergin tolgaergin changed the title feat(remote-cache): honour the Registry's request limits and its 100 MiB artifact cap fix(remote-cache): wait out the Registry's request limits and honour its 100 MiB artifact cap Oct 11, 2026
tolgaergin and others added 7 commits October 11, 2026 05:53
…s in MiB

An upload refused with a short Retry-After (5 s at most), as the
Registry answers when the account has as many uploads in flight as it
may or another upload holds its cache, now waits and goes once more
instead of being dropped and pausing other uploads. A longer wait is a
spent limit and pauses uploads as before.

Size limits read "100 MiB" and "500 MiB", as the Registry and the docs
state them. The size check and the upload are their own methods, with
unit tests for the upload cap, a streamed download past the cap, and
the retry.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rnings for real limits

- An upload retried after a short refusal waits up to half the wait
  again, by its artifact's key, so uploads refused together don't all
  retry together. It doesn't go again if uploads were paused while it
  waited.
- A pause of 5 seconds or less no longer uses up its kind's one
  warning, which a longer limit later needs.
- Tests pin the 5-second boundary, the re-check and the spread.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…leave a conflicting upload to the other

A refusal of 5 seconds or less gets its own once-a-run notice, apart from
the warning for a spent limit, including an upload dropped because uploads
were paused while it waited to retry. An upload answered 409, which another
upload of the same artifact is storing, ends quietly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… wait, and note busy skips in LPM's own words

A lookup refused for 5 seconds or less goes once more after the wait, plus
its share of the spread, as uploads do, unless lookups were paused
meanwhile; a cache hit is worth the wait. The busy notice is a skipped row
in LPM's own words, and only for requests not already paused by a spent
limit, whose warning covers them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… retries per run

The busy notice was held back while a pause had more than 5 seconds
left, so a spent limit's last seconds read as a busy server. The CLI now
records which pauses are spent limits when it takes them. Retry jitter
hashes the artifact's key with a seed drawn each run, so parallel CI jobs
refused on the same key don't retry together. Tests cover nothing paused
while a lookup waits, the lookup's 5-second boundary and spread, and the
busy notice when a lookup's retry is cut short.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An upload skipped because a lookup's short pause also pauses uploads
now gets the once-a-run busy notice; a spent limit keeps its warning
instead, and a read-only client stays silent. Lookups skipped while
paused take the same path. Retry jitter is tested to hold per key for
the run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e run

A 403 on an upload stops uploads for the run: the token can't write
there, and every later upload would only spend the server's limit for
refused requests. A 403 on a lookup stops lookups and uploads.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant