[SCH-7477] cancel in-flight prewarm before releasing leases on shutdown - #119
Merged
bpapillon merged 1 commit intoSep 17, 2026
Conversation
shutdown() stopped the sweeper and released the per-process leases, but left the fire-and-forget prewarm work running. A prewarm acquire that landed after release_all_local_leases() listed the store installed a lease nothing was left to release, holding its credits against the company balance until server-side expiry. The load-bearing part is tracking the single-flight acquire. Cancelling a prewarm cancels its asyncio.shield, not the task underneath, and drops the registry entry as it unwinds, so the acquire was tracked nowhere at all. It now joins the manager's drain set alongside the registry, which dedupes concurrent callers and has a different lifetime. Shutdown cancels the tracked prewarms, drains the manager (bounded by SHUTDOWN_DRAIN_TIMEOUT, since a shutdown that hangs is worse than a hold the server expires), and only then releases. Cancel and drain run for a shared backend too; only the release stays gated on it. _spawn is a no-op past stop(), and _prewarm_one returns early while the client is shutting down, for a caller awaiting prewarm() directly. Only the in-memory store releases on shutdown, so this is single-process deployments.
schematic-bot
approved these changes
Sep 17, 2026
bpapillon
deleted the
orbit/sch-7477-sdk-shutdown-does-not-cancel-in-flight-prewarm-before
branch
September 17, 2026 00:08
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
shutdown()stopped the sweeper and released the per-process leases, but left the fire-and-forget prewarm work running. A prewarm acquire that landed afterrelease_all_local_leases()listed the store installed a lease nothing was left to release, so its credits stayed held against the company balance until the server expired it — five minutes by default.The load-bearing part is tracking the single-flight acquire.
_single_flightregisters the task and thenawait asyncio.shield(task), so cancelling a prewarm cancels the shield, not the acquire underneath, and thefinallydrops the registry entry as the caller unwinds. The acquire was therefore tracked nowhere at all and went on to call_lease_store.replace()behind the release. It now joins the manager's drain set alongside the registry entry: the registry dedupes concurrent callers, the drain set is what a close waits on, and they have different lifetimes.shutdown()now goes stop → cancel → drain → release. It cancels the prewarms in_background_tasks, awaitsLeaseManager.drain()(bounded by a new internalSHUTDOWN_DRAIN_TIMEOUTof 5s, since a shutdown that hangs is worse than a hold the server expires inDEFAULT_LEASE_DURATION), and only then releases. Cancel and drain run for a shared backend too — that work must not outlive the client — while the release stays gated onnot self._lease_backend_shared._spawnis a no-op paststop(), soextend_in_backgroundstarts nothing new, and_prewarm_onereturns early while the client is shutting down, which covers a caller awaitingclient.prewarm(...)directly rather than throughidentify(prewarm=...).Only the in-memory store releases on shutdown, so this is single-process deployments; a Redis-backed store still releases nothing.
Before, with an acquire on the wire when shutdown begins:
After, the lease is drained, found, and released —
release_credit_leaseis awaited forlse_1and the store is empty.Notes, not in the diff:
close()inschematic-node) is not here — one repo per change. It needs aclosingflagresolveCompanyIdWithWait's poll loop checks, a tracked set for the voidedthis.prewarm(...)inidentify, and adrain()onCreditLeaseManagerawaitinginflightAcquire/inflightExtendplus the voided release and background extend. Promises are not cancellable, so Node drains where Python cancels.check()racingshutdown()can still install a lease after the release. Gatingacquire_if_neededon_stoppedwould change what a mid-shutdown check returns (fail-closed denies), which is a call for a human and not what this ticket claims.conformance/SPEC.md's manager-level ops have no prewarm and no client-close op, so covering this there means a spec change shared with schematic-node.https://linear.app/schematic/issue/SCH-7477/sdk-shutdown-does-not-cancel-in-flight-prewarm-before-releasing-leases