Skip to content

[SCH-7477] cancel in-flight prewarm before releasing leases on shutdown - #119

Merged
bpapillon merged 1 commit into
mainfrom
orbit/sch-7477-sdk-shutdown-does-not-cancel-in-flight-prewarm-before
Sep 17, 2026
Merged

bpapillon merged 1 commit into
mainfrom
orbit/sch-7477-sdk-shutdown-does-not-cancel-in-flight-prewarm-before

Conversation

@schematic-orbit

Copy link
Copy Markdown
Contributor

shutdown() stopped the sweeper and released the per-process leases, but left the fire-and-forget prewarm work running. A prewarm acquire that landed after release_all_local_leases() listed the store installed a lease nothing was left to release, so its credits stayed held against the company balance until the server expired it — five minutes by default.

The load-bearing part is tracking the single-flight acquire. _single_flight registers the task and then await asyncio.shield(task), so cancelling a prewarm cancels the shield, not the acquire underneath, and the finally drops the registry entry as the caller unwinds. The acquire was therefore tracked nowhere at all and went on to call _lease_store.replace() behind the release. It now joins the manager's drain set alongside the registry entry: the registry dedupes concurrent callers, the drain set is what a close waits on, and they have different lifetimes.

shutdown() now goes stop → cancel → drain → release. It cancels the prewarms in _background_tasks, awaits LeaseManager.drain() (bounded by a new internal SHUTDOWN_DRAIN_TIMEOUT of 5s, since a shutdown that hangs is worse than a hold the server expires in DEFAULT_LEASE_DURATION), and only then releases. Cancel and drain run for a shared backend too — that work must not outlive the client — while the release stays gated on not self._lease_backend_shared. _spawn is a no-op past stop(), so extend_in_background starts nothing new, and _prewarm_one returns early while the client is shutting down, which covers a caller awaiting client.prewarm(...) directly rather than through identify(prewarm=...).

Only the in-memory store releases on shutdown, so this is single-process deployments; a Redis-backed store still releases nothing.

Before, with an acquire on the wire when shutdown begins:

>       assert client._lease_store.list_leases() == []
E       AssertionError: assert [LeaseState(lease_id='lse_1', company_id='co_1', ...)] == []
E         Left contains one more item

After, the lease is drained, found, and released — release_credit_lease is awaited for lse_1 and the store is empty.

Notes, not in the diff:

  • The Node half of the ticket (close() in schematic-node) is not here — one repo per change. It needs a closing flag resolveCompanyIdWithWait's poll loop checks, a tracked set for the voided this.prewarm(...) in identify, and a drain() on CreditLeaseManager awaiting inflightAcquire/inflightExtend plus the voided release and background extend. Promises are not cancellable, so Node drains where Python cancels.
  • A check() racing shutdown() can still install a lease after the release. Gating acquire_if_needed on _stopped would change what a mid-shutdown check returns (fail-closed denies), which is a call for a human and not what this ticket claims.
  • No conformance vector: conformance/SPEC.md's manager-level ops have no prewarm and no client-close op, so covering this there means a spec change shared with schematic-node.

https://linear.app/schematic/issue/SCH-7477/sdk-shutdown-does-not-cancel-in-flight-prewarm-before-releasing-leases

shutdown() stopped the sweeper and released the per-process leases, but
left the fire-and-forget prewarm work running. A prewarm acquire that
landed after release_all_local_leases() listed the store installed a
lease nothing was left to release, holding its credits against the
company balance until server-side expiry.

The load-bearing part is tracking the single-flight acquire. Cancelling
a prewarm cancels its asyncio.shield, not the task underneath, and drops
the registry entry as it unwinds, so the acquire was tracked nowhere at
all. It now joins the manager's drain set alongside the registry, which
dedupes concurrent callers and has a different lifetime.

Shutdown cancels the tracked prewarms, drains the manager (bounded by
SHUTDOWN_DRAIN_TIMEOUT, since a shutdown that hangs is worse than a hold
the server expires), and only then releases. Cancel and drain run for a
shared backend too; only the release stays gated on it. _spawn is a
no-op past stop(), and _prewarm_one returns early while the client is
shutting down, for a caller awaiting prewarm() directly.

Only the in-memory store releases on shutdown, so this is single-process
deployments.
@bpapillon bpapillon self-assigned this Sep 16, 2026
@bpapillon
bpapillon marked this pull request as ready for review September 16, 2026 23:48
@bpapillon
bpapillon requested a review from a team as a code owner September 16, 2026 23:48
@bpapillon
bpapillon merged commit cbe3c49 into main Sep 17, 2026
6 checks passed
@bpapillon
bpapillon deleted the orbit/sch-7477-sdk-shutdown-does-not-cancel-in-flight-prewarm-before branch September 17, 2026 00:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants