rule create and rule improve give up after a single transient failure while they poll a submitted request. Retrying means running the command again, which submits a new generation request even though the first one is still running on the service.
Line numbers are against origin/main at 416a2db, under packages/cli/src/.
Current behavior
awaitRequest() in rules/generate.ts:69 polls every 15 seconds (POLL_INTERVAL_MS, :38). The first unavailable outcome throws NETWORK_ERROR (:109-111). That covers a network failure, a 408, a 429, or a 5xx. Since #462 the message adds "This is usually temporary; try again in a few minutes."
There is no way to resume polling a request, so "try again" means running rule create or rule improve again (commands/rules.ts:343, :490). That submits a second request while the first one may still finish.
Fetching the rule once generation succeeds has the same shape. servedOrThrow() (rules/generate.ts:252) throws on the first unavailable, after the rule has already been generated.
#462 deliberately left this out because it changes behavior. It made these outcomes carry retryable but kept them failing on the first error.
Expected
- While polling, an
unavailable outcome with retryable: true is tolerated and the next poll goes ahead. The command fails only when the failures persist, for example after N consecutive failures or a time budget. The limit should also apply to 429, so the CLI does not hammer a rate-limited service.
- A non-retryable
unavailable (an undocumented 4xx or a malformed body), an error, or unauthorized still fails immediately, as now.
- Fetching the rule after generation tolerates transient failures the same way.
- If the command still gives up, the message says the request may still complete and names the request id. It should not imply that resubmitting is the only option.
Open question
Should there be a way to resume a request by id, such as rule create --resume <requestId>? That would make giving up safe even with retries in place, but it is a larger change and could be a separate issue.
Refs #454
rule createandrule improvegive up after a single transient failure while they poll a submitted request. Retrying means running the command again, which submits a new generation request even though the first one is still running on the service.Line numbers are against
origin/mainat416a2db, underpackages/cli/src/.Current behavior
awaitRequest()inrules/generate.ts:69polls every 15 seconds (POLL_INTERVAL_MS,:38). The firstunavailableoutcome throwsNETWORK_ERROR(:109-111). That covers a network failure, a408, a429, or a5xx. Since #462 the message adds "This is usually temporary; try again in a few minutes."There is no way to resume polling a request, so "try again" means running
rule createorrule improveagain (commands/rules.ts:343,:490). That submits a second request while the first one may still finish.Fetching the rule once generation succeeds has the same shape.
servedOrThrow()(rules/generate.ts:252) throws on the firstunavailable, after the rule has already been generated.#462 deliberately left this out because it changes behavior. It made these outcomes carry
retryablebut kept them failing on the first error.Expected
unavailableoutcome withretryable: trueis tolerated and the next poll goes ahead. The command fails only when the failures persist, for example after N consecutive failures or a time budget. The limit should also apply to429, so the CLI does not hammer a rate-limited service.unavailable(an undocumented4xxor a malformed body), anerror, orunauthorizedstill fails immediately, as now.Open question
Should there be a way to resume a request by id, such as
rule create --resume <requestId>? That would make giving up safe even with retries in place, but it is a larger change and could be a separate issue.Refs #454