fix(rules): ride out transient failures while polling, and resume a request by id - #468
Conversation
…equest by id `rule create` and `rule improve` gave up on the first transient failure while polling a submitted request or fetching the rules it produced, and the only retry was running the command again, which submitted a second request while the first could still finish. A retryable `unavailable` (network failure, 408, 429, 5xx) is now retried on the next poll, up to 8 in a row; any other answer resets the count, and a non-retryable one still fails at once. Both commands take `--resume <requestId>` to pick up a request without resubmitting it, and every give-up names the request id, says it may still complete, and gives that command.
|
@claude /review |
|
Claude finished @theCodeDrift's task in 30m 18s —— View job Review completeReview mode: incremental — read 0 prior review thread(s) before reviewing.
Findings[New] One finding posted inline on Everything else checked out
Per the review instructions, build/lint/tests were not run locally — CI reports those separately. |
…--resume command The service's requestId was accepted as any object field and printed inside the --resume command the recipes tell an agent to run, so server text could reach a command line. Submit and iterate now accept a body only when requestId is a UUID, sharing one check with --resume's own validation; anything else is an invalid response body.
Thanks for the review. Its one finding, the unsanitized server — AI Coding Agent |
rule createandrule improvegave up on the first transient failure while polling a submitted request. The only retry was running the command again, which submitted a second generation request while the first could still finish. Fetching the rules after generation had the same problem, and a give-up there threw away a generation that had already succeeded.What changes
unavailableoutcome markedretryable(network failure,408,429,5xx) is reported as progress and retried after the 15s poll interval. The command fails only after 8 in a row (about 2 minutes). Any other answer resets the count. A non-retryableunavailable(undocumented4xx, malformed body), a documentederror, a refusal, or a401still fails on the first occurrence.--resume <requestId>on both commands polls an already-submitted request and fetches, verifies, and writes what it produced, without submitting anything. It refuses--fromand any value that is not a UUID withINVALID_INPUT, before calling the service. This also recovers a generated rule whose fetch gave up, without generating it again.--resumecommand. The fetch give-up says no rules were written, and no longer says "try again", since that would regenerate the rule.create-remote-rulev5,improve-rulev7): onNETWORK_ERROR, the recipe now says to run the named--resumecommand, and to confirm with the user before re-running with--from, which submits a new request. TherequestIdnotes say--resumeis the one command that takes it back.openspec/specs/cli-rules/spec.mdby hand, with no OpenSpec change: transient-failure tolerance, and--resume.Tests
test/rule-poll-transient.test.ts(12 tests) drives the real command against the v2 stub. The stub can now return a queue of HTTP statuses, network failures, andbuildinganswers. The tests cover:401;--resumeafter a polling give-up and after a fetch give-up (exactly onePOSTacross both runs);improve --resumemaking no iterate call;INVALID_INPUTrefusals.Setting the budget to 1 (the old behavior) fails every retry test.
Fixes #466