Conversation
The Chat service sends an SSE heartbeat every 15 seconds, but the reader never checked for it. When the connection stalled without closing, the reader blocked in reader.read() forever and the pending reply kept showing its last phase with no sign that nothing was arriving. Each attempt now aborts after 45 seconds without any byte (three missed heartbeats) and resumes from its cursor through the existing bounded retry. While it reconnects the pending reply shows a local phase that never moves the cursor. A caller abort still ends the stream without a retry. Signed-off-by: song <liusongstep@gmail.com>
A stub stream sends one event and then goes silent. The reader must resume from that event's cursor, report the reconnect phase, complete the Turn, and still end immediately on a caller abort. Signed-off-by: song <liusongstep@gmail.com>
There was a problem hiding this comment.
English verdict: REQUEST_CHANGES
Reviewed exact head: b49e9e121ffa33f778b48c2ec4e9f33164e8f45d.
阻塞项:[P2] 取消发生在重试间隙或调用前时仍发起新 GET;[P2] watchdog 四次耗尽返回原始 AbortError,丢失现有 wrapper 的 typed recovery 信息。定位:attempt signal、第四次异常抛出。
动机
这个问题值得修:SSE 连接已收到 headers 或第一帧,随后既没有数据也没有 heartbeat,原 reader 会一直等。用户需要恢复的是已经接受的同一个 Turn,从最后 cursor 继续读取,不是再提交一轮。可用恢复还应保留真实的用户取消语义,并在有限重试失败后保留既有 wrapper 的同一 Turn typed recovery 信息。head 确实恢复了第一次静默连接,但取消与耗尽错误边界有实际回归,因此当前完整目标尚未交付。
改动思路
改动放在现有 TypeScript streamChatTurn owner:每次 fetch/read 用自己的 AbortController,默认四十五秒无字节则中止连接;复用原来的最多四次 retry 与 after cursor,并发送没有 event_id 的本地重连 phase,避免推进服务端 cursor。heartbeat 字节会续期。复用现有 reader 是合理的,既不需要新 provider,也不需要新增持久化状态。问题在于同一个 abortAttempt 同时服务 watchdog 与父 signal,而父 signal 的既有 aborted 状态和最终错误类型没有完整传递到新控制器、退避以及现有 wrapper。
具体改动
完整 PR 三个文件,+85/-1:chat.ts 的三十行生产增量、五十四行薄 reader smoke 和一个 npm 命令注册。它是既有 caller 的默认 deadline 变更,stallTimeoutMs 可选参数不表示 default-off 或只对 opt-in 用户生效。
关键代码讲解
- streamChatTurn:入口仍是 eventsUrl、onEvent 与可选父 AbortSignal;每轮创建 attempt,再监听父 abort,以 attempt.signal 发起 GET。监听只接收未来事件,不能重放调用前或退避期间已经发生的 abort。这会使新一轮 GET 脱离已取消的父生命周期。
- armStallTimer:先在 fetch 前启动,再在 reader.read 前重置,因此覆盖无 headers 和无后续字节两种静默;timer 最终清理。无字节只是传输故障,不意味着后台 Turn 被用户中断。id-empty 本地 phase 不更新 cursor,这个正向行为已验证。
- receiveChatTurnStreaming:既有 wrapper 仅在错误是 ChatApiError 且父 signal 未取消时,补充 reconnectable 与 Session/Turn/events 信息。第四次 watchdog 的原始 DOMException/AbortError 不满足该分支,所以新故障类型失去既有恢复信息。
- Dashboard failure projection:既有 catch 会把 pending 设为 false,展示 error.message,并计算 payload.reconnectable 对应的消息标记。进一步追到现有 WorkspaceTimelineItem 映射后确认:该标记没有传给 PersonalWorkspace,当前没有已接通的重连按钮,base/head 都如此。它没有吞异常或让 pending 永远为 true;这项旧 UI 缺口不归因于本 PR。
对主干的风险
[P2] 取消边界:用相同独立 fixture、实际 HTTP server、原生 fetch/ReadableStream 和生产导出测试。第一 GET 返回 503,三十毫秒后取消父 signal;原退避二百五十毫秒保持不变。base 只有一次请求并以 AbortError 结束;head 在取消后仍发起第二次 GET,收到 terminal 并 resolve。预先 aborted 的输入,base 零请求,head 一请求并交付 terminal。read 已进行时取消在两边都能结束,说明 native smoke 的同步 onEvent abort 正例不足以证明整个取消契约。
最小修改是每次 attempt 之前检查并传播父 aborted 状态,使 backoff 可取消,且取消后不再开启连接或交付 callback。请加入调用前、退避中与读取中的三个取消反例,而不是改变后台 Turn 的中断权限。
[P2] 重试耗尽:连续四个无字节连接、父 signal 始终未取消,生产 resume wrapper 返回原始 AbortError,reconnectable=false,没有 session_id/turn_id/events_url 恢复 metadata;普通四次 HTTP503 则仍保留 typed reconnectable 信息。这是新增 watchdog 分支与既有 public wrapper 错误契约不匹配。实际 React/Ego 也确认同一 Turn、after=1、四次 GET 后 pending 已结算并显示原始 BodyStreamBuffer abort;但“没有重连按钮”是现有 Timeline 映射未消费 reconnect 标记的旧问题,不作为本 PR blocker,也不要求此 PR 顺便新增按钮。最小修改是区分 watchdog 传输超时与真实用户 abort,并在 bounded exhaustion 使用既有 ChatApiError/session/turn/events 恢复契约。请验证生产 resume wrapper 的完整错误 payload,而不只断言 reader 会停止。
语义与 CI 对齐
当前共享契约分别是“父取消后停止传输活动”以及“非用户故障有限失败后保留同一 Turn 的 typed recovery 信息”。新 controller/timeout 分支违反了这两项,不需要另起 RFC 或词汇。修复后重跑 npm run smoke:chat-stream-stall、npm run smoke:chat-turn-acceptance-retry,并补上上述退避/预取消和四次静默 wrapper 错误契约反例;不要把旧 UI affordance 缺口冒充 PR 回归。
已亲自运行 exact-head npm run build(TS/Vite/source 与 packaged Chat assets)、新 stall smoke、既有 acceptance-retry smoke,均通过。独立八场景通过真实 HTTP/native stream 检查:首帧后静默与无 headers 能恢复;heartbeat 连续到达时只用一个连接;cursor 不被本地 phase 推进;普通 503 仍 bounded/typed。三个取消/耗尽 oracle 在 head 失败。仅将 watchdog 四十五秒缩到七十毫秒(HTTP)/八十毫秒(UI)作为测试时钟,原 retry 二百五十/五百/一千毫秒未改;base 静默用例仅由测试安全 abort 收尾,不能当作旧产品恢复能力。本轮没有查询或等待远端 CI,也不因 unrelated 红灯拒绝。
我的整体评价
归因更正:本评审已原地修订,明确区分新增 watchdog 的 typed error 缺失与未改动的 PersonalWorkspace 重连动作缺口;后者不要求本 PR 修复。
reader 的放置、同一 Turn cursor 与有限 retry 方向正确,问题规模与三十行生产机制相称;请求修改不是代码量或新架构要求。相邻 future-facing pass 建议在原 owner 用明确的取消/timeout原因和 existing typed error 统一监督,避免再增加平行 busy flag、transport wrapper 或协议版本;评审-only 没有代作者改代码。long_horizon 的已取消工作收敛、user_experience 的取消后不再交付陈旧响应存在已复现回归;非用户 timeout 的 typed wrapper 契约也需保住。保留第一处静默恢复与 heartbeat 正向收益,补齐真实取消和 typed exhaustion 后可继续重审。此证据仅覆盖隔离合成 HTTP 和真实客户端/UI,不声称 live 模型、长时间 soak 或完整仓库套件通过;本轮不修复、合并或升级。
Goal And Delivered Outcome
streamChatTurnnever checked for it. When the connection stalled without closing, the reader blocked inreader.read()forever; the pending reply kept its last phase (for exampleAgent 已开始处理) with no sign that nothing was arriving.Agent 已开始处理indefinitely; after it switches to连接中断,正在重连…once 45 seconds pass without any byte, and when the service resumes the reader continues from its cursor and completes the reply without duplicating it.main.Scope And Continuation
aftercursor. The reconnect notice is a localagent.phaseevent with no event id, so it never moves the cursor. A caller abort still ends the stream without a retry.Validation
staticpassednpx tsc --noEmitinapps/presentation/dashboard.unitpassednpm run smoke:chat-stream-stall: a stub stream sends one event then goes silent; the reader must resume with?after=<that event>, emit the reconnect phase, complete the Turn, and still end at once (no retry) on a caller abort.regression_paritypassedreal_entrypointpassedloopx serve-status+loopx chaton an isolated synthetic registry with a stub Codex app-server; the Chat process was paused with SIGSTOP 1.5 s into a Turn: 20 s →Agent 已开始处理, 50 s →连接中断,正在重连…; after SIGCONT at 75 s the full answer rendered within 1 s, once.integrationpassedmainalso pass here (all stream Turns through this reader).npm run smoke:chat-turn-acceptance-retrypasses.integrationnot_runnpm run smoke:chat-routeand six browser scenarios (goal-draft,capability-scope,steward-group-trigger,conversation-input,automation-cadence,steward-model-settings) already fail on a cleanmaincheckout.streamChatTurn; covered by the unit smoke, the real-service freeze and the browser scenarios. Not covered: a proxy that sends partial bytes slower than one byte per 45 s would still be treated as alive, which matches the heartbeat contract.Frontend / Visual Evidence
连接中断,正在重连…until the stream resumes or the bounded retry gives up with the existingAgent 事件流连接已断开。error.Type of Change
LoopX Area
Technical Direction
Shared-authority RFC fixture impact
N/A
Boundary Checklist
none.Signed-off-bytrailer (git commit -s).Future-facing refactor pass: considered a shared timeout option in
requestJson; kept the timer local to the SSE reader because only a stream has a heartbeat contract to measure against.