Skip to content

fix(chat): reconnect a Turn stream that stops delivering bytes - #5262

Open
songoow wants to merge 2 commits into
loopx-project:mainfrom
songoow:codex/chat-stream-stall-watchdog
Open

songoow wants to merge 2 commits into
loopx-project:mainfrom
songoow:codex/chat-stream-stall-watchdog

Conversation

@songoow

@songoow songoow commented Sep 28, 2026

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

  • Outcome basis / optional anchor: reproduced defect (no issue).
  • Goal/source and gap: the Chat service sends an SSE heartbeat every 15 seconds while a Turn runs, but streamChatTurn never checked for it. When the connection stalled without closing, the reader blocked in reader.read() forever; the pending reply kept its last phase (for example Agent 已开始处理) with no sign that nothing was arriving.
  • Observable before → after: with the Chat service process frozen mid-Turn, before the reply showed Agent 已开始处理 indefinitely; after it switches to 连接中断,正在重连… once 45 seconds pass without any byte, and when the service resumes the reader continues from its cursor and completes the reply without duplicating it.
  • Issue/task and intended base: none; base main.

Scope And Continuation

  • Completed scope and remaining work: complete within this scope. Each connection attempt aborts after 45 seconds without any byte (three missed heartbeats, response headers included) and goes through the existing bounded retry with the after cursor. The reconnect notice is a local agent.phase event with no event id, so it never moves the cursor. A caller abort still ends the stream without a retry.
  • Slice boundary / successor: none. The phase label follows the existing Chinese-only labels in this module; localizing stream phases belongs to the typed phase-code work, not this fix.

Validation

  • Tested revision: b49e9e1
  • Run state: finished
  • Input classes: synthetic
Check kind Result Public-safe evidence / limitation
static passed npx tsc --noEmit in apps/presentation/dashboard.
unit passed New npm run smoke:chat-stream-stall: a stub stream sends one event then goes silent; the reader must resume with ?after=<that event>, emit the reconnect phase, complete the Turn, and still end at once (no retry) on a caller abort.
regression_parity passed Failing-before check: with the source change reverted the smoke never settles (Node exits 13 on an unsettled top-level await).
real_entrypoint passed loopx serve-status + loopx chat on an isolated synthetic registry with a stub Codex app-server; the Chat process was paused with SIGSTOP 1.5 s into a Turn: 20 s → Agent 已开始处理, 50 s → 连接中断,正在重连…; after SIGCONT at 75 s the full answer rendered within 1 s, once.
integration passed 16 Personal Workspace browser scenarios that pass on main also pass here (all stream Turns through this reader). npm run smoke:chat-turn-acceptance-retry passes.
integration not_run npm run smoke:chat-route and six browser scenarios (goal-draft, capability-scope, steward-group-trigger, conversation-input, automation-cadence, steward-model-settings) already fail on a clean main checkout.
  • Coverage and gaps: the changed path is the per-attempt stall timer and reconnect notice in streamChatTurn; covered by the unit smoke, the real-service freeze and the browser scenarios. Not covered: a proxy that sends partial bytes slower than one byte per 45 s would still be treated as alive, which matches the heartbeat contract.

Frontend / Visual Evidence

  • UI impact: changed
  • Before: a stalled Turn keeps showing its last phase forever.
  • After: after 45 s of silence the pending reply reads 连接中断,正在重连… until the stream resumes or the bounded retry gives up with the existing Agent 事件流连接已断开。 error.
  • States and viewports shown: desktop 1440×900 reconnecting state (screenshots available on request; not attachable from the CLI).
  • Source data: synthetic
  • Attention review: no new element; the existing pending-phase line now tells the truth about the connection.

Type of Change

  • Bug fix
  • Test update

LoopX Area

  • Public docs or presentation surface (README, protocols, dashboard)

Technical Direction

  • Direction / acceptance reference, when applicable: Operator surface and IM integration.

Shared-authority RFC fixture impact

N/A

Boundary Checklist

  • Neither the diff nor this PR body/comments/attachments disclose private state, credentials, raw traces or verifier output, internal links, or local machine paths.
  • I did not duplicate maintainer-owned benchmark work unless a maintainer split out a public issue for it.
  • I kept the change scoped to the linked issue/task.
  • I completed the visual evidence section for UI changes, or marked UI impact none.
  • Every commit includes a DCO Signed-off-by trailer (git commit -s).

Future-facing refactor pass: considered a shared timeout option in requestJson; kept the timer local to the SSE reader because only a stream has a heartbeat contract to measure against.

The Chat service sends an SSE heartbeat every 15 seconds, but the reader
never checked for it. When the connection stalled without closing, the
reader blocked in reader.read() forever and the pending reply kept showing
its last phase with no sign that nothing was arriving.

Each attempt now aborts after 45 seconds without any byte (three missed
heartbeats) and resumes from its cursor through the existing bounded
retry. While it reconnects the pending reply shows a local phase that never
moves the cursor. A caller abort still ends the stream without a retry.

Signed-off-by: song <liusongstep@gmail.com>
A stub stream sends one event and then goes silent. The reader must resume
from that event's cursor, report the reconnect phase, complete the Turn,
and still end immediately on a caller abort.

Signed-off-by: song <liusongstep@gmail.com>

@huangruiteng huangruiteng left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

English verdict: REQUEST_CHANGES
Reviewed exact head: b49e9e121ffa33f778b48c2ec4e9f33164e8f45d.

阻塞项:[P2] 取消发生在重试间隙或调用前时仍发起新 GET;[P2] watchdog 四次耗尽返回原始 AbortError,丢失现有 wrapper 的 typed recovery 信息。定位:attempt signal、第四次异常抛出。

动机

这个问题值得修:SSE 连接已收到 headers 或第一帧,随后既没有数据也没有 heartbeat,原 reader 会一直等。用户需要恢复的是已经接受的同一个 Turn,从最后 cursor 继续读取,不是再提交一轮。可用恢复还应保留真实的用户取消语义,并在有限重试失败后保留既有 wrapper 的同一 Turn typed recovery 信息。head 确实恢复了第一次静默连接,但取消与耗尽错误边界有实际回归,因此当前完整目标尚未交付。

改动思路

改动放在现有 TypeScript streamChatTurn owner:每次 fetch/read 用自己的 AbortController,默认四十五秒无字节则中止连接;复用原来的最多四次 retry 与 after cursor,并发送没有 event_id 的本地重连 phase,避免推进服务端 cursor。heartbeat 字节会续期。复用现有 reader 是合理的,既不需要新 provider,也不需要新增持久化状态。问题在于同一个 abortAttempt 同时服务 watchdog 与父 signal,而父 signal 的既有 aborted 状态和最终错误类型没有完整传递到新控制器、退避以及现有 wrapper。

具体改动

完整 PR 三个文件,+85/-1:chat.ts 的三十行生产增量、五十四行薄 reader smoke 和一个 npm 命令注册。它是既有 caller 的默认 deadline 变更,stallTimeoutMs 可选参数不表示 default-off 或只对 opt-in 用户生效。

关键代码讲解

  1. streamChatTurn:入口仍是 eventsUrl、onEvent 与可选父 AbortSignal;每轮创建 attempt,再监听父 abort,以 attempt.signal 发起 GET。监听只接收未来事件,不能重放调用前或退避期间已经发生的 abort。这会使新一轮 GET 脱离已取消的父生命周期。
  2. armStallTimer:先在 fetch 前启动,再在 reader.read 前重置,因此覆盖无 headers 和无后续字节两种静默;timer 最终清理。无字节只是传输故障,不意味着后台 Turn 被用户中断。id-empty 本地 phase 不更新 cursor,这个正向行为已验证。
  3. receiveChatTurnStreaming:既有 wrapper 仅在错误是 ChatApiError 且父 signal 未取消时,补充 reconnectable 与 Session/Turn/events 信息。第四次 watchdog 的原始 DOMException/AbortError 不满足该分支,所以新故障类型失去既有恢复信息。
  4. Dashboard failure projection:既有 catch 会把 pending 设为 false,展示 error.message,并计算 payload.reconnectable 对应的消息标记。进一步追到现有 WorkspaceTimelineItem 映射后确认:该标记没有传给 PersonalWorkspace,当前没有已接通的重连按钮,base/head 都如此。它没有吞异常或让 pending 永远为 true;这项旧 UI 缺口不归因于本 PR。

对主干的风险

[P2] 取消边界:用相同独立 fixture、实际 HTTP server、原生 fetch/ReadableStream 和生产导出测试。第一 GET 返回 503,三十毫秒后取消父 signal;原退避二百五十毫秒保持不变。base 只有一次请求并以 AbortError 结束;head 在取消后仍发起第二次 GET,收到 terminal 并 resolve。预先 aborted 的输入,base 零请求,head 一请求并交付 terminal。read 已进行时取消在两边都能结束,说明 native smoke 的同步 onEvent abort 正例不足以证明整个取消契约。

最小修改是每次 attempt 之前检查并传播父 aborted 状态,使 backoff 可取消,且取消后不再开启连接或交付 callback。请加入调用前、退避中与读取中的三个取消反例,而不是改变后台 Turn 的中断权限。

[P2] 重试耗尽:连续四个无字节连接、父 signal 始终未取消,生产 resume wrapper 返回原始 AbortError,reconnectable=false,没有 session_id/turn_id/events_url 恢复 metadata;普通四次 HTTP503 则仍保留 typed reconnectable 信息。这是新增 watchdog 分支与既有 public wrapper 错误契约不匹配。实际 React/Ego 也确认同一 Turn、after=1、四次 GET 后 pending 已结算并显示原始 BodyStreamBuffer abort;但“没有重连按钮”是现有 Timeline 映射未消费 reconnect 标记的旧问题,不作为本 PR blocker,也不要求此 PR 顺便新增按钮。最小修改是区分 watchdog 传输超时与真实用户 abort,并在 bounded exhaustion 使用既有 ChatApiError/session/turn/events 恢复契约。请验证生产 resume wrapper 的完整错误 payload,而不只断言 reader 会停止。

语义与 CI 对齐

当前共享契约分别是“父取消后停止传输活动”以及“非用户故障有限失败后保留同一 Turn 的 typed recovery 信息”。新 controller/timeout 分支违反了这两项,不需要另起 RFC 或词汇。修复后重跑 npm run smoke:chat-stream-stall、npm run smoke:chat-turn-acceptance-retry,并补上上述退避/预取消和四次静默 wrapper 错误契约反例;不要把旧 UI affordance 缺口冒充 PR 回归。

已亲自运行 exact-head npm run build(TS/Vite/source 与 packaged Chat assets)、新 stall smoke、既有 acceptance-retry smoke,均通过。独立八场景通过真实 HTTP/native stream 检查:首帧后静默与无 headers 能恢复;heartbeat 连续到达时只用一个连接;cursor 不被本地 phase 推进;普通 503 仍 bounded/typed。三个取消/耗尽 oracle 在 head 失败。仅将 watchdog 四十五秒缩到七十毫秒(HTTP)/八十毫秒(UI)作为测试时钟,原 retry 二百五十/五百/一千毫秒未改;base 静默用例仅由测试安全 abort 收尾,不能当作旧产品恢复能力。本轮没有查询或等待远端 CI,也不因 unrelated 红灯拒绝。

我的整体评价

归因更正:本评审已原地修订,明确区分新增 watchdog 的 typed error 缺失与未改动的 PersonalWorkspace 重连动作缺口;后者不要求本 PR 修复。

reader 的放置、同一 Turn cursor 与有限 retry 方向正确,问题规模与三十行生产机制相称;请求修改不是代码量或新架构要求。相邻 future-facing pass 建议在原 owner 用明确的取消/timeout原因和 existing typed error 统一监督,避免再增加平行 busy flag、transport wrapper 或协议版本;评审-only 没有代作者改代码。long_horizon 的已取消工作收敛、user_experience 的取消后不再交付陈旧响应存在已复现回归;非用户 timeout 的 typed wrapper 契约也需保住。保留第一处静默恢复与 heartbeat 正向收益,补齐真实取消和 typed exhaustion 后可继续重审。此证据仅覆盖隔离合成 HTTP 和真实客户端/UI,不声称 live 模型、长时间 soak 或完整仓库套件通过;本轮不修复、合并或升级。

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants