Skip to content

[Bug]: OpenCode prompt admission times out on large attachments - turn declared failed while it actually completes #14011

Description

@imluckii

Environment

  • T3 Code: 0.0.43-nightly.20260927.2344
  • OpenCode: 1.18.32 (remote opencode serve on the same host)
  • Hardware: 1-core VPS (this makes the race much easier to hit)
  • OS: Linux

Summary

When a prompt with large attachments (9 screenshots, ~1080x2400 each) is sent to an OpenCode thread, T3 Code's prompt-admission confirmation window expires before OpenCode can even acknowledge the message. T3 then declares the turn failed and emits the error banner, while OpenCode continues processing the turn to completion. The user sees a hard failure for a turn that actually succeeds.

Steps to reproduce

  1. Connect T3 Code to an OpenCode instance on modest hardware (1 core).
  2. Send a prompt with 9 image attachments.
  3. Within ~15s, the app shows:

Runtime error
OpenCode accepted the prompt, but T3 Code could not confirm its message or session status. The cleanup abort also failed: Operation timed out after '1s'

  1. The OpenCode turn keeps running and completes normally (verified server-side: loop exits at step 27, files edited, session completes).

Timeline from a live repro (server logs, UTC)

22:16:24.118  session created (prompt submitted)
22:16:25.7-8  OpenCode receives the 9 image parts
22:16:32.669  resized image 1/9
22:17:22.809  resized image 9/9        <- 50 seconds of ingest
22:17:28.765  agent loop step=0        <- processing starts 64s after send

T3 Code's admission recovery (schedulePromptAdmissionRecovery in apps/server/src/provider/Layers/OpenCodeAdapter.ts) polls session.message with a hardcoded Effect.timeout("1 second") per attempt, up to 5 attempts with 250ms→2s backoff (~14s worst case). The window expired at ~22:16:40 while OpenCode was still resizing image 2/9. failPromptAdmissionRecovery then calls session.abort, also with a hardcoded 1s timeout, which also times out - producing the "cleanup abort also failed" string - and the session is marked error with turn.completed { state: "failed" }.

Meanwhile the abort never actually reached OpenCode, which is the only reason the turn survived. A faster machine where the abort lands within 1s would kill a turn that was about to succeed anyway.

Problems

  1. Fixed deadlines don't scale with payload. The 1 second literals at lines ~861/1293/1384/1418 assume tiny prompts. Attachment ingest time is unbounded from T3's perspective.
  2. A failed cleanup abort declares the turn dead. If session.abort times out, the correct inference is "OpenCode is busy, turn may be alive" - T3 should re-poll session.status / session.message instead of emitting a terminal failure.
  3. No warning escalation. There is already a runtime.warning event ("OpenCode turn completion is waiting for session status") for the idle path; the admission path jumps straight to terminal failure.

Relation to existing PRs

Expected

Admission confirmation should survive slow-but-progressing ingestion: deadline proportional to recent event flow (the SSE stream was actively delivering file/using resized image events the whole time - that's strong liveness evidence the probes could use), and failed cleanup abort should trigger re-confirmation, not terminal failure.

Actual

Terminal "Runtime error" banner + session marked error in the app while the agent works and lands the changes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions