Skip to content

fix(server): decode streamed SSE text from the cumulative id list - #105

Open
ahrazzle wants to merge 1 commit into
Edge0-AI:mainfrom
ahrazzle:fix/sse-incremental-decode
Open

ahrazzle wants to merge 1 commit into
Edge0-AI:mainfrom
ahrazzle:fix/sse-incremental-decode

Conversation

@ahrazzle

Copy link
Copy Markdown
Contributor

What changed

Streamed SSE responses now decode the full id list cumulatively and emit only the newly completed suffix, holding back partial multi-byte characters until they resolve.

Root cause

The streaming path decoded each token id in isolation (src/edge0/server/app.py:83). A byte-level BPE token carrying part of a multi-byte character streamed as U+FFFD while the non-streaming path returned correct text.

Changes

  • Added _incremental_suffix helper to extract the new suffix from a cumulative decode
  • Modified _chat_stream to maintain a running list of seen ids
  • Flush held-back partials at end of generation so streamed equals non-streaming text

Validation

Live edge0-8b server run (stdlib transport), prompt asking the model to echo 🌊🏖️🦀🍣:

path text
base, streaming 9 U+FFFD, no emoji
head, streaming 🌊🏖️🦀🍣 (0 U+FFFD)

The non-streaming path decodes the whole id list and was already correct.

Test: tests/test_server.py::test_chat_stream_incremental_decode_of_split_characters

  • Red on base (base tree + head test file)
  • Green on head (66 passed, 1 skipped)

Limits

Common CJK text in the shipped vocabulary is not affected. Corruption needs characters that split across byte tokens (emoji, rare CJK, some Korean).

O(n²) cumulative re-decode per request. This is the minimal fix. Throughput was not measured.

Automated posting by agentic team with human oversight.

The streaming path decoded every token id on its own
(`decode_tokens(engine, [tid])`) and shipped the result as the SSE
delta. A byte-level BPE token is a byte fragment, not a character, so
any multi-byte UTF-8 character split across ids decoded to U+FFFD and
the streamed reply was corrupted, while the non-streaming path — which
decodes the whole id list — returned correct text.

Keep the ids seen so far, decode the cumulative list, and emit only the
newly completed suffix; a trailing replacement character is held back
until the next id resolves it, so a partial multi-byte character is
never emitted, and the tail is flushed when generation ends so the
concatenated deltas equal the non-streaming text exactly.

Add a regression test that streams an id sequence whose UTF-8 bytes are
split across ids through the handler and asserts the deltas match the
non-streaming text with no U+FFFD.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant