Files
CLIProxyAPI/internal/runtime/executor
Viggo95 cb73cd9936 fix(codex): buffer keepalive and empty item announcements during codex bootstrap
The bootstrap buffer released the stream on the first frame outside the
handshake allow-list. Upstream sends `keepalive` heartbeats and
`response.output_item.added` before the first token, so the downstream headers
were committed while nothing had been produced, and an overload rejection
arriving afterwards could no longer be retried on another credential.

Extend the allow-list with `keepalive`, and with `response.output_item.added`,
`response.content_part.added` and `response.reasoning_summary_part.added` when
they announce something that is still empty. `response.output_item.added` is
accepted only for `message`, `reasoning`, `function_call` and `custom_tool_call`
items, and only while their content, summary, arguments or input is empty: every
other item type stands for a server-side operation that may already be running -
a `web_search_call` is announced with status "in_progress" and its `searching`
event follows immediately - and failing the attempt over after one would run
that operation again on another credential. Part announcements are matched on a
closed list of textual part types for the same reason.

The list stays closed. "Nothing has happened yet" cannot be derived from the
absence of a TTFT token, because TTFT deliberately ignores server-side tool
traffic such as `response.shell_call_output_content.delta` and its `.done`
counterpart, so an unrecognised frame, an unrecognised item type and an
unrecognised part type all release the stream exactly as before.

Holding those frames also required fixing the bound on how much may be held.
codexBootstrapMaxBufferedEvents was enforced with len(bufferedChunks), and a
chunk count bounds only the downstream formats that render every upstream frame:
the OpenAI Chat Completions and Gemini translators return zero chunks for a frame
they do not recognise, which is true of every frame added here. Count frames read
from the upstream instead, and cap what a bootstrap accumulates over the upstream
frames and the chunks they translate into. The SSE scanner walks physical lines
and an event can arrive as one line (": keepalive"), two, or three
(event:/data:/blank), so the frame budget is sized for the widest framing rather
than derived from any separator. The websocket loop counts a message as soon as
it is read, before the branches that skip non-text and whitespace-only messages,
so a peer sending only frames the loop skips cannot hold the downstream headers
open; the message that exhausts the budget is still processed and delivered
rather than dropped. Both caps are checked before a frame is admitted, so one
oversized frame cannot be taken on the strength of an empty buffer. This widens
the window where one frame rendered one chunk - the websocket path effectively
moves from 16 frames to 48.

An empty `response.incomplete` seen while buffering is now delivered in-stream
instead of failing the attempt over, matching the websocket executor and the
documented contract that only overload and rate-limit rejections trigger
failover.

`isCodexHandshakeMetadataEvent` is renamed to `isCodexBootstrapBufferableEvent`
because the list is no longer only handshake metadata. Both config doc sites
describe what is held, that heartbeats are held too, that SSE also holds the
lines around a `data:` frame, and that the hold is bounded by frames and bytes
rather than by wall-clock time.
2026-09-12 22:27:44 +08:00
..