- Send X-Codex-Routing-Hint ("model=<slug>" plus ";tier=<service_tier>")
on ChatGPT-backend requests over HTTP, compact and websocket handshakes,
matching openai/codex rust-v0.155.0 build_routing_hint_header. Translated
requests previously carried service_tier=priority only in the body.
- Derive the hint from the final upstream body, replacing a hint a native
client forwarded, so it names the model and tier actually sent.
- Keep operator precedence: an auth header rule that resolves to a value
wins, a dynamic rule that resolves to nothing falls back to the derived
hint, and models.json override_header still applies last.
- Leave API-key requests unchanged, and keep websocket reuse as native Codex
does: an open connection keeps its handshake hint when the tier changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Add `RequestID` (unique execution UUID) and `TraceID` (inbound request ID) fields to usage records and plugin API
- Generate unique UUIDs per model attempt in `UsageReporter` while linking to parent trace context
- Expose `execution_id` and `trace_id` in Redis queue usage payloads while preserving legacy `request_id`
- Add header definitions and ordering for per-turn control, timing, inline tools, clear-at, dangerous tool use, thinking binding/resumption, and prompt caching evict betas
- Dynamically assemble beta headers based on model capabilities and request body fields matching Claude Code 2.1.280 behavior
Closes: #6054
- Update default Claude CLI version to 2.1.280 in cloaking and device profile
- Add `mid-conversation-tool-changes-2026-07-01` beta header support for supported models
Closes: #6054
- Convert `input_video` and `video_url` content parts to chat completion video parts
- Preserve video URL and processing options for upstream handling
Closes: #6037
Read Gemini interactions scanner lines directly because emitted frames are already cloned before reuse, and allocate each Kimi responses passthrough chunk once so it stays owned without a second copy.
Skip the second compatibility translation for Claude, Gemini, and Gemini
interactions when the baseline and working payloads share a backing array
and no plugin hooks are installed. Keep independent buffers, the existing
hook order, and native interactions copies.
- Wait for bootstrap initialization before advancing mock clock in Codex WebSocket tests
- Reorder data channel closures in OpenAI responses stream error tests
- Drain stream chunks in byte-cap release test to prevent dangling goroutines
Closes: #5981
- Implement `streamCodexDuplex` to handle bidirectional WebSocket interactions for Codex
- Support steering continuations (`response.steer`) and queued creations (`response.create`, `response.append`)
- Preserve response settings and state transitions across steering updates and automatic successors
- Isolate connection-level failures with `codexDuplexConnectionError` to prevent unnecessary credential cooldowns
Closes: #5968
- Normalize boolean `true` subschemas across root, properties, items, definitions, and dependencies into empty object schemas `{}`
- Include `dependencies` in schema definition traversal and normalization
- Strip unsupported schema keywords including `additionalItems`, `unevaluatedProperties`, `unevaluatedItems`, and `contentSchema`
Closes: #3551
- Buffer text deltas arriving after tool calls in streaming execution and flush them after closing active tool slots
- Prioritize tool call deltas over content text deltas during Connect frame processing
- Separate pre-tool and post-tool text parts in non-streaming responses to preserve step ordering
- Correct thought step index resolution when closing thoughts during content chunk emission
Closes: #5951
- Track declared passthrough MCP tools in `claudeMCPAliasResolver`
- Recover hybrid tool names when the model prepends a virtual server prefix or replaces the caller's server with the virtual server
- Ensure client-tool alias matches take precedence over passthrough recovery and reject ambiguous passthrough tool matches
Closes: #5949
- Add `use-max-completion-tokens` setting to OpenAI compatibility model configuration
- Provide token normalization between `max_tokens` and `max_completion_tokens` based on model preference
Closes: #5939
- Only detect `server_tool_use` for advisor tool invocations in Claude conversation history
- Prevent client-side tools named `advisor` using standard `tool_use` from improperly triggering advisor cloaking restrictions
Closes: #5934
- Resolve model `is_compat` status via model info, config index, or model entries in Codex executor
- Skip wiping `reasoning.content` and stripping reasoning IDs during Responses sanitization when `is_compat` is enabled
- Pass resolved compat flag across standard, streaming, compact, and WebSocket Codex request flows
Closes: #5930
- Set expected upstream model for Devin in both streaming and non-streaming paths to avoid false positive substitution warnings on intentional mappings
- Account for line overhead and enforce max lines per stream event in StreamResponseModelObserver, dropping overflowed events until the event boundary
- Add unit tests for Devin intentional mappings and stream observer bounded memory behavior
- observe response model from terminal sourceEvent in Meta non-stream multi-event SSE responses
- record expected upstream model in UsageReporter to prevent false substitution warnings on Kimi canonical mappings
- use bounded stream observer to extract response models across image stream chunk boundaries
- add regression tests for Meta SSE non-stream, Kimi model mappings, and chunked image streams
- Devin: extract authentic upstream model name from parsed Usage.ModelName
instead of synthesized interactions JSON or binary Connect frames. Keep
response model empty when upstream does not report one.
- Gemini: support interaction.model and event_type terminal semantics
in response model extractors for Gemini Interactions streaming.
- Add unit tests for Devin and Gemini Interactions response model observability.
- Add `AlignOpenAIToolCallMessages` to reorder tool results immediately after the matching assistant tool calls while preserving original content and numeric precision.
- Prevent deferred message reordering and call ID guessing when tool outputs are incomplete, duplicate, or missing IDs.
- Normalize translated requests after applying summary configuration in Codex multi-agent execution.
Closes: #5925
Codex upstreams can silently serve a different model than the one requested
(HTTP 200, with response.model naming the substitute). The proxy kept no record
of it: nothing logged, nothing reported, only the pass-through response body.
- Add Record.ResponseModel to sdk/cliproxy/usage, aligned with the existing
ResponseServiceTier field, and emit it from the redis usage queue as the
optional response_model payload field alongside response_service_tier. Only
the record for the requested model carries it: additional-model records
(image generation tool usage) describe a side model the upstream response
never refers to, and would otherwise look like a substitution downstream.
- Add internal/runtime/executor/helps/response_model.go with
extractCodexResponseModelEvent (SSE frames and raw JSON, restricted to the events
that embed the authoritative response object, rejecting non-string and
oversized upstream model names) and IsCodexModelSubstituted (both sides
trimmed, lower-cased and stripped of thinking suffixes, dated aliases such as
gpt-5.6-terra-2026-05-13 accepted in either direction).
- UsageReporter records the served model on the event path and emits the WARN
when the attempt publishes its usage record, so no logging work happens
before the first event is forwarded. Repeats are throttled per
(auth id, requested model, served model) with a 10 minute window, because on
an affected credential every request is substituted and an unthrottled
warning would mirror the whole request volume into the logs. The credential
is labelled auth_index=<index> only: codex credential file names embed the
account e-mail, which must not be written to the logs at request rate.
Coverage, by entry point. The served model is observed on the HTTP streaming
path (both the bootstrap-buffered handshake and the streaming goroutine), the
HTTP non-streaming Execute loop, the websocket streaming and non-streaming
paths, and the two /responses-shaped image entry points. The remaining codex
entry points cannot report it and are therefore left alone: executeCompact
(/responses/compact answers with a compaction object that has no event type and
no response.model), the two direct image endpoints (/images/generations and
/images/edits answer in the Images API shape and stream image_generation.*
events), and CountTokens (counts locally with tiktoken, never reaching an
upstream).
TokenAccountingSchemaVersion is not bumped: it versions the token accounting
contract (token breakdown semantics), and this change only adds an optional
non-token field that leaves existing consumers and all token math untouched.
Note: response_model ships with the usage record and is the counting source;
the WARN is a throttled alerting signal and must not be used to count
substitutions.
Tests: table-driven unit tests for both helpers over real model ids, reporter
tests covering the published record, the single throttled warning, the absence
of account identifiers in it, concurrent observation and publishing under
-race, the throttle window and its entry bound, an executor-level guard for the
observeCodexTokenEvent wiring and the per-model records, plus a redisqueue
payload assertion for response_model. gofmt, go vet, go test -race on the
touched packages and go test ./... are clean.
- Restore client-defined `web_search` tool names from their alias across SSE, websocket, and buffered execution responses.
- Preserve namespaced tool calls when matching and restoring tool names.
Closes: #5923
- Skip parsing the `Retry-After` header when a rate limit rejection is overage-only to prevent global credential cooldowns.
- Allow exponential backoff to handle model recovery while keeping shared subscription windows available for other models.
Closes: #5920
- Generalize forced hosted tool choice handling across both image generation and web search.
- Normalize forced web search tool choices to string mode and isolate the target tool in the tools list.
- Strip hosted web search from mixed `allowed_tools` definitions.
- Skip native `x_search` injection when a hosted tool is exclusively required.
Closes: #5916
- Add `claude.model-level-cooling` configuration to scope rate limit cooldowns to the requested model.
- Treat overage-only and spend cap rejections as model-scoped when shared subscription windows remain healthy.
- Propagate model-level cooling settings into streaming, token counting, and direct execution error classifiers.
Closes: #5915
- Match tool results against pending tool calls and downgrade unmatched results to user messages.
- Prevent downgraded orphaned tool results from consuming images intended for user turns.
- Unwrap protocol wrapper envelopes and extract structured text parts while preserving arbitrary business JSON.
- Provide a placeholder for empty or whitespace-only tool results.
Closes: #5911
- Track and aggregate tool calls by call ID instead of slot index in streaming and buffered execution.
- Support raw arguments from invalid JSON fields for custom tool calls.
- Parse usage field 4 as cache write tokens instead of adding to prompt tokens.
- Align client metadata with the default client name and drop deprecated tag 28.
Closes: #5910
- Extract and attach images from tool results to corresponding tool prompts by call ID.
- Preserve structured business JSON and raw objects in function result content.
- Prepend image headers to tool prompt content when images are present.
Closes: #5893
- Filter out `automation_update` tools and sanitize tool descriptions in Devin wire requests and logs.
- Strip additional Codex prompt directives from system messages.
- Support `children` field fallback when collecting namespace tools.
- Replace tool image placeholders with omission markers for compatibility with text-only upstream models.
- Strip synthetic image relay notices and image parts from user messages.
Closes: #5884
- Buffer content and tool call deltas while thinking is active so late-arriving thinking signatures can be attached before closing thinking blocks.
- Ensure pending actions and open steps are properly flushed and closed on stream completion or trailer errors.
Closes: #5873
- Add incremental merging for Gemini `groundingMetadata` with chunk index remapping and query deduplication.
- Implement rune offset mapping across multipart messages for accurate `url_citation` annotations.
- Manage full streaming lifecycle for web search calls, emitting `searching`, `completed`, and output item done events.
- Stream incremental citation annotations via `response.output_text.annotation.added` events.
- Support `web_search_preview_2025_03_11` as a recognized web search tool type.