- Pass resolved `baseModel` directly to `applyCodexRoutingHint` instead of reading model from the upstream body
- Ensure the routing hint header accurately reflects the resolved model across HTTP and WebSocket executions
- Send X-Codex-Routing-Hint ("model=<slug>" plus ";tier=<service_tier>")
on ChatGPT-backend requests over HTTP, compact and websocket handshakes,
matching openai/codex rust-v0.155.0 build_routing_hint_header. Translated
requests previously carried service_tier=priority only in the body.
- Derive the hint from the final upstream body, replacing a hint a native
client forwarded, so it names the model and tier actually sent.
- Keep operator precedence: an auth header rule that resolves to a value
wins, a dynamic rule that resolves to nothing falls back to the derived
hint, and models.json override_header still applies last.
- Leave API-key requests unchanged, and keep websocket reuse as native Codex
does: an open connection keeps its handshake hint when the tier changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Resolve model `is_compat` status via model info, config index, or model entries in Codex executor
- Skip wiping `reasoning.content` and stripping reasoning IDs during Responses sanitization when `is_compat` is enabled
- Pass resolved compat flag across standard, streaming, compact, and WebSocket Codex request flows
Closes: #5930
Codex upstreams can silently serve a different model than the one requested
(HTTP 200, with response.model naming the substitute). The proxy kept no record
of it: nothing logged, nothing reported, only the pass-through response body.
- Add Record.ResponseModel to sdk/cliproxy/usage, aligned with the existing
ResponseServiceTier field, and emit it from the redis usage queue as the
optional response_model payload field alongside response_service_tier. Only
the record for the requested model carries it: additional-model records
(image generation tool usage) describe a side model the upstream response
never refers to, and would otherwise look like a substitution downstream.
- Add internal/runtime/executor/helps/response_model.go with
extractCodexResponseModelEvent (SSE frames and raw JSON, restricted to the events
that embed the authoritative response object, rejecting non-string and
oversized upstream model names) and IsCodexModelSubstituted (both sides
trimmed, lower-cased and stripped of thinking suffixes, dated aliases such as
gpt-5.6-terra-2026-05-13 accepted in either direction).
- UsageReporter records the served model on the event path and emits the WARN
when the attempt publishes its usage record, so no logging work happens
before the first event is forwarded. Repeats are throttled per
(auth id, requested model, served model) with a 10 minute window, because on
an affected credential every request is substituted and an unthrottled
warning would mirror the whole request volume into the logs. The credential
is labelled auth_index=<index> only: codex credential file names embed the
account e-mail, which must not be written to the logs at request rate.
Coverage, by entry point. The served model is observed on the HTTP streaming
path (both the bootstrap-buffered handshake and the streaming goroutine), the
HTTP non-streaming Execute loop, the websocket streaming and non-streaming
paths, and the two /responses-shaped image entry points. The remaining codex
entry points cannot report it and are therefore left alone: executeCompact
(/responses/compact answers with a compaction object that has no event type and
no response.model), the two direct image endpoints (/images/generations and
/images/edits answer in the Images API shape and stream image_generation.*
events), and CountTokens (counts locally with tiktoken, never reaching an
upstream).
TokenAccountingSchemaVersion is not bumped: it versions the token accounting
contract (token breakdown semantics), and this change only adds an optional
non-token field that leaves existing consumers and all token math untouched.
Note: response_model ships with the usage record and is the counting source;
the WARN is a throttled alerting signal and must not be used to count
substitutions.
Tests: table-driven unit tests for both helpers over real model ids, reporter
tests covering the published record, the single throttled warning, the absence
of account identifiers in it, concurrent observation and publishing under
-race, the throttle window and its entry bound, an executor-level guard for the
observeCodexTokenEvent wiring and the per-model records, plus a redisqueue
payload assertion for response_model. gofmt, go vet, go test -race on the
touched packages and go test ./... are clean.
- Initialize `util.SessionIDResolver` to resolve session IDs from request contexts, metadata, and headers.
- Ensure canonical session metadata is injected into execution options across execution flows.
- Propagate session context in executors to support `$CPA-SESSION-ID` expansion in custom headers.
- Sync cleared or updated session identities back to execution contexts during conductor execution.
Closes: #5690
- Add `model-level-cooling` configuration option to Codex settings.
- Scope `usage_limit_reached` quota cooldowns to the requested model instead of the entire credential when enabled.
- Propagate model-level cooling checks across HTTP, SSE, and WebSocket execution paths.
Closes: #5619
- Normalize complex constant `oneOf` and `anyOf` tool parameter schemas into equivalent enums to prevent upstream aborts.
- Escape property keys containing dots and colons during schema updates to prevent invalid path splitting.
- Identify terminal `response.incomplete` events with zero output tokens and no content as upstream failures.
Closes: #5551
- Propagate request headers into custom-header resolution for OpenAI/Gemini/XAI/Codex execution and websocket flows.
- Resolve auth `header:` values like `$ABC` from incoming request headers at request time and omit headers when no value is available.
- Add documentation for the dynamic custom-header behavior in `config.example.yaml`.
Closes: #5053
- add shared `EnsureResponsesUsageDetails` helper to patch `usage` objects with:
- `output_tokens_details.reasoning_tokens = 0`
- `input_tokens_details.cached_tokens = 0`
- for both plain JSON and SSE `data:` frames, including multi-line frames
- apply the helper to OpenAI Response format outputs in non-stream and stream paths across executors/plugins so translated payloads consistently include required usage details
- update websocket/completion payload builders to emit default `usage` detail fields for prewarm/finish responses
Closes: #4985
- Added `is-compat` model metadata plumbing from config through executor and helpers, including hash computation.
- Introduced a compatibility-aware translation path (`TranslateRequestWithAPIKeyModelCompatibility`) and wired it into Claude/Gemini/Codex/Interactions request flows.
- Updated Claude message sanitization/translation behavior to keep empty-thinking compatibility blocks (including signatures) when `is-compat` is enabled, while keeping default behavior unchanged.
- Replaced instances of `thinking.ApplyThinking` with `helps.ApplyRequestThinking` across all executors for consistency.
- Updated `applyGeminiInteractionsThinking` to accept `cliproxyexecutor.Request` and `Options`.
- Centralized logic for request thinking application to `helps` package for improved maintainability.
Closes: #4618
- Removed deprecated interceptor and executor-related methods, including `callRequestInterceptor`, `callResponseInterceptor`, and `callStreamChunkInterceptor`.
- Consolidated unused logic and pruned redundant imports to streamline `adapters.go`.
- No functional changes.