Commit Graph

14 Commits

Author SHA1 Message Date
Luis Pater
dd3b657b5d fix(codex): use resolved base model for routing hint
- Pass resolved `baseModel` directly to `applyCodexRoutingHint` instead of reading model from the upstream body
- Ensure the routing hint header accurately reflects the resolved model across HTTP and WebSocket executions
2026-09-24 13:02:16 +08:00
SamGu-NRX
b97f71da0c fix(codex): send the ChatGPT routing hint native Codex sends
- Send X-Codex-Routing-Hint ("model=<slug>" plus ";tier=<service_tier>")
  on ChatGPT-backend requests over HTTP, compact and websocket handshakes,
  matching openai/codex rust-v0.155.0 build_routing_hint_header. Translated
  requests previously carried service_tier=priority only in the body.
- Derive the hint from the final upstream body, replacing a hint a native
  client forwarded, so it names the model and tier actually sent.
- Keep operator precedence: an auth header rule that resolves to a value
  wins, a dynamic rule that resolves to nothing falls back to the derived
  hint, and models.json override_header still applies last.
- Leave API-key requests unchanged, and keep websocket reuse as native Codex
  does: an open connection keeps its handshake hint when the tier changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 19:56:53 -05:00
Luis Pater
81d6ba7746 fix(codex): preserve reasoning content and IDs for compat models in responses
- Resolve model `is_compat` status via model info, config index, or model entries in Codex executor
- Skip wiping `reasoning.content` and stripping reasoning IDs during Responses sanitization when `is_compat` is enabled
- Pass resolved compat flag across standard, streaming, compact, and WebSocket Codex request flows

Closes: #5930
2026-09-18 23:02:54 +08:00
Viggo95
25f40d8cf8 feat(codex): record upstream response model and warn on silent model substitution
Codex upstreams can silently serve a different model than the one requested
(HTTP 200, with response.model naming the substitute). The proxy kept no record
of it: nothing logged, nothing reported, only the pass-through response body.

- Add Record.ResponseModel to sdk/cliproxy/usage, aligned with the existing
  ResponseServiceTier field, and emit it from the redis usage queue as the
  optional response_model payload field alongside response_service_tier. Only
  the record for the requested model carries it: additional-model records
  (image generation tool usage) describe a side model the upstream response
  never refers to, and would otherwise look like a substitution downstream.
- Add internal/runtime/executor/helps/response_model.go with
  extractCodexResponseModelEvent (SSE frames and raw JSON, restricted to the events
  that embed the authoritative response object, rejecting non-string and
  oversized upstream model names) and IsCodexModelSubstituted (both sides
  trimmed, lower-cased and stripped of thinking suffixes, dated aliases such as
  gpt-5.6-terra-2026-05-13 accepted in either direction).
- UsageReporter records the served model on the event path and emits the WARN
  when the attempt publishes its usage record, so no logging work happens
  before the first event is forwarded. Repeats are throttled per
  (auth id, requested model, served model) with a 10 minute window, because on
  an affected credential every request is substituted and an unthrottled
  warning would mirror the whole request volume into the logs. The credential
  is labelled auth_index=<index> only: codex credential file names embed the
  account e-mail, which must not be written to the logs at request rate.

Coverage, by entry point. The served model is observed on the HTTP streaming
path (both the bootstrap-buffered handshake and the streaming goroutine), the
HTTP non-streaming Execute loop, the websocket streaming and non-streaming
paths, and the two /responses-shaped image entry points. The remaining codex
entry points cannot report it and are therefore left alone: executeCompact
(/responses/compact answers with a compaction object that has no event type and
no response.model), the two direct image endpoints (/images/generations and
/images/edits answer in the Images API shape and stream image_generation.*
events), and CountTokens (counts locally with tiktoken, never reaching an
upstream).

TokenAccountingSchemaVersion is not bumped: it versions the token accounting
contract (token breakdown semantics), and this change only adds an optional
non-token field that leaves existing consumers and all token math untouched.

Note: response_model ships with the usage record and is the counting source;
the WARN is a throttled alerting signal and must not be used to count
substitutions.

Tests: table-driven unit tests for both helpers over real model ids, reporter
tests covering the published record, the single throttled warning, the absence
of account identifiers in it, concurrent observation and publishing under
-race, the throttle window and its entry bound, an executor-level guard for the
observeCodexTokenEvent wiring and the per-model records, plus a redisqueue
payload assertion for response_model. gofmt, go vet, go test -race on the
touched packages and go test ./... are clean.
2026-09-18 12:31:52 +08:00
Luis Pater
f702bc1ac2 feat(codex): preserve native fidelity for responses-lite requests
- Detect native responses-lite requests via headers and client metadata.
- Skip instructions normalization and synthetic session cloaking for native requests.
- Preserve upstream completion output during websocket response forwarding.

Closes: #5780
2026-09-13 15:05:03 +08:00
Luis Pater
3bf787fc1d feat(auth): propagate canonical session id for custom header expansion
- Initialize `util.SessionIDResolver` to resolve session IDs from request contexts, metadata, and headers.
- Ensure canonical session metadata is injected into execution options across execution flows.
- Propagate session context in executors to support `$CPA-SESSION-ID` expansion in custom headers.
- Sync cleared or updated session identities back to execution contexts during conductor execution.

Closes: #5690
2026-09-10 01:43:37 +08:00
Luis Pater
b064b832e2 feat(codex): support model-level quota cooling
- Add `model-level-cooling` configuration option to Codex settings.
- Scope `usage_limit_reached` quota cooldowns to the requested model instead of the entire credential when enabled.
- Propagate model-level cooling checks across HTTP, SSE, and WebSocket execution paths.

Closes: #5619
2026-09-09 11:57:58 +08:00
Luis Pater
bf20b999de fix(codex): simplify complex tool schema unions and detect empty incomplete responses
- Normalize complex constant `oneOf` and `anyOf` tool parameter schemas into equivalent enums to prevent upstream aborts.
- Escape property keys containing dots and colons during schema updates to prevent invalid path splitting.
- Identify terminal `response.incomplete` events with zero output tokens and no content as upstream failures.

Closes: #5551
2026-09-07 20:09:29 +08:00
Luis Pater
e424bfad00 feat(executor): support $-based custom headers from downstream request headers
- Propagate request headers into custom-header resolution for OpenAI/Gemini/XAI/Codex execution and websocket flows.
- Resolve auth `header:` values like `$ABC` from incoming request headers at request time and omit headers when no value is available.
- Add documentation for the dynamic custom-header behavior in `config.example.yaml`.

Closes: #5053
2026-08-18 19:33:50 +08:00
Luis Pater
e0b4956242 fix(openai): ensure Responses usage includes token detail fields
- add shared `EnsureResponsesUsageDetails` helper to patch `usage` objects with:
  - `output_tokens_details.reasoning_tokens = 0`
  - `input_tokens_details.cached_tokens = 0`
  - for both plain JSON and SSE `data:` frames, including multi-line frames
- apply the helper to OpenAI Response format outputs in non-stream and stream paths across executors/plugins so translated payloads consistently include required usage details
- update websocket/completion payload builders to emit default `usage` detail fields for prewarm/finish responses

Closes: #4985
2026-08-15 15:00:19 +08:00
Luis Pater
dcee14dd3c feat(compat): preserve compat-mode thinking/signature blocks for API-key models
- Added `is-compat` model metadata plumbing from config through executor and helpers, including hash computation.
- Introduced a compatibility-aware translation path (`TranslateRequestWithAPIKeyModelCompatibility`) and wired it into Claude/Gemini/Codex/Interactions request flows.
- Updated Claude message sanitization/translation behavior to keep empty-thinking compatibility blocks (including signatures) when `is-compat` is enabled, while keeping default behavior unchanged.
2026-08-06 17:19:24 +08:00
Luis Pater
e5ea945ed9 feat(codex): add model-level is-compat flag to rewrite MultiAgentV2 agent_message for Responses-compatible endpoints
Closes: #4801
2026-08-06 04:49:28 +08:00
Luis Pater
f32291436a refactor(executor): consolidate thinking.ApplyThinking into helps.ApplyRequestThinking
- Replaced instances of `thinking.ApplyThinking` with `helps.ApplyRequestThinking` across all executors for consistency.
- Updated `applyGeminiInteractionsThinking` to accept `cliproxyexecutor.Request` and `Options`.
- Centralized logic for request thinking application to `helps` package for improved maintainability.

Closes: #4618
2026-07-29 14:14:18 +08:00
Luis Pater
fe4ae4989c chore(pluginhost): refactor and remove unused interceptors and executor methods
- Removed deprecated interceptor and executor-related methods, including `callRequestInterceptor`, `callResponseInterceptor`, and `callStreamChunkInterceptor`.
- Consolidated unused logic and pruned redundant imports to streamline `adapters.go`.
- No functional changes.
2026-07-26 14:31:45 +08:00