- Parse and pass `quality` in xAI image generation and edit requests
- Extract `quality` field from JSON payloads and multipart form data
- Include `quality` in the upstream xAI image request payload when provided
Closes: #6157
- Share request payload reference during session identity enrichment when original request is unset
- Avoid cloning request payload when interceptors do not mutate the body
- Pass request body directly to handler and conductor interceptors
Closes: #6101
- Marshal Codex client catalog as compact single-line JSON without HTML escaping
- Use compact base instructions for non-template models to stay within the 1MiB payload limit
- Keep `apply_patch_tool_type`, `upgrade`, and `availability_nux` present as null instead of deleting them
- Add `trusted-proxies` configuration option with IP and CIDR validation
- Apply trusted proxies to Gin engine to properly resolve forwarded client IPs
- Record `ResolvedClientIP` in request metadata and usage queue details
Closes: #6024
- Allow prewarm requests to be handled locally after an existing request turn
- Reset transcript root for mid-connection prewarms that omit `previous_response_id`
- Reuse local prewarm evaluation flag across request normalization and response handling
Closes: #6006
- Wait for bootstrap initialization before advancing mock clock in Codex WebSocket tests
- Reorder data channel closures in OpenAI responses stream error tests
- Drain stream chunks in byte-cap release test to prevent dangling goroutines
Closes: #5981
- Implement `streamCodexDuplex` to handle bidirectional WebSocket interactions for Codex
- Support steering continuations (`response.steer`) and queued creations (`response.create`, `response.append`)
- Preserve response settings and state transitions across steering updates and automatic successors
- Isolate connection-level failures with `codexDuplexConnectionError` to prevent unnecessary credential cooldowns
Closes: #5968
- Consolidate stream error code and error type derivations into paired error classes
- Map `StatusRequestTimeout` to `server_error` instead of `invalid_request_error` so interrupted streams remain retryable by clients
Closes: #5931
- Detect Devin models by ID prefix, type, ownership, model registry metadata, or provider.
- Append `(Devin)` to model display names in Codex client responses when not already present.
- Set model `type` to `devin` in home Codex model formatting when served by the Devin provider.
Grok Imagine already supports the 9:20 phone ratio, but the OpenAI image
handler whitelist dropped it and fell back to 1:1. Accept 9:20 from
aspect_ratio and from size, and keep unknown ratios on the existing
fallback.
- Introduce `authUnavailableError` with detailed recovery hints, status codes, and retry-after headers.
- Refactor scheduler logic to incorporate local retry deadlines and classify cooldown states without altering quota assessments.
- Refine error propagation to properly handle upstream errors and enrich context for retryable scenarios.
- Add comprehensive tests for retry deadlines, scheduler behavior, and error handling during credential unavailability.
- Ensure embedded recovery hints in HTTP headers for improved client-side handling and diagnostics.
Closes: #5842
- Add `ForcedProvider` and `AuthID` fields to `HostModelExecutionRequest` and `ModelExecutionRequest`.
- Pin credentials via `WithPinnedAuthID` and forward forced provider options during model execution and streaming.
- Forward forced provider and auth ID from plugin host callbacks to model execution handlers.
Closes: #5814
- Verify observed compaction response replay and retention across subsequent turns.
- Test plugin route overrides, provider auth pinning, and home runtime auth handling during compaction.
Closes: #5811
- Add `WriteModelListResponse` to `BaseAPIHandler` to apply plugin interceptors and record request lifecycles for model catalog responses.
- Update OpenAI, Claude, Gemini, Grok, and Codex model listing endpoints to route responses through the unified interceptor helper.
Closes: #5742
- Resolve canonical templates using metadata model IDs for model aliases and prefixed routes.
- Apply descriptions, base instructions, and thinking support overrides to matched templates.
- Restrict protocol capabilities and reasoning levels based on provider support.
Closes: #5699
- Track trailing carriage returns across chunk boundaries in `sseJSONValidationState`.
- Strip leading newline in subsequent chunks to prevent duplicate newline insertion from split CRLF sequences.
- Reset trailing carriage return state upon stream completion.
Closes: #5657
- Format streaming error payloads with nested error objects matching official OpenAI Responses SSE specifications.
- Extract and propagate sequence numbers from upstream terminal events and framer states.
- Use `json.Number` to prevent precision loss for large integers and token metrics.
- Sanitize sensitive keys recursively across nested error objects without dropping custom fields.
- Track pending synthetic prewarm response IDs to merge warmup inputs into subsequent delta followups.
- Normalize transcript replacements when followups do not reference the prewarm parent response ID.
- Validate that the `input` field is an array for `response.create` requests.
- Allow `function_call_output` items without a `call_id` when a non-empty tool name is present.
Closes: #5631
- Register builtin model definitions for `gpt-image-2.5`, `gpt-image-2.5-flare`, and `gpt-image-2.5-sunburst`.
- Update OpenAI image handlers and request routing to recognize GPT Image 2.5 models.
- Support direct image generation and edit execution for GPT Image 2.5 variants in the Codex executor.
- Apply client visibility overrides to hide new builtin image models where appropriate.
- Introduce `IsTerminalAuthError` and `NewTerminalAuthError` to identify permanent upstream authentication failures.
- Return terminal auth errors from candidate selection and scheduling when all available credentials fail with unauthorized errors.
- Support `BuildErrorResponseBodyWithError` to format terminal upstream auth errors as non-retryable `upstream_authentication_required` responses.
- Propagate terminal error classifications and the `retryable` field across HTTP and WebSocket response handlers.
Closes: #5645
- Initialize `util.SessionIDResolver` to resolve session IDs from request contexts, metadata, and headers.
- Ensure canonical session metadata is injected into execution options across execution flows.
- Propagate session context in executors to support `$CPA-SESSION-ID` expansion in custom headers.
- Sync cleared or updated session identities back to execution contexts during conductor execution.
Closes: #5690
- Reuse coresession.ExtractSessionInfo across HTTP headers and request payloads to unify canonical session prefix namespaces with the scheduler.
- Extract hierarchical session identities in two phases: initial extraction from request headers on entry, and authoritative deep extraction once request payloads and metadata are available.
- Support Claude Code multi-level subagents (X-Claude-Code-Agent-Id, metadata.agent_id) and Codex thread fork lineages.
- Propagate SessionID and ParentSessionID across ClientRequestMetadata, UsageReporter, and coreusage.Record without root_session_id.
- Include session_id and parent_session_id in queuedUsageDetail for Home LPushUsage forwarding and Redis consumption with self-loop guards.
- Add comprehensive test coverage for canonical headers, body extraction, ghost parent elimination, and self-referential loop guards.
- Bump plugin schema version to 5 and introduce `SchemaVersionStreamChunkOmitHistory`.
- Omit `HistoryChunks` on payload stream chunks for schema version 5+ to avoid per-chunk cloning and serialization overhead.
- Conditionally accumulate and clone history chunks only when legacy plugins with schema version < 5 are active.
Closes: #5451
- Add `writePing` to responses websocket writer to emit Ping control frames.
- Send periodic keep-alive Ping frames based on streaming configuration during response forwarding.
- Reset keep-alive interval upon receiving data chunks and abort session if ping write fails.
Closes: #5413
- Add `codex.orphan-delegation-compatibility` configuration option and mirror it to SDK configuration.
- Convert orphan Codex delegation outputs into standard user messages for requests with `X-Openai-Subagent: collab_spawn`.
- Integrate orphan delegation rewriting into OpenAI responses request handling pipeline.
Closes: #5401
- Attach candidate upstream errors as causes to scheduler availability and cooldown errors.
- Introduce `errorWithCause` wrapper to extract and display upstream error summaries.
- Enrich auth selection error messages with upstream details and preserve model cooldown errors.
- Forward `Retry-After` response headers for model cooldown errors in Claude Code handler.
Closes: #5365
- Add `ParsePluginExecutorResponseUsage` to extract token usage from non-streaming plugin responses across Claude, Gemini, Interactions, Antigravity, and OpenAI/Codex protocols.
- Add `ObservePluginExecutorStreamUsage` to observe and aggregate token usage across streaming chunks.
Closes: #5340
- Filter `max` and `ultra` reasoning effort levels for Codex client versions prior to `0.144.0`.
- Extract and forward the `client_version` query parameter across model catalog response handlers.
- Add dotted version parsing and comparison utilities to verify extended reasoning level compatibility.
Closes: #5262
- Introduce `WebSocketResponseObserver` capability and bump plugin ABI schema version to 4.
- Forward upstream WebSocket response frames from Codex and xAI executors to configured observers.
- Wire `WebSocketResponseObserver` across API handlers and plugin host dispatchers.
Closes: #5248
- Add compact-specific error classification to mark transient/non-credential failures as availability-neutral instead of triggering cooldown penalties.
- Stop auth fallback immediately on compact request-fault errors (e.g., bad/not-found/unsupported request errors) and return the upstream compact error.
- Preserve existing cooldown behavior for auth/credential faults (`401`, `403`, `429`) while allowing non-auth compact failures to fail fast without tainting normal traffic routing.
Closes: #5031
After remote compaction succeeds, the normalized previous request still contains the consumed compaction_trigger. The next WebSocket-to-HTTP merge replays that trigger before the compaction output and current input, causing the upstream 400 error.
Remove compaction_trigger only from previous request items when the previous response contains compaction or compaction_summary. Add a two-step regression test covering trigger normalization followed by compact replay.
Fixes#5041
- Return buffered pending stream errors from `errChan` when the data channel closes in streaming handlers.
- Emit error response instead of silently falling back to normal SSE completion paths.
- Replace Claude’s local `pendingClaudeStreamError` usage with shared `handlers.PendingStreamError`.
Closes: #4710
- add shared `EnsureResponsesUsageDetails` helper to patch `usage` objects with:
- `output_tokens_details.reasoning_tokens = 0`
- `input_tokens_details.cached_tokens = 0`
- for both plain JSON and SSE `data:` frames, including multi-line frames
- apply the helper to OpenAI Response format outputs in non-stream and stream paths across executors/plugins so translated payloads consistently include required usage details
- update websocket/completion payload builders to emit default `usage` detail fields for prewarm/finish responses
Closes: #4985
- Added `grok-imagine-image-2.0` as a first-class xAI image base model across validation, canonicalization, and routing checks.
- Registered the model in built-in model definitions so it appears in model metadata.
- Updated image model allowlists and request validation error messaging to include the new model.
- Marked the new model as hidden in client visibility override handling.