292 Commits

Author SHA1 Message Date
Luis Pater
4b99d1d6f8 feat(openai): support quality parameter for xAI image requests
- Parse and pass `quality` in xAI image generation and edit requests
- Extract `quality` field from JSON payloads and multipart form data
- Include `quality` in the upstream xAI image request payload when provided

Closes: #6157
2026-09-28 06:22:38 +08:00
Luis Pater
31f4cfab3f feat: migrate CLIProxyAPI to version 8 2026-09-27 14:42:26 +08:00
hkfires
d9922f0589 feat(config): add v8 configuration API
Introduce the v8 management configuration tree with migration-aware reads and
writes, grouped upstream API-key providers, and independent v8 operational
routes.

Preserve legacy configuration compatibility while ensuring successful v8 writes
migrate layouts, redact JSON TURN secrets, and keep OAuth-only provider
settings from affecting API-key credentials.
2026-09-27 14:16:27 +08:00
Luis Pater
c9a06ba6f2 fix(interceptors): avoid cloning request body for read-only interceptors
- Share request payload reference during session identity enrichment when original request is unset
- Avoid cloning request payload when interceptors do not mutate the body
- Pass request body directly to handler and conductor interceptors

Closes: #6101
2026-09-25 00:34:17 +08:00
Luis Pater
673131f574 fix(codex): compact client model catalog and preserve required fields
- Marshal Codex client catalog as compact single-line JSON without HTML escaping
- Use compact base instructions for non-template models to stay within the 1MiB payload limit
- Keep `apply_patch_tool_type`, `upgrade`, and `availability_nux` present as null instead of deleting them
2026-09-23 08:54:52 +08:00
Luis Pater
a962b77d43 feat(server): support trusted proxies configuration for client IP resolution
- Add `trusted-proxies` configuration option with IP and CIDR validation
- Apply trusted proxies to Gin engine to properly resolve forwarded client IPs
- Record `ResolvedClientIP` in request metadata and usage queue details

Closes: #6024
2026-09-23 04:26:45 +08:00
Luis Pater
d582067c06 feat(executor): support execution-scoped request proxy overrides
- Add context helpers to manage execution-scoped proxy overrides
- Prioritize request proxy over credential and global proxy settings across HTTP and uTLS clients
- Propagate request proxy in conductor execution and plugin host adapters
- Isolate Codex WebSocket connection reuse by proxy endpoint

Closes: #6013
2026-09-22 03:54:34 +08:00
Luis Pater
dd013f9e29 fix(responses): filter upstream private and telemetry events in SSE streams
- Filter internal `responsesapi.` telemetry and private `codex.` events from SSE output
- Preserve Codex response metadata for official Codex clients while dropping rate limits
- Discard frames matching filtered event names or payload types in the SSE framer

Closes: #6007
2026-09-21 21:35:48 +08:00
Luis Pater
7b6fafce1b fix(responses): support mid-connection prewarm in websocket handler
- Allow prewarm requests to be handled locally after an existing request turn
- Reset transcript root for mid-connection prewarms that omit `previous_response_id`
- Reuse local prewarm evaluation flag across request normalization and response handling

Closes: #6006
2026-09-21 20:13:05 +08:00
Luis Pater
563865e77a test: fix timing synchronization and channel races in streaming tests
- Wait for bootstrap initialization before advancing mock clock in Codex WebSocket tests
- Reorder data channel closures in OpenAI responses stream error tests
- Drain stream chunks in byte-cap release test to prevent dangling goroutines

Closes: #5981
2026-09-21 00:39:28 +08:00
Luis Pater
42c9680eee feat(executor): support duplex streaming for codex websockets
- Implement `streamCodexDuplex` to handle bidirectional WebSocket interactions for Codex
- Support steering continuations (`response.steer`) and queued creations (`response.create`, `response.append`)
- Preserve response settings and state transitions across steering updates and automatic successors
- Isolate connection-level failures with `codexDuplexConnectionError` to prevent unnecessary credential cooldowns

Closes: #5968
2026-09-20 00:02:10 +08:00
Luis Pater
cb62a6748b fix(openai): classify stream request timeout as server error
- Consolidate stream error code and error type derivations into paired error classes
- Map `StatusRequestTimeout` to `server_error` instead of `invalid_request_error` so interrupted streams remain retryable by clients

Closes: #5931
2026-09-18 23:10:19 +08:00
Luis Pater
05391d7b72 Merge pull request #5919 from Fesaluo/feat/xai-imagine-aspect-ratio-9-20
feat(xai): allow grok imagine aspect_ratio 9:20 and 20:9
2026-09-18 12:49:11 +08:00
Luis Pater
28743473c1 feat(codex): append (Devin) suffix to Devin model display names
- Detect Devin models by ID prefix, type, ownership, model registry metadata, or provider.
- Append `(Devin)` to model display names in Codex client responses when not already present.
- Set model `type` to `devin` in home Codex model formatting when served by the Devin provider.
2026-09-18 10:37:52 +08:00
Fesaluo
f049e00b76 feat(xai): also allow grok imagine aspect_ratio 20:9
Grok Imagine already supports the 20:9 wide phone ratio. Accept it
from aspect_ratio and from size, matching the 9:20 whitelist change.
2026-09-18 04:04:24 +08:00
Fesaluo
75bd6a60eb feat(xai): allow grok imagine aspect_ratio 9:20
Grok Imagine already supports the 9:20 phone ratio, but the OpenAI image
handler whitelist dropped it and fell back to 1:1. Accept 9:20 from
aspect_ratio and from size, and keep unknown ratios on the existing
fallback.
2026-09-18 03:58:11 +08:00
Luis Pater
6724a95851 feat(auth): enhance retry logic with enriched error handling and tests
- Introduce `authUnavailableError` with detailed recovery hints, status codes, and retry-after headers.
- Refactor scheduler logic to incorporate local retry deadlines and classify cooldown states without altering quota assessments.
- Refine error propagation to properly handle upstream errors and enrich context for retryable scenarios.
- Add comprehensive tests for retry deadlines, scheduler behavior, and error handling during credential unavailability.
- Ensure embedded recovery hints in HTTP headers for improved client-side handling and diagnostics.

Closes: #5842
2026-09-16 00:13:37 +08:00
Luis Pater
748d576731 feat(pluginhost): propagate forced provider and auth ID in host model execution
- Add `ForcedProvider` and `AuthID` fields to `HostModelExecutionRequest` and `ModelExecutionRequest`.
- Pin credentials via `WithPinnedAuthID` and forward forced provider options during model execution and streaming.
- Forward forced provider and auth ID from plugin host callbacks to model execution handlers.

Closes: #5814
2026-09-14 21:59:32 +08:00
Luis Pater
ca929459f9 test(openai): add tests for responses websocket compaction replay and routing
- Verify observed compaction response replay and retention across subsequent turns.
- Test plugin route overrides, provider auth pinning, and home runtime auth handling during compaction.

Closes: #5811
2026-09-14 19:36:46 +08:00
Luis Pater
f702bc1ac2 feat(codex): preserve native fidelity for responses-lite requests
- Detect native responses-lite requests via headers and client metadata.
- Skip instructions normalization and synthetic session cloaking for native requests.
- Preserve upstream completion output during websocket response forwarding.

Closes: #5780
2026-09-13 15:05:03 +08:00
Luis Pater
5b2785617d feat(plugins): expose model list responses to plugin interceptors
- Add `WriteModelListResponse` to `BaseAPIHandler` to apply plugin interceptors and record request lifecycles for model catalog responses.
- Update OpenAI, Claude, Gemini, Grok, and Codex model listing endpoints to route responses through the unified interceptor helper.

Closes: #5742
2026-09-12 00:04:11 +08:00
Luis Pater
8f23ad0291 fix(codex): inherit template metadata for model aliases and restrict provider capabilities
- Resolve canonical templates using metadata model IDs for model aliases and prefixed routes.
- Apply descriptions, base instructions, and thinking support overrides to matched templates.
- Restrict protocol capabilities and reasoning levels based on provider support.

Closes: #5699
2026-09-11 01:42:30 +08:00
Luis Pater
638ed7e1cc fix(stream): handle split CRLF across chunk boundaries in SSE validation
- Track trailing carriage returns across chunk boundaries in `sseJSONValidationState`.
- Strip leading newline in subsequent chunks to prevent duplicate newline insertion from split CRLF sequences.
- Reset trailing carriage return state upon stream completion.

Closes: #5657
2026-09-10 17:53:25 +08:00
Luis Pater
259130863d fix(openai): preserve nested error details and sequence numbers in responses stream
- Format streaming error payloads with nested error objects matching official OpenAI Responses SSE specifications.
- Extract and propagate sequence numbers from upstream terminal events and framer states.
- Use `json.Number` to prevent precision loss for large integers and token metrics.
- Sanitize sensitive keys recursively across nested error objects without dropping custom fields.
2026-09-10 11:58:36 +08:00
Luis Pater
bd03aabcf1 fix(openai): preserve prewarm input and allow named tool outputs in responses websocket
- Track pending synthetic prewarm response IDs to merge warmup inputs into subsequent delta followups.
- Normalize transcript replacements when followups do not reference the prewarm parent response ID.
- Validate that the `input` field is an array for `response.create` requests.
- Allow `function_call_output` items without a `call_id` when a non-empty tool name is present.

Closes: #5631
2026-09-10 11:20:49 +08:00
Luis Pater
d1a024e940 feat(codex): add support for gpt-image-2.5 models
- Register builtin model definitions for `gpt-image-2.5`, `gpt-image-2.5-flare`, and `gpt-image-2.5-sunburst`.
- Update OpenAI image handlers and request routing to recognize GPT Image 2.5 models.
- Support direct image generation and edit execution for GPT Image 2.5 variants in the Codex executor.
- Apply client visibility overrides to hide new builtin image models where appropriate.
2026-09-10 10:40:53 +08:00
Luis Pater
aedc9e6a39 fix(auth): classify terminal upstream auth failures as non-retryable errors
- Introduce `IsTerminalAuthError` and `NewTerminalAuthError` to identify permanent upstream authentication failures.
- Return terminal auth errors from candidate selection and scheduling when all available credentials fail with unauthorized errors.
- Support `BuildErrorResponseBodyWithError` to format terminal upstream auth errors as non-retryable `upstream_authentication_required` responses.
- Propagate terminal error classifications and the `retryable` field across HTTP and WebSocket response handlers.

Closes: #5645
2026-09-10 10:02:08 +08:00
Luis Pater
3bf787fc1d feat(auth): propagate canonical session id for custom header expansion
- Initialize `util.SessionIDResolver` to resolve session IDs from request contexts, metadata, and headers.
- Ensure canonical session metadata is injected into execution options across execution flows.
- Propagate session context in executors to support `$CPA-SESSION-ID` expansion in custom headers.
- Sync cleared or updated session identities back to execution contexts during conductor execution.

Closes: #5690
2026-09-10 01:43:37 +08:00
sususu
390589159e feat(session): enhance harness hierarchy recognition and deduplicate selector extraction 2026-09-08 17:04:58 +08:00
sususu
580df36423 feat(usage): propagate session and parent session hierarchy to usage reporting queue
- Reuse coresession.ExtractSessionInfo across HTTP headers and request payloads to unify canonical session prefix namespaces with the scheduler.
- Extract hierarchical session identities in two phases: initial extraction from request headers on entry, and authoritative deep extraction once request payloads and metadata are available.
- Support Claude Code multi-level subagents (X-Claude-Code-Agent-Id, metadata.agent_id) and Codex thread fork lineages.
- Propagate SessionID and ParentSessionID across ClientRequestMetadata, UsageReporter, and coreusage.Record without root_session_id.
- Include session_id and parent_session_id in queuedUsageDetail for Home LPushUsage forwarding and Redis consumption with self-loop guards.
- Add comprehensive test coverage for canonical headers, body extraction, ghost parent elimination, and self-referential loop guards.
2026-09-06 10:45:32 +08:00
Luis Pater
649a8bdb6f feat(plugin): omit stream chunk history on payload chunks for schema v5
- Bump plugin schema version to 5 and introduce `SchemaVersionStreamChunkOmitHistory`.
- Omit `HistoryChunks` on payload stream chunks for schema version 5+ to avoid per-chunk cloning and serialization overhead.
- Conditionally accumulate and clone history chunks only when legacy plugins with schema version < 5 are active.

Closes: #5451
2026-09-04 01:19:50 +08:00
Luis Pater
2a6b87aca0 feat(openai): send periodic ping control frames during responses websocket streaming
- Add `writePing` to responses websocket writer to emit Ping control frames.
- Send periodic keep-alive Ping frames based on streaming configuration during response forwarding.
- Reset keep-alive interval upon receiving data chunks and abort session if ping write fails.

Closes: #5413
2026-09-03 21:27:23 +08:00
Luis Pater
291cfb87ef feat(codex): support orphan delegation compatibility via orphan-delegation-compatibility
- Add `codex.orphan-delegation-compatibility` configuration option and mirror it to SDK configuration.
- Convert orphan Codex delegation outputs into standard user messages for requests with `X-Openai-Subagent: collab_spawn`.
- Integrate orphan delegation rewriting into OpenAI responses request handling pipeline.

Closes: #5401
2026-09-03 20:21:24 +08:00
Luis Pater
f2e2d713b2 fix(auth): propagate upstream error causes in auth selection failures
- Attach candidate upstream errors as causes to scheduler availability and cooldown errors.
- Introduce `errorWithCause` wrapper to extract and display upstream error summaries.
- Enrich auth selection error messages with upstream details and preserve model cooldown errors.
- Forward `Retry-After` response headers for model cooldown errors in Claude Code handler.

Closes: #5365
2026-09-01 08:15:51 +08:00
Luis Pater
9721d9939e feat(usage): track streaming execution state in usage records
- Add `Stream` field to usage records, context helpers, and Redis queue payloads.
- Propagate streaming mode across execution handlers, conductors, and usage reporters.

Closes: #5361
2026-08-31 20:29:19 +08:00
Luis Pater
d31b15916d feat(executor): support token usage parsing for plugin executors
- Add `ParsePluginExecutorResponseUsage` to extract token usage from non-streaming plugin responses across Claude, Gemini, Interactions, Antigravity, and OpenAI/Codex protocols.
- Add `ObservePluginExecutorStreamUsage` to observe and aggregate token usage across streaming chunks.

Closes: #5340
2026-08-30 14:51:08 +08:00
hkfires
9a2201c36a fix(auth): forward Home unauthorized upstream errors
Stop refreshing Home-owned OAuth credentials after upstream 401s.
Preserve marked upstream response bodies for direct responses, usage
records, request logs, and websocket handshake failures.
2026-08-29 12:50:44 +08:00
Luis Pater
1cc72b9d13 fix(codex): filter extended reasoning levels for older client versions
- Filter `max` and `ultra` reasoning effort levels for Codex client versions prior to `0.144.0`.
- Extract and forward the `client_version` query parameter across model catalog response handlers.
- Add dotted version parsing and comparison utilities to verify extended reasoning level compatibility.

Closes: #5262
2026-08-27 16:51:33 +08:00
Luis Pater
4b5f1eab25 feat(plugin): support observing upstream websocket response events
- Introduce `WebSocketResponseObserver` capability and bump plugin ABI schema version to 4.
- Forward upstream WebSocket response frames from Codex and xAI executors to configured observers.
- Wire `WebSocketResponseObserver` across API handlers and plugin host dispatchers.

Closes: #5248
2026-08-27 05:30:19 +08:00
Luis Pater
80de901550 fix: preserve multi-reference video durations 2026-08-25 12:44:00 +08:00
Luis Pater
1d5b7612c6 fix(cliproxy): add protocol-aware plugin executor usage parsing for response and streaming payloads
Closes: #5122
2026-08-21 13:05:54 +08:00
Luis Pater
ec105dac94 fix(cliproxy): handle responses/compact auth cooldowns and fallback semantics
- Add compact-specific error classification to mark transient/non-credential failures as availability-neutral instead of triggering cooldown penalties.
- Stop auth fallback immediately on compact request-fault errors (e.g., bad/not-found/unsupported request errors) and return the upstream compact error.
- Preserve existing cooldown behavior for auth/credential faults (`401`, `403`, `429`) while allowing non-auth compact failures to fail fast without tainting normal traffic routing.

Closes: #5031
2026-08-19 22:05:28 +08:00
DragonFSKY
45c90e8d0e fix(websocket): drop consumed compaction triggers
After remote compaction succeeds, the normalized previous request still contains the consumed compaction_trigger. The next WebSocket-to-HTTP merge replays that trigger before the compaction output and current input, causing the upstream 400 error.

Remove compaction_trigger only from previous request items when the previous response contains compaction or compaction_summary. Add a two-step regression test covering trigger normalization followed by compact replay.

Fixes #5041
2026-08-18 04:33:00 +08:00
Luis Pater
92f03e68e3 fix(claude,gemini,openai): preserve upstream stream errors when no data payload is emitted
- Return buffered pending stream errors from `errChan` when the data channel closes in streaming handlers.
- Emit error response instead of silently falling back to normal SSE completion paths.
- Replace Claude’s local `pendingClaudeStreamError` usage with shared `handlers.PendingStreamError`.

Closes: #4710
2026-08-16 04:29:51 +08:00
Luis Pater
e0b4956242 fix(openai): ensure Responses usage includes token detail fields
- add shared `EnsureResponsesUsageDetails` helper to patch `usage` objects with:
  - `output_tokens_details.reasoning_tokens = 0`
  - `input_tokens_details.cached_tokens = 0`
  - for both plain JSON and SSE `data:` frames, including multi-line frames
- apply the helper to OpenAI Response format outputs in non-stream and stream paths across executors/plugins so translated payloads consistently include required usage details
- update websocket/completion payload builders to emit default `usage` detail fields for prewarm/finish responses

Closes: #4985
2026-08-15 15:00:19 +08:00
Luis Pater
810d4dddb3 Merge pull request #4360 from router-for-me/perf/skip-inactive-request-interceptors 2026-08-15 05:12:18 +08:00
Luis Pater
8b02fedec2 Merge pull request #4928 from ramapitecusment/codex/websocket-transcript-allocations-v2
perf(openai): reduce websocket transcript merge allocations
2026-08-15 04:58:38 +08:00
Luis Pater
db35b91e2a feat(openai): add xAI Grok Imagine Image 2.0 image model support
- Added `grok-imagine-image-2.0` as a first-class xAI image base model across validation, canonicalization, and routing checks.
- Registered the model in built-in model definitions so it appears in model metadata.
- Updated image model allowlists and request validation error messaging to include the new model.
- Marked the new model as hidden in client visibility override handling.
2026-08-13 14:37:29 +08:00
Luis Pater
75d2c4a4b4 fix(openai): avoid JSON copies in websocket responses tool-call repair path
Closes: #4925
2026-08-12 20:32:10 +08:00
Ramapitecus
baa11ed6dd fix(openai): match websocket item metadata case-insensitively 2026-08-12 16:52:16 +05:00