Commit Graph

1095 Commits

Author SHA1 Message Date
SamGu-NRX
b97f71da0c fix(codex): send the ChatGPT routing hint native Codex sends
- Send X-Codex-Routing-Hint ("model=<slug>" plus ";tier=<service_tier>")
  on ChatGPT-backend requests over HTTP, compact and websocket handshakes,
  matching openai/codex rust-v0.155.0 build_routing_hint_header. Translated
  requests previously carried service_tier=priority only in the body.
- Derive the hint from the final upstream body, replacing a hint a native
  client forwarded, so it names the model and tier actually sent.
- Keep operator precedence: an auth header rule that resolves to a value
  wins, a dynamic rule that resolves to nothing falls back to the derived
  hint, and models.json override_header still applies last.
- Leave API-key requests unchanged, and keep websocket reuse as native Codex
  does: an open connection keeps its handshake hint when the tier changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 19:56:53 -05:00
Luis Pater
2fe9932bb9 feat(usage): track execution request ID and trace ID across usage records
- Add `RequestID` (unique execution UUID) and `TraceID` (inbound request ID) fields to usage records and plugin API
- Generate unique UUIDs per model attempt in `UsageReporter` while linking to parent trace context
- Expose `execution_id` and `trace_id` in Redis queue usage payloads while preserving legacy `request_id`
2026-09-23 19:31:39 +08:00
Luis Pater
bd584a7523 feat(claude): support Claude Code 2.1.280 feature-gated betas
- Add header definitions and ordering for per-turn control, timing, inline tools, clear-at, dangerous tool use, thinking binding/resumption, and prompt caching evict betas
- Dynamically assemble beta headers based on model capabilities and request body fields matching Claude Code 2.1.280 behavior

Closes: #6054
2026-09-23 08:11:33 +08:00
Luis Pater
779bf317e0 feat(claude): bump claude cli baseline to 2.1.280 and support tool changes beta
- Update default Claude CLI version to 2.1.280 in cloaking and device profile
- Add `mid-conversation-tool-changes-2026-07-01` beta header support for supported models

Closes: #6054
2026-09-23 05:07:49 +08:00
Luis Pater
f351924f42 feat(codex): add support for disable codex cloaking per credential
Closes: #6034
2026-09-22 22:39:20 +08:00
Luis Pater
4b5adbbe9a feat(openai): support video input in responses request translation
- Convert `input_video` and `video_url` content parts to chat completion video parts
- Preserve video URL and processing options for upstream handling

Closes: #6037
2026-09-22 21:38:10 +08:00
Luis Pater
320100ecf7 fix(codex): strip octal NUL pattern escapes from tool schemas
- Strip octal NUL (`\0`) escapes in regex patterns to prevent rejection by strict upstream JSON Schema validators
- Update tool schema normalization fast-path to check for serialized `\0` escape sequences

Closes: #6025
2026-09-22 18:25:48 +08:00
Luis Pater
a26cf2a8c2 perf(executor): avoid discarded copies of Gemini and Kimi stream lines
Read Gemini interactions scanner lines directly because emitted frames are already cloned before reuse, and allocate each Kimi responses passthrough chunk once so it stays owned without a second copy.
2026-09-22 08:32:35 +08:00
Luis Pater
f6d582cb68 Merge pull request #5678 from cookerpapa/perf/codex-stream-copy
perf(codex): avoid discarded copies of SSE data lines
2026-09-22 08:27:31 +08:00
Luis Pater
639b7f1126 perf(executor): reuse identical Claude and Gemini request translations
Skip the second compatibility translation for Claude, Gemini, and Gemini
interactions when the baseline and working payloads share a backing array
and no plugin hooks are installed. Keep independent buffers, the existing
hook order, and native interactions copies.
2026-09-22 05:16:25 +08:00
Luis Pater
5083f65037 Merge pull request #5990 from jack-5566/fix/responses-request-cache-pr
perf: cache Responses request validation and index tool declarations
2026-09-22 04:51:07 +08:00
Luis Pater
7f63f8c8d2 Merge pull request #5680 from cookerpapa/perf/codex-tool-schema-batch
perf(codex): batch tool schema rewrites
2026-09-22 04:23:03 +08:00
Luis Pater
d582067c06 feat(executor): support execution-scoped request proxy overrides
- Add context helpers to manage execution-scoped proxy overrides
- Prioritize request proxy over credential and global proxy settings across HTTP and uTLS clients
- Propagate request proxy in conductor execution and plugin host adapters
- Isolate Codex WebSocket connection reuse by proxy endpoint

Closes: #6013
2026-09-22 03:54:34 +08:00
sususu
bf44a7f892 style: gofmt codex bootstrap and cache-control tests 2026-09-22 00:18:05 +08:00
sususu
ed751ea08a fix(claude): gate fallback-credit beta on fallback tokens or explicit fallbacks 2026-09-21 10:17:25 +08:00
Luis Pater
fd5cd228a8 test(executor): allow non-negative TTFT in xAI usage record assertions
- Accept zero TTFT duration to accommodate sub-tick loopback responses
- Add unit test verifying zero TTFT validation in usage records

Closes: #5996
2026-09-21 10:15:52 +08:00
Luis Pater
a5ab69521f feat: add full support for kimi.ai oauth and runtime execution 2026-09-21 04:13:44 +08:00
Luis Pater
563865e77a test: fix timing synchronization and channel races in streaming tests
- Wait for bootstrap initialization before advancing mock clock in Codex WebSocket tests
- Reorder data channel closures in OpenAI responses stream error tests
- Drain stream chunks in byte-cap release test to prevent dangling goroutines

Closes: #5981
2026-09-21 00:39:28 +08:00
jack
8a6a39684d perf: index Responses tools and cache stream request validation 2026-09-20 23:10:57 +09:00
sususu
c52ca7bd4e docs(claude): clarify patch version floor rationale for native passthrough 2026-09-20 19:26:48 +08:00
camy-x
83a4913aa4 fix(claude): accept newer patch releases as native clients (#5820) 2026-09-20 19:15:51 +08:00
Luis Pater
42c9680eee feat(executor): support duplex streaming for codex websockets
- Implement `streamCodexDuplex` to handle bidirectional WebSocket interactions for Codex
- Support steering continuations (`response.steer`) and queued creations (`response.create`, `response.append`)
- Preserve response settings and state transitions across steering updates and automatic successors
- Isolate connection-level failures with `codexDuplexConnectionError` to prevent unnecessary credential cooldowns

Closes: #5968
2026-09-20 00:02:10 +08:00
Luis Pater
c93978c4ea fix(schema): normalize true boolean subschemas and strip unsupported keywords
- Normalize boolean `true` subschemas across root, properties, items, definitions, and dependencies into empty object schemas `{}`
- Include `dependencies` in schema definition traversal and normalization
- Strip unsupported schema keywords including `additionalItems`, `unevaluatedProperties`, `unevaluatedItems`, and `contentSchema`

Closes: #3551
2026-09-19 03:07:33 +08:00
Luis Pater
b6d1f050af fix(executor): buffer post-tool text to order devin tool calls before assistant response
- Buffer text deltas arriving after tool calls in streaming execution and flush them after closing active tool slots
- Prioritize tool call deltas over content text deltas during Connect frame processing
- Separate pre-tool and post-tool text parts in non-streaming responses to preserve step ordering
- Correct thought step index resolution when closing thoughts during content chunk emission

Closes: #5951
2026-09-19 02:41:17 +08:00
Luis Pater
22392c537d fix(executor): restore hybrid passthrough mcp tool names in claude oauth
- Track declared passthrough MCP tools in `claudeMCPAliasResolver`
- Recover hybrid tool names when the model prepends a virtual server prefix or replaces the caller's server with the virtual server
- Ensure client-tool alias matches take precedence over passthrough recovery and reject ambiguous passthrough tool matches

Closes: #5949
2026-09-19 02:13:11 +08:00
Luis Pater
690f4f3116 feat(executor): support use-max-completion-tokens for openai compatibility models
- Add `use-max-completion-tokens` setting to OpenAI compatibility model configuration
- Provide token normalization between `max_tokens` and `max_completion_tokens` based on model preference

Closes: #5939
2026-09-19 01:54:05 +08:00
Luis Pater
f86a33f721 fix(executor): restrict claude advisor tool check to server tool use
- Only detect `server_tool_use` for advisor tool invocations in Claude conversation history
- Prevent client-side tools named `advisor` using standard `tool_use` from improperly triggering advisor cloaking restrictions

Closes: #5934
2026-09-18 23:32:00 +08:00
Luis Pater
81d6ba7746 fix(codex): preserve reasoning content and IDs for compat models in responses
- Resolve model `is_compat` status via model info, config index, or model entries in Codex executor
- Skip wiping `reasoning.content` and stripping reasoning IDs during Responses sanitization when `is_compat` is enabled
- Pass resolved compat flag across standard, streaming, compact, and WebSocket Codex request flows

Closes: #5930
2026-09-18 23:02:54 +08:00
Luis Pater
e84e248c51 fix(executor): record expected Devin upstream model and bound stream observer memory
- Set expected upstream model for Devin in both streaming and non-streaming paths to avoid false positive substitution warnings on intentional mappings
- Account for line overhead and enforce max lines per stream event in StreamResponseModelObserver, dropping overflowed events until the event boundary
- Add unit tests for Devin intentional mappings and stream observer bounded memory behavior
2026-09-18 22:08:35 +08:00
Luis Pater
cde7d57e44 fix(executor): robust response model observability across meta, kimi, and openai-compat streams
- observe response model from terminal sourceEvent in Meta non-stream multi-event SSE responses
- record expected upstream model in UsageReporter to prevent false substitution warnings on Kimi canonical mappings
- use bounded stream observer to extract response models across image stream chunk boundaries
- add regression tests for Meta SSE non-stream, Kimi model mappings, and chunked image streams
2026-09-18 21:53:19 +08:00
Luis Pater
f8467f07dc fix(executor): record authentic Devin and Gemini Interactions response models
- Devin: extract authentic upstream model name from parsed Usage.ModelName
  instead of synthesized interactions JSON or binary Connect frames. Keep
  response model empty when upstream does not report one.
- Gemini: support interaction.model and event_type terminal semantics
  in response model extractors for Gemini Interactions streaming.
- Add unit tests for Devin and Gemini Interactions response model observability.
2026-09-18 21:32:59 +08:00
Luis Pater
e9463ff5a7 feat(executor): extend response model recording and substitution warnings to all providers 2026-09-18 21:22:43 +08:00
Luis Pater
cd5af08e31 Merge pull request #5928 from Viggo95/feat/codex-response-model-observability 2026-09-18 21:05:35 +08:00
Luis Pater
cc545cbf90 fix(openai): align tool call messages and preserve ordering on ambiguous outputs
- Add `AlignOpenAIToolCallMessages` to reorder tool results immediately after the matching assistant tool calls while preserving original content and numeric precision.
- Prevent deferred message reordering and call ID guessing when tool outputs are incomplete, duplicate, or missing IDs.
- Normalize translated requests after applying summary configuration in Codex multi-agent execution.

Closes: #5925
2026-09-18 12:32:21 +08:00
Viggo95
25f40d8cf8 feat(codex): record upstream response model and warn on silent model substitution
Codex upstreams can silently serve a different model than the one requested
(HTTP 200, with response.model naming the substitute). The proxy kept no record
of it: nothing logged, nothing reported, only the pass-through response body.

- Add Record.ResponseModel to sdk/cliproxy/usage, aligned with the existing
  ResponseServiceTier field, and emit it from the redis usage queue as the
  optional response_model payload field alongside response_service_tier. Only
  the record for the requested model carries it: additional-model records
  (image generation tool usage) describe a side model the upstream response
  never refers to, and would otherwise look like a substitution downstream.
- Add internal/runtime/executor/helps/response_model.go with
  extractCodexResponseModelEvent (SSE frames and raw JSON, restricted to the events
  that embed the authoritative response object, rejecting non-string and
  oversized upstream model names) and IsCodexModelSubstituted (both sides
  trimmed, lower-cased and stripped of thinking suffixes, dated aliases such as
  gpt-5.6-terra-2026-05-13 accepted in either direction).
- UsageReporter records the served model on the event path and emits the WARN
  when the attempt publishes its usage record, so no logging work happens
  before the first event is forwarded. Repeats are throttled per
  (auth id, requested model, served model) with a 10 minute window, because on
  an affected credential every request is substituted and an unthrottled
  warning would mirror the whole request volume into the logs. The credential
  is labelled auth_index=<index> only: codex credential file names embed the
  account e-mail, which must not be written to the logs at request rate.

Coverage, by entry point. The served model is observed on the HTTP streaming
path (both the bootstrap-buffered handshake and the streaming goroutine), the
HTTP non-streaming Execute loop, the websocket streaming and non-streaming
paths, and the two /responses-shaped image entry points. The remaining codex
entry points cannot report it and are therefore left alone: executeCompact
(/responses/compact answers with a compaction object that has no event type and
no response.model), the two direct image endpoints (/images/generations and
/images/edits answer in the Images API shape and stream image_generation.*
events), and CountTokens (counts locally with tiktoken, never reaching an
upstream).

TokenAccountingSchemaVersion is not bumped: it versions the token accounting
contract (token breakdown semantics), and this change only adds an optional
non-token field that leaves existing consumers and all token math untouched.

Note: response_model ships with the usage record and is the counting source;
the WARN is a throttled alerting signal and must not be used to count
substitutions.

Tests: table-driven unit tests for both helpers over real model ids, reporter
tests covering the published record, the single throttled warning, the absence
of account identifiers in it, concurrent observation and publishing under
-race, the throttle window and its entry bound, an executor-level guard for the
observeCodexTokenEvent wiring and the per-model records, plus a redisqueue
payload assertion for response_model. gofmt, go vet, go test -race on the
touched packages and go test ./... are clean.
2026-09-18 12:31:52 +08:00
Luis Pater
660a5800e7 fix(xai): restore aliased client web search tool name in responses
- Restore client-defined `web_search` tool names from their alias across SSE, websocket, and buffered execution responses.
- Preserve namespaced tool calls when matching and restoring tool names.

Closes: #5923
2026-09-18 09:57:06 +08:00
Luis Pater
1cce932573 fix(claude): skip retry-after header on overage-only rejections
- Skip parsing the `Retry-After` header when a rate limit rejection is overage-only to prevent global credential cooldowns.
- Allow exponential backoff to handle model recovery while keeping shared subscription windows available for other models.

Closes: #5920
2026-09-18 09:15:48 +08:00
Luis Pater
c616193a6c fix(xai): unify forced hosted tool choice normalization
- Generalize forced hosted tool choice handling across both image generation and web search.
- Normalize forced web search tool choices to string mode and isolate the target tool in the tools list.
- Strip hosted web search from mixed `allowed_tools` definitions.
- Skip native `x_search` injection when a hosted tool is exclusively required.

Closes: #5916
2026-09-18 02:20:41 +08:00
Luis Pater
44eaef0009 feat(claude): support model-level cooling and scope overage rate limits
- Add `claude.model-level-cooling` configuration to scope rate limit cooldowns to the requested model.
- Treat overage-only and spend cap rejections as model-scoped when shared subscription windows remain healthy.
- Propagate model-level cooling settings into streaming, token counting, and direct execution error classifiers.

Closes: #5915
2026-09-18 01:55:36 +08:00
Luis Pater
b6fe4f20c4 fix(devin): handle orphaned tool results and normalize function result payloads
- Match tool results against pending tool calls and downgrade unmatched results to user messages.
- Prevent downgraded orphaned tool results from consuming images intended for user turns.
- Unwrap protocol wrapper envelopes and extract structured text parts while preserving arbitrary business JSON.
- Provide a placeholder for empty or whitespace-only tool results.

Closes: #5911
2026-09-18 01:34:38 +08:00
Luis Pater
9e10db53ad fix(devin): aggregate tool calls by id and track cache write tokens
- Track and aggregate tool calls by call ID instead of slot index in streaming and buffered execution.
- Support raw arguments from invalid JSON fields for custom tool calls.
- Parse usage field 4 as cache write tokens instead of adding to prompt tokens.
- Align client metadata with the default client name and drop deprecated tag 28.

Closes: #5910
2026-09-18 01:04:43 +08:00
Luis Pater
64c9433fd2 fix(devin): support images in tool results
- Extract and attach images from tool results to corresponding tool prompts by call ID.
- Preserve structured business JSON and raw objects in function result content.
- Prepend image headers to tool prompt content when images are present.

Closes: #5893
2026-09-17 20:00:46 +08:00
Luis Pater
8c664b2fed Merge pull request #5896 from router-for-me/translator
feat(translator): preserve model metadata in requests
2026-09-17 13:52:35 +08:00
Luis Pater
ad088a8795 fix(devin): filter automation update tools and sanitize tool descriptions
- Filter out `automation_update` tools and sanitize tool descriptions in Devin wire requests and logs.
- Strip additional Codex prompt directives from system messages.
- Support `children` field fallback when collecting namespace tools.
2026-09-17 13:51:26 +08:00
hkfires
a9e92b8145 feat(translator): preserve model metadata in requests 2026-09-17 13:10:47 +08:00
Luis Pater
311efcb3a2 test(executor): use metaUserAgent constant in meta executor test
Closes: #5885
2026-09-17 10:01:49 +08:00
Luis Pater
c4982e846e fix(executor): strip relayed tool result images for text-only models
- Replace tool image placeholders with omission markers for compatibility with text-only upstream models.
- Strip synthetic image relay notices and image parts from user messages.

Closes: #5884
2026-09-17 09:57:13 +08:00
Luis Pater
c2bb91d2cb fix(devin): buffer content deltas to handle late thinking signatures
- Buffer content and tool call deltas while thinking is active so late-arriving thinking signatures can be attached before closing thinking blocks.
- Ensure pending actions and open steps are properly flushed and closed on stream completion or trailer errors.

Closes: #5873
2026-09-17 04:49:55 +08:00
Luis Pater
4613cfd44e Merge pull request #5870 from avabbbb/fix/devin-high-demand-rate-limit
fix(devin): classify high-demand errors as rate limits
2026-09-17 03:59:25 +08:00
Luis Pater
7fcbdf8896 feat(translator): enhance web search streaming and citation mapping in OpenAI Responses
- Add incremental merging for Gemini `groundingMetadata` with chunk index remapping and query deduplication.
- Implement rune offset mapping across multipart messages for accurate `url_citation` annotations.
- Manage full streaming lifecycle for web search calls, emitting `searching`, `completed`, and output item done events.
- Stream incremental citation annotations via `response.output_text.annotation.added` events.
- Support `web_search_preview_2025_03_11` as a recognized web search tool type.
2026-09-17 03:51:48 +08:00