Commit Graph

2949 Commits

Author SHA1 Message Date
Luis Pater
cd5af08e31 Merge pull request #5928 from Viggo95/feat/codex-response-model-observability 2026-09-18 21:05:35 +08:00
Luis Pater
cc545cbf90 fix(openai): align tool call messages and preserve ordering on ambiguous outputs
- Add `AlignOpenAIToolCallMessages` to reorder tool results immediately after the matching assistant tool calls while preserving original content and numeric precision.
- Prevent deferred message reordering and call ID guessing when tool outputs are incomplete, duplicate, or missing IDs.
- Normalize translated requests after applying summary configuration in Codex multi-agent execution.

Closes: #5925
2026-09-18 12:32:21 +08:00
Viggo95
25f40d8cf8 feat(codex): record upstream response model and warn on silent model substitution
Codex upstreams can silently serve a different model than the one requested
(HTTP 200, with response.model naming the substitute). The proxy kept no record
of it: nothing logged, nothing reported, only the pass-through response body.

- Add Record.ResponseModel to sdk/cliproxy/usage, aligned with the existing
  ResponseServiceTier field, and emit it from the redis usage queue as the
  optional response_model payload field alongside response_service_tier. Only
  the record for the requested model carries it: additional-model records
  (image generation tool usage) describe a side model the upstream response
  never refers to, and would otherwise look like a substitution downstream.
- Add internal/runtime/executor/helps/response_model.go with
  extractCodexResponseModelEvent (SSE frames and raw JSON, restricted to the events
  that embed the authoritative response object, rejecting non-string and
  oversized upstream model names) and IsCodexModelSubstituted (both sides
  trimmed, lower-cased and stripped of thinking suffixes, dated aliases such as
  gpt-5.6-terra-2026-05-13 accepted in either direction).
- UsageReporter records the served model on the event path and emits the WARN
  when the attempt publishes its usage record, so no logging work happens
  before the first event is forwarded. Repeats are throttled per
  (auth id, requested model, served model) with a 10 minute window, because on
  an affected credential every request is substituted and an unthrottled
  warning would mirror the whole request volume into the logs. The credential
  is labelled auth_index=<index> only: codex credential file names embed the
  account e-mail, which must not be written to the logs at request rate.

Coverage, by entry point. The served model is observed on the HTTP streaming
path (both the bootstrap-buffered handshake and the streaming goroutine), the
HTTP non-streaming Execute loop, the websocket streaming and non-streaming
paths, and the two /responses-shaped image entry points. The remaining codex
entry points cannot report it and are therefore left alone: executeCompact
(/responses/compact answers with a compaction object that has no event type and
no response.model), the two direct image endpoints (/images/generations and
/images/edits answer in the Images API shape and stream image_generation.*
events), and CountTokens (counts locally with tiktoken, never reaching an
upstream).

TokenAccountingSchemaVersion is not bumped: it versions the token accounting
contract (token breakdown semantics), and this change only adds an optional
non-token field that leaves existing consumers and all token math untouched.

Note: response_model ships with the usage record and is the counting source;
the WARN is a throttled alerting signal and must not be used to count
substitutions.

Tests: table-driven unit tests for both helpers over real model ids, reporter
tests covering the published record, the single throttled warning, the absence
of account identifiers in it, concurrent observation and publishing under
-race, the throttle window and its entry bound, an executor-level guard for the
observeCodexTokenEvent wiring and the per-model records, plus a redisqueue
payload assertion for response_model. gofmt, go vet, go test -race on the
touched packages and go test ./... are clean.
2026-09-18 12:31:52 +08:00
Luis Pater
859c486512 fix(codex): normalize and support ultrafast service tier
- Normalize `service_tier` by trimming whitespace and matching case-insensitively.
- Map `fast` to `priority` and preserve `ultrafast` service tier.
- Strip unsupported service tier values and non-string types before forwarding.

Closes: #5924
2026-09-18 11:21:55 +08:00
Luis Pater
28743473c1 feat(codex): append (Devin) suffix to Devin model display names
- Detect Devin models by ID prefix, type, ownership, model registry metadata, or provider.
- Append `(Devin)` to model display names in Codex client responses when not already present.
- Set model `type` to `devin` in home Codex model formatting when served by the Devin provider.
2026-09-18 10:37:52 +08:00
Luis Pater
660a5800e7 fix(xai): restore aliased client web search tool name in responses
- Restore client-defined `web_search` tool names from their alias across SSE, websocket, and buffered execution responses.
- Preserve namespaced tool calls when matching and restoring tool names.

Closes: #5923
2026-09-18 09:57:06 +08:00
Luis Pater
3662d1535a fix(codex): strip item-level and tool output prompt cache breakpoints
- Strip `prompt_cache_breakpoint` from `output` arrays in function call outputs.
- Remove item-level `prompt_cache_breakpoint` fields from input items.

Closes: #5922
2026-09-18 09:29:06 +08:00
Luis Pater
1cce932573 fix(claude): skip retry-after header on overage-only rejections
- Skip parsing the `Retry-After` header when a rate limit rejection is overage-only to prevent global credential cooldowns.
- Allow exponential backoff to handle model recovery while keeping shared subscription windows available for other models.

Closes: #5920
2026-09-18 09:15:48 +08:00
Luis Pater
c616193a6c fix(xai): unify forced hosted tool choice normalization
- Generalize forced hosted tool choice handling across both image generation and web search.
- Normalize forced web search tool choices to string mode and isolate the target tool in the tools list.
- Strip hosted web search from mixed `allowed_tools` definitions.
- Skip native `x_search` injection when a hosted tool is exclusively required.

Closes: #5916
2026-09-18 02:20:41 +08:00
Luis Pater
44eaef0009 feat(claude): support model-level cooling and scope overage rate limits
- Add `claude.model-level-cooling` configuration to scope rate limit cooldowns to the requested model.
- Treat overage-only and spend cap rejections as model-scoped when shared subscription windows remain healthy.
- Propagate model-level cooling settings into streaming, token counting, and direct execution error classifiers.

Closes: #5915
2026-09-18 01:55:36 +08:00
Luis Pater
b6fe4f20c4 fix(devin): handle orphaned tool results and normalize function result payloads
- Match tool results against pending tool calls and downgrade unmatched results to user messages.
- Prevent downgraded orphaned tool results from consuming images intended for user turns.
- Unwrap protocol wrapper envelopes and extract structured text parts while preserving arbitrary business JSON.
- Provide a placeholder for empty or whitespace-only tool results.

Closes: #5911
2026-09-18 01:34:38 +08:00
Luis Pater
9e10db53ad fix(devin): aggregate tool calls by id and track cache write tokens
- Track and aggregate tool calls by call ID instead of slot index in streaming and buffered execution.
- Support raw arguments from invalid JSON fields for custom tool calls.
- Parse usage field 4 as cache write tokens instead of adding to prompt tokens.
- Align client metadata with the default client name and drop deprecated tag 28.

Closes: #5910
2026-09-18 01:04:43 +08:00
Luis Pater
0b55053944 fix(management): reject unresolved token placeholders in api-call
- Return an error when the `$TOKEN$` placeholder cannot be resolved or the credential for `auth_index` is missing.
- Ensure `$TOKEN$` substitutions in headers and request payloads fail fast instead of proceeding with empty values.
- Support resolving and refreshing provider OAuth tokens for API calls.

Closes: #5838
2026-09-17 23:43:04 +08:00
Luis Pater
b715526add feat(plugin): support scheduling across priorities
- Add `SchedulerAcrossPriorities` capability to allow plugin schedulers to receive candidates across all priority tiers.
- Propagate scheduler priority preferences through the plugin host and RPC capabilities.
- Update auth manager selection logic to supply candidates across all priorities when the scheduler opts in.

Closes: #5894
2026-09-17 21:21:56 +08:00
Luis Pater
afba07ba26 fix(config): preserve plugin configurations when saving yaml
- Replace the plugin configs subtree directly instead of merging to prevent zero values, booleans, and empty collections from being pruned as defaults.
- Ensure stale plugin keys and removed configuration entries are properly cleaned up while preserving comments.

Closes: #5907
2026-09-17 21:01:31 +08:00
Luis Pater
64c9433fd2 fix(devin): support images in tool results
- Extract and attach images from tool results to corresponding tool prompts by call ID.
- Preserve structured business JSON and raw objects in function result content.
- Prepend image headers to tool prompt content when images are present.

Closes: #5893
2026-09-17 20:00:46 +08:00
camy-x
76ac75e68a feat(management): paginate auth file listings 2026-09-17 14:37:36 +08:00
Luis Pater
8c664b2fed Merge pull request #5896 from router-for-me/translator
feat(translator): preserve model metadata in requests
2026-09-17 13:52:35 +08:00
Luis Pater
ad088a8795 fix(devin): filter automation update tools and sanitize tool descriptions
- Filter out `automation_update` tools and sanitize tool descriptions in Devin wire requests and logs.
- Strip additional Codex prompt directives from system messages.
- Support `children` field fallback when collecting namespace tools.
2026-09-17 13:51:26 +08:00
hkfires
a9e92b8145 feat(translator): preserve model metadata in requests 2026-09-17 13:10:47 +08:00
Luis Pater
f668ac417d fix(schema): strip unsupported schema identifier keywords
- Remove `id`, `$anchor`, `$vocabulary`, `$dynamicRef`, and `$dynamicAnchor` when cleaning JSON schemas for Gemini.

Closes: #5888
2026-09-17 10:10:50 +08:00
Luis Pater
311efcb3a2 test(executor): use metaUserAgent constant in meta executor test
Closes: #5885
2026-09-17 10:01:49 +08:00
Luis Pater
c4982e846e fix(executor): strip relayed tool result images for text-only models
- Replace tool image placeholders with omission markers for compatibility with text-only upstream models.
- Strip synthetic image relay notices and image parts from user messages.

Closes: #5884
2026-09-17 09:57:13 +08:00
Luis Pater
c8ec91fbd9 Merge pull request #5509 from trukhinyuri/fix/codex-exec-user-agent
fix(codex): recognize codex_exec user agent for multi-agent v2 optimization
2026-09-17 08:16:05 +08:00
Luis Pater
b681a1e0f7 feat(translator): wrap demoted mid-session system messages in system-reminder envelope
- Add reusable `SystemReminderText` helper to format directives within `<system-reminder>` envelopes.
- Wrap demoted mid-session system and developer messages in Gemini and Antigravity OpenAI request translators so upstream models treat them as system directives.

Closes: #5878
2026-09-17 05:22:17 +08:00
Luis Pater
c2bb91d2cb fix(devin): buffer content deltas to handle late thinking signatures
- Buffer content and tool call deltas while thinking is active so late-arriving thinking signatures can be attached before closing thinking blocks.
- Ensure pending actions and open steps are properly flushed and closed on stream completion or trailer errors.

Closes: #5873
2026-09-17 04:49:55 +08:00
Luis Pater
77820cb2f4 feat(translator): support alternative reasoning fields in OpenAI Claude translation
- Extract reasoning text sequentially from `reasoning_content`, `reasoning`, and `reasoning_details`.
- Unify reasoning extraction across streaming deltas and non-streaming message conversions.

Closes: #5872
2026-09-17 04:17:30 +08:00
Luis Pater
4613cfd44e Merge pull request #5870 from avabbbb/fix/devin-high-demand-rate-limit
fix(devin): classify high-demand errors as rate limits
2026-09-17 03:59:25 +08:00
Luis Pater
7fcbdf8896 feat(translator): enhance web search streaming and citation mapping in OpenAI Responses
- Add incremental merging for Gemini `groundingMetadata` with chunk index remapping and query deduplication.
- Implement rune offset mapping across multipart messages for accurate `url_citation` annotations.
- Manage full streaming lifecycle for web search calls, emitting `searching`, `completed`, and output item done events.
- Stream incremental citation annotations via `response.output_text.annotation.added` events.
- Support `web_search_preview_2025_03_11` as a recognized web search tool type.
2026-09-17 03:51:48 +08:00
Luis Pater
8c6d4dcde8 Merge pull request #5862 from sususu98/feat/antigravity-websearch
feat(antigravity): support web search translation and URL resolution in OpenAI Responses
2026-09-16 20:21:46 +08:00
Luis Pater
7c32971b91 fix(executor): prevent stream failure on client disconnect after claude completion
- Break stream scan loops immediately when upstream completion is reached.
- Skip scanner error handling and cancellation checks when `upstreamCompleted` is true.

Closes: #5866
2026-09-16 20:00:27 +08:00
Luis Pater
44f8343f30 Merge pull request #5867 from avabbbb/fix/devin-swe-1-6-slow
fix(devin): support swe-1-6-slow model variant
2026-09-16 19:36:05 +08:00
bekkilove
d44901f91c fix(devin): classify high-demand errors as rate limits
Devin upstream occasionally encodes transient capacity failures using
Connect code permission_denied with message containing 'high demand'.
CPA's ParseDevinTrailerError currently maps every permission_denied to
HTTP 403. The auth cooldown manager interprets 403 as a 30-minute model
permission cooldown, keeping a recovered model locally unavailable.

This fix narrowly reclassifies the observed high-demand variant as
HTTP 429 (Too Many Requests), so it enters the quota/retry cooldown
path instead of the long permission denial path. Genuine
permission_denied errors (model access denied, plan entitlement denied,
etc.) remain HTTP 403.

Tests: 4 new cases in TestParseDevinTrailerError covering transient
high-demand (429), genuine permission error (403), resource_exhausted
unchanged (429), and case-insensitive matching. All existing tests
pass with no regressions.
2026-09-16 19:24:08 +08:00
Luis Pater
7def842554 feat(translator): preserve tool names in tool result messages
- Track tool names by tool use ID from assistant tool call blocks.
- Populate the `name` field on converted tool result messages when a matching ID exists.

Closes: #5859
2026-09-16 18:10:02 +08:00
sususu
ef63d2e7fa feat(antigravity): support web search translation and URL resolution in OpenAI Responses
- Add bidirectional web search translation between OpenAI Responses API and Gemini/Antigravity
- Map Responses web_search tool to Antigravity web_search requestType envelope and googleSearch
- Map Google groundingMetadata to Responses web_search_call output item and url_citation annotations
- Buffer streaming text deltas while awaiting groundingMetadata so web_search_call strictly precedes message in SSE events and response.completed.output
- Derive search stream mode from effective translated request (requestRawJSON) to accurately support model aliases and rewrites
- Calculate streaming and non-streaming URL citation Unicode character (rune) offsets on full accumulated text, eliminating multi-byte CJK truncation and clamping
- Prioritize models.json native_capabilities.web_search explicit false as absolute veto before checking dynamic Antigravity probe capability
- Isolate Antigravity web search gating to Antigravity-specific model capabilities
- Suppress native googleSearch in Antigravity chat fallback when tools are mixed with function declarations
- Support Responses allowed_tools tool_choice containing web search and concatenate multi-part text queries
- Resolve Vertex Search grounding redirect URLs to target destination URLs in Antigravity executor
- Add comprehensive unit test coverage for stream/non-stream translation, late grounding, CJK offsets, model aliases, mixed tools, allowed_tools, and URL resolution
2026-09-16 16:56:31 +08:00
bekkilove
dea4ce8aeb fix(devin): support swe-1-6-slow model variant 2026-09-16 16:49:49 +08:00
Luis Pater
e3e97ad9be feat(translator): introduce namespace-aware tool name capping and collision handling
- Add logic to cap long OpenAI Responses tool names to 64 characters while retaining the most identifying portion.
- Implement disambiguation for truncation-induced collisions by appending suffixes to ensure unique tool names.
- Prevent truncated names from overlapping with namespace-less local names or previously declared aliases.
- Refactor declaration processing logic to consistently apply capping and disambiguation across requests and replayed history.
- Extend test coverage to validate correct capping, disambiguation, and namespace recovery for tool calls in streaming and non-streaming scenarios.

Closes: #5856
2026-09-16 11:29:50 +08:00
Luis Pater
e54a8e97ff feat(translator): enhance cache usage handling and input token computation logic
- Refactor input token computation to account for inclusive and exclusive cache scenarios, supporting fields like `input_tokens`, `total_input_tokens`, and `prompt_tokens`.
- Add support for distinguishing and handling `cache_read_input_tokens` and `cache_creation_input_tokens` separately.
- Update Claude translator logic to ensure correct token usage attribution, including fallback mechanisms for missing fields.
- Extend test coverage to validate various cache hit, partial, and write scenarios across streaming and non-streaming modes.

Closes: #5854
2026-09-16 11:14:47 +08:00
Luis Pater
6f908cbcff feat(translator): enhance finish reason handling and tool call translations
- Refine finish reason assignment logic in OpenAI interactions to account for incomplete states (`length`, `content_filter`) and tool call indices.
- Implement changes to normalize tool call indexing to ensure contiguous 0-based indexing.
- Add `status` and `incomplete_details` fields to OpenAI response payloads for enriched status handling.
- Improve logic for handling `generation_config` fields during interactions request conversion.
- Add robust defaulting for missing usage metrics in response payloads.
- Ensure comprehensive test coverage for all enhanced scenarios.

Closes: #5851
2026-09-16 10:24:47 +08:00
sususu
f51d3ae93b fix(interactions): strip invalid id from function_result and redundant call_id from function_call
- Remove unexpected id parameter from function_result in Claude and OpenAI Chat Interactions request translators to comply with Google Interactions API schema and fix HTTP 400 (Unknown parameter 'id' at 'input[N]').
- Remove redundant call_id parameter from function_call in Claude Interactions request translator to match schema requirements.
- Propagate is_error flag between Claude tool_result and Interactions function_result.
- Align Devin executor to prioritize call_id for function_result steps.
- Add regression tests covering parameter schemas and end-to-end executor request generation.

Closes: #5828
2026-09-16 00:31:31 +08:00
Luis Pater
8335eac731 feat(config): add Meta OAuth aliasing and error rules support
- Introduced `meta` as a supported OAuth channel for model aliasing and request-scoped error rules.
- Added alias config to rename `muse-spark-1.3` as `muse-latest` with enforced mapping and fork behavior.
- Extended error rules for `meta` to classify "context_length_exceeded" responses with status `400` as `stop` actions.
- Updated examples and synthesizer logic to merge exclusions and apply Meta-specific mechanisms.
2026-09-15 20:08:12 +08:00
Luis Pater
65348b9594 feat(executor): add native Meta (Muse Code) provider integration
- Introduced `MetaExecutor` to support native Meta (Muse Code) API operations, including `Execute`, `ExecuteStream`, and `CountTokens`.
- Added logic to handle Meta-specific `responses` endpoint, including request preparation, enriched authentication, header management, and response translation.
- Implemented token counting functionality with custom `CountTokens` logic.
- Updated `meta_executor` to process streamed and non-streamed responses, maintaining compatibility with the Codex schema.
- Added utility functions to handle events, errors, and enriched metadata common to Meta API requests.
- Extended `models.json` to include Meta's updated configuration with additional thinking levels like "minimal" and "max."
- Introduced comprehensive tests to validate Meta-specific execution, streaming, header assignments, and response formatting.
2026-09-15 18:12:49 +08:00
Luis Pater
42ca5d3412 Merge PR #5502 feat/meta-provider into cpa/muse
Bring in native Meta (Muse Code) provider support while keeping the
existing Devin integration and original Meta commit history.
2026-09-15 12:05:40 +08:00
Luis Pater
772c63c8e2 feat(translator): refine tool call argument validation for finish reason logic
- Introduce `hasValidToolCallArguments` to validate tool call arguments in OpenAI response translation.
- Update `effectiveOpenAIFinishReason` logic to consider argument validity for determining finish reasons like `tool_calls` or `length`.
- Extend finish reason assignment during response parsing to ensure robust handling of invalid or truncated tool call arguments.
- Add comprehensive test coverage for various tool call argument scenarios and finish reason outputs.

Closes: #5836
2026-09-15 10:51:46 +08:00
Luis Pater
512c453e09 Merge pull request #5803 from anpicasso/feat/plugin-quota-summary
feat(plugin-quota): add typed summary metrics
2026-09-15 08:35:47 +08:00
Luis Pater
7bbfeaf8a7 feat(executor): promote reasoning content as summary and sanitize inputs
- Introduce `promoteOpenAIResponsesReasoningTextToSummary` to move `reasoning_text` parts from content to summary when the summary is empty.
- Clear `reasoning.content` to comply with Codex schema constraints (`maxItems: 0`).
- Implement safeguards for handling cleartext reasoning, preserving valid encrypted content, and stripping invalid `encrypted_content`.
- Add comprehensive tests to validate behavior for promoting reasoning texts, preserving existing summaries, and ensuring sanitization rules.

Closes: #5825
2026-09-15 03:42:08 +08:00
Luis Pater
8c984672a6 feat(translator): handle orphan function outputs as user text in OpenAI and Gemini translators
- Introduce logic to convert orphaned or unpaired `function_call_output` into user text instead of mismatched tool responses.
- Add `buildOpenAIResponsesStandaloneToolOutputTextParts` and `appendStandaloneResponsesToolOutputAsUser` helpers for managing orphaned outputs.
- Update OpenAI, Gemini, and Antigravity translators to ensure unmatched tool outputs are surfaced correctly.
- Add comprehensive test cases for both Gemini and OpenAI translator workflows to validate behavior for orphan function outputs and explicit call IDs.

Closes: #5831
2026-09-15 03:26:12 +08:00
Luis Pater
1fac8cc0c8 feat(executor): strip unsupported id fields from Gemini interactions payload
- Introduce `sanitizeGeminiInteractionsUnsupportedInputIDs` to remove `id` fields from Gemini interactions `input` and `content` items.
- Update execution logic to apply the sanitization process before upstreaming payloads.
- Add tests to ensure payloads are correctly sanitized and validate `call_id` pairing for Gemini interactions.

Closes: #5828
2026-09-15 02:21:45 +08:00
Luis Pater
e3cbe437d0 feat(auth): add tests to ensure async operations don’t block unrelated actions
- Add tests to verify that runtime hooks, auth modifications, and related updates do not block on unrelated operations such as antigravity probes or plugin virtual models.
- Introduce detailed scenarios, e.g., stale disables, batch operations, and conflict handling during concurrent auth updates.
- Refactor locking mechanisms to avoid unnecessary blocking during hooks and model registrations.

Fixed: #5813
Closes: #5773
2026-09-15 01:23:26 +08:00
Luis Pater
748d576731 feat(pluginhost): propagate forced provider and auth ID in host model execution
- Add `ForcedProvider` and `AuthID` fields to `HostModelExecutionRequest` and `ModelExecutionRequest`.
- Pin credentials via `WithPinnedAuthID` and forward forced provider options during model execution and streaming.
- Forward forced provider and auth ID from plugin host callbacks to model execution handlers.

Closes: #5814
2026-09-14 21:59:32 +08:00