Commit Graph

228 Commits

Author SHA1 Message Date
sususu98
f7c738dad3 fix(claude): align 2.1.280 fingerprint and thinking visibility (#6096)
* fix(claude): align Opus 5.5 cloak and preserve native CLI hints

* fix(claude): recognize the 2.1.280 Haiku title helper

The native title request now appends server-side fallback, fallback credit,
and cache diagnosis after the structured-output beta. Keep the helper
allowlist exact so this sequence passes through without cloaking.

* fix(claude): align thinking visibility and Messages passthrough

* fix(claude): preserve thinking visibility through compatibility paths

* fix(claude): harden OAuth beta, session alignment, and 2.1.280 helper headers
2026-09-24 15:26:39 +08:00
Luis Pater
2fe9932bb9 feat(usage): track execution request ID and trace ID across usage records
- Add `RequestID` (unique execution UUID) and `TraceID` (inbound request ID) fields to usage records and plugin API
- Generate unique UUIDs per model attempt in `UsageReporter` while linking to parent trace context
- Expose `execution_id` and `trace_id` in Redis queue usage payloads while preserving legacy `request_id`
2026-09-23 19:31:39 +08:00
Luis Pater
779bf317e0 feat(claude): bump claude cli baseline to 2.1.280 and support tool changes beta
- Update default Claude CLI version to 2.1.280 in cloaking and device profile
- Add `mid-conversation-tool-changes-2026-07-01` beta header support for supported models

Closes: #6054
2026-09-23 05:07:49 +08:00
Luis Pater
320100ecf7 fix(codex): strip octal NUL pattern escapes from tool schemas
- Strip octal NUL (`\0`) escapes in regex patterns to prevent rejection by strict upstream JSON Schema validators
- Update tool schema normalization fast-path to check for serialized `\0` escape sequences

Closes: #6025
2026-09-22 18:25:48 +08:00
Luis Pater
639b7f1126 perf(executor): reuse identical Claude and Gemini request translations
Skip the second compatibility translation for Claude, Gemini, and Gemini
interactions when the baseline and working payloads share a backing array
and no plugin hooks are installed. Keep independent buffers, the existing
hook order, and native interactions copies.
2026-09-22 05:16:25 +08:00
Luis Pater
5083f65037 Merge pull request #5990 from jack-5566/fix/responses-request-cache-pr
perf: cache Responses request validation and index tool declarations
2026-09-22 04:51:07 +08:00
Luis Pater
7f63f8c8d2 Merge pull request #5680 from cookerpapa/perf/codex-tool-schema-batch
perf(codex): batch tool schema rewrites
2026-09-22 04:23:03 +08:00
Luis Pater
d582067c06 feat(executor): support execution-scoped request proxy overrides
- Add context helpers to manage execution-scoped proxy overrides
- Prioritize request proxy over credential and global proxy settings across HTTP and uTLS clients
- Propagate request proxy in conductor execution and plugin host adapters
- Isolate Codex WebSocket connection reuse by proxy endpoint

Closes: #6013
2026-09-22 03:54:34 +08:00
Luis Pater
a5ab69521f feat: add full support for kimi.ai oauth and runtime execution 2026-09-21 04:13:44 +08:00
jack
8a6a39684d perf: index Responses tools and cache stream request validation 2026-09-20 23:10:57 +09:00
sususu
c52ca7bd4e docs(claude): clarify patch version floor rationale for native passthrough 2026-09-20 19:26:48 +08:00
camy-x
83a4913aa4 fix(claude): accept newer patch releases as native clients (#5820) 2026-09-20 19:15:51 +08:00
Luis Pater
690f4f3116 feat(executor): support use-max-completion-tokens for openai compatibility models
- Add `use-max-completion-tokens` setting to OpenAI compatibility model configuration
- Provide token normalization between `max_tokens` and `max_completion_tokens` based on model preference

Closes: #5939
2026-09-19 01:54:05 +08:00
Luis Pater
e84e248c51 fix(executor): record expected Devin upstream model and bound stream observer memory
- Set expected upstream model for Devin in both streaming and non-streaming paths to avoid false positive substitution warnings on intentional mappings
- Account for line overhead and enforce max lines per stream event in StreamResponseModelObserver, dropping overflowed events until the event boundary
- Add unit tests for Devin intentional mappings and stream observer bounded memory behavior
2026-09-18 22:08:35 +08:00
Luis Pater
cde7d57e44 fix(executor): robust response model observability across meta, kimi, and openai-compat streams
- observe response model from terminal sourceEvent in Meta non-stream multi-event SSE responses
- record expected upstream model in UsageReporter to prevent false substitution warnings on Kimi canonical mappings
- use bounded stream observer to extract response models across image stream chunk boundaries
- add regression tests for Meta SSE non-stream, Kimi model mappings, and chunked image streams
2026-09-18 21:53:19 +08:00
Luis Pater
f8467f07dc fix(executor): record authentic Devin and Gemini Interactions response models
- Devin: extract authentic upstream model name from parsed Usage.ModelName
  instead of synthesized interactions JSON or binary Connect frames. Keep
  response model empty when upstream does not report one.
- Gemini: support interaction.model and event_type terminal semantics
  in response model extractors for Gemini Interactions streaming.
- Add unit tests for Devin and Gemini Interactions response model observability.
2026-09-18 21:32:59 +08:00
Luis Pater
e9463ff5a7 feat(executor): extend response model recording and substitution warnings to all providers 2026-09-18 21:22:43 +08:00
Luis Pater
cd5af08e31 Merge pull request #5928 from Viggo95/feat/codex-response-model-observability 2026-09-18 21:05:35 +08:00
Luis Pater
cc545cbf90 fix(openai): align tool call messages and preserve ordering on ambiguous outputs
- Add `AlignOpenAIToolCallMessages` to reorder tool results immediately after the matching assistant tool calls while preserving original content and numeric precision.
- Prevent deferred message reordering and call ID guessing when tool outputs are incomplete, duplicate, or missing IDs.
- Normalize translated requests after applying summary configuration in Codex multi-agent execution.

Closes: #5925
2026-09-18 12:32:21 +08:00
Viggo95
25f40d8cf8 feat(codex): record upstream response model and warn on silent model substitution
Codex upstreams can silently serve a different model than the one requested
(HTTP 200, with response.model naming the substitute). The proxy kept no record
of it: nothing logged, nothing reported, only the pass-through response body.

- Add Record.ResponseModel to sdk/cliproxy/usage, aligned with the existing
  ResponseServiceTier field, and emit it from the redis usage queue as the
  optional response_model payload field alongside response_service_tier. Only
  the record for the requested model carries it: additional-model records
  (image generation tool usage) describe a side model the upstream response
  never refers to, and would otherwise look like a substitution downstream.
- Add internal/runtime/executor/helps/response_model.go with
  extractCodexResponseModelEvent (SSE frames and raw JSON, restricted to the events
  that embed the authoritative response object, rejecting non-string and
  oversized upstream model names) and IsCodexModelSubstituted (both sides
  trimmed, lower-cased and stripped of thinking suffixes, dated aliases such as
  gpt-5.6-terra-2026-05-13 accepted in either direction).
- UsageReporter records the served model on the event path and emits the WARN
  when the attempt publishes its usage record, so no logging work happens
  before the first event is forwarded. Repeats are throttled per
  (auth id, requested model, served model) with a 10 minute window, because on
  an affected credential every request is substituted and an unthrottled
  warning would mirror the whole request volume into the logs. The credential
  is labelled auth_index=<index> only: codex credential file names embed the
  account e-mail, which must not be written to the logs at request rate.

Coverage, by entry point. The served model is observed on the HTTP streaming
path (both the bootstrap-buffered handshake and the streaming goroutine), the
HTTP non-streaming Execute loop, the websocket streaming and non-streaming
paths, and the two /responses-shaped image entry points. The remaining codex
entry points cannot report it and are therefore left alone: executeCompact
(/responses/compact answers with a compaction object that has no event type and
no response.model), the two direct image endpoints (/images/generations and
/images/edits answer in the Images API shape and stream image_generation.*
events), and CountTokens (counts locally with tiktoken, never reaching an
upstream).

TokenAccountingSchemaVersion is not bumped: it versions the token accounting
contract (token breakdown semantics), and this change only adds an optional
non-token field that leaves existing consumers and all token math untouched.

Note: response_model ships with the usage record and is the counting source;
the WARN is a throttled alerting signal and must not be used to count
substitutions.

Tests: table-driven unit tests for both helpers over real model ids, reporter
tests covering the published record, the single throttled warning, the absence
of account identifiers in it, concurrent observation and publishing under
-race, the throttle window and its entry bound, an executor-level guard for the
observeCodexTokenEvent wiring and the per-model records, plus a redisqueue
payload assertion for response_model. gofmt, go vet, go test -race on the
touched packages and go test ./... are clean.
2026-09-18 12:31:52 +08:00
Luis Pater
1cce932573 fix(claude): skip retry-after header on overage-only rejections
- Skip parsing the `Retry-After` header when a rate limit rejection is overage-only to prevent global credential cooldowns.
- Allow exponential backoff to handle model recovery while keeping shared subscription windows available for other models.

Closes: #5920
2026-09-18 09:15:48 +08:00
Luis Pater
44eaef0009 feat(claude): support model-level cooling and scope overage rate limits
- Add `claude.model-level-cooling` configuration to scope rate limit cooldowns to the requested model.
- Treat overage-only and spend cap rejections as model-scoped when shared subscription windows remain healthy.
- Propagate model-level cooling settings into streaming, token counting, and direct execution error classifiers.

Closes: #5915
2026-09-18 01:55:36 +08:00
Luis Pater
b6fe4f20c4 fix(devin): handle orphaned tool results and normalize function result payloads
- Match tool results against pending tool calls and downgrade unmatched results to user messages.
- Prevent downgraded orphaned tool results from consuming images intended for user turns.
- Unwrap protocol wrapper envelopes and extract structured text parts while preserving arbitrary business JSON.
- Provide a placeholder for empty or whitespace-only tool results.

Closes: #5911
2026-09-18 01:34:38 +08:00
Luis Pater
9e10db53ad fix(devin): aggregate tool calls by id and track cache write tokens
- Track and aggregate tool calls by call ID instead of slot index in streaming and buffered execution.
- Support raw arguments from invalid JSON fields for custom tool calls.
- Parse usage field 4 as cache write tokens instead of adding to prompt tokens.
- Align client metadata with the default client name and drop deprecated tag 28.

Closes: #5910
2026-09-18 01:04:43 +08:00
Luis Pater
8c664b2fed Merge pull request #5896 from router-for-me/translator
feat(translator): preserve model metadata in requests
2026-09-17 13:52:35 +08:00
Luis Pater
ad088a8795 fix(devin): filter automation update tools and sanitize tool descriptions
- Filter out `automation_update` tools and sanitize tool descriptions in Devin wire requests and logs.
- Strip additional Codex prompt directives from system messages.
- Support `children` field fallback when collecting namespace tools.
2026-09-17 13:51:26 +08:00
hkfires
a9e92b8145 feat(translator): preserve model metadata in requests 2026-09-17 13:10:47 +08:00
Luis Pater
c4982e846e fix(executor): strip relayed tool result images for text-only models
- Replace tool image placeholders with omission markers for compatibility with text-only upstream models.
- Strip synthetic image relay notices and image parts from user messages.

Closes: #5884
2026-09-17 09:57:13 +08:00
Luis Pater
4613cfd44e Merge pull request #5870 from avabbbb/fix/devin-high-demand-rate-limit
fix(devin): classify high-demand errors as rate limits
2026-09-17 03:59:25 +08:00
bekkilove
d44901f91c fix(devin): classify high-demand errors as rate limits
Devin upstream occasionally encodes transient capacity failures using
Connect code permission_denied with message containing 'high demand'.
CPA's ParseDevinTrailerError currently maps every permission_denied to
HTTP 403. The auth cooldown manager interprets 403 as a 30-minute model
permission cooldown, keeping a recovered model locally unavailable.

This fix narrowly reclassifies the observed high-demand variant as
HTTP 429 (Too Many Requests), so it enters the quota/retry cooldown
path instead of the long permission denial path. Genuine
permission_denied errors (model access denied, plan entitlement denied,
etc.) remain HTTP 403.

Tests: 4 new cases in TestParseDevinTrailerError covering transient
high-demand (429), genuine permission error (403), resource_exhausted
unchanged (429), and case-insensitive matching. All existing tests
pass with no regressions.
2026-09-16 19:24:08 +08:00
bekkilove
dea4ce8aeb fix(devin): support swe-1-6-slow model variant 2026-09-16 16:49:49 +08:00
sususu
4c331bb953 fix(devin): unwrap repeated field 28 groups, merge partial field 7 usage, and harden APICall escaping 2026-09-14 09:32:05 +08:00
sususu
b4749cb204 fix(devin): parse field 8 header submessages, accumulate field 4 prompt tokens, and add field 28 usage fallback 2026-09-14 09:19:51 +08:00
sususu
30b2ac8996 refactor(devin): deduplicate auth credentials extraction, filter sparse tool calls, and optimize model lookup 2026-09-13 23:34:43 +08:00
sususu
fb2c1c1afa fix(devin): transport reuse, interleaved stream steps, strings.Builder panic, and updater URL 2026-09-13 23:34:43 +08:00
sususu
926450e87b fix(devin): dynamic catalog-driven chat_model_uid resolution and effort clamping 2026-09-13 23:34:43 +08:00
sususu
1604cb0334 fix(devin): normalize upstream internal errors to 502 Bad Gateway 2026-09-13 23:34:43 +08:00
sususu
98b106f0e8 fix(devin): store transient quota metrics strictly in Quota.Signals and keep Metadata static 2026-09-13 23:34:43 +08:00
sususu
6a239f5715 fix(devin): propagate stream chunk errors, fix truncated protobuf infinite loop, respect ctx roundtripper, and limit auth response body reads 2026-09-13 23:34:43 +08:00
sususu
86de823daa perf(devin): cache Devin HTTP transports per proxy URL to reuse connection pools 2026-09-13 23:34:43 +08:00
sususu
6c7d2d57f7 perf(devin): cache sensitive word regex matcher, preallocate request bytes and frame decompression buffer 2026-09-13 23:34:43 +08:00
sususu
5d0c77cf3f fix(devin): enforce Connect-RPC EOS trailer invariant, validate frame flag, and bind OAuth callback to ctx 2026-09-13 23:34:43 +08:00
sususu
a5ea971f35 fix(devin): sort streaming tool call stop events and normalize prompt CRLF 2026-09-13 23:34:43 +08:00
sususu
f1f5506c0b fix(devin): align wire protocol, harden streaming, and resolve multi-turn tool/signature parity
- Wire parity: align Connect-RPC Sentry-Trace, User-Agent suppression, float32 double pattern, and dynamic 732-char hex device fingerprint
- Session ordinal & cache: implement process-scoped Field 15.2 with bounded LRU (5000 entries) and Field 15.4=14 user boundary; prioritize stable session_id over previous_interaction_id to preserve prompt caching
- Streaming robustness: unblock hung TCP reads on client cancellation via context watcher; accurately propagate stream read errors and trailer errors instead of swallowing truncated frames
- Thought signature & reasoning: emit raw delta signatures directly in active thought steps; eliminate redundant tail base64 re-encoding; ensure 1:1 assistant signature and thinking alignment across multi-turn history
- Tool call de-multiplexing: route parallel tool calls by tc.Index in both streaming step events and non-streaming aggregations
- Security & transport: escape OAuth callback error HTML against reflected XSS, enforce strict state validation, and isolate Devin HTTP transport with tr.Clone()
2026-09-13 23:34:43 +08:00
sususu
bf06746d42 feat(devin): map model aliases and swe-1-7/haiku/sonnet/gpt-4-1 UIDs 2026-09-13 23:34:43 +08:00
sususu
2683ec201d feat(cmd): add fetch_devin_models CLI tool for dynamic model catalog extraction 2026-09-13 23:34:43 +08:00
sususu
469aa3678f feat(devin): parse protobuf timestamp and harden partial failure logging 2026-09-13 23:34:43 +08:00
sususu
16cb6c0b02 feat(devin): add symmetric decoded upstream response in request log 2026-09-13 23:34:43 +08:00
sususu
0c2351bb89 feat(devin): support none thinking level for glm-5-2 2026-09-13 23:34:43 +08:00
sususu
308e5ad3b1 feat(devin): add deepseek-v4-flash and deepseek-v4-1-flash models 2026-09-13 23:34:43 +08:00