* fix(claude): align Opus 5.5 cloak and preserve native CLI hints
* fix(claude): recognize the 2.1.280 Haiku title helper
The native title request now appends server-side fallback, fallback credit,
and cache diagnosis after the structured-output beta. Keep the helper
allowlist exact so this sequence passes through without cloaking.
* fix(claude): align thinking visibility and Messages passthrough
* fix(claude): preserve thinking visibility through compatibility paths
* fix(claude): harden OAuth beta, session alignment, and 2.1.280 helper headers
- Add `RequestID` (unique execution UUID) and `TraceID` (inbound request ID) fields to usage records and plugin API
- Generate unique UUIDs per model attempt in `UsageReporter` while linking to parent trace context
- Expose `execution_id` and `trace_id` in Redis queue usage payloads while preserving legacy `request_id`
- Update default Claude CLI version to 2.1.280 in cloaking and device profile
- Add `mid-conversation-tool-changes-2026-07-01` beta header support for supported models
Closes: #6054
Skip the second compatibility translation for Claude, Gemini, and Gemini
interactions when the baseline and working payloads share a backing array
and no plugin hooks are installed. Keep independent buffers, the existing
hook order, and native interactions copies.
- Add `use-max-completion-tokens` setting to OpenAI compatibility model configuration
- Provide token normalization between `max_tokens` and `max_completion_tokens` based on model preference
Closes: #5939
- Set expected upstream model for Devin in both streaming and non-streaming paths to avoid false positive substitution warnings on intentional mappings
- Account for line overhead and enforce max lines per stream event in StreamResponseModelObserver, dropping overflowed events until the event boundary
- Add unit tests for Devin intentional mappings and stream observer bounded memory behavior
- observe response model from terminal sourceEvent in Meta non-stream multi-event SSE responses
- record expected upstream model in UsageReporter to prevent false substitution warnings on Kimi canonical mappings
- use bounded stream observer to extract response models across image stream chunk boundaries
- add regression tests for Meta SSE non-stream, Kimi model mappings, and chunked image streams
- Devin: extract authentic upstream model name from parsed Usage.ModelName
instead of synthesized interactions JSON or binary Connect frames. Keep
response model empty when upstream does not report one.
- Gemini: support interaction.model and event_type terminal semantics
in response model extractors for Gemini Interactions streaming.
- Add unit tests for Devin and Gemini Interactions response model observability.
- Add `AlignOpenAIToolCallMessages` to reorder tool results immediately after the matching assistant tool calls while preserving original content and numeric precision.
- Prevent deferred message reordering and call ID guessing when tool outputs are incomplete, duplicate, or missing IDs.
- Normalize translated requests after applying summary configuration in Codex multi-agent execution.
Closes: #5925
Codex upstreams can silently serve a different model than the one requested
(HTTP 200, with response.model naming the substitute). The proxy kept no record
of it: nothing logged, nothing reported, only the pass-through response body.
- Add Record.ResponseModel to sdk/cliproxy/usage, aligned with the existing
ResponseServiceTier field, and emit it from the redis usage queue as the
optional response_model payload field alongside response_service_tier. Only
the record for the requested model carries it: additional-model records
(image generation tool usage) describe a side model the upstream response
never refers to, and would otherwise look like a substitution downstream.
- Add internal/runtime/executor/helps/response_model.go with
extractCodexResponseModelEvent (SSE frames and raw JSON, restricted to the events
that embed the authoritative response object, rejecting non-string and
oversized upstream model names) and IsCodexModelSubstituted (both sides
trimmed, lower-cased and stripped of thinking suffixes, dated aliases such as
gpt-5.6-terra-2026-05-13 accepted in either direction).
- UsageReporter records the served model on the event path and emits the WARN
when the attempt publishes its usage record, so no logging work happens
before the first event is forwarded. Repeats are throttled per
(auth id, requested model, served model) with a 10 minute window, because on
an affected credential every request is substituted and an unthrottled
warning would mirror the whole request volume into the logs. The credential
is labelled auth_index=<index> only: codex credential file names embed the
account e-mail, which must not be written to the logs at request rate.
Coverage, by entry point. The served model is observed on the HTTP streaming
path (both the bootstrap-buffered handshake and the streaming goroutine), the
HTTP non-streaming Execute loop, the websocket streaming and non-streaming
paths, and the two /responses-shaped image entry points. The remaining codex
entry points cannot report it and are therefore left alone: executeCompact
(/responses/compact answers with a compaction object that has no event type and
no response.model), the two direct image endpoints (/images/generations and
/images/edits answer in the Images API shape and stream image_generation.*
events), and CountTokens (counts locally with tiktoken, never reaching an
upstream).
TokenAccountingSchemaVersion is not bumped: it versions the token accounting
contract (token breakdown semantics), and this change only adds an optional
non-token field that leaves existing consumers and all token math untouched.
Note: response_model ships with the usage record and is the counting source;
the WARN is a throttled alerting signal and must not be used to count
substitutions.
Tests: table-driven unit tests for both helpers over real model ids, reporter
tests covering the published record, the single throttled warning, the absence
of account identifiers in it, concurrent observation and publishing under
-race, the throttle window and its entry bound, an executor-level guard for the
observeCodexTokenEvent wiring and the per-model records, plus a redisqueue
payload assertion for response_model. gofmt, go vet, go test -race on the
touched packages and go test ./... are clean.
- Skip parsing the `Retry-After` header when a rate limit rejection is overage-only to prevent global credential cooldowns.
- Allow exponential backoff to handle model recovery while keeping shared subscription windows available for other models.
Closes: #5920
- Add `claude.model-level-cooling` configuration to scope rate limit cooldowns to the requested model.
- Treat overage-only and spend cap rejections as model-scoped when shared subscription windows remain healthy.
- Propagate model-level cooling settings into streaming, token counting, and direct execution error classifiers.
Closes: #5915
- Match tool results against pending tool calls and downgrade unmatched results to user messages.
- Prevent downgraded orphaned tool results from consuming images intended for user turns.
- Unwrap protocol wrapper envelopes and extract structured text parts while preserving arbitrary business JSON.
- Provide a placeholder for empty or whitespace-only tool results.
Closes: #5911
- Track and aggregate tool calls by call ID instead of slot index in streaming and buffered execution.
- Support raw arguments from invalid JSON fields for custom tool calls.
- Parse usage field 4 as cache write tokens instead of adding to prompt tokens.
- Align client metadata with the default client name and drop deprecated tag 28.
Closes: #5910
- Filter out `automation_update` tools and sanitize tool descriptions in Devin wire requests and logs.
- Strip additional Codex prompt directives from system messages.
- Support `children` field fallback when collecting namespace tools.
- Replace tool image placeholders with omission markers for compatibility with text-only upstream models.
- Strip synthetic image relay notices and image parts from user messages.
Closes: #5884
Devin upstream occasionally encodes transient capacity failures using
Connect code permission_denied with message containing 'high demand'.
CPA's ParseDevinTrailerError currently maps every permission_denied to
HTTP 403. The auth cooldown manager interprets 403 as a 30-minute model
permission cooldown, keeping a recovered model locally unavailable.
This fix narrowly reclassifies the observed high-demand variant as
HTTP 429 (Too Many Requests), so it enters the quota/retry cooldown
path instead of the long permission denial path. Genuine
permission_denied errors (model access denied, plan entitlement denied,
etc.) remain HTTP 403.
Tests: 4 new cases in TestParseDevinTrailerError covering transient
high-demand (429), genuine permission error (403), resource_exhausted
unchanged (429), and case-insensitive matching. All existing tests
pass with no regressions.
- Wire parity: align Connect-RPC Sentry-Trace, User-Agent suppression, float32 double pattern, and dynamic 732-char hex device fingerprint
- Session ordinal & cache: implement process-scoped Field 15.2 with bounded LRU (5000 entries) and Field 15.4=14 user boundary; prioritize stable session_id over previous_interaction_id to preserve prompt caching
- Streaming robustness: unblock hung TCP reads on client cancellation via context watcher; accurately propagate stream read errors and trailer errors instead of swallowing truncated frames
- Thought signature & reasoning: emit raw delta signatures directly in active thought steps; eliminate redundant tail base64 re-encoding; ensure 1:1 assistant signature and thinking alignment across multi-turn history
- Tool call de-multiplexing: route parallel tool calls by tc.Index in both streaming step events and non-streaming aggregations
- Security & transport: escape OAuth callback error HTML against reflected XSS, enforce strict state validation, and isolate Devin HTTP transport with tr.Clone()