- Wire parity: align Connect-RPC Sentry-Trace, User-Agent suppression, float32 double pattern, and dynamic 732-char hex device fingerprint
- Session ordinal & cache: implement process-scoped Field 15.2 with bounded LRU (5000 entries) and Field 15.4=14 user boundary; prioritize stable session_id over previous_interaction_id to preserve prompt caching
- Streaming robustness: unblock hung TCP reads on client cancellation via context watcher; accurately propagate stream read errors and trailer errors instead of swallowing truncated frames
- Thought signature & reasoning: emit raw delta signatures directly in active thought steps; eliminate redundant tail base64 re-encoding; ensure 1:1 assistant signature and thinking alignment across multi-turn history
- Tool call de-multiplexing: route parallel tool calls by tc.Index in both streaming step events and non-streaming aggregations
- Security & transport: escape OAuth callback error HTML against reflected XSS, enforce strict state validation, and isolate Devin HTTP transport with tr.Clone()
- Register static model definitions for devin/gemini-3-8-flash (1M context, Google) and devin/grok-4-6 (500k context, xAI).
- Support gemini38Efforts (low/medium/high) and grok46Efforts (low/medium/high/xhigh) in ResolveDevinChatModelUID, supporting both colon and parenthesis suffix parsing.
- Recognize Gemini Tink thought signatures (AY-prefix / 0x01 Tink header) in detectSignatureType and parseSignatureBytes.
- Add unit tests for both models in registry and devin_models.
- Implement Connect-RPC GetUserStatus serialization and response parsing in internal/auth/devin/user_status.go.
- Extract user email, plan, username, user_id, team_id, org_id, daily/weekly quota percentages, and reset timestamps.
- Wire user status into DevinExecutor.Refresh to update auth metadata and Quota.Signals.
- Add devin to ProviderSupportsQuotaObservation so CPA management endpoints surface quota observations.
- Enrich Devin OAuth login flow with user status, email, and quota information, and add CSRF state verification.
- Support base_url override in DevinAuthService for mock testing and custom gateways.
- Add example payload filter rules in `config.example.yaml` for stripping tools from both flat and nested `additional_tools` Codex request payloads.
Closes: #5792
- Define a managed Claude beta set to distinguish proxy-governed betas from caller extensions.
- Forward unmanaged caller betas on direct Anthropic endpoints to support newer client features.
Closes: #5738
- Avoid acquiring write mutex in ping handlers so keepalive pongs reply immediately during active writes.
- Stream payload messages larger than 32KB in chunks using writer streams.
- Support ephemeral websocket sessions for sessionless execution and enrich disconnect logs with session kind and terminal event context.
Closes: #5734
- Add kimi-k2.8 and kimi-k2.8-code with 1M context, 64k completion tokens, low/high/max thinking, and zero_allowed support to models.json.
- Remap K2.8 aliases (kimi-k2.8, k2.8, kimi-k2.8-code, k2.8-code, and -preview variants) to upstream canonical kimi-for-coding.
- Normalize temperature for Kimi upstream to prevent 400 errors (strict 0.6 for disabled thinking, 1.0 for enabled).
- Align zero_allowed: true across K2.8 and K3 models based on live upstream verification of thinking.type=disabled.
- Add test coverage for model normalization, thinking replay family, temperature stripping, and Claude effort=max preservation.
- Derive billing fingerprint message text from the first user message to prevent prompt cache invalidation across turns.
- Centralize continuity tag resolution to maintain stable prompt IDs and request references across requests.
Closes: #5730
- Inspect `input` items for the latest `configuration_update` reasoning effort before falling back to top-level configuration.
- Route `codex`, `xai`, and `openai-response` providers through usage-specific thinking config extraction.
- Support `none`, `auto`, and explicit reasoning level modes from in-turn updates.
Closes: #5714
- Add `BaseURL` field to usage records and host auth file entries.
- Extract `base_url` from auth attributes and metadata during usage reporting.
- Propagate `base_url` through plugin usage adapters and runtime auth callbacks.
Closes: #5693
- Synthesize function response parts for interrupted or missing OpenAI Responses tool calls to maintain strict Gemini call-response pairing.
- Preserve function response ordering matching pending call IDs across parallel and partial tool execution turns.
- Degrade gracefully to the original request payload when Antigravity reasoning replay breaks Gemini function call pairing.
Closes: #5682
- Map upstream `model_not_found` errors to HTTP 404 before evaluating generic invalid request types in Codex terminal error handling.
- Prevent treating structured model not found responses as client request faults to preserve credential rotation.
- Recognize model access denial errors to apply model-level cooldown and failover.
- Respect `disable_cooling` configuration during model-level cooldown processing.
Closes: #5635
- Broaden pattern matching for Codex model capacity errors.
- Classify model capacity rejections as overload bootstrap failures to enable failover.
Closes: #5634
- Format streaming error payloads with nested error objects matching official OpenAI Responses SSE specifications.
- Extract and propagate sequence numbers from upstream terminal events and framer states.
- Use `json.Number` to prevent precision loss for large integers and token metrics.
- Sanitize sensitive keys recursively across nested error objects without dropping custom fields.
- Retain `cache_control` blocks with 1h TTL and `extended-cache-ttl` beta header when explicitly requested by subagents.
- Detect 1h TTL configuration from request payloads and incoming Anthropic-Beta headers.
- Ensure `extended-cache-ttl` beta header is preserved or injected when 1h TTL is present.
Closes: #5629
- Register builtin model definitions for `gpt-image-2.5`, `gpt-image-2.5-flare`, and `gpt-image-2.5-sunburst`.
- Update OpenAI image handlers and request routing to recognize GPT Image 2.5 models.
- Support direct image generation and edit execution for GPT Image 2.5 variants in the Codex executor.
- Apply client visibility overrides to hide new builtin image models where appropriate.
- Initialize `util.SessionIDResolver` to resolve session IDs from request contexts, metadata, and headers.
- Ensure canonical session metadata is injected into execution options across execution flows.
- Propagate session context in executors to support `$CPA-SESSION-ID` expansion in custom headers.
- Sync cleared or updated session identities back to execution contexts during conductor execution.
Closes: #5690
- Check for \u in fast-path check to prevent JSON Unicode escapes from bypassing inspection.
- Inspect regex keys under patternProperties and drop keys with unsupported Unicode property escapes.
- Add tests covering Unicode escape representations (\u005c, \u0070, \u0050) and patternProperties keys.
- Add HasUnsupportedUnicodePropertyEscape in internal/util to detect \p{...} / \P{...} escapes that fail Python re compilation.
- Strip incompatible pattern attributes during tool parameter normalization in codex/claude and openai/claude translators.
- Provide schema-aware fallback stripping in codex executor helps to protect downstream Codex requests without mutating non-schema user data.
- Export unified schema keyword lists in internal/util to eliminate duplication.
- Add comprehensive unit tests covering Artifact fixtures, lookaheads, and user data preservation.
Closes: #5644
- Add `model-level-cooling` configuration option to Codex settings.
- Scope `usage_limit_reached` quota cooldowns to the requested model instead of the entire credential when enabled.
- Propagate model-level cooling checks across HTTP, SSE, and WebSocket execution paths.
Closes: #5619
- Emit an OpenAI-compatible trailing usage chunk with an empty choices array on `message_stop`.
- Include `cache_write_tokens` in prompt token details for OpenAI response translations.
- Parse usage from `message.usage` and buffer Claude stream usage across chunks to merge input and output token counts.
- Ensure observed streaming usage details are published on completion or stream failure.
Closes: #5617
Repair placeholder functionResponse names with part-local sjson and one
Index splice of the request body, matching the sibling content-edit path
instead of one full-body sjson.SetBytes per repaired name.
- Add shortMode, idleConnTimeout, and maxIdleConnsPerHost to antigravityTransportKey.
- Prevent stale transport reuse during hot-reload race windows under load.
- Add TestAntigravityTransportKeySeparatesPoolSettingsAcrossReload regression test.
- Route `openai-response` source format requests to dedicated streaming and non-streaming responses handlers.
- Add upstream URL resolution for Kimi Responses API endpoints based on auth attributes.
- Normalize `openai-response` provider format to codex during thinking parameter resolution.
Closes: #5400#5571