- Add incremental merging for Gemini `groundingMetadata` with chunk index remapping and query deduplication.
- Implement rune offset mapping across multipart messages for accurate `url_citation` annotations.
- Manage full streaming lifecycle for web search calls, emitting `searching`, `completed`, and output item done events.
- Stream incremental citation annotations via `response.output_text.annotation.added` events.
- Support `web_search_preview_2025_03_11` as a recognized web search tool type.
- Break stream scan loops immediately when upstream completion is reached.
- Skip scanner error handling and cancellation checks when `upstreamCompleted` is true.
Closes: #5866
Devin upstream occasionally encodes transient capacity failures using
Connect code permission_denied with message containing 'high demand'.
CPA's ParseDevinTrailerError currently maps every permission_denied to
HTTP 403. The auth cooldown manager interprets 403 as a 30-minute model
permission cooldown, keeping a recovered model locally unavailable.
This fix narrowly reclassifies the observed high-demand variant as
HTTP 429 (Too Many Requests), so it enters the quota/retry cooldown
path instead of the long permission denial path. Genuine
permission_denied errors (model access denied, plan entitlement denied,
etc.) remain HTTP 403.
Tests: 4 new cases in TestParseDevinTrailerError covering transient
high-demand (429), genuine permission error (403), resource_exhausted
unchanged (429), and case-insensitive matching. All existing tests
pass with no regressions.
- Add bidirectional web search translation between OpenAI Responses API and Gemini/Antigravity
- Map Responses web_search tool to Antigravity web_search requestType envelope and googleSearch
- Map Google groundingMetadata to Responses web_search_call output item and url_citation annotations
- Buffer streaming text deltas while awaiting groundingMetadata so web_search_call strictly precedes message in SSE events and response.completed.output
- Derive search stream mode from effective translated request (requestRawJSON) to accurately support model aliases and rewrites
- Calculate streaming and non-streaming URL citation Unicode character (rune) offsets on full accumulated text, eliminating multi-byte CJK truncation and clamping
- Prioritize models.json native_capabilities.web_search explicit false as absolute veto before checking dynamic Antigravity probe capability
- Isolate Antigravity web search gating to Antigravity-specific model capabilities
- Suppress native googleSearch in Antigravity chat fallback when tools are mixed with function declarations
- Support Responses allowed_tools tool_choice containing web search and concatenate multi-part text queries
- Resolve Vertex Search grounding redirect URLs to target destination URLs in Antigravity executor
- Add comprehensive unit test coverage for stream/non-stream translation, late grounding, CJK offsets, model aliases, mixed tools, allowed_tools, and URL resolution
- Refine finish reason assignment logic in OpenAI interactions to account for incomplete states (`length`, `content_filter`) and tool call indices.
- Implement changes to normalize tool call indexing to ensure contiguous 0-based indexing.
- Add `status` and `incomplete_details` fields to OpenAI response payloads for enriched status handling.
- Improve logic for handling `generation_config` fields during interactions request conversion.
- Add robust defaulting for missing usage metrics in response payloads.
- Ensure comprehensive test coverage for all enhanced scenarios.
Closes: #5851
- Remove unexpected id parameter from function_result in Claude and OpenAI Chat Interactions request translators to comply with Google Interactions API schema and fix HTTP 400 (Unknown parameter 'id' at 'input[N]').
- Remove redundant call_id parameter from function_call in Claude Interactions request translator to match schema requirements.
- Propagate is_error flag between Claude tool_result and Interactions function_result.
- Align Devin executor to prioritize call_id for function_result steps.
- Add regression tests covering parameter schemas and end-to-end executor request generation.
Closes: #5828
- Introduced `MetaExecutor` to support native Meta (Muse Code) API operations, including `Execute`, `ExecuteStream`, and `CountTokens`.
- Added logic to handle Meta-specific `responses` endpoint, including request preparation, enriched authentication, header management, and response translation.
- Implemented token counting functionality with custom `CountTokens` logic.
- Updated `meta_executor` to process streamed and non-streamed responses, maintaining compatibility with the Codex schema.
- Added utility functions to handle events, errors, and enriched metadata common to Meta API requests.
- Extended `models.json` to include Meta's updated configuration with additional thinking levels like "minimal" and "max."
- Introduced comprehensive tests to validate Meta-specific execution, streaming, header assignments, and response formatting.
- Introduce `promoteOpenAIResponsesReasoningTextToSummary` to move `reasoning_text` parts from content to summary when the summary is empty.
- Clear `reasoning.content` to comply with Codex schema constraints (`maxItems: 0`).
- Implement safeguards for handling cleartext reasoning, preserving valid encrypted content, and stripping invalid `encrypted_content`.
- Add comprehensive tests to validate behavior for promoting reasoning texts, preserving existing summaries, and ensuring sanitization rules.
Closes: #5825
- Introduce `sanitizeGeminiInteractionsUnsupportedInputIDs` to remove `id` fields from Gemini interactions `input` and `content` items.
- Update execution logic to apply the sanitization process before upstreaming payloads.
- Add tests to ensure payloads are correctly sanitized and validate `call_id` pairing for Gemini interactions.
Closes: #5828
- Add `stream-bootstrap-timeout` configuration (defaulting to 0/unlimited, recommended 20s behind reverse proxies) to bound how long early handshake or trickled events may hold response headers.
- Release stream buffering into normal in-stream delivery once the time budget is exhausted on both SSE and WebSocket executors.
- Deliver post-timeout overload and status-bearing errors in-stream rather than triggering credential failover, preventing latency doubling on long reasoning turns.
- Provide thread-safe mock clock test harness and comprehensive unit tests covering timeout release, unlimited default, disabled ceilings, and post-timeout error delivery.
- Introduce `SetStringWithoutHTMLEscape` to prevent escaping of HTML characters (`<`, `>`, `&`) in tool call arguments and argument deltas across translators.
- Handle sequential tool calls in Devin executor that share the same stream index but have distinct tool call IDs.
- Wire parity: align Connect-RPC Sentry-Trace, User-Agent suppression, float32 double pattern, and dynamic 732-char hex device fingerprint
- Session ordinal & cache: implement process-scoped Field 15.2 with bounded LRU (5000 entries) and Field 15.4=14 user boundary; prioritize stable session_id over previous_interaction_id to preserve prompt caching
- Streaming robustness: unblock hung TCP reads on client cancellation via context watcher; accurately propagate stream read errors and trailer errors instead of swallowing truncated frames
- Thought signature & reasoning: emit raw delta signatures directly in active thought steps; eliminate redundant tail base64 re-encoding; ensure 1:1 assistant signature and thinking alignment across multi-turn history
- Tool call de-multiplexing: route parallel tool calls by tc.Index in both streaming step events and non-streaming aggregations
- Security & transport: escape OAuth callback error HTML against reflected XSS, enforce strict state validation, and isolate Devin HTTP transport with tr.Clone()
- Register static model definitions for devin/gemini-3-8-flash (1M context, Google) and devin/grok-4-6 (500k context, xAI).
- Support gemini38Efforts (low/medium/high) and grok46Efforts (low/medium/high/xhigh) in ResolveDevinChatModelUID, supporting both colon and parenthesis suffix parsing.
- Recognize Gemini Tink thought signatures (AY-prefix / 0x01 Tink header) in detectSignatureType and parseSignatureBytes.
- Add unit tests for both models in registry and devin_models.