- Generalize forced hosted tool choice handling across both image generation and web search.
- Normalize forced web search tool choices to string mode and isolate the target tool in the tools list.
- Strip hosted web search from mixed `allowed_tools` definitions.
- Skip native `x_search` injection when a hosted tool is exclusively required.
Closes: #5916
- Add `claude.model-level-cooling` configuration to scope rate limit cooldowns to the requested model.
- Treat overage-only and spend cap rejections as model-scoped when shared subscription windows remain healthy.
- Propagate model-level cooling settings into streaming, token counting, and direct execution error classifiers.
Closes: #5915
- Match tool results against pending tool calls and downgrade unmatched results to user messages.
- Prevent downgraded orphaned tool results from consuming images intended for user turns.
- Unwrap protocol wrapper envelopes and extract structured text parts while preserving arbitrary business JSON.
- Provide a placeholder for empty or whitespace-only tool results.
Closes: #5911
- Track and aggregate tool calls by call ID instead of slot index in streaming and buffered execution.
- Support raw arguments from invalid JSON fields for custom tool calls.
- Parse usage field 4 as cache write tokens instead of adding to prompt tokens.
- Align client metadata with the default client name and drop deprecated tag 28.
Closes: #5910
- Extract and attach images from tool results to corresponding tool prompts by call ID.
- Preserve structured business JSON and raw objects in function result content.
- Prepend image headers to tool prompt content when images are present.
Closes: #5893
- Filter out `automation_update` tools and sanitize tool descriptions in Devin wire requests and logs.
- Strip additional Codex prompt directives from system messages.
- Support `children` field fallback when collecting namespace tools.
- Replace tool image placeholders with omission markers for compatibility with text-only upstream models.
- Strip synthetic image relay notices and image parts from user messages.
Closes: #5884
- Buffer content and tool call deltas while thinking is active so late-arriving thinking signatures can be attached before closing thinking blocks.
- Ensure pending actions and open steps are properly flushed and closed on stream completion or trailer errors.
Closes: #5873
- Add incremental merging for Gemini `groundingMetadata` with chunk index remapping and query deduplication.
- Implement rune offset mapping across multipart messages for accurate `url_citation` annotations.
- Manage full streaming lifecycle for web search calls, emitting `searching`, `completed`, and output item done events.
- Stream incremental citation annotations via `response.output_text.annotation.added` events.
- Support `web_search_preview_2025_03_11` as a recognized web search tool type.
- Break stream scan loops immediately when upstream completion is reached.
- Skip scanner error handling and cancellation checks when `upstreamCompleted` is true.
Closes: #5866
Devin upstream occasionally encodes transient capacity failures using
Connect code permission_denied with message containing 'high demand'.
CPA's ParseDevinTrailerError currently maps every permission_denied to
HTTP 403. The auth cooldown manager interprets 403 as a 30-minute model
permission cooldown, keeping a recovered model locally unavailable.
This fix narrowly reclassifies the observed high-demand variant as
HTTP 429 (Too Many Requests), so it enters the quota/retry cooldown
path instead of the long permission denial path. Genuine
permission_denied errors (model access denied, plan entitlement denied,
etc.) remain HTTP 403.
Tests: 4 new cases in TestParseDevinTrailerError covering transient
high-demand (429), genuine permission error (403), resource_exhausted
unchanged (429), and case-insensitive matching. All existing tests
pass with no regressions.
- Add bidirectional web search translation between OpenAI Responses API and Gemini/Antigravity
- Map Responses web_search tool to Antigravity web_search requestType envelope and googleSearch
- Map Google groundingMetadata to Responses web_search_call output item and url_citation annotations
- Buffer streaming text deltas while awaiting groundingMetadata so web_search_call strictly precedes message in SSE events and response.completed.output
- Derive search stream mode from effective translated request (requestRawJSON) to accurately support model aliases and rewrites
- Calculate streaming and non-streaming URL citation Unicode character (rune) offsets on full accumulated text, eliminating multi-byte CJK truncation and clamping
- Prioritize models.json native_capabilities.web_search explicit false as absolute veto before checking dynamic Antigravity probe capability
- Isolate Antigravity web search gating to Antigravity-specific model capabilities
- Suppress native googleSearch in Antigravity chat fallback when tools are mixed with function declarations
- Support Responses allowed_tools tool_choice containing web search and concatenate multi-part text queries
- Resolve Vertex Search grounding redirect URLs to target destination URLs in Antigravity executor
- Add comprehensive unit test coverage for stream/non-stream translation, late grounding, CJK offsets, model aliases, mixed tools, allowed_tools, and URL resolution
- Refine finish reason assignment logic in OpenAI interactions to account for incomplete states (`length`, `content_filter`) and tool call indices.
- Implement changes to normalize tool call indexing to ensure contiguous 0-based indexing.
- Add `status` and `incomplete_details` fields to OpenAI response payloads for enriched status handling.
- Improve logic for handling `generation_config` fields during interactions request conversion.
- Add robust defaulting for missing usage metrics in response payloads.
- Ensure comprehensive test coverage for all enhanced scenarios.
Closes: #5851
- Remove unexpected id parameter from function_result in Claude and OpenAI Chat Interactions request translators to comply with Google Interactions API schema and fix HTTP 400 (Unknown parameter 'id' at 'input[N]').
- Remove redundant call_id parameter from function_call in Claude Interactions request translator to match schema requirements.
- Propagate is_error flag between Claude tool_result and Interactions function_result.
- Align Devin executor to prioritize call_id for function_result steps.
- Add regression tests covering parameter schemas and end-to-end executor request generation.
Closes: #5828
- Introduced `MetaExecutor` to support native Meta (Muse Code) API operations, including `Execute`, `ExecuteStream`, and `CountTokens`.
- Added logic to handle Meta-specific `responses` endpoint, including request preparation, enriched authentication, header management, and response translation.
- Implemented token counting functionality with custom `CountTokens` logic.
- Updated `meta_executor` to process streamed and non-streamed responses, maintaining compatibility with the Codex schema.
- Added utility functions to handle events, errors, and enriched metadata common to Meta API requests.
- Extended `models.json` to include Meta's updated configuration with additional thinking levels like "minimal" and "max."
- Introduced comprehensive tests to validate Meta-specific execution, streaming, header assignments, and response formatting.
- Introduce `promoteOpenAIResponsesReasoningTextToSummary` to move `reasoning_text` parts from content to summary when the summary is empty.
- Clear `reasoning.content` to comply with Codex schema constraints (`maxItems: 0`).
- Implement safeguards for handling cleartext reasoning, preserving valid encrypted content, and stripping invalid `encrypted_content`.
- Add comprehensive tests to validate behavior for promoting reasoning texts, preserving existing summaries, and ensuring sanitization rules.
Closes: #5825
- Introduce `sanitizeGeminiInteractionsUnsupportedInputIDs` to remove `id` fields from Gemini interactions `input` and `content` items.
- Update execution logic to apply the sanitization process before upstreaming payloads.
- Add tests to ensure payloads are correctly sanitized and validate `call_id` pairing for Gemini interactions.
Closes: #5828
- Add `stream-bootstrap-timeout` configuration (defaulting to 0/unlimited, recommended 20s behind reverse proxies) to bound how long early handshake or trickled events may hold response headers.
- Release stream buffering into normal in-stream delivery once the time budget is exhausted on both SSE and WebSocket executors.
- Deliver post-timeout overload and status-bearing errors in-stream rather than triggering credential failover, preventing latency doubling on long reasoning turns.
- Provide thread-safe mock clock test harness and comprehensive unit tests covering timeout release, unlimited default, disabled ceilings, and post-timeout error delivery.
- Introduce `SetStringWithoutHTMLEscape` to prevent escaping of HTML characters (`<`, `>`, `&`) in tool call arguments and argument deltas across translators.
- Handle sequential tool calls in Devin executor that share the same stream index but have distinct tool call IDs.
- Wire parity: align Connect-RPC Sentry-Trace, User-Agent suppression, float32 double pattern, and dynamic 732-char hex device fingerprint
- Session ordinal & cache: implement process-scoped Field 15.2 with bounded LRU (5000 entries) and Field 15.4=14 user boundary; prioritize stable session_id over previous_interaction_id to preserve prompt caching
- Streaming robustness: unblock hung TCP reads on client cancellation via context watcher; accurately propagate stream read errors and trailer errors instead of swallowing truncated frames
- Thought signature & reasoning: emit raw delta signatures directly in active thought steps; eliminate redundant tail base64 re-encoding; ensure 1:1 assistant signature and thinking alignment across multi-turn history
- Tool call de-multiplexing: route parallel tool calls by tc.Index in both streaming step events and non-streaming aggregations
- Security & transport: escape OAuth callback error HTML against reflected XSS, enforce strict state validation, and isolate Devin HTTP transport with tr.Clone()