Commit Graph

981 Commits

Author SHA1 Message Date
sususu
d115fe2c45 refactor(devin): align sensitive-words with antigravity to config.yaml only 2026-09-13 23:34:43 +08:00
sususu
1b6948513d feat(devin): support sensitive-words in auth json metadata and attributes 2026-09-13 23:34:43 +08:00
sususu
f5247e496f fix(devin): restrict sensitive word obfuscation strictly to system prompt only 2026-09-13 23:34:43 +08:00
sususu
02fd1bde78 feat(devin): restrict glm-5-2 to free tier, remove static swe-1-7-lightning, and harden cloak 2026-09-13 23:34:43 +08:00
sususu
eed249072d feat(devin): prefix all Devin model IDs with devin/ namespace 2026-09-13 23:34:43 +08:00
sususu
cbe800aa28 feat(devin): bind upstream session_id and cascade_id to CPA canonical session 2026-09-13 23:34:43 +08:00
sususu
f94752762b feat(devin): add Devin/Cognition provider integration and CLI OAuth
Implement the full Devin/Cognition Connect-RPC provider support across all CPA endpoints (/v1/chat/completions, /v1/messages, /v1/responses), complete with binary protobuf wire framing, streaming tools/arguments delta handling, thinking/reasoning replay, and CLI OAuth authentication.

Key highlights:
- Wire Protocol & Streaming:
  * Implemented Connect-RPC uncompressed 5-byte framing (0x00 + 4-byte length + protobuf) for ApiServerService/GetChatMessage.
  * Implemented Devin protobuf encoder/decoder in internal/runtime/executor/helps/devin_wire.go, including ClientMetadata, prompts, tools, completion_config, and multimodal image handling (Prompt Field #10).
  * Stream frame consumption via interactions protocol, correctly mapping arguments_delta and tracking multiple sequential tool calls (currentToolCallActive).
  * Streaming thought summary and sealed.v1 signature deltas targeting the thinking step.

- Model Registration & Thinking Clamping:
  * Registered static fallback models in model_definitions.go (swe-2, claude-fable-5-1, gpt-6-astra, swe-1-7-lightning, glm-5-2, glm-5-3).
  * Configured ThinkingSupport with discrete levels per model family.
  * Implemented CPA-standard nearest-neighbor clamping for thinking levels (minimal/low -> medium, xhigh -> max for swe-2).
  * Mapped thinking effort to Devin upstream model UID (e.g. swe-2-medium, swe-2-high, swe-2-max).

- Sensitive Words & System Prompt Sanitization:
  * Added devin.sensitive-words configuration in internal/config/config_types.go and config.go, matching Antigravity conventions.
  * Supported zero-width space (\u200b) obfuscation in prompts, tools, and system instructions via SensitiveWordMatcher.
  * Stripped Claude Code billing headers (x-anthropic-billing-header:) and CLI identity signatures from system instructions and tool descriptions to avoid upstream content filter rejections.

- Signature Compatibility:
  * Added SignatureProviderSWE = "swe" recognizing sealed.v1.* reasoning signatures in internal/signature/provider_compatibility.go.
  * Propagated reasoning.encrypted_content on Responses API and thinking.signature on Messages API.

- Authentication:
  * Implemented Devin PKCE OAuth flow with loopback callback server and headless manual token/code paste (--no-browser).
  * Registered Devin authenticator in SDK and CLI (-devin-login flag).
  * Integrated with management OAuth session endpoints and credentials manager.
2026-09-13 23:34:43 +08:00
Luis Pater
94d6eb535e docs(config): document payload filter examples for codex tools
- Add example payload filter rules in `config.example.yaml` for stripping tools from both flat and nested `additional_tools` Codex request payloads.

Closes: #5792
2026-09-13 22:12:28 +08:00
Luis Pater
f702bc1ac2 feat(codex): preserve native fidelity for responses-lite requests
- Detect native responses-lite requests via headers and client metadata.
- Skip instructions normalization and synthetic session cloaking for native requests.
- Preserve upstream completion output during websocket response forwarding.

Closes: #5780
2026-09-13 15:05:03 +08:00
Luis Pater
e696ea47c5 feat(codex): forward X-Codex-Turn-State header in executor requests
- Ensure `X-Codex-Turn-State` is preserved and forwarded from incoming client requests.

Closes: #5778
2026-09-13 14:42:26 +08:00
Luis Pater
d9b8fdb77f fix(claude): forward unmanaged caller betas on direct anthropic requests
- Define a managed Claude beta set to distinguish proxy-governed betas from caller extensions.
- Forward unmanaged caller betas on direct Anthropic endpoints to support newer client features.

Closes: #5738
2026-09-11 23:44:21 +08:00
Luis Pater
b5ba02c2e3 fix(websockets): prevent keepalive pong starvation during large payload writes
- Avoid acquiring write mutex in ping handlers so keepalive pongs reply immediately during active writes.
- Stream payload messages larger than 32KB in chunks using writer streams.
- Support ephemeral websocket sessions for sessionless execution and enrich disconnect logs with session kind and terminal event context.

Closes: #5734
2026-09-11 19:53:45 +08:00
sususu
8bd67f3338 feat(kimi): add Kimi K2.8 model definitions, normalization, and temperature guard
- Add kimi-k2.8 and kimi-k2.8-code with 1M context, 64k completion tokens, low/high/max thinking, and zero_allowed support to models.json.
- Remap K2.8 aliases (kimi-k2.8, k2.8, kimi-k2.8-code, k2.8-code, and -preview variants) to upstream canonical kimi-for-coding.
- Normalize temperature for Kimi upstream to prevent 400 errors (strict 0.6 for disabled thinking, 1.0 for enabled).
- Align zero_allowed: true across K2.8 and K3 models based on live upstream verification of thinking.type=disabled.
- Add test coverage for model normalization, thinking replay family, temperature stripping, and Claude effort=max preservation.
2026-09-11 17:01:00 +08:00
Luis Pater
377c315fd7 fix(claude): anchor billing fingerprint to initial turn for cloaked cache stability
- Derive billing fingerprint message text from the first user message to prevent prompt cache invalidation across turns.
- Centralize continuity tag resolution to maintain stable prompt IDs and request references across requests.

Closes: #5730
2026-09-11 16:59:57 +08:00
Luis Pater
4edf9d1dd6 fix(thinking): extract configuration_update reasoning effort in codex usage reporting
- Inspect `input` items for the latest `configuration_update` reasoning effort before falling back to top-level configuration.
- Route `codex`, `xai`, and `openai-response` providers through usage-specific thinking config extraction.
- Support `none`, `auto`, and explicit reasoning level modes from in-turn updates.

Closes: #5714
2026-09-11 12:15:57 +08:00
sususu
8461b4e91d chore(codex): update codex user-agent to 0.154.0 2026-09-11 00:30:24 +08:00
Luis Pater
c8f723e0fb feat(usage): propagate upstream base_url across usage records and plugin auth
- Add `BaseURL` field to usage records and host auth file entries.
- Extract `base_url` from auth attributes and metadata during usage reporting.
- Propagate `base_url` through plugin usage adapters and runtime auth callbacks.

Closes: #5693
2026-09-11 00:25:17 +08:00
Luis Pater
b8e6ec0ae7 fix(gemini): preserve function call pairing for interrupted calls and replay failures
- Synthesize function response parts for interrupted or missing OpenAI Responses tool calls to maintain strict Gemini call-response pairing.
- Preserve function response ordering matching pending call IDs across parallel and partial tool execution turns.
- Degrade gracefully to the original request payload when Antigravity reasoning replay breaks Gemini function call pairing.

Closes: #5682
2026-09-10 20:01:39 +08:00
Luis Pater
dde250f1c3 fix(auth): enable model cooldown and rotation for model not found errors
- Map upstream `model_not_found` errors to HTTP 404 before evaluating generic invalid request types in Codex terminal error handling.
- Prevent treating structured model not found responses as client request faults to preserve credential rotation.
- Recognize model access denial errors to apply model-level cooldown and failover.
- Respect `disable_cooling` configuration during model-level cooldown processing.

Closes: #5635
2026-09-10 12:35:07 +08:00
Luis Pater
3ae9093da8 fix(codex): treat model capacity errors as bootstrap overload failures
- Broaden pattern matching for Codex model capacity errors.
- Classify model capacity rejections as overload bootstrap failures to enable failover.

Closes: #5634
2026-09-10 12:06:01 +08:00
Luis Pater
259130863d fix(openai): preserve nested error details and sequence numbers in responses stream
- Format streaming error payloads with nested error objects matching official OpenAI Responses SSE specifications.
- Extract and propagate sequence numbers from upstream terminal events and framer states.
- Use `json.Number` to prevent precision loss for large integers and token metrics.
- Sanitize sensitive keys recursively across nested error objects without dropping custom fields.
2026-09-10 11:58:36 +08:00
Luis Pater
6a73f39627 fix(claude): preserve 1h cache ttl and beta header for subagent requests
- Retain `cache_control` blocks with 1h TTL and `extended-cache-ttl` beta header when explicitly requested by subagents.
- Detect 1h TTL configuration from request payloads and incoming Anthropic-Beta headers.
- Ensure `extended-cache-ttl` beta header is preserved or injected when 1h TTL is present.

Closes: #5629
2026-09-10 10:58:34 +08:00
Luis Pater
d1a024e940 feat(codex): add support for gpt-image-2.5 models
- Register builtin model definitions for `gpt-image-2.5`, `gpt-image-2.5-flare`, and `gpt-image-2.5-sunburst`.
- Update OpenAI image handlers and request routing to recognize GPT Image 2.5 models.
- Support direct image generation and edit execution for GPT Image 2.5 variants in the Codex executor.
- Apply client visibility overrides to hide new builtin image models where appropriate.
2026-09-10 10:40:53 +08:00
Luis Pater
3bf787fc1d feat(auth): propagate canonical session id for custom header expansion
- Initialize `util.SessionIDResolver` to resolve session IDs from request contexts, metadata, and headers.
- Ensure canonical session metadata is injected into execution options across execution flows.
- Propagate session context in executors to support `$CPA-SESSION-ID` expansion in custom headers.
- Sync cleared or updated session identities back to execution contexts during conductor execution.

Closes: #5690
2026-09-10 01:43:37 +08:00
Luis Pater
db011593ca Merge pull request #5677 from sususu98/fix/codex-tool-schema-pattern
fix(translator): strip unsupported unicode property escape patterns from tool schemas
2026-09-09 18:22:16 +08:00
sususu
37ce368c50 fix(schema): inspect patternProperties keys and avoid Unicode escape fast-path bypass
- Check for \u in fast-path check to prevent JSON Unicode escapes from bypassing inspection.
- Inspect regex keys under patternProperties and drop keys with unsupported Unicode property escapes.
- Add tests covering Unicode escape representations (\u005c, \u0070, \u0050) and patternProperties keys.
2026-09-09 18:03:39 +08:00
sususu
e56abd56f1 fix(translator): strip unsupported unicode property escape patterns from tool schemas
- Add HasUnsupportedUnicodePropertyEscape in internal/util to detect \p{...} / \P{...} escapes that fail Python re compilation.
- Strip incompatible pattern attributes during tool parameter normalization in codex/claude and openai/claude translators.
- Provide schema-aware fallback stripping in codex executor helps to protect downstream Codex requests without mutating non-schema user data.
- Export unified schema keyword lists in internal/util to eliminate duplication.
- Add comprehensive unit tests covering Artifact fixtures, lookaheads, and user data preservation.

Closes: #5644
2026-09-09 17:49:13 +08:00
Luis Pater
b064b832e2 feat(codex): support model-level quota cooling
- Add `model-level-cooling` configuration option to Codex settings.
- Scope `usage_limit_reached` quota cooldowns to the requested model instead of the entire credential when enabled.
- Propagate model-level cooling checks across HTTP, SSE, and WebSocket execution paths.

Closes: #5619
2026-09-09 11:57:58 +08:00
Luis Pater
a59b1764e7 fix(claude): emit trailing usage chunk and aggregate stream usage
- Emit an OpenAI-compatible trailing usage chunk with an empty choices array on `message_stop`.
- Include `cache_write_tokens` in prompt token details for OpenAI response translations.
- Parse usage from `message.usage` and buffer Claude stream usage across chunks to merge input and output token counts.
- Ensure observed streaming usage details are published on completion or stream failure.

Closes: #5617
2026-09-09 10:48:05 +08:00
rome-xi
d8f2dceef7 perf(antigravity): batch functionResponse name repairs
Repair placeholder functionResponse names with part-local sjson and one
Index splice of the request body, matching the sibling content-edit path
instead of one full-body sjson.SetBytes per repaired name.
2026-09-08 18:32:26 +08:00
sususu
390589159e feat(session): enhance harness hierarchy recognition and deduplicate selector extraction 2026-09-08 17:04:58 +08:00
sususu
68dd99d56f fix(antigravity): include resolved pool settings in transport cache key
- Add shortMode, idleConnTimeout, and maxIdleConnsPerHost to antigravityTransportKey.
- Prevent stale transport reuse during hot-reload race windows under load.
- Add TestAntigravityTransportKeySeparatesPoolSettingsAcrossReload regression test.
2026-09-08 11:29:05 +08:00
Luis Pater
d4146bde12 feat(kimi): support openai responses api
- Route `openai-response` source format requests to dedicated streaming and non-streaming responses handlers.
- Add upstream URL resolution for Kimi Responses API endpoints based on auth attributes.
- Normalize `openai-response` provider format to codex during thinking parameter resolution.

Closes: #5400 #5571
2026-09-08 10:52:27 +08:00
이현민
35a4723872 fix(claude): attach helper request IDs according to the upstream base
2.1.258 attaches x-client-request-id only when the base URL is
api.anthropic.com. Since the helper check made this header optional, a
helper that arrived without one gained a freshly generated UUID on a
custom upstream. Attach it only for a first-party upstream or when the
caller actually sent one, and otherwise keep the header absent as the
real client would.
2026-09-08 10:34:05 +08:00
이현민
280b96acea fix(claude): align beta assembly and Haiku helper transport with the measured 2.1.258
advanced-tool-use-2025-11-20 is sent only while tool search or another
advanced tool-use feature is on the wire, or when the caller asks for it;
2.1.258 no longer attaches it to plain tool declarations (measured: 158
inline tools, no beta). A caller-supplied afk-mode-2026-01-31 is forwarded
between fast-mode and extended-cache-ttl and is an insertion boundary for
the advisor beta.

2.1.258 Haiku helpers offer the same full compression set as the main
thread and never send X-Stainless-Async. x-client-request-id is attached
only when the client's base URL is api.anthropic.com, so the helper
transport check accepts an empty value or a valid UUID and rejects only a
malformed one.
2026-09-08 10:34:05 +08:00
Luis Pater
ba7e55836d fix(codex): recognize retryable server errors for bootstrap failover
- Check `error.message` and `message` for retry advice on upstream `server_error` responses.
- Treat server errors indicating the request can be retried as eligible overload bootstrap failures.
2026-09-08 03:49:10 +08:00
Luis Pater
bf20b999de fix(codex): simplify complex tool schema unions and detect empty incomplete responses
- Normalize complex constant `oneOf` and `anyOf` tool parameter schemas into equivalent enums to prevent upstream aborts.
- Escape property keys containing dots and colons during schema updates to prevent invalid path splitting.
- Identify terminal `response.incomplete` events with zero output tokens and no content as upstream failures.

Closes: #5551
2026-09-07 20:09:29 +08:00
sususu
d5397905f0 fix(antigravity): default to short connections and harden connection pool lifecycle (fixes #5494)
- Configure upstream connection pool under antigravity.connection-pool with enabled: false by default.
- In short connection mode, set MaxIdleConnsPerHost = -1 with DisableKeepAlives = false, ensuring immediate TCP termination after response body completion without leaking Connection: close request headers.
- When pooling is explicitly enabled (enabled: true), cap idle-conn-timeout at 210s (leaving a 30s safety buffer below Google Frontend's 240s Keep-Alive cutoff) and default max-idle-conns-per-host to 2 (bounded at 100).
- Refactor TransportCache to execute CloseIdleConnections outside the mutex lock during LRU eviction and matching closes.
- Proactively evict and close idle connections on 429 quota exhaustion across Execute, ExecuteStream, and CountTokens.
- Wire hot-reload diff detection and server reload purge hooks for graceful transport pool updates.
2026-09-07 19:15:39 +08:00
Luis Pater
82f4f370df fix(codex): restore dotted collaboration tool names in multi-agent v2
- Detect `collaboration-optimize.` tool conflicts during multi-agent v2 request optimization.
- Restore dot-prefixed tool calls back to the collaboration namespace and base tool name in responses.

Closes: #5524
2026-09-07 19:12:46 +08:00
Luis Pater
5dc428f392 fix(gemini): append trailing user turn for requests ending with model content
- Ensure Gemini, AI Studio, Vertex, and Antigravity requests end with a user turn to prevent upstream errors when trailing with model content.
- Preserve trailing turns that contain a function response and skip trailing adjustments for `countTokens` actions or Claude models.
- Introduce boundary normalization helpers to validate both leading and trailing turn roles.

Closes: #5358
2026-09-06 23:41:12 +08:00
Luis Pater
0e85eb46f3 fix(xai): support compaction fallback from payload input or previous response id
- Fall back to request payload input items or `previous_response_id` during websocket compaction when transcript snapshot is empty.
- Preserve `previous_response_id` in prepared request body for xAI execution.

Closes: #5205
2026-09-06 22:16:06 +08:00
Luis Pater
70f4560452 feat(antigravity): add conversation compaction support and capsule encryption
- Add payload detection and preparation helpers for compaction triggers and summary generation.
- Implement AES-GCM sealing and unsealing to store and restore opaque compaction capsules across turns.
- Support expanding compaction capsules into developer context for downstream turns.
- Provide summary text extraction across Gemini, Claude, and OpenAI response formats.
- Add builders for compaction non-stream JSON responses and SSE stream frames.

Closes: #5094
2026-09-06 21:51:47 +08:00
Luis Pater
c76dfd4e0e chore(codex): update codex user-agent to 0.153.3
- Update default Codex executor user-agent and model override headers to `codex-tui/0.153.3`.

Closes: #5521
2026-09-06 15:27:15 +08:00
sususu
580df36423 feat(usage): propagate session and parent session hierarchy to usage reporting queue
- Reuse coresession.ExtractSessionInfo across HTTP headers and request payloads to unify canonical session prefix namespaces with the scheduler.
- Extract hierarchical session identities in two phases: initial extraction from request headers on entry, and authoritative deep extraction once request payloads and metadata are available.
- Support Claude Code multi-level subagents (X-Claude-Code-Agent-Id, metadata.agent_id) and Codex thread fork lineages.
- Propagate SessionID and ParentSessionID across ClientRequestMetadata, UsageReporter, and coreusage.Record without root_session_id.
- Include session_id and parent_session_id in queuedUsageDetail for Home LPushUsage forwarding and Redis consumption with self-loop guards.
- Add comprehensive test coverage for canonical headers, body extraction, ghost parent elimination, and self-referential loop guards.
2026-09-06 10:45:32 +08:00
Luis Pater
5ab0bca040 fix(codex): scope usage limit errors to credentials and parse flexible quota resets
- Mark Codex usage limit errors as credential-scoped across HTTP and WebSocket executors.
- Support both top-level and nested error structures with case-insensitive matching when parsing retry-after resets.
- Propagate prevalidated candidate context to session affinity and built-in selectors during auth selection.

Closes: #5529
2026-09-06 06:20:20 +08:00
Luis Pater
7c2f6ce0d1 fix(claude): avoid mid-conversation system splicing for advisor calls or results
- Detect advisor tool calls and advisor results in conversation history.
- Preserve forwarded system prompt blocks in top-level system instead of splicing them into message turns or prepending reminders.
- Prevent message index shifts that break layout bindings for encrypted advisor results and trigger upstream 400 errors.

Closes: #5470
2026-09-06 05:55:42 +08:00
Luis Pater
9dfddd613d fix(aistudio): normalize thinking level to uppercase
- Normalize `generationConfig.thinkingConfig.thinkingLevel` to canonical uppercase enum values (`MINIMAL`, `LOW`, `MEDIUM`, `HIGH`).
- Prevent upstream HTTP 400 invalid argument errors caused by case-sensitive validation.

Closes: #5481
2026-09-06 04:14:41 +08:00
rome-xi
acf919ce50 perf(antigravity): batch reasoning replay mutations 2026-09-04 18:53:21 +08:00
sususu
4a5ab534f8 feat(claude): harden probe and helper request classification, diagnostics isolation, and late cloaking
- Require exactly one non-reminder text block matching quota/test/probe/./Hi for probe request matching.
- Identify title helper requests via expanded session title patterns.
- Implement post-payload bidirectional probe reclassification: strip CPA diagnostics and billing tags on probes, while restoring continuity if declassified.
- Preserve caller-owned and payload-supplied diagnostics on probe requests.
- Gate late sensitive-word obfuscation strictly on cloaked requests in Execute and ExecuteStream.
- Add comprehensive end-to-end tests for probe classification, diagnostics isolation, and late cloaking.
2026-09-04 15:23:31 +08:00
sususu
de4aa60028 feat(claude): add Fable 5.1 reporting outcomes block and post-payload reconciliation
- Inject '# Reporting outcomes' system block for Fable 5.1 / Mythos 5.1 models matching Claude Code 2.1.258 tengu_dapper_lagoon behavior.
- Support both string-format and block-array system prompts for Fable reporting injection.
- Add reconcileClaudeCodeFableModelAfterPayload to reconcile model additions when payload rules override model or thinking.
- Reliably re-add reporting block if omitted after payload rules modify system blocks.
- Drop injected thinking.display on Fable models when payload changes thinking type from adaptive to disabled.
- Enhance payload rule path tracking with ancestor target matching in ApplyPayloadConfigWithTrackedPaths.
2026-09-04 15:23:31 +08:00