Commit Graph

998 Commits

Author SHA1 Message Date
sususu
f1f5506c0b fix(devin): align wire protocol, harden streaming, and resolve multi-turn tool/signature parity
- Wire parity: align Connect-RPC Sentry-Trace, User-Agent suppression, float32 double pattern, and dynamic 732-char hex device fingerprint
- Session ordinal & cache: implement process-scoped Field 15.2 with bounded LRU (5000 entries) and Field 15.4=14 user boundary; prioritize stable session_id over previous_interaction_id to preserve prompt caching
- Streaming robustness: unblock hung TCP reads on client cancellation via context watcher; accurately propagate stream read errors and trailer errors instead of swallowing truncated frames
- Thought signature & reasoning: emit raw delta signatures directly in active thought steps; eliminate redundant tail base64 re-encoding; ensure 1:1 assistant signature and thinking alignment across multi-turn history
- Tool call de-multiplexing: route parallel tool calls by tc.Index in both streaming step events and non-streaming aggregations
- Security & transport: escape OAuth callback error HTML against reflected XSS, enforce strict state validation, and isolate Devin HTTP transport with tr.Clone()
2026-09-13 23:34:43 +08:00
sususu
85ddf3aeb5 fix(devin): calculate total_input_tokens and total_tokens correctly 2026-09-13 23:34:43 +08:00
sususu
bf06746d42 feat(devin): map model aliases and swe-1-7/haiku/sonnet/gpt-4-1 UIDs 2026-09-13 23:34:43 +08:00
sususu
ca664c6ede feat(devin): clamp maxTokens to model MaxCompletionTokens 2026-09-13 23:34:43 +08:00
sususu
2683ec201d feat(cmd): add fetch_devin_models CLI tool for dynamic model catalog extraction 2026-09-13 23:34:43 +08:00
sususu
469aa3678f feat(devin): parse protobuf timestamp and harden partial failure logging 2026-09-13 23:34:43 +08:00
sususu
16cb6c0b02 feat(devin): add symmetric decoded upstream response in request log 2026-09-13 23:34:43 +08:00
sususu
59df75d20c docs(devin): update DevinExecutor comment with verbatim reconstructed cloud system prompt 2026-09-13 23:34:43 +08:00
sususu
0c2351bb89 feat(devin): support none thinking level for glm-5-2 2026-09-13 23:34:43 +08:00
sususu
308e5ad3b1 feat(devin): add deepseek-v4-flash and deepseek-v4-1-flash models 2026-09-13 23:34:43 +08:00
sususu
8a3770710e docs(devin): document cloud-side system instructions baseline in DevinExecutor 2026-09-13 23:34:43 +08:00
sususu
2caab7dbf9 feat(devin): enhance request-log with intermediate interactions and decoded upstream body 2026-09-13 23:34:43 +08:00
sususu
ea2f29feec feat(devin): add devin/gemini-3-8-flash and devin/grok-4-6 model definitions and signature recognition
- Register static model definitions for devin/gemini-3-8-flash (1M context, Google) and devin/grok-4-6 (500k context, xAI).
- Support gemini38Efforts (low/medium/high) and grok46Efforts (low/medium/high/xhigh) in ResolveDevinChatModelUID, supporting both colon and parenthesis suffix parsing.
- Recognize Gemini Tink thought signatures (AY-prefix / 0x01 Tink header) in detectSignatureType and parseSignatureBytes.
- Add unit tests for both models in registry and devin_models.
2026-09-13 23:34:43 +08:00
sususu
5b8e3821b1 fix(devin): strip system prompt lines matching configured sensitive words to evade unicode normalization bypass 2026-09-13 23:34:43 +08:00
sususu
7b5741c639 feat(devin): support credential quota and seat status query via GetUserStatus
- Implement Connect-RPC GetUserStatus serialization and response parsing in internal/auth/devin/user_status.go.
- Extract user email, plan, username, user_id, team_id, org_id, daily/weekly quota percentages, and reset timestamps.
- Wire user status into DevinExecutor.Refresh to update auth metadata and Quota.Signals.
- Add devin to ProviderSupportsQuotaObservation so CPA management endpoints surface quota observations.
- Enrich Devin OAuth login flow with user status, email, and quota information, and add CSRF state verification.
- Support base_url override in DevinAuthService for mock testing and custom gateways.
2026-09-13 23:34:43 +08:00
sususu
c0b76c2d09 refactor(devin): keep sensitive words strictly external in config.yaml without hardcoding 2026-09-13 23:34:43 +08:00
sususu
c0b86059c4 fix(devin): sanitize claude subagent identity and emoji directives to prevent content policy 403 2026-09-13 23:34:43 +08:00
sususu
d115fe2c45 refactor(devin): align sensitive-words with antigravity to config.yaml only 2026-09-13 23:34:43 +08:00
sususu
1b6948513d feat(devin): support sensitive-words in auth json metadata and attributes 2026-09-13 23:34:43 +08:00
sususu
f5247e496f fix(devin): restrict sensitive word obfuscation strictly to system prompt only 2026-09-13 23:34:43 +08:00
sususu
02fd1bde78 feat(devin): restrict glm-5-2 to free tier, remove static swe-1-7-lightning, and harden cloak 2026-09-13 23:34:43 +08:00
sususu
eed249072d feat(devin): prefix all Devin model IDs with devin/ namespace 2026-09-13 23:34:43 +08:00
sususu
cbe800aa28 feat(devin): bind upstream session_id and cascade_id to CPA canonical session 2026-09-13 23:34:43 +08:00
sususu
f94752762b feat(devin): add Devin/Cognition provider integration and CLI OAuth
Implement the full Devin/Cognition Connect-RPC provider support across all CPA endpoints (/v1/chat/completions, /v1/messages, /v1/responses), complete with binary protobuf wire framing, streaming tools/arguments delta handling, thinking/reasoning replay, and CLI OAuth authentication.

Key highlights:
- Wire Protocol & Streaming:
  * Implemented Connect-RPC uncompressed 5-byte framing (0x00 + 4-byte length + protobuf) for ApiServerService/GetChatMessage.
  * Implemented Devin protobuf encoder/decoder in internal/runtime/executor/helps/devin_wire.go, including ClientMetadata, prompts, tools, completion_config, and multimodal image handling (Prompt Field #10).
  * Stream frame consumption via interactions protocol, correctly mapping arguments_delta and tracking multiple sequential tool calls (currentToolCallActive).
  * Streaming thought summary and sealed.v1 signature deltas targeting the thinking step.

- Model Registration & Thinking Clamping:
  * Registered static fallback models in model_definitions.go (swe-2, claude-fable-5-1, gpt-6-astra, swe-1-7-lightning, glm-5-2, glm-5-3).
  * Configured ThinkingSupport with discrete levels per model family.
  * Implemented CPA-standard nearest-neighbor clamping for thinking levels (minimal/low -> medium, xhigh -> max for swe-2).
  * Mapped thinking effort to Devin upstream model UID (e.g. swe-2-medium, swe-2-high, swe-2-max).

- Sensitive Words & System Prompt Sanitization:
  * Added devin.sensitive-words configuration in internal/config/config_types.go and config.go, matching Antigravity conventions.
  * Supported zero-width space (\u200b) obfuscation in prompts, tools, and system instructions via SensitiveWordMatcher.
  * Stripped Claude Code billing headers (x-anthropic-billing-header:) and CLI identity signatures from system instructions and tool descriptions to avoid upstream content filter rejections.

- Signature Compatibility:
  * Added SignatureProviderSWE = "swe" recognizing sealed.v1.* reasoning signatures in internal/signature/provider_compatibility.go.
  * Propagated reasoning.encrypted_content on Responses API and thinking.signature on Messages API.

- Authentication:
  * Implemented Devin PKCE OAuth flow with loopback callback server and headless manual token/code paste (--no-browser).
  * Registered Devin authenticator in SDK and CLI (-devin-login flag).
  * Integrated with management OAuth session endpoints and credentials manager.
2026-09-13 23:34:43 +08:00
Luis Pater
94d6eb535e docs(config): document payload filter examples for codex tools
- Add example payload filter rules in `config.example.yaml` for stripping tools from both flat and nested `additional_tools` Codex request payloads.

Closes: #5792
2026-09-13 22:12:28 +08:00
Luis Pater
f702bc1ac2 feat(codex): preserve native fidelity for responses-lite requests
- Detect native responses-lite requests via headers and client metadata.
- Skip instructions normalization and synthetic session cloaking for native requests.
- Preserve upstream completion output during websocket response forwarding.

Closes: #5780
2026-09-13 15:05:03 +08:00
Luis Pater
e696ea47c5 feat(codex): forward X-Codex-Turn-State header in executor requests
- Ensure `X-Codex-Turn-State` is preserved and forwarded from incoming client requests.

Closes: #5778
2026-09-13 14:42:26 +08:00
Luis Pater
d9b8fdb77f fix(claude): forward unmanaged caller betas on direct anthropic requests
- Define a managed Claude beta set to distinguish proxy-governed betas from caller extensions.
- Forward unmanaged caller betas on direct Anthropic endpoints to support newer client features.

Closes: #5738
2026-09-11 23:44:21 +08:00
Luis Pater
b5ba02c2e3 fix(websockets): prevent keepalive pong starvation during large payload writes
- Avoid acquiring write mutex in ping handlers so keepalive pongs reply immediately during active writes.
- Stream payload messages larger than 32KB in chunks using writer streams.
- Support ephemeral websocket sessions for sessionless execution and enrich disconnect logs with session kind and terminal event context.

Closes: #5734
2026-09-11 19:53:45 +08:00
sususu
8bd67f3338 feat(kimi): add Kimi K2.8 model definitions, normalization, and temperature guard
- Add kimi-k2.8 and kimi-k2.8-code with 1M context, 64k completion tokens, low/high/max thinking, and zero_allowed support to models.json.
- Remap K2.8 aliases (kimi-k2.8, k2.8, kimi-k2.8-code, k2.8-code, and -preview variants) to upstream canonical kimi-for-coding.
- Normalize temperature for Kimi upstream to prevent 400 errors (strict 0.6 for disabled thinking, 1.0 for enabled).
- Align zero_allowed: true across K2.8 and K3 models based on live upstream verification of thinking.type=disabled.
- Add test coverage for model normalization, thinking replay family, temperature stripping, and Claude effort=max preservation.
2026-09-11 17:01:00 +08:00
Luis Pater
377c315fd7 fix(claude): anchor billing fingerprint to initial turn for cloaked cache stability
- Derive billing fingerprint message text from the first user message to prevent prompt cache invalidation across turns.
- Centralize continuity tag resolution to maintain stable prompt IDs and request references across requests.

Closes: #5730
2026-09-11 16:59:57 +08:00
Luis Pater
4edf9d1dd6 fix(thinking): extract configuration_update reasoning effort in codex usage reporting
- Inspect `input` items for the latest `configuration_update` reasoning effort before falling back to top-level configuration.
- Route `codex`, `xai`, and `openai-response` providers through usage-specific thinking config extraction.
- Support `none`, `auto`, and explicit reasoning level modes from in-turn updates.

Closes: #5714
2026-09-11 12:15:57 +08:00
sususu
8461b4e91d chore(codex): update codex user-agent to 0.154.0 2026-09-11 00:30:24 +08:00
Luis Pater
c8f723e0fb feat(usage): propagate upstream base_url across usage records and plugin auth
- Add `BaseURL` field to usage records and host auth file entries.
- Extract `base_url` from auth attributes and metadata during usage reporting.
- Propagate `base_url` through plugin usage adapters and runtime auth callbacks.

Closes: #5693
2026-09-11 00:25:17 +08:00
Luis Pater
b8e6ec0ae7 fix(gemini): preserve function call pairing for interrupted calls and replay failures
- Synthesize function response parts for interrupted or missing OpenAI Responses tool calls to maintain strict Gemini call-response pairing.
- Preserve function response ordering matching pending call IDs across parallel and partial tool execution turns.
- Degrade gracefully to the original request payload when Antigravity reasoning replay breaks Gemini function call pairing.

Closes: #5682
2026-09-10 20:01:39 +08:00
Luis Pater
dde250f1c3 fix(auth): enable model cooldown and rotation for model not found errors
- Map upstream `model_not_found` errors to HTTP 404 before evaluating generic invalid request types in Codex terminal error handling.
- Prevent treating structured model not found responses as client request faults to preserve credential rotation.
- Recognize model access denial errors to apply model-level cooldown and failover.
- Respect `disable_cooling` configuration during model-level cooldown processing.

Closes: #5635
2026-09-10 12:35:07 +08:00
Luis Pater
3ae9093da8 fix(codex): treat model capacity errors as bootstrap overload failures
- Broaden pattern matching for Codex model capacity errors.
- Classify model capacity rejections as overload bootstrap failures to enable failover.

Closes: #5634
2026-09-10 12:06:01 +08:00
Luis Pater
259130863d fix(openai): preserve nested error details and sequence numbers in responses stream
- Format streaming error payloads with nested error objects matching official OpenAI Responses SSE specifications.
- Extract and propagate sequence numbers from upstream terminal events and framer states.
- Use `json.Number` to prevent precision loss for large integers and token metrics.
- Sanitize sensitive keys recursively across nested error objects without dropping custom fields.
2026-09-10 11:58:36 +08:00
Luis Pater
6a73f39627 fix(claude): preserve 1h cache ttl and beta header for subagent requests
- Retain `cache_control` blocks with 1h TTL and `extended-cache-ttl` beta header when explicitly requested by subagents.
- Detect 1h TTL configuration from request payloads and incoming Anthropic-Beta headers.
- Ensure `extended-cache-ttl` beta header is preserved or injected when 1h TTL is present.

Closes: #5629
2026-09-10 10:58:34 +08:00
Luis Pater
d1a024e940 feat(codex): add support for gpt-image-2.5 models
- Register builtin model definitions for `gpt-image-2.5`, `gpt-image-2.5-flare`, and `gpt-image-2.5-sunburst`.
- Update OpenAI image handlers and request routing to recognize GPT Image 2.5 models.
- Support direct image generation and edit execution for GPT Image 2.5 variants in the Codex executor.
- Apply client visibility overrides to hide new builtin image models where appropriate.
2026-09-10 10:40:53 +08:00
Luis Pater
3bf787fc1d feat(auth): propagate canonical session id for custom header expansion
- Initialize `util.SessionIDResolver` to resolve session IDs from request contexts, metadata, and headers.
- Ensure canonical session metadata is injected into execution options across execution flows.
- Propagate session context in executors to support `$CPA-SESSION-ID` expansion in custom headers.
- Sync cleared or updated session identities back to execution contexts during conductor execution.

Closes: #5690
2026-09-10 01:43:37 +08:00
Luis Pater
db011593ca Merge pull request #5677 from sususu98/fix/codex-tool-schema-pattern
fix(translator): strip unsupported unicode property escape patterns from tool schemas
2026-09-09 18:22:16 +08:00
sususu
37ce368c50 fix(schema): inspect patternProperties keys and avoid Unicode escape fast-path bypass
- Check for \u in fast-path check to prevent JSON Unicode escapes from bypassing inspection.
- Inspect regex keys under patternProperties and drop keys with unsupported Unicode property escapes.
- Add tests covering Unicode escape representations (\u005c, \u0070, \u0050) and patternProperties keys.
2026-09-09 18:03:39 +08:00
sususu
e56abd56f1 fix(translator): strip unsupported unicode property escape patterns from tool schemas
- Add HasUnsupportedUnicodePropertyEscape in internal/util to detect \p{...} / \P{...} escapes that fail Python re compilation.
- Strip incompatible pattern attributes during tool parameter normalization in codex/claude and openai/claude translators.
- Provide schema-aware fallback stripping in codex executor helps to protect downstream Codex requests without mutating non-schema user data.
- Export unified schema keyword lists in internal/util to eliminate duplication.
- Add comprehensive unit tests covering Artifact fixtures, lookaheads, and user data preservation.

Closes: #5644
2026-09-09 17:49:13 +08:00
Luis Pater
b064b832e2 feat(codex): support model-level quota cooling
- Add `model-level-cooling` configuration option to Codex settings.
- Scope `usage_limit_reached` quota cooldowns to the requested model instead of the entire credential when enabled.
- Propagate model-level cooling checks across HTTP, SSE, and WebSocket execution paths.

Closes: #5619
2026-09-09 11:57:58 +08:00
Luis Pater
a59b1764e7 fix(claude): emit trailing usage chunk and aggregate stream usage
- Emit an OpenAI-compatible trailing usage chunk with an empty choices array on `message_stop`.
- Include `cache_write_tokens` in prompt token details for OpenAI response translations.
- Parse usage from `message.usage` and buffer Claude stream usage across chunks to merge input and output token counts.
- Ensure observed streaming usage details are published on completion or stream failure.

Closes: #5617
2026-09-09 10:48:05 +08:00
rome-xi
d8f2dceef7 perf(antigravity): batch functionResponse name repairs
Repair placeholder functionResponse names with part-local sjson and one
Index splice of the request body, matching the sibling content-edit path
instead of one full-body sjson.SetBytes per repaired name.
2026-09-08 18:32:26 +08:00
sususu
390589159e feat(session): enhance harness hierarchy recognition and deduplicate selector extraction 2026-09-08 17:04:58 +08:00
sususu
68dd99d56f fix(antigravity): include resolved pool settings in transport cache key
- Add shortMode, idleConnTimeout, and maxIdleConnsPerHost to antigravityTransportKey.
- Prevent stale transport reuse during hot-reload race windows under load.
- Add TestAntigravityTransportKeySeparatesPoolSettingsAcrossReload regression test.
2026-09-08 11:29:05 +08:00
Luis Pater
d4146bde12 feat(kimi): support openai responses api
- Route `openai-response` source format requests to dedicated streaming and non-streaming responses handlers.
- Add upstream URL resolution for Kimi Responses API endpoints based on auth attributes.
- Normalize `openai-response` provider format to codex during thinking parameter resolution.

Closes: #5400 #5571
2026-09-08 10:52:27 +08:00