Commit Graph

3779 Commits

Author SHA1 Message Date
sususu
eed249072d feat(devin): prefix all Devin model IDs with devin/ namespace 2026-09-13 23:34:43 +08:00
sususu
cbe800aa28 feat(devin): bind upstream session_id and cascade_id to CPA canonical session 2026-09-13 23:34:43 +08:00
sususu
f94752762b feat(devin): add Devin/Cognition provider integration and CLI OAuth
Implement the full Devin/Cognition Connect-RPC provider support across all CPA endpoints (/v1/chat/completions, /v1/messages, /v1/responses), complete with binary protobuf wire framing, streaming tools/arguments delta handling, thinking/reasoning replay, and CLI OAuth authentication.

Key highlights:
- Wire Protocol & Streaming:
  * Implemented Connect-RPC uncompressed 5-byte framing (0x00 + 4-byte length + protobuf) for ApiServerService/GetChatMessage.
  * Implemented Devin protobuf encoder/decoder in internal/runtime/executor/helps/devin_wire.go, including ClientMetadata, prompts, tools, completion_config, and multimodal image handling (Prompt Field #10).
  * Stream frame consumption via interactions protocol, correctly mapping arguments_delta and tracking multiple sequential tool calls (currentToolCallActive).
  * Streaming thought summary and sealed.v1 signature deltas targeting the thinking step.

- Model Registration & Thinking Clamping:
  * Registered static fallback models in model_definitions.go (swe-2, claude-fable-5-1, gpt-6-astra, swe-1-7-lightning, glm-5-2, glm-5-3).
  * Configured ThinkingSupport with discrete levels per model family.
  * Implemented CPA-standard nearest-neighbor clamping for thinking levels (minimal/low -> medium, xhigh -> max for swe-2).
  * Mapped thinking effort to Devin upstream model UID (e.g. swe-2-medium, swe-2-high, swe-2-max).

- Sensitive Words & System Prompt Sanitization:
  * Added devin.sensitive-words configuration in internal/config/config_types.go and config.go, matching Antigravity conventions.
  * Supported zero-width space (\u200b) obfuscation in prompts, tools, and system instructions via SensitiveWordMatcher.
  * Stripped Claude Code billing headers (x-anthropic-billing-header:) and CLI identity signatures from system instructions and tool descriptions to avoid upstream content filter rejections.

- Signature Compatibility:
  * Added SignatureProviderSWE = "swe" recognizing sealed.v1.* reasoning signatures in internal/signature/provider_compatibility.go.
  * Propagated reasoning.encrypted_content on Responses API and thinking.signature on Messages API.

- Authentication:
  * Implemented Devin PKCE OAuth flow with loopback callback server and headless manual token/code paste (--no-browser).
  * Registered Devin authenticator in SDK and CLI (-devin-login flag).
  * Integrated with management OAuth session endpoints and credentials manager.
2026-09-13 23:34:43 +08:00
Supra4E8C
d23ba5ee05 Merge pull request #5795 from router-for-me/feat/management-auth-cooldowns
feat(cooldowns): add cooldown snapshot feature for management auth files
2026-09-13 22:17:48 +08:00
Luis Pater
94d6eb535e docs(config): document payload filter examples for codex tools
- Add example payload filter rules in `config.example.yaml` for stripping tools from both flat and nested `additional_tools` Codex request payloads.

Closes: #5792
2026-09-13 22:12:28 +08:00
Supra4E8C
1ca975dfc0 feat(cooldowns): add cooldown snapshot feature for management auth files 2026-09-13 18:50:21 +08:00
Luis Pater
f702bc1ac2 feat(codex): preserve native fidelity for responses-lite requests
- Detect native responses-lite requests via headers and client metadata.
- Skip instructions normalization and synthetic session cloaking for native requests.
- Preserve upstream completion output during websocket response forwarding.

Closes: #5780
2026-09-13 15:05:03 +08:00
Luis Pater
e696ea47c5 feat(codex): forward X-Codex-Turn-State header in executor requests
- Ensure `X-Codex-Turn-State` is preserved and forwarded from incoming client requests.

Closes: #5778
2026-09-13 14:42:26 +08:00
Luis Pater
ac02da6c05 docs(readme): update APIKEY.FUN sponsor link
- Update sponsor registration URL domain from `apikey.fun` to `apikey.fan` across README files.
v7.2.159
2026-09-13 01:59:34 +08:00
Luis Pater
3c3938feb1 feat(plugins): forward query parameters as metadata in auth provider start login
- Accept optional metadata parameters in SDK and internal plugin host `StartLogin` methods.
- Clone and pass metadata to the auth provider's `AuthLoginStartRequest`.
- Convert management auth URL query parameters into metadata before initiating plugin login flows.

Closes: #5760
2026-09-12 20:05:34 +08:00
Luis Pater
b192f6550c feat(management): add plugin quota and declarative probe endpoints
- Add endpoints to list quota providers and fetch or reset credential quotas via plugins.
- Support declarative metadata quota probes with token substitution and response mapping.
- Clear core routing quota state when provider quota reset succeeds.

Closes: #5752
2026-09-12 19:07:43 +08:00
Luis Pater
e30de3d561 perf(antigravity): cache and deduplicate model capability probe requests
- Deduplicate concurrent capability probe requests using singleflight.
- Cache capability hints with TTL and apply backoff for transient failures.
- Track authentication failures per account to avoid poisoning shared endpoint caches.
- Restrict default model capability base URL to the daily endpoint.

Closes: #5749
2026-09-12 14:06:44 +08:00
Luis Pater
5b2785617d feat(plugins): expose model list responses to plugin interceptors
- Add `WriteModelListResponse` to `BaseAPIHandler` to apply plugin interceptors and record request lifecycles for model catalog responses.
- Update OpenAI, Claude, Gemini, Grok, and Codex model listing endpoints to route responses through the unified interceptor helper.

Closes: #5742
v7.2.158
2026-09-12 00:04:11 +08:00
Luis Pater
d9b8fdb77f fix(claude): forward unmanaged caller betas on direct anthropic requests
- Define a managed Claude beta set to distinguish proxy-governed betas from caller extensions.
- Forward unmanaged caller betas on direct Anthropic endpoints to support newer client features.

Closes: #5738
2026-09-11 23:44:21 +08:00
Luis Pater
456d4c371b fix(auth): clear unauthorized cooldowns on credential changes and sync codex plan type
- Reset unauthorized errors and model cooldowns in lifecycle updates when credentials change.
- Sync `plan_type` attribute from metadata or JWT `id_token` in auth file handlers and synthesizer.
- Invoke `postAuthPersistHook` after auth file upload and field patch operations.

Closes: #5736
2026-09-11 23:31:48 +08:00
Luis Pater
137db720c5 Merge pull request #5746 from smartconnect365/fix/gemini-responses-output-tokens
fix(gemini): include thought tokens in OpenAI responses usage
2026-09-11 23:29:53 +08:00
Luis Pater
9781fe4681 Merge pull request #5744 from smartconnect365/fix/5733
fix(gemini): include thought tokens in OpenAI completion usage
2026-09-11 20:02:36 +08:00
Luis Pater
b5ba02c2e3 fix(websockets): prevent keepalive pong starvation during large payload writes
- Avoid acquiring write mutex in ping handlers so keepalive pongs reply immediately during active writes.
- Stream payload messages larger than 32KB in chunks using writer streams.
- Support ephemeral websocket sessions for sessionless execution and enrich disconnect logs with session kind and terminal event context.

Closes: #5734
2026-09-11 19:53:45 +08:00
charles
c2562d8a5e fix(gemini): include thought tokens in OpenAI responses usage 2026-09-11 19:35:53 +08:00
charles
cd1e8e10c0 fix(gemini): include thought tokens in OpenAI completion usage
Map Gemini thoughtsTokenCount into OpenAI-compatible completion_tokens in
addition to completion_tokens_details.reasoning_tokens. This keeps
prompt_tokens + completion_tokens aligned with total_tokens when Gemini returns
reasoning/thought tokens without visible candidate output.
2026-09-11 19:23:41 +08:00
Supra4E8C
ae8f1f8be5 feat(readme): add RapidProxy sponsorship details and logo in Chinese and Japanese 2026-09-11 18:35:47 +08:00
Supra4E8C
e88cd94767 feat(readme): add RapidProxy sponsorship details and logo 2026-09-11 18:32:32 +08:00
sususu
8bd67f3338 feat(kimi): add Kimi K2.8 model definitions, normalization, and temperature guard
- Add kimi-k2.8 and kimi-k2.8-code with 1M context, 64k completion tokens, low/high/max thinking, and zero_allowed support to models.json.
- Remap K2.8 aliases (kimi-k2.8, k2.8, kimi-k2.8-code, k2.8-code, and -preview variants) to upstream canonical kimi-for-coding.
- Normalize temperature for Kimi upstream to prevent 400 errors (strict 0.6 for disabled thinking, 1.0 for enabled).
- Align zero_allowed: true across K2.8 and K3 models based on live upstream verification of thinking.type=disabled.
- Add test coverage for model normalization, thinking replay family, temperature stripping, and Claude effort=max preservation.
2026-09-11 17:01:00 +08:00
Luis Pater
377c315fd7 fix(claude): anchor billing fingerprint to initial turn for cloaked cache stability
- Derive billing fingerprint message text from the first user message to prevent prompt cache invalidation across turns.
- Centralize continuity tag resolution to maintain stable prompt IDs and request references across requests.

Closes: #5730
2026-09-11 16:59:57 +08:00
sususu
75ce635293 feat(signature): validate Claude CAQS reasoning signatures
Support Claude Envelope v4 reasoning signatures (encoded base64 prefix "CAQS",
EnvelopeVersion 4, Channel ID 17), observed on models such as claude-fable-5-1
and claude-fable-5-1-max:
- Extract signature bytes from Container Field 5 when absent in Channel Field 5
  for EnvelopeVersion >= 4
- Omit mandatory plaintext model_text (Channel Field 6) validation for
  EnvelopeVersion >= 4
- Validate BlockKind is either "thinking" or "narration" for EnvelopeVersion >= 4
- Differentiate CAQS compatibility reason strings when model_text is omitted
- Add test fixtures and regression/rejection test cases for CAQS thinking and narration
2026-09-11 12:20:23 +08:00
Luis Pater
4edf9d1dd6 fix(thinking): extract configuration_update reasoning effort in codex usage reporting
- Inspect `input` items for the latest `configuration_update` reasoning effort before falling back to top-level configuration.
- Route `codex`, `xai`, and `openai-response` providers through usage-specific thinking config extraction.
- Support `none`, `auto`, and explicit reasoning level modes from in-turn updates.

Closes: #5714
2026-09-11 12:15:57 +08:00
sususu
54776fb3a0 fix(signature): preserve Gemini 3 server-side tool thought signatures (#5652)
- Recognize Gemini 3 server-side tool invocation protobuf envelopes (field 1 varint + field 2 Tink ciphertext) via isLikelyGeminiToolInvocationPayload in gemini_validation.
- Early-skip sanitizing toolCall/toolResponse and tool_call/tool_response parts in SanitizeGeminiRequestThoughtSignatures and geminiContentsThoughtSignaturesNeedSanitize, preserving upstream-issued signatures per the Gemini API contract.
- Maintain existing sanitization policies: synthesize skip_thought_signature_validator only on the first functionCall when missing/incompatible, keep sibling functionCalls unsigned, and drop unrecognized foreign signatures on normal text parts to prevent upstream 400 Corrupted thought signature errors.
- Add comprehensive unit tests covering toolCall/toolResponse preservation, snake_case variants, negative malformed envelope rejection, mixed toolCall and unsigned functionCall turns, and tool-block compatibility in ValidateGeminiFunctionCallPairing.

Closes #5652
2026-09-11 09:48:48 +08:00
Luis Pater
d1702fdffd fix(auth): drop stale auth updates using monotonic watcher revisions
- Track monotonic watcher revisions across persisted auth updates to filter out out-of-order events.
- Validate registration epochs before applying auth updates and deletions to prevent stale state overwrites.
- Synchronize auth status patches through post-persist hooks using detached background contexts.
- Guard auth status modifications with a dedicated handler mutex.

Closes: #5729
2026-09-11 08:18:51 +08:00
Luis Pater
5c80a01ca9 fix(translator): support audio transcription parts in gemini response translation
- Fall back to `audioTranscription.text` when explicit text content is absent in both streaming and non-streaming responses.
- Simplify non-streaming message content assignment using accumulated text content.

Closes: #5722
2026-09-11 07:05:25 +08:00
Luis Pater
942bda9979 fix(translator): preserve unknown thinking signatures in claude to codex compat translation
- Retain assistant thinking blocks with unknown-format string signatures when compatibility translation is enabled.
- Fall back to using the raw signature when provider detection classifies the signature provider as unknown.

Closes: #5710
2026-09-11 04:13:27 +08:00
Luis Pater
fc96a87fa6 feat(auth): support organization-hashed claude credentials and legacy migration
- Disambiguate Claude credential filenames using organization and account UUID hashes to keep multiple organizations distinct.
- Migrate legacy Claude credentials during login and save flows while preserving existing metadata and deleting obsolete files.
- Introduce `WithAuthCreationIntent` context policy across token stores to allow creating missing disabled credentials during login and migration.
- Preserve existing `disabled` status during auth metadata merges when not explicitly specified.

Closes: #5709
2026-09-11 03:58:23 +08:00
Luis Pater
4cd17293a3 fix(translator): relay claude tool result images as user messages for openai
- Extract image parts from Claude tool results and replay them in a user message following tool messages.
- Merge relayed images into existing user message content when present within the same turn.
- Insert a placeholder text for image-only tool results to keep tool message content non-empty.

Closes: #5707
2026-09-11 02:52:04 +08:00
Luis Pater
2912516cea feat(auth): support execution result policy before quota and cooldown mutations
- Introduce `ResultPolicy` interface and adapter to inspect and mutate execution results.
- Apply result policy in `MarkResult` prior to in-memory quota mutation, cooldown persistence, and hook invocation.
- Expose result policy configuration across auth `Manager`, SDK `Builder`, and `Service`.

Closes: #5705
2026-09-11 02:19:22 +08:00
Luis Pater
4efcac7966 fix(translator): support incomplete status and response.incomplete event in responses translation
- Handle `response.incomplete` events as terminal interaction completion events during stream translation.
- Propagate upstream response `status` to interaction response and completed event payloads.

Closes: #5702
2026-09-11 02:03:50 +08:00
Luis Pater
8f23ad0291 fix(codex): inherit template metadata for model aliases and restrict provider capabilities
- Resolve canonical templates using metadata model IDs for model aliases and prefixed routes.
- Apply descriptions, base instructions, and thinking support overrides to matched templates.
- Restrict protocol capabilities and reasoning levels based on provider support.

Closes: #5699
2026-09-11 01:42:30 +08:00
sususu
8461b4e91d chore(codex): update codex user-agent to 0.154.0 2026-09-11 00:30:24 +08:00
Luis Pater
c8f723e0fb feat(usage): propagate upstream base_url across usage records and plugin auth
- Add `BaseURL` field to usage records and host auth file entries.
- Extract `base_url` from auth attributes and metadata during usage reporting.
- Propagate `base_url` through plugin usage adapters and runtime auth callbacks.

Closes: #5693
2026-09-11 00:25:17 +08:00
Luis Pater
5a07045e6d fix(translator): ignore fco output item IDs when extracting responses call ID
- Exclude IDs with `fco_` prefix in `ExtractResponsesCallID` to prevent treating output item IDs as tool call identifiers.
- Ensure tool call outputs without explicit call IDs correctly pair with pending function calls via fallback.

Closes: #5692
2026-09-10 20:34:05 +08:00
Luis Pater
6dce78673f fix(auth): bound force refresh all concurrency using worker pool
- Limit concurrent credential refreshes in `ForceRefreshAll` using a worker pool bounded by `AuthAutoRefreshWorkers`.
- Centralize refresh worker pool size resolution in `refreshWorkers`.
- Check context cancellation prior to refreshing to fast-fail queued credentials.

Closes: #5687
2026-09-10 20:21:11 +08:00
Luis Pater
b8e6ec0ae7 fix(gemini): preserve function call pairing for interrupted calls and replay failures
- Synthesize function response parts for interrupted or missing OpenAI Responses tool calls to maintain strict Gemini call-response pairing.
- Preserve function response ordering matching pending call IDs across parallel and partial tool execution turns.
- Degrade gracefully to the original request payload when Antigravity reasoning replay breaks Gemini function call pairing.

Closes: #5682
2026-09-10 20:01:39 +08:00
Luis Pater
9fad505505 fix(auth): treat cloudflare 520-526 origin errors as transient upstream failures
- Exclude HTTP 5xx status codes from Cloudflare challenge classification to avoid treating origin errors as challenges.
- Tighten Cloudflare challenge detection pattern to require challenge indicators instead of generic HTML tags.
- Include HTTP 520-526 status codes in transient error cooldown handling across auth and model states.
- Support upstream `RetryAfter` hints when calculating recoverable failure cooldown durations.

Closes: #5681
2026-09-10 19:31:08 +08:00
Luis Pater
fd3e662376 fix(antigravity): keep active block open on empty text parts in claude stream
- Only transition and close the previous content block when text parts are non-empty.
- Prevent prematurely emitting `content_block_stop` on active blocks like thinking blocks when encountering empty text parts.

Closes: #5674
2026-09-10 19:07:48 +08:00
Luis Pater
c8ecb4f3c9 fix(openai): process reasoning deltas before content in responses stream
- Process `reasoning_content` deltas before message `content` in streaming chat completion translation.
- Ensure reasoning output items receive incremental text before message creation closes reasoning.

Closes: #5659
2026-09-10 18:59:11 +08:00
Luis Pater
09a29bd345 docs(readme): remove RunAPI sponsor
- Remove RunAPI sponsor entries from English, Chinese, and Japanese README files.
v7.2.157
2026-09-10 18:41:46 +08:00
Luis Pater
638ed7e1cc fix(stream): handle split CRLF across chunk boundaries in SSE validation
- Track trailing carriage returns across chunk boundaries in `sseJSONValidationState`.
- Strip leading newline in subsequent chunks to prevent duplicate newline insertion from split CRLF sequences.
- Reset trailing carriage return state upon stream completion.

Closes: #5657
2026-09-10 17:53:25 +08:00
Luis Pater
4dce5f3a2b fix(translator): ignore null and empty finish reasons in openai to gemini response
- Check that `finish_reason` is a non-empty string before mapping to Gemini `finishReason`.
- Prevent chunks or messages with `null` or empty `finish_reason` from emitting unexpected completion statuses.

Closes: #5651
2026-09-10 17:37:15 +08:00
Luis Pater
dde250f1c3 fix(auth): enable model cooldown and rotation for model not found errors
- Map upstream `model_not_found` errors to HTTP 404 before evaluating generic invalid request types in Codex terminal error handling.
- Prevent treating structured model not found responses as client request faults to preserve credential rotation.
- Recognize model access denial errors to apply model-level cooldown and failover.
- Respect `disable_cooling` configuration during model-level cooldown processing.

Closes: #5635
2026-09-10 12:35:07 +08:00
Luis Pater
3ae9093da8 fix(codex): treat model capacity errors as bootstrap overload failures
- Broaden pattern matching for Codex model capacity errors.
- Classify model capacity rejections as overload bootstrap failures to enable failover.

Closes: #5634
2026-09-10 12:06:01 +08:00
Luis Pater
259130863d fix(openai): preserve nested error details and sequence numbers in responses stream
- Format streaming error payloads with nested error objects matching official OpenAI Responses SSE specifications.
- Extract and propagate sequence numbers from upstream terminal events and framer states.
- Use `json.Number` to prevent precision loss for large integers and token metrics.
- Sanitize sensitive keys recursively across nested error objects without dropping custom fields.
2026-09-10 11:58:36 +08:00
Luis Pater
bd03aabcf1 fix(openai): preserve prewarm input and allow named tool outputs in responses websocket
- Track pending synthetic prewarm response IDs to merge warmup inputs into subsequent delta followups.
- Normalize transcript replacements when followups do not reference the prewarm parent response ID.
- Validate that the `input` field is an array for `response.create` requests.
- Allow `function_call_output` items without a `call_id` when a non-empty tool name is present.

Closes: #5631
2026-09-10 11:20:49 +08:00