Commit Graph

218 Commits

Author SHA1 Message Date
Luis Pater
44eaef0009 feat(claude): support model-level cooling and scope overage rate limits
- Add `claude.model-level-cooling` configuration to scope rate limit cooldowns to the requested model.
- Treat overage-only and spend cap rejections as model-scoped when shared subscription windows remain healthy.
- Propagate model-level cooling settings into streaming, token counting, and direct execution error classifiers.

Closes: #5915
2026-09-18 01:55:36 +08:00
Luis Pater
77820cb2f4 feat(translator): support alternative reasoning fields in OpenAI Claude translation
- Extract reasoning text sequentially from `reasoning_content`, `reasoning`, and `reasoning_details`.
- Unify reasoning extraction across streaming deltas and non-streaming message conversions.

Closes: #5872
2026-09-17 04:17:30 +08:00
Luis Pater
8335eac731 feat(config): add Meta OAuth aliasing and error rules support
- Introduced `meta` as a supported OAuth channel for model aliasing and request-scoped error rules.
- Added alias config to rename `muse-spark-1.3` as `muse-latest` with enforced mapping and fork behavior.
- Extended error rules for `meta` to classify "context_length_exceeded" responses with status `400` as `stop` actions.
- Updated examples and synthesizer logic to merge exclusions and apply Meta-specific mechanisms.
2026-09-15 20:08:12 +08:00
Luis Pater
42ca5d3412 Merge PR #5502 feat/meta-provider into cpa/muse
Bring in native Meta (Muse Code) provider support while keeping the
existing Devin integration and original Meta commit history.
2026-09-15 12:05:40 +08:00
sususu98
9b52a49927 Merge pull request #5796 from router-for-me/feat/ai-gateway-discovery
feat(discovery): introduce mDNS/DNS-SD LAN service advertising and CLI discovery
2026-09-14 17:07:32 +08:00
sususu
9e847e596e fix(discovery): uniquify instance names without a startup browse
Always advertise CPA-<ID> or <service-name>-<ID> so LAN instance names
are unique from local state instead of probing mDNS at process start.
2026-09-14 16:35:57 +08:00
sususu
6e307553f4 feat(codex): add optional time ceiling for stream bootstrap buffering
- Add `stream-bootstrap-timeout` configuration (defaulting to 0/unlimited, recommended 20s behind reverse proxies) to bound how long early handshake or trickled events may hold response headers.
- Release stream buffering into normal in-stream delivery once the time budget is exhausted on both SSE and WebSocket executors.
- Deliver post-timeout overload and status-bearing errors in-stream rather than triggering credential failover, preventing latency doubling on long reasoning turns.
- Provide thread-safe mock clock test harness and comprehensive unit tests covering timeout release, unlimited default, disabled ceilings, and post-timeout error delivery.
2026-09-14 14:25:45 +08:00
sususu98
bb20fa2d5e Merge pull request #5724 from Viggo95/fix/codex-bootstrap-buffer-noncontent-events
fix(codex): keep bootstrap buffer open for events that carry no output
2026-09-14 12:27:00 +08:00
sususu
c1b7c91f2f fix(discovery): honor caller cancellation and reload endpoint 2026-09-14 12:22:22 +08:00
sususu
3428110d49 feat(discovery): add LAN gateway discovery
Advertise _ai-gateway._tcp with API-protocol DNS-SD subtypes, refresh
on address changes, uniquify colliding instance names, and filter
discovery to physical LAN interfaces.
2026-09-14 10:54:46 +08:00
sususu
f94752762b feat(devin): add Devin/Cognition provider integration and CLI OAuth
Implement the full Devin/Cognition Connect-RPC provider support across all CPA endpoints (/v1/chat/completions, /v1/messages, /v1/responses), complete with binary protobuf wire framing, streaming tools/arguments delta handling, thinking/reasoning replay, and CLI OAuth authentication.

Key highlights:
- Wire Protocol & Streaming:
  * Implemented Connect-RPC uncompressed 5-byte framing (0x00 + 4-byte length + protobuf) for ApiServerService/GetChatMessage.
  * Implemented Devin protobuf encoder/decoder in internal/runtime/executor/helps/devin_wire.go, including ClientMetadata, prompts, tools, completion_config, and multimodal image handling (Prompt Field #10).
  * Stream frame consumption via interactions protocol, correctly mapping arguments_delta and tracking multiple sequential tool calls (currentToolCallActive).
  * Streaming thought summary and sealed.v1 signature deltas targeting the thinking step.

- Model Registration & Thinking Clamping:
  * Registered static fallback models in model_definitions.go (swe-2, claude-fable-5-1, gpt-6-astra, swe-1-7-lightning, glm-5-2, glm-5-3).
  * Configured ThinkingSupport with discrete levels per model family.
  * Implemented CPA-standard nearest-neighbor clamping for thinking levels (minimal/low -> medium, xhigh -> max for swe-2).
  * Mapped thinking effort to Devin upstream model UID (e.g. swe-2-medium, swe-2-high, swe-2-max).

- Sensitive Words & System Prompt Sanitization:
  * Added devin.sensitive-words configuration in internal/config/config_types.go and config.go, matching Antigravity conventions.
  * Supported zero-width space (\u200b) obfuscation in prompts, tools, and system instructions via SensitiveWordMatcher.
  * Stripped Claude Code billing headers (x-anthropic-billing-header:) and CLI identity signatures from system instructions and tool descriptions to avoid upstream content filter rejections.

- Signature Compatibility:
  * Added SignatureProviderSWE = "swe" recognizing sealed.v1.* reasoning signatures in internal/signature/provider_compatibility.go.
  * Propagated reasoning.encrypted_content on Responses API and thinking.signature on Messages API.

- Authentication:
  * Implemented Devin PKCE OAuth flow with loopback callback server and headless manual token/code paste (--no-browser).
  * Registered Devin authenticator in SDK and CLI (-devin-login flag).
  * Integrated with management OAuth session endpoints and credentials manager.
2026-09-13 23:34:43 +08:00
Luis Pater
94d6eb535e docs(config): document payload filter examples for codex tools
- Add example payload filter rules in `config.example.yaml` for stripping tools from both flat and nested `additional_tools` Codex request payloads.

Closes: #5792
2026-09-13 22:12:28 +08:00
Viggo95
cb73cd9936 fix(codex): buffer keepalive and empty item announcements during codex bootstrap
The bootstrap buffer released the stream on the first frame outside the
handshake allow-list. Upstream sends `keepalive` heartbeats and
`response.output_item.added` before the first token, so the downstream headers
were committed while nothing had been produced, and an overload rejection
arriving afterwards could no longer be retried on another credential.

Extend the allow-list with `keepalive`, and with `response.output_item.added`,
`response.content_part.added` and `response.reasoning_summary_part.added` when
they announce something that is still empty. `response.output_item.added` is
accepted only for `message`, `reasoning`, `function_call` and `custom_tool_call`
items, and only while their content, summary, arguments or input is empty: every
other item type stands for a server-side operation that may already be running -
a `web_search_call` is announced with status "in_progress" and its `searching`
event follows immediately - and failing the attempt over after one would run
that operation again on another credential. Part announcements are matched on a
closed list of textual part types for the same reason.

The list stays closed. "Nothing has happened yet" cannot be derived from the
absence of a TTFT token, because TTFT deliberately ignores server-side tool
traffic such as `response.shell_call_output_content.delta` and its `.done`
counterpart, so an unrecognised frame, an unrecognised item type and an
unrecognised part type all release the stream exactly as before.

Holding those frames also required fixing the bound on how much may be held.
codexBootstrapMaxBufferedEvents was enforced with len(bufferedChunks), and a
chunk count bounds only the downstream formats that render every upstream frame:
the OpenAI Chat Completions and Gemini translators return zero chunks for a frame
they do not recognise, which is true of every frame added here. Count frames read
from the upstream instead, and cap what a bootstrap accumulates over the upstream
frames and the chunks they translate into. The SSE scanner walks physical lines
and an event can arrive as one line (": keepalive"), two, or three
(event:/data:/blank), so the frame budget is sized for the widest framing rather
than derived from any separator. The websocket loop counts a message as soon as
it is read, before the branches that skip non-text and whitespace-only messages,
so a peer sending only frames the loop skips cannot hold the downstream headers
open; the message that exhausts the budget is still processed and delivered
rather than dropped. Both caps are checked before a frame is admitted, so one
oversized frame cannot be taken on the strength of an empty buffer. This widens
the window where one frame rendered one chunk - the websocket path effectively
moves from 16 frames to 48.

An empty `response.incomplete` seen while buffering is now delivered in-stream
instead of failing the attempt over, matching the websocket executor and the
documented contract that only overload and rate-limit rejections trigger
failover.

`isCodexHandshakeMetadataEvent` is renamed to `isCodexBootstrapBufferableEvent`
because the list is no longer only handshake metadata. Both config doc sites
describe what is held, that heartbeats are held too, that SSE also holds the
lines around a `data:` frame, and that the hold is bounded by frames and bytes
rather than by wall-clock time.
2026-09-12 22:27:44 +08:00
Kenny
c6e076ff41 Merge dev into feat/meta-provider and preserve both registry tests 2026-09-10 14:03:29 +00:00
Luis Pater
6dce78673f fix(auth): bound force refresh all concurrency using worker pool
- Limit concurrent credential refreshes in `ForceRefreshAll` using a worker pool bounded by `AuthAutoRefreshWorkers`.
- Centralize refresh worker pool size resolution in `refreshWorkers`.
- Check context cancellation prior to refreshing to fast-fail queued credentials.

Closes: #5687
2026-09-10 20:21:11 +08:00
Luis Pater
9fad505505 fix(auth): treat cloudflare 520-526 origin errors as transient upstream failures
- Exclude HTTP 5xx status codes from Cloudflare challenge classification to avoid treating origin errors as challenges.
- Tighten Cloudflare challenge detection pattern to require challenge indicators instead of generic HTML tags.
- Include HTTP 520-526 status codes in transient error cooldown handling across auth and model states.
- Support upstream `RetryAfter` hints when calculating recoverable failure cooldown durations.

Closes: #5681
2026-09-10 19:31:08 +08:00
Luis Pater
b064b832e2 feat(codex): support model-level quota cooling
- Add `model-level-cooling` configuration option to Codex settings.
- Scope `usage_limit_reached` quota cooldowns to the requested model instead of the entire credential when enabled.
- Propagate model-level cooling checks across HTTP, SSE, and WebSocket execution paths.

Closes: #5619
2026-09-09 11:57:58 +08:00
이현민
280b96acea fix(claude): align beta assembly and Haiku helper transport with the measured 2.1.258
advanced-tool-use-2025-11-20 is sent only while tool search or another
advanced tool-use feature is on the wire, or when the caller asks for it;
2.1.258 no longer attaches it to plain tool declarations (measured: 158
inline tools, no beta). A caller-supplied afk-mode-2026-01-31 is forwarded
between fast-mode and extended-cache-ttl and is an insertion boundary for
the advisor beta.

2.1.258 Haiku helpers offer the same full compression set as the main
thread and never send X-Stainless-Async. x-client-request-id is attached
only when the client's base URL is api.anthropic.com, so the helper
transport check accepts an empty value or a valid UUID and rejects only a
malformed one.
2026-09-08 10:34:05 +08:00
sususu
d5397905f0 fix(antigravity): default to short connections and harden connection pool lifecycle (fixes #5494)
- Configure upstream connection pool under antigravity.connection-pool with enabled: false by default.
- In short connection mode, set MaxIdleConnsPerHost = -1 with DisableKeepAlives = false, ensuring immediate TCP termination after response body completion without leaking Connection: close request headers.
- When pooling is explicitly enabled (enabled: true), cap idle-conn-timeout at 210s (leaving a 30s safety buffer below Google Frontend's 240s Keep-Alive cutoff) and default max-idle-conns-per-host to 2 (bounded at 100).
- Refactor TransportCache to execute CloseIdleConnections outside the mutex lock during LRU eviction and matching closes.
- Proactively evict and close idle connections on 429 quota exhaustion across Execute, ExecuteStream, and CountTokens.
- Wire hot-reload diff detection and server reload purge hooks for graceful transport pool updates.
2026-09-07 19:15:39 +08:00
Kenny
21aa46b603 chore(meta): keep provider PR scoped to Meta support 2026-09-05 04:23:06 +00:00
Kenny
54d4f4c019 feat(meta): add native Meta (Muse Code) provider and OAuth device flow 2026-09-04 13:41:24 +00:00
Luis Pater
ba2cdea3b9 fix(translator/claude): handle incomplete status and terminal state on max_tokens
- Map Claude `max_tokens` stop reason to `response.incomplete` with `max_output_tokens` incomplete details.
- Propagate `incomplete` status to trailing output items across streaming and non-streaming responses.
- Defer and track output item finalization to ensure correct status and arguments when turns are truncated.

Closes: #5439
2026-09-04 00:39:18 +08:00
Luis Pater
291cfb87ef feat(codex): support orphan delegation compatibility via orphan-delegation-compatibility
- Add `codex.orphan-delegation-compatibility` configuration option and mirror it to SDK configuration.
- Convert orphan Codex delegation outputs into standard user messages for requests with `X-Openai-Subagent: collab_spawn`.
- Integrate orphan delegation rewriting into OpenAI responses request handling pipeline.

Closes: #5401
2026-09-03 20:21:24 +08:00
Luis Pater
d577e630b1 feat(routing): make subagent session affinity configurable via session-affinity-subagents
- Add routing.session-affinity-subagents (defaulting to true) to RoutingConfig.
- Unify subagent parent credential inheritance across all providers by removing hardcoded provider/model blacklists.
- When session-affinity-subagents is false, isolate subagents and distribute them via the fallback selector.
- Ensure changing session-affinity-subagents is a no-op when session-affinity is false.
- Preserve alias isolation and failure isolation invariants.

Closes: #5417
2026-09-03 17:05:35 +08:00
sususu
df7e04ea28 fix(claude): upgrade default Claude Code baseline and fingerprint to 2.1.258
- Bump default Claude Code version baseline from 2.1.220 to 2.1.258 to resolve upstream 400 version gate error
- Update Stainless SDK package version to 0.112.1 and runtime version to v26.3.0
- Refresh default device profile, billing header build hashes, and cloaking signatures
- Decouple unconfirmed client OS/Arch test assertions from host runner platform
- Update config.example.yaml and test suites across executor and helps packages
2026-09-03 10:21:13 +08:00
hkfires
e04d620cc1 feat(auth): normalize credential metadata keys
Canonicalize legacy config-style credential keys across stores,
management handlers, plugin auth, and file synthesis while preserving
explicit canonical values. Expose per-auth request_retry in auth file
management and add max-retry-credentials management routes.
2026-08-22 12:01:06 +08:00
hkfires
0a14eb70ce feat(auth): add retry round credential filtering 2026-08-22 01:20:02 +08:00
hkfires
601ca43090 feat(auth): add credential retry round contract
Redefine request-retry as additional credential retry rounds and
enforce max-retry-credentials per round. Home dispatch now carries
excluded and pinned auth constraints, supports remote retry limits, and
propagates cooldown retry-after metadata across exhausted rounds.

Move Antigravity upstream retries under conductor ownership to avoid
double-consuming retry attempts. Update configuration comments and add
coverage for Home retry rounds, cooldown handling, pinned credentials,
and legacy dispatcher compatibility.
2026-08-22 01:19:58 +08:00
sususu98
4b9d404fb0 feat(codex): add opt-in stream bootstrap buffering and overload failover (#5115)
* feat(codex): add opt-in stream bootstrap buffering

The upstream smuggles capacity rejections into an HTTP 200 stream. The
handshake events arrive normally and only a later event carries
{"error":{"type":"service_unavailable_error","code":
"server_is_overloaded"}}. By then the executor has already handed the
first chunk downstream, the response is committed, and the conductor can
no longer retry on another credential, so the request fails even though
other credentials were available.

When codex.stream-bootstrap-buffering is enabled the executor holds back
the handshake events until it can tell whether the stream carries real
output or a rejection. An overload rejection then fails the attempt
before any chunk is delivered, letting the conductor retry on another
credential; every other terminal failure is flushed in order and
delivered in-stream exactly as before.

Detection uses an event-type allow-list rather than a fixed count. On the
websocket transport codex.rate_limits and codex.response.metadata arrive
before response.created, making the first generated event the fifth
frame, so a small counter would release the stream before the rejection
is visible. Buffering is bounded and hitting the bound degrades to the
original unbuffered behaviour.

Two details are load-bearing. The error must be returned synchronously:
delivering it as the first stream chunk makes ExecuteStream downgrade it
into a committed 200 and the status is lost. And the websocket path must
not signal an upstream disconnect for a rejection it intends to retry,
because the downstream handler closes the client connection on that
signal and the retry would have nowhere to deliver.

The 503 status is produced only on this path rather than in the shared
codexTerminalFailureStatus mapping, so disabling the feature restores the
previous behaviour exactly, including cooldown classification and
retry-after parsing.

Defaults to false: response headers are withheld until generation
starts, which can trip client or reverse-proxy read timeouts.

* test(codex): pin bootstrap overload failover through the conductor

Executor-level tests cannot show what the client finally receives. These
exercise ExecuteStream end to end to pin three properties that are easy
to regress:

- consecutive overloaded credentials are skipped until one serves the
  request, and retries are capped by max-retry-credentials rather than
  multiplying with request-retry
- exhausting the pool surfaces the upstream status instead of a
  committed 200 stream
- with buffering disabled the rejection stays an in-stream error on a
  committed stream, which is the behaviour the feature must preserve

The third case also documents why the executor returns its error
synchronously: an error arriving as the first stream chunk is wrapped and
downgraded into a committed 200, silently losing the status.
2026-08-20 21:31:53 +08:00
Luis Pater
9dc51b1f87 feat(cliproxy): add OAuth request-scoped error rules support
- Add `oauth-request-scoped-errors` configuration with normalization, sanitization, and YAML management persistence/hot-reload hooks.
- Route request-scoped error classification to use per-provider rules only for OAuth auth entries.
- Add config diff reporting and management CRUD endpoints for `oauth-request-scoped-errors` (get/put/patch/delete) with input sanitization.

Closes: #5085
2026-08-20 02:55:49 +08:00
Luis Pater
e424bfad00 feat(executor): support $-based custom headers from downstream request headers
- Propagate request headers into custom-header resolution for OpenAI/Gemini/XAI/Codex execution and websocket flows.
- Resolve auth `header:` values like `$ABC` from incoming request headers at request time and omit headers when no value is available.
- Add documentation for the dynamic custom-header behavior in `config.example.yaml`.

Closes: #5053
2026-08-18 19:33:50 +08:00
sususu
aec70dfec4 fix(claude): gate CCH signing by upstream origin and validate fingerprint-profile
Claude Code 2.1.220 through 2.1.234 emit the cch attribution only for
firstParty on api.anthropic.com and for vertex; every other backend sends
the billing header unsigned. CPA had dropped its endpoint check, so an
opted-in API key signed a per-request hash on any gateway and could bust
that gateway's prompt cache.

- Restore the endpoint gate in claudeCCHSigningEnabled: a real Claude OAuth
  credential still signs on every upstream, because a downstream Claude Code
  pointed at CPA cannot produce that value itself, while a claude-code-cli
  API key signs only on api.anthropic.com or Vertex
- Drop the unused origin parameter from Claude fingerprint policy resolution
  and restore the original resolveClaudeWirePolicy signature; the wire profile
  follows the credential and only CCH follows the origin
- Add config.NormalizeClaudeFingerprintProfile / ValidateClaudeFingerprintProfile
  as the single source of truth for fingerprint-profile values
- Reject unknown fingerprint-profile values in the Management API, and warn
  once per distinct value at request time instead of on every resolution,
  which previously logged about four warnings per request for one typo
- Preserve unrecognized values through config sanitization so rewriting a
  config file never discards operator input
- Update config.example.yaml and tests for the origin-scoped CCH behavior
2026-08-18 18:06:40 +08:00
sususu98
f1b0431c77 feat(claude): add fingerprint-profile=claude-code-cli for API keys and delegated providers (#5047)
* feat(config): add fingerprint-profile to Claude keys and auth JSON

- Add FingerprintProfile to ClaudeKey configuration struct and normalizer
- Track fingerprint-profile in config diff
- Map fingerprint-profile / fingerprint_profile to auth attributes in file and config synthesizers
- Support fingerprint-profile in Management API PatchClaudeKey and normalization
- Add Claude billing attribution string manipulation utilities in internal/util
- Document fingerprint-profile options in config.example.yaml

* feat(claude): add fingerprint policy and request-local CLI identity

- Centralize Claude fingerprint policy resolution in claude_fingerprint_policy.go
- Support stable Claude CLI identity synthesis (UUIDv5 account_uuid and SHA-256 device_id)
  seeded from API keys or stable OAuth IDs, keeping access tokens isolated
- Warn on unrecognized fingerprint-profile values

* feat(claude): apply CLI fingerprint to Messages and keep API keys caller-owned

- Wire centralized fingerprint policy into Claude and Kimi executors
- Keep first-party Anthropic API keys and delegated providers caller-owned by default
- Apply Claude Code CLI wire profile (betas, metadata, diagnostics, MCP aliases)
  when fingerprint-profile=claude-code-cli is configured
- Strictly align CCH signing with native Claude Code 2.1.220: only first-party
  api.anthropic.com and Vertex sign dynamic CCH; third-party gateways and Kimi
  receive billing header without cch= to avoid prompt cache busting
- Respect caller-owned count_tokens bodies by default while aligning CLI shape on opt-in
- Fall back to CLIProxyAPI/<version> User-Agent when caller sends no UA in caller-owned mode
- Scope custom operator header overrides accurately in caller-owned mode
- Add comprehensive test coverage for policy resolution, gateway opt-in, Kimi, and token counting
2026-08-18 17:29:26 +08:00
hkfires
5bffd1514f Refactor cooling management across configuration and handlers
- Changed `DisableCooling` from a boolean to a pointer in various config types to allow explicit inheritance.
- Updated tests to reflect the new pointer usage for `DisableCooling`.
- Enhanced the `BuildConfigChangeDetails` function to handle optional boolean changes for `DisableCooling`.
- Added new tests to ensure proper handling of cooling overrides in configurations.
- Refactored the `SetQuotaCooldownDisabled` function and related logic to clarify the purpose of cooldown management.
- Introduced new tests for cooling override precedence in the auth manager.
- Ensured that all relevant handlers and synthesizers correctly manage the `DisableCooling` setting.
2026-08-17 10:54:07 +08:00
Luis Pater
aa10847e12 feat(auth): add request-scoped error action handling in conductor
- Add request-scoped error rule extraction from auth metadata or runtime provider config (including OpenAI compatibility fallback)
- Match rules by HTTP/status-code plus error body substring or regex patterns
- Support `stop`, `stop-and-cooldown`, `continue`, `continue-and-cooldown` actions with normalized validation
- Apply matched actions to execution results via request-scoped vs force-cooldown error codes and stop/continue flow control
- Introduce request-stop error wrappers/helpers for matching and unwrapping scoped stop state

Closes: #5006
2026-08-16 21:39:38 +08:00
Luis Pater
6f2cea9484 feat(config): add per-credential request-retry override support
Closes: #4931
2026-08-13 01:50:47 +08:00
Luis Pater
0a95fa62a1 feat(compat): preserve Claude thinking/tool-call content for is-compat OpenAI compatibility models
- Add `is-compat` support to OpenAI compatibility model config, capabilities, hashing, and example config.
- Propagate `IsCompat` through API-key model resolution and switch OpenAI-compat executor translation to compatibility-aware routing.
- Keep Claude assistant thinking content in compatibility mode while keeping default behavior unchanged when `is-compat` is disabled.

Closes: #4776
2026-08-07 06:26:33 +08:00
Luis Pater
dcee14dd3c feat(compat): preserve compat-mode thinking/signature blocks for API-key models
- Added `is-compat` model metadata plumbing from config through executor and helpers, including hash computation.
- Introduced a compatibility-aware translation path (`TranslateRequestWithAPIKeyModelCompatibility`) and wired it into Claude/Gemini/Codex/Interactions request flows.
- Updated Claude message sanitization/translation behavior to keep empty-thinking compatibility blocks (including signatures) when `is-compat` is enabled, while keeping default behavior unchanged.
2026-08-06 17:19:24 +08:00
sususu
65fc536e6e fix(auth): preserve session affinity across priorities
Keep a healthy cached session binding even when a higher-priority credential recovers. Continue applying credential priority to cold selection, plugin scheduling, and true failover.

Resolve the across-priority candidate list in a single availability pass and narrow the highest-priority tier from the same cloned auths, so the selection path does not scan or deep-clone candidates twice. Document that an established binding outranks credential priority.

Fixes #4808
2026-08-06 10:29:57 +08:00
Luis Pater
e5ea945ed9 feat(codex): add model-level is-compat flag to rewrite MultiAgentV2 agent_message for Responses-compatible endpoints
Closes: #4801
2026-08-06 04:49:28 +08:00
Luis Pater
42eef103d6 feat(antigravity): obfuscate sensitive words in system instructions
Closes: #4696 #4723 #4732
2026-08-05 00:27:32 +08:00
Luis Pater
7fe8473766 feat(codex): prepare multi-agent v2 tool definitions at the Responses boundary for Codex clients 2026-08-03 22:22:15 +08:00
sususu
903e41b6dc docs(claude): clarify exact fingerprint baseline 2026-08-03 15:25:17 +08:00
sususu
f63a925d15 fix(claude): replay measured OAuth wire, Fast and diagnostic profiles
Align the remaining measured OAuth wire profiles, including the ordered
connection writer in internal/httpwire that reproduces the observed header
sequence, and the refresh/profile response shapes in internal/auth/claude.

Replay the measured Fast path and keep diagnostic continuity across cloaked and
native requests.

Preserve the native direct token-counting shape so a caller that reaches
count_tokens itself is not reshaped into the cloaked form.

Scope cloak dates to the credential's timezone rather than the host's, so
currentDate matches what the real client would have sent for that account.
2026-08-03 14:47:26 +08:00
sususu
a2933c7737 fix(claude): scope Anthropic beta and count_tokens policies
Claude Code builds Anthropic-Beta per request instead of sending a fixed list.
Captured from an isolated 2.1.220 profile pointed at api.anthropic.com through
a local proxy, over two rounds covering 11 model IDs and the [1m] variants:

  constant  claude-code, interleaved-thinking, redact-thinking,
            thinking-token-count, context-management, prompt-caching-scope
  tools     advanced-tool-use-2025-11-20 only when tools are declared
  model     mid-conversation-system-2026-04-07 only on models that accept a
            role=system turn
  [1m]      context-1m-2025-08-07, directly after claude-code-20250219 rather
            than at the end
  trailing  effort-2025-11-24, then server-side-fallback-2026-06-01

claude-sonnet-5 emits mid-conversation-system-2026-04-07, so it accepts a
role=system turn and must not sit in the legacy reminder whitelist.

count_tokens does not reuse the inference fingerprint. Running /context in an
interactive session issues 37 identical calls, which made the endpoint
observable for the first time: four betas only, and 21 headers rather than 22
because X-Stainless-Timeout is absent. The profile is selected from the request
path so no call site has to thread another flag.
2026-08-03 14:47:26 +08:00
sususu
ef89c6a69d fix(claude): reconstruct cloaked system prompts like the real client
Preserve a cloaked caller's own system prompt instead of discarding it, place
it as a mid-conversation system turn on models that accept one, and route the
remaining legacy models through system reminders.

Scope the legacy reminder whitelist to official model IDs. claude-opus-4-6-thinking
was dropped: Anthropic publishes no -thinking IDs, that one belongs to the
antigravity provider in models.json and is served by a different executor, so it
can never reach ClaudeExecutor cloaking. Keeping it implied that synthetic
suffixes are normalized here, which they are not, since thinking.ParseSuffix only
strips parenthesis suffixes. The map is anchored to the "claude" provider block
plus Anthropic's bare and "-latest" aliases, and now covers claude-opus-4-7.
2026-08-03 14:47:26 +08:00
sususu
f3e25ab2ba feat(claude): align OAuth wire identity and TLS with Claude Code 2.1.220
Detect confirmed CLI, sdk-cli and VSCode callers before mutation so native
software, system, tool, cache and beta shapes pass through, while unconfirmed
OAuth clients receive a coherent minimum CLI identity.

Persist each Claude OAuth credential's upstream account metadata and one stable
device ID, derive one stable session per agent conversation, and keep body and
header identity synchronized across Messages, streaming and count_tokens.

Alias every cloaked third-party custom tool through caller-stable opaque MCP
names and restore declarations, choices, history, references, non-stream
responses and SSE events without changing tool ownership.

Implement the Claude Code 2.1.220 CCH algorithm over the final serialized
request bytes, align currentDate and first-user cache layout, update the
official beta/header baseline, and use upstream count_tokens for OAuth and
first-party Anthropic credentials.

Match the 2.1.220 TLS ClientHello so the transport fingerprint agrees with the
identity the request now claims, and document the CLI defaults and automatic
OAuth signing / tool alias behaviour in config.example.yaml.
2026-08-03 14:47:26 +08:00
Luis Pater
13435c93d2 feat(executor): normalize OpenAI tool results for text-only compatibility models
- Add OpenAI-compat executor-time normalization for `tool` message content, converting non-string tool results to plain text and replacing image parts with a clear unsupported marker for models configured with text-only `input-modalities`.
- Apply normalization in both regular and streaming OpenAI-compat execution paths before prompt-cache processing.
- Update config example to document `[text]` `input-modalities` for upstreams that reject multimodal tool-result content.

Closes: #4737
2026-08-03 03:37:24 +08:00
Luis Pater
a303fd869b feat(codex): support max-context-length overrides for configured models
- Add optional `max-context-length` model configuration across supported provider model types and expose it via `GetMaxContextLength`.
- Propagate the override into model metadata so Codex/client model catalog responses honor the configured value (`context_window`, `max_context_window`, and `max_context_length`).
- Update example configuration with documented usage of the new option.

Closes: #4728
2026-08-03 03:01:43 +08:00
sususu
91561df0b1 fix(codex): align websocket cloaking headers 2026-08-01 13:01:33 +08:00