- Add example payload filter rules in `config.example.yaml` for stripping tools from both flat and nested `additional_tools` Codex request payloads.
Closes: #5792
- Add endpoints to list quota providers and fetch or reset credential quotas via plugins.
- Support declarative metadata quota probes with token substitution and response mapping.
- Clear core routing quota state when provider quota reset succeeds.
Closes: #5752
- Deduplicate concurrent capability probe requests using singleflight.
- Cache capability hints with TTL and apply backoff for transient failures.
- Track authentication failures per account to avoid poisoning shared endpoint caches.
- Restrict default model capability base URL to the daily endpoint.
Closes: #5749
- Add `WriteModelListResponse` to `BaseAPIHandler` to apply plugin interceptors and record request lifecycles for model catalog responses.
- Update OpenAI, Claude, Gemini, Grok, and Codex model listing endpoints to route responses through the unified interceptor helper.
Closes: #5742
- Define a managed Claude beta set to distinguish proxy-governed betas from caller extensions.
- Forward unmanaged caller betas on direct Anthropic endpoints to support newer client features.
Closes: #5738
- Reset unauthorized errors and model cooldowns in lifecycle updates when credentials change.
- Sync `plan_type` attribute from metadata or JWT `id_token` in auth file handlers and synthesizer.
- Invoke `postAuthPersistHook` after auth file upload and field patch operations.
Closes: #5736
- Avoid acquiring write mutex in ping handlers so keepalive pongs reply immediately during active writes.
- Stream payload messages larger than 32KB in chunks using writer streams.
- Support ephemeral websocket sessions for sessionless execution and enrich disconnect logs with session kind and terminal event context.
Closes: #5734
Map Gemini thoughtsTokenCount into OpenAI-compatible completion_tokens in
addition to completion_tokens_details.reasoning_tokens. This keeps
prompt_tokens + completion_tokens aligned with total_tokens when Gemini returns
reasoning/thought tokens without visible candidate output.
- Add kimi-k2.8 and kimi-k2.8-code with 1M context, 64k completion tokens, low/high/max thinking, and zero_allowed support to models.json.
- Remap K2.8 aliases (kimi-k2.8, k2.8, kimi-k2.8-code, k2.8-code, and -preview variants) to upstream canonical kimi-for-coding.
- Normalize temperature for Kimi upstream to prevent 400 errors (strict 0.6 for disabled thinking, 1.0 for enabled).
- Align zero_allowed: true across K2.8 and K3 models based on live upstream verification of thinking.type=disabled.
- Add test coverage for model normalization, thinking replay family, temperature stripping, and Claude effort=max preservation.
- Derive billing fingerprint message text from the first user message to prevent prompt cache invalidation across turns.
- Centralize continuity tag resolution to maintain stable prompt IDs and request references across requests.
Closes: #5730
Support Claude Envelope v4 reasoning signatures (encoded base64 prefix "CAQS",
EnvelopeVersion 4, Channel ID 17), observed on models such as claude-fable-5-1
and claude-fable-5-1-max:
- Extract signature bytes from Container Field 5 when absent in Channel Field 5
for EnvelopeVersion >= 4
- Omit mandatory plaintext model_text (Channel Field 6) validation for
EnvelopeVersion >= 4
- Validate BlockKind is either "thinking" or "narration" for EnvelopeVersion >= 4
- Differentiate CAQS compatibility reason strings when model_text is omitted
- Add test fixtures and regression/rejection test cases for CAQS thinking and narration
- Inspect `input` items for the latest `configuration_update` reasoning effort before falling back to top-level configuration.
- Route `codex`, `xai`, and `openai-response` providers through usage-specific thinking config extraction.
- Support `none`, `auto`, and explicit reasoning level modes from in-turn updates.
Closes: #5714
- Recognize Gemini 3 server-side tool invocation protobuf envelopes (field 1 varint + field 2 Tink ciphertext) via isLikelyGeminiToolInvocationPayload in gemini_validation.
- Early-skip sanitizing toolCall/toolResponse and tool_call/tool_response parts in SanitizeGeminiRequestThoughtSignatures and geminiContentsThoughtSignaturesNeedSanitize, preserving upstream-issued signatures per the Gemini API contract.
- Maintain existing sanitization policies: synthesize skip_thought_signature_validator only on the first functionCall when missing/incompatible, keep sibling functionCalls unsigned, and drop unrecognized foreign signatures on normal text parts to prevent upstream 400 Corrupted thought signature errors.
- Add comprehensive unit tests covering toolCall/toolResponse preservation, snake_case variants, negative malformed envelope rejection, mixed toolCall and unsigned functionCall turns, and tool-block compatibility in ValidateGeminiFunctionCallPairing.
Closes#5652
- Track monotonic watcher revisions across persisted auth updates to filter out out-of-order events.
- Validate registration epochs before applying auth updates and deletions to prevent stale state overwrites.
- Synchronize auth status patches through post-persist hooks using detached background contexts.
- Guard auth status modifications with a dedicated handler mutex.
Closes: #5729
- Fall back to `audioTranscription.text` when explicit text content is absent in both streaming and non-streaming responses.
- Simplify non-streaming message content assignment using accumulated text content.
Closes: #5722
- Retain assistant thinking blocks with unknown-format string signatures when compatibility translation is enabled.
- Fall back to using the raw signature when provider detection classifies the signature provider as unknown.
Closes: #5710
- Disambiguate Claude credential filenames using organization and account UUID hashes to keep multiple organizations distinct.
- Migrate legacy Claude credentials during login and save flows while preserving existing metadata and deleting obsolete files.
- Introduce `WithAuthCreationIntent` context policy across token stores to allow creating missing disabled credentials during login and migration.
- Preserve existing `disabled` status during auth metadata merges when not explicitly specified.
Closes: #5709
- Extract image parts from Claude tool results and replay them in a user message following tool messages.
- Merge relayed images into existing user message content when present within the same turn.
- Insert a placeholder text for image-only tool results to keep tool message content non-empty.
Closes: #5707
- Introduce `ResultPolicy` interface and adapter to inspect and mutate execution results.
- Apply result policy in `MarkResult` prior to in-memory quota mutation, cooldown persistence, and hook invocation.
- Expose result policy configuration across auth `Manager`, SDK `Builder`, and `Service`.
Closes: #5705
- Resolve canonical templates using metadata model IDs for model aliases and prefixed routes.
- Apply descriptions, base instructions, and thinking support overrides to matched templates.
- Restrict protocol capabilities and reasoning levels based on provider support.
Closes: #5699
- Add `BaseURL` field to usage records and host auth file entries.
- Extract `base_url` from auth attributes and metadata during usage reporting.
- Propagate `base_url` through plugin usage adapters and runtime auth callbacks.
Closes: #5693
- Exclude IDs with `fco_` prefix in `ExtractResponsesCallID` to prevent treating output item IDs as tool call identifiers.
- Ensure tool call outputs without explicit call IDs correctly pair with pending function calls via fallback.
Closes: #5692
- Limit concurrent credential refreshes in `ForceRefreshAll` using a worker pool bounded by `AuthAutoRefreshWorkers`.
- Centralize refresh worker pool size resolution in `refreshWorkers`.
- Check context cancellation prior to refreshing to fast-fail queued credentials.
Closes: #5687
- Synthesize function response parts for interrupted or missing OpenAI Responses tool calls to maintain strict Gemini call-response pairing.
- Preserve function response ordering matching pending call IDs across parallel and partial tool execution turns.
- Degrade gracefully to the original request payload when Antigravity reasoning replay breaks Gemini function call pairing.
Closes: #5682
- Exclude HTTP 5xx status codes from Cloudflare challenge classification to avoid treating origin errors as challenges.
- Tighten Cloudflare challenge detection pattern to require challenge indicators instead of generic HTML tags.
- Include HTTP 520-526 status codes in transient error cooldown handling across auth and model states.
- Support upstream `RetryAfter` hints when calculating recoverable failure cooldown durations.
Closes: #5681
- Only transition and close the previous content block when text parts are non-empty.
- Prevent prematurely emitting `content_block_stop` on active blocks like thinking blocks when encountering empty text parts.
Closes: #5674
- Track trailing carriage returns across chunk boundaries in `sseJSONValidationState`.
- Strip leading newline in subsequent chunks to prevent duplicate newline insertion from split CRLF sequences.
- Reset trailing carriage return state upon stream completion.
Closes: #5657
- Check that `finish_reason` is a non-empty string before mapping to Gemini `finishReason`.
- Prevent chunks or messages with `null` or empty `finish_reason` from emitting unexpected completion statuses.
Closes: #5651
- Map upstream `model_not_found` errors to HTTP 404 before evaluating generic invalid request types in Codex terminal error handling.
- Prevent treating structured model not found responses as client request faults to preserve credential rotation.
- Recognize model access denial errors to apply model-level cooldown and failover.
- Respect `disable_cooling` configuration during model-level cooldown processing.
Closes: #5635
- Broaden pattern matching for Codex model capacity errors.
- Classify model capacity rejections as overload bootstrap failures to enable failover.
Closes: #5634
- Format streaming error payloads with nested error objects matching official OpenAI Responses SSE specifications.
- Extract and propagate sequence numbers from upstream terminal events and framer states.
- Use `json.Number` to prevent precision loss for large integers and token metrics.
- Sanitize sensitive keys recursively across nested error objects without dropping custom fields.
- Track pending synthetic prewarm response IDs to merge warmup inputs into subsequent delta followups.
- Normalize transcript replacements when followups do not reference the prewarm parent response ID.
- Validate that the `input` field is an array for `response.create` requests.
- Allow `function_call_output` items without a `call_id` when a non-empty tool name is present.
Closes: #5631