- Register static model definitions for devin/gemini-3-8-flash (1M context, Google) and devin/grok-4-6 (500k context, xAI).
- Support gemini38Efforts (low/medium/high) and grok46Efforts (low/medium/high/xhigh) in ResolveDevinChatModelUID, supporting both colon and parenthesis suffix parsing.
- Recognize Gemini Tink thought signatures (AY-prefix / 0x01 Tink header) in detectSignatureType and parseSignatureBytes.
- Add unit tests for both models in registry and devin_models.
- Implement Connect-RPC GetUserStatus serialization and response parsing in internal/auth/devin/user_status.go.
- Extract user email, plan, username, user_id, team_id, org_id, daily/weekly quota percentages, and reset timestamps.
- Wire user status into DevinExecutor.Refresh to update auth metadata and Quota.Signals.
- Add devin to ProviderSupportsQuotaObservation so CPA management endpoints surface quota observations.
- Enrich Devin OAuth login flow with user status, email, and quota information, and add CSRF state verification.
- Support base_url override in DevinAuthService for mock testing and custom gateways.
- Add example payload filter rules in `config.example.yaml` for stripping tools from both flat and nested `additional_tools` Codex request payloads.
Closes: #5792
- Add endpoints to list quota providers and fetch or reset credential quotas via plugins.
- Support declarative metadata quota probes with token substitution and response mapping.
- Clear core routing quota state when provider quota reset succeeds.
Closes: #5752
- Deduplicate concurrent capability probe requests using singleflight.
- Cache capability hints with TTL and apply backoff for transient failures.
- Track authentication failures per account to avoid poisoning shared endpoint caches.
- Restrict default model capability base URL to the daily endpoint.
Closes: #5749
- Add `WriteModelListResponse` to `BaseAPIHandler` to apply plugin interceptors and record request lifecycles for model catalog responses.
- Update OpenAI, Claude, Gemini, Grok, and Codex model listing endpoints to route responses through the unified interceptor helper.
Closes: #5742
- Define a managed Claude beta set to distinguish proxy-governed betas from caller extensions.
- Forward unmanaged caller betas on direct Anthropic endpoints to support newer client features.
Closes: #5738
- Reset unauthorized errors and model cooldowns in lifecycle updates when credentials change.
- Sync `plan_type` attribute from metadata or JWT `id_token` in auth file handlers and synthesizer.
- Invoke `postAuthPersistHook` after auth file upload and field patch operations.
Closes: #5736
- Avoid acquiring write mutex in ping handlers so keepalive pongs reply immediately during active writes.
- Stream payload messages larger than 32KB in chunks using writer streams.
- Support ephemeral websocket sessions for sessionless execution and enrich disconnect logs with session kind and terminal event context.
Closes: #5734
Map Gemini thoughtsTokenCount into OpenAI-compatible completion_tokens in
addition to completion_tokens_details.reasoning_tokens. This keeps
prompt_tokens + completion_tokens aligned with total_tokens when Gemini returns
reasoning/thought tokens without visible candidate output.
- Add kimi-k2.8 and kimi-k2.8-code with 1M context, 64k completion tokens, low/high/max thinking, and zero_allowed support to models.json.
- Remap K2.8 aliases (kimi-k2.8, k2.8, kimi-k2.8-code, k2.8-code, and -preview variants) to upstream canonical kimi-for-coding.
- Normalize temperature for Kimi upstream to prevent 400 errors (strict 0.6 for disabled thinking, 1.0 for enabled).
- Align zero_allowed: true across K2.8 and K3 models based on live upstream verification of thinking.type=disabled.
- Add test coverage for model normalization, thinking replay family, temperature stripping, and Claude effort=max preservation.
- Derive billing fingerprint message text from the first user message to prevent prompt cache invalidation across turns.
- Centralize continuity tag resolution to maintain stable prompt IDs and request references across requests.
Closes: #5730
Support Claude Envelope v4 reasoning signatures (encoded base64 prefix "CAQS",
EnvelopeVersion 4, Channel ID 17), observed on models such as claude-fable-5-1
and claude-fable-5-1-max:
- Extract signature bytes from Container Field 5 when absent in Channel Field 5
for EnvelopeVersion >= 4
- Omit mandatory plaintext model_text (Channel Field 6) validation for
EnvelopeVersion >= 4
- Validate BlockKind is either "thinking" or "narration" for EnvelopeVersion >= 4
- Differentiate CAQS compatibility reason strings when model_text is omitted
- Add test fixtures and regression/rejection test cases for CAQS thinking and narration
- Inspect `input` items for the latest `configuration_update` reasoning effort before falling back to top-level configuration.
- Route `codex`, `xai`, and `openai-response` providers through usage-specific thinking config extraction.
- Support `none`, `auto`, and explicit reasoning level modes from in-turn updates.
Closes: #5714
- Recognize Gemini 3 server-side tool invocation protobuf envelopes (field 1 varint + field 2 Tink ciphertext) via isLikelyGeminiToolInvocationPayload in gemini_validation.
- Early-skip sanitizing toolCall/toolResponse and tool_call/tool_response parts in SanitizeGeminiRequestThoughtSignatures and geminiContentsThoughtSignaturesNeedSanitize, preserving upstream-issued signatures per the Gemini API contract.
- Maintain existing sanitization policies: synthesize skip_thought_signature_validator only on the first functionCall when missing/incompatible, keep sibling functionCalls unsigned, and drop unrecognized foreign signatures on normal text parts to prevent upstream 400 Corrupted thought signature errors.
- Add comprehensive unit tests covering toolCall/toolResponse preservation, snake_case variants, negative malformed envelope rejection, mixed toolCall and unsigned functionCall turns, and tool-block compatibility in ValidateGeminiFunctionCallPairing.
Closes#5652