Files
Illia Polosukhin 764e586717 feat(engine): LLM council via per-call model override in CodeAct (#2320)
* feat(engine): LLM council via per-call model override in CodeAct

Extend `llm_query()` and `llm_query_batched()` with a `model=` (and
`models=` for the batched variant) keyword so CodeAct can route
individual sub-queries to specific LLMs. The "LLM council" pattern
becomes a skill — the agent broadcasts the same prompt across a
parallel array of models and synthesizes the responses — with no new
tool, dispatch path, or capability boundary.

- Add `model: Option<String>` to `LlmCallConfig`; thread it through
  `LlmBridgeAdapter` onto `CompletionRequest.model` /
  `ToolCompletionRequest.model` so providers that honor per-request
  overrides (NEAR AI, Anthropic OAuth, GitHub Copilot, Bedrock) pick
  it up. Other providers fall back to their configured model.
- `handle_llm_query` extracts a `model` arg; `handle_llm_query_batched`
  accepts either `model="..."` (broadcast) or `models=[...]` (parallel
  array, length-validated against `prompts`).
- `__llm_complete__` host fn extracts `model` from explicit_config so
  the Python orchestrator can also forward it.
- Update CodeAct preamble docs so the agent sees the new parameters.
- Add `skills/llm-council/SKILL.md` with the council pattern,
  recommended NEAR AI model line-ups, and a synthesis example.

Tests: 5 new scripting tests (model kwarg forwarding, default `None`,
`models=` broadcast, single-`model=` broadcast, length-mismatch error)
and 2 new bridge tests (config.model → CompletionRequest.model on both
the no-tools and with-tools paths).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): address llm-council review feedback

Address three review comments on the LLM council PR:

1. Test coverage for the orchestrator entry point. Add two tests
   driving `handle_llm_complete` directly with explicit_config
   containing `model` (and a control case without it). Closes the
   "test through the caller, not just the helper" gap — the previous
   tests only exercised `handle_llm_query`, leaving the parallel
   `__llm_complete__` host fn path unverified.

2. Loud failure on non-string entries in `models=[...]`. Previously
   `models=[1, 2]` was silently coerced via `monty_to_string` to
   `["1", "2"]`. Now returns `TypeError` with the offending value,
   matching the existing length-mismatch handling style.

3. `None` slots in `models=[...]` are no longer backfilled by the
   singular `model=` kwarg. Each slot is authoritative: a `None`
   means "no override for this prompt" (use the configured default).
   Mixing the two would have been surprising — the docs now spell
   out the contract explicitly. Add a regression test that passes
   both `models=[None, "gpt-4o"]` and `model="claude-..."` and
   asserts the None slot stays None.

Also update `skills/llm-council/SKILL.md` to default to a 4-model
council of `anthropic/claude-opus-4-6`, `google/gemini-3-pro`,
`zai-org/GLM-latest`, `openai/gpt-5.4`. Per-call provider errors
already flow through the existing `Ok(Err(e))` arm as
`"Error: ..."` strings — the batch never fails as a whole, so
unavailable models just surface in their own slot.

Tests: 4 new (2 orchestrator, 2 scripting), all 346 engine unit
tests pass, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): strict optional-string parsing for llm_query model= kwarg

Address Copilot review on PR #2320. Four comments, all valid:

1 & 2. `model` and `single_model` were extracted via `extract_string_arg`,
   which calls `monty_to_string` — that coerces `MontyObject::None` to the
   literal string "None" and stringifies non-string values (ints become
   "1", etc.). So `llm_query(prompt="hi", model=None)` would silently
   route every call to a bogus model ID called "None".

   Add a strict `extract_optional_string_kwarg` helper that returns
   `Ok(None)` for missing/`None`, `Ok(Some(s))` for strings, and a
   `TypeError` for anything else. Use it in both `handle_llm_query` and
   `handle_llm_query_batched` for the `model=` kwarg. Regression tests
   cover: `model=None` → no override, `model=<int>` → TypeError, and the
   same two cases on the batched path.

3. The `models=` list-type error message said "list of strings" but we
   accept `None` entries. Updated to "list of str or None".

4. SKILL.md claimed the batched call "never raises". It does — for
   argument validation errors (wrong types, length mismatch). Clarified
   that per-model failures return as `"Error: ..."` strings, but
   argument validation still raises.

Tests: 4 new regression tests, all 4687 main-crate and 358 engine tests
pass, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): positional args for llm_query_batched + correct provider docs

Address two review comments from serrrfirat on PR #2320.

1. `llm_query_batched` silently dropped positional `context`/`model`/
   `models` args. The documented signature is
   `llm_query_batched(prompts, context=None, model=None, models=None)`,
   but the extractors were hardcoded to kwargs only (`&[]` for args).
   A call like `llm_query_batched(prompts, None, "gpt-4o")` routed to
   the default model — silent contract violation.

   Thread the real `args` slice into each extractor with the documented
   positional indices: context=1, model=2, models=3. Positional
   `MontyObject::None` at any of those slots now correctly means "no
   override". Added 3 regression tests:
   - `llm_query_batched_honors_positional_context_and_model`
   - `llm_query_batched_honors_positional_models_list`
   - `llm_query_batched_positional_none_for_models_is_no_override`

2. SKILL.md claimed Bedrock honors per-request model overrides, but
   `bedrock.rs::complete()` unconditionally uses
   `self.current_model_id()` and ignores `request.model`. Also, the
   default 4-model prefixed lineup (`anthropic/...`, `google/...`,
   `openai/...`) only works on aggregator backends like NEAR AI — a
   direct Anthropic OAuth or Copilot provider honors `set_model` but
   can only switch between models within its own vendor.

   Rewrite the SKILL.md preamble with a provider capability table
   (dropping Bedrock from "honors it" and adding cross-vendor routing
   as a separate column), and add per-backend default lineups:
   NEAR AI (prefixed cross-vendor), Anthropic OAuth (Anthropic tiers
   only), Copilot (Copilot-exposed models). For backends that don't
   honor `model=` at all (Bedrock, raw OpenAI/Ollama/Tinfoil), the
   skill now instructs the agent to tell the user and fall back to a
   single-model answer.

Tests: 3 new regression tests, all 4687 main-crate and 361 engine tests
pass, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 09:23:05 +03:00
..