feat(engine-v2): Phase 4 cost tracking + Phase 6 mission lifecycle acceptance (#2660)

* feat(engine-v2): Phase 4 cost tracking + Phase 6 mission lifecycle acceptance

Closes two gaps blocking v2 engine becoming the default:

**Phase 4 — token + cost accounting**

- Delete orphaned `crates/ironclaw_engine/src/executor/compaction.rs`
  (176 lines). The Python orchestrator (`default.py::compact_if_needed`)
  has owned compaction policy since #1557; the Rust module had no
  callers anywhere in the workspace.
- Wire `cost_usd` in `LlmBridgeAdapter` by calling
  `LlmProvider::calculate_cost()` at both the no-tools and with-tools
  response paths. The engine's `Thread::total_cost_usd` accumulator and
  `max_budget_usd` gates were already plumbed — only the adapter was
  hardcoding 0.0.
- Persist `total_cost_usd` through `ThreadArchiveSummary` round-trip in
  `store_adapter.rs`. Previously, rehydrating an archived thread
  silently dropped the cost to 0.0. `#[serde(default)]` keeps existing
  archive files deserializing cleanly.

**Phase 6 — mission lifecycle acceptance**

Three new integration tests in `bridge/effect_adapter.rs` driving
`execute_action()` end-to-end (per `.claude/rules/testing.md` "Test
Through the Caller"):

- `mission_full_lifecycle_via_execute_action` — create → list → complete
  → list, asserting the `Completed` status surfaces through
  `mission_list` after `mission_complete`.
- `mission_fire_returns_thread_id_for_manual_cadence_via_execute_action`
  — fresh manual mission fires successfully and returns a UUID thread_id
  rather than `not_fired`.
- `mission_list_returns_all_user_missions_via_execute_action` — all
  three created missions appear in `mission_list` output.

**Regression tests for cost wiring**

Three new tests in `bridge/llm_adapter.rs`:

- `complete_no_tools_populates_cost_usd_through_adapter`
- `complete_with_tools_populates_cost_usd_through_adapter`
- `complete_routes_subcalls_through_cheap_provider_for_cost` — pins that
  `depth > 0` is priced with the cheap provider, not the primary.

Coordinated with in-flight work: skipped paths owned by #2504 (auth
E2E), #2631 (paused-lease resume), #2570 (mission re-fire), #2549
(mission_get), #2452 (tool_calls persistence), #2621 (replay snapshot).

Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features` (0 warnings), `cargo test -p ironclaw_engine` (409
passed), `cargo test -p ironclaw --lib` (5079 passed).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(engine-v2): surface engine capability actions to LLM via available_actions

Fixes the gap called out in the PR body: `EffectBridgeAdapter::available_actions`
was only enumerating v1 `ToolRegistry` tools + latent OAuth actions, so
engine-native capabilities like `missions` never appeared in the LLM's
tools list even when a thread held an active lease for them. The LLM
was therefore unable to call `mission_create` / `mission_list` / etc.
via structured tool calls; the only ways to drive missions were CodeAct
Python calls (which relied on the same `known_actions` set and hit the
same gap) or `/routine` slash commands falling through to v1.

Wire `CapabilityRegistry` into the adapter and iterate active leases to
surface every leased, engine-registered capability action. Respects
lease grant scope — a lease granting only `mission_list` does not leak
`mission_create`. Skips the `"tools"` capability since that lease is
already reconciled from the v1 path.

Router wires the shared `Arc<CapabilityRegistry>` to both the adapter
and `ThreadManager` at setup.

Three new regression tests:
- `available_actions_surfaces_leased_mission_capability`
- `available_actions_respects_partial_lease_grant`
- `available_actions_omits_capability_without_lease`

Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features` (0 warnings), `cargo test -p ironclaw_engine` (409
passed), `cargo test -p ironclaw --lib` (5082 passed).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(engine-v2): close review gaps — archive round-trip, v1/engine merge, defensive filters

Addresses gaps raised in PR #2660 review:

- **Consolidate `use ironclaw_engine::{...}`** into a single grouped
  import in `effect_adapter.rs` (was split across two statements).

- **Apply `is_v1_only_tool` / `is_v1_auth_tool` filters** to the engine
  capability path in `available_actions`. Defensive guardrail: a future
  engine capability that registers an action under a v1-denylisted name
  (`create_job`, `tool_auth`, ...) must not bypass the v2-isolation
  filters by virtue of coming through a different capability registry.

- **`ThreadArchiveSummary` serialization round-trip tests** in
  `store_adapter.rs`:
    - `archive_summary_preserves_total_cost_usd_through_round_trip` —
      pins the regression the PR fixed (cost silently zeroed on
      rehydration).
    - `archive_summary_handles_legacy_json_without_total_cost_usd_field`
      — pins `#[serde(default)]` back-compat for archive files written
      before this PR.

- **`available_actions` combined advertising tests** in
  `effect_adapter.rs`:
    - `available_actions_merges_v1_tools_with_engine_capabilities` —
      v1 tool + mission capability both surface on one call.
    - `available_actions_filters_v1_denylisted_names_from_engine_capabilities`
      — pins the new defensive filter.

- **`cost_usd_from` subscription-billed-provider test** in
  `llm_adapter.rs`:
    - `complete_with_subscription_billed_provider_yields_zero_cost` —
      zero `cost_per_token` round-trips to exactly `0.0`, no NaN/Inf.

Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features` (0 warnings), `cargo test -p ironclaw_engine` (409
passed), `cargo test -p ironclaw --lib` (5087 passed, +5 new).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(engine-v2): price cache tokens correctly in LlmBridgeAdapter

Addresses PR #2660 review (gemini-code-assist + Copilot, L23/L115/L189):
`cost_usd_from` only priced `input_tokens + output_tokens`, ignoring
`cache_read_input_tokens` and `cache_creation_input_tokens`. For
providers with prompt caching (Anthropic, OpenAI), this undercounted
input cost and silently neutered the `max_budget_usd` gate.

Extend the helper to mirror the canonical formula in
`src/agent/cost_guard.rs::CostGuard::record_llm_call`:

  uncached_input = input_tokens - (cache_read + cache_write)
  cache_read_cost  = input_rate * cache_read  / cache_read_discount()
  cache_write_cost = input_rate * cache_write * cache_write_multiplier()
  cost = input_rate * uncached_input
       + cache_read_cost
       + cache_write_cost
       + output_rate * output_tokens

All `LlmProvider` implementations already supply `cache_read_discount()`
(default 1, Anthropic 10, OpenAI 2) and `cache_write_multiplier()`
(default 1, Anthropic 1.25 for 5m / 2.0 for 1h) through the decorator
chain, so no trait surgery is required.

Regression test: `complete_prices_cache_tokens_with_discount_and_multiplier`
uses Anthropic Sonnet 5m-TTL rates, exercises a 10k-input / 2k-read /
1k-write / 500-output response, and pins the correct total ($0.03285)
against the old naive $0.0375 that would have undercounted ~14%.

Verified: `cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features` (0 warnings), `cargo test -p ironclaw_engine` (435
passed), `cargo test -p ironclaw --lib` (5136 passed).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Illia Polosukhin
2026-04-19 21:25:12 +09:00
committed by GitHub
parent 4104e87b65
commit 43d6fc16b0
6 changed files with 984 additions and 184 deletions

View File

@@ -1,176 +0,0 @@
//! Context compaction and token counting.
//!
//! When message history approaches the model's context limit, compaction
//! asks the LLM to summarize progress and resets the history. This follows
//! the official RLM pattern (compaction at 85% of context limit).
use std::sync::Arc;
use tracing::debug;
use crate::traits::llm::{LlmBackend, LlmCallConfig};
use crate::types::error::EngineError;
use crate::types::message::{MessageRole, ThreadMessage};
use crate::types::step::{LlmResponse, TokenUsage};
/// Characters per token estimate when no tokenizer is available.
/// Conservative estimate (official RLM uses 4).
const CHARS_PER_TOKEN: usize = 4;
/// Estimate token count for a list of messages.
///
/// Uses character length / `CHARS_PER_TOKEN` as a rough estimate.
/// The official RLM uses tiktoken when available; we use this fallback
/// since we don't depend on a Python tokenizer.
pub fn estimate_tokens(messages: &[ThreadMessage]) -> usize {
let total_chars: usize = messages
.iter()
.map(|m| {
m.content.len() + m.action_name.as_ref().map_or(0, |n| n.len()) + 4 // overhead per message (role token, delimiters)
})
.sum();
total_chars.div_ceil(CHARS_PER_TOKEN)
}
/// Check if compaction should be triggered.
///
/// Returns `true` when estimated token count exceeds `threshold_pct` of
/// the model's context limit.
pub fn should_compact(
messages: &[ThreadMessage],
model_context_limit: usize,
threshold_pct: f64,
) -> bool {
let tokens = estimate_tokens(messages);
let threshold = (model_context_limit as f64 * threshold_pct) as usize;
tokens >= threshold
}
/// The compaction prompt sent to the LLM.
const COMPACTION_PROMPT: &str = "\
Summarize your progress so far in a concise but complete way. Include:
1. What you have accomplished
2. Key intermediate results and variable values
3. What still needs to be done
4. Any errors encountered and how they were handled
Preserve all information needed to continue the task. Be specific about data values.";
/// Compact the message history by asking the LLM to summarize.
///
/// Returns the new (shorter) message list and the token usage from the
/// summarization call. The original messages are replaced with:
/// `[system_prompt, summary, continuation_note]`
///
/// The full original messages are returned separately so the caller can
/// store them (e.g., in a `history` variable or event log).
pub async fn compact_messages(
messages: &[ThreadMessage],
llm: &Arc<dyn LlmBackend>,
compaction_count: u32,
) -> Result<CompactionResult, EngineError> {
// Build a summarization request from existing messages + prompt
let mut summarize_messages = messages.to_vec();
summarize_messages.push(ThreadMessage::user(COMPACTION_PROMPT.to_string()));
let config = LlmCallConfig {
force_text: true,
..LlmCallConfig::default()
};
let output = llm.complete(&summarize_messages, &[], &config).await?;
let summary_text = match output.response {
LlmResponse::Text(t) => t,
LlmResponse::ActionCalls { content, .. } | LlmResponse::Code { content, .. } => {
content.unwrap_or_else(|| "[compaction produced no summary]".into())
}
};
// Preserve the system prompt (first message if it's a system message)
let system_msg = messages
.iter()
.find(|m| m.role == MessageRole::System)
.cloned();
// Build compacted history
let mut compacted = Vec::new();
if let Some(sys) = system_msg {
compacted.push(sys);
}
compacted.push(ThreadMessage::assistant(summary_text.clone()));
compacted.push(ThreadMessage::user(format!(
"Your conversation has been compacted {n} time(s). \
The summary above captures your progress. Continue working on the task.",
n = compaction_count + 1,
)));
let tokens_before = estimate_tokens(messages);
let tokens_after = estimate_tokens(&compacted);
debug!(
tokens_before,
tokens_after,
compaction_count = compaction_count + 1,
"context compacted"
);
Ok(CompactionResult {
compacted_messages: compacted,
summary: summary_text,
tokens_used: output.usage,
tokens_before,
tokens_after,
})
}
/// Result of a compaction operation.
pub struct CompactionResult {
/// The new (shorter) message list.
pub compacted_messages: Vec<ThreadMessage>,
/// The summary text produced by the LLM.
pub summary: String,
/// Tokens used by the summarization LLM call.
pub tokens_used: TokenUsage,
/// Estimated token count before compaction.
pub tokens_before: usize,
/// Estimated token count after compaction.
pub tokens_after: usize,
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn estimate_tokens_empty() {
assert_eq!(estimate_tokens(&[]), 0);
}
#[test]
fn estimate_tokens_basic() {
let msgs = vec![
ThreadMessage::system("Hello world"), // 11 chars + 4 overhead = 15 / 4 = 3.75
ThreadMessage::user("Hi"), // 2 chars + 4 = 6 / 4 = 1.5
];
let tokens = estimate_tokens(&msgs);
// (11+4 + 2+4) / 4 = 21/4 = 5.25 → 6 (ceiling)
assert!(tokens > 0);
assert!(tokens < 100);
}
#[test]
fn should_compact_below_threshold() {
let msgs = vec![ThreadMessage::user("short message")];
assert!(!should_compact(&msgs, 128_000, 0.85));
}
#[test]
fn should_compact_above_threshold() {
// Create a message large enough to trigger compaction at low limit
let big = "x".repeat(1000);
let msgs = vec![ThreadMessage::user(big)];
// 1000 chars / 4 = 250 tokens. Context limit 200, threshold 85% = 170
assert!(should_compact(&msgs, 200, 0.85));
}
}

View File

@@ -5,7 +5,6 @@
//! - [`context`] — context building for LLM calls
//! - [`intent`] — tool intent nudge detection
pub mod compaction;
pub mod context;
pub mod loop_engine;
pub mod orchestrator;