mirror of
https://github.com/nearai/ironclaw.git
synced 2026-09-03 08:06:01 +08:00
* v2 architecture phase 1 * feat(engine): Phase 2 — execution loop, capability system, thread runtime Add the core execution engine to ironclaw_engine crate: - CapabilityRegistry: register/get/list capabilities and actions - LeaseManager: async lease lifecycle (grant, check, consume, revoke, expire) - PolicyEngine: deterministic effect-level allow/deny/approve - ThreadTree: parent-child relationship tracking - ThreadSignal/ThreadOutcome: inter-thread messaging via mpsc - ThreadManager: spawn threads as tokio tasks, stop, inject messages, join - ExecutionLoop: core loop replacing run_agentic_loop() with signals, context building, LLM calls, action execution, and event recording - Structured executor (Tier 0): lease lookup → policy check → effect execution - Tool intent nudge detection - MemoryStore + RetrievalEngine stubs for Phase 4 - Full 8-phase architecture plan in docs/plans/ - CLAUDE.md spec for the engine crate 74 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 3 — Monty Python executor with RLM pattern Add CodeAct execution (Tier 1) using the Monty embedded Python interpreter, following the Recursive Language Model (RLM) pattern from arXiv:2512.24601. Key additions: - executor/scripting.rs: Monty integration with FunctionCall-based tool dispatch, catch_unwind panic safety, resource limits (30s, 64MB, 1M allocs) - LlmResponse::Code variant + ExecutionTier::Scripting - Context-as-variables (RLM 3.4): thread messages, goal, step_number, previous_results injected as Python variables — LLM context stays lean while code accesses data selectively - llm_query(prompt, context) (RLM 3.5): recursive subagent calls from within Python code — results stored as variables, not injected into parent's attention window (symbolic composition) - Compact output metadata between code steps instead of full stdout - MontyObject ↔ serde_json::Value bidirectional conversion - Updated architecture plan with RLM design principles 74 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): RLM best-practices enhancements from cross-reference analysis Cross-referenced our implementation against the official RLM (alexzhang13/rlm), fast-rlm (avbiswas/fast-rlm), and Prime Intellect's verifiers implementation. Key enhancements: - FINAL(answer) / FINAL_VAR(name): explicit termination pattern matching all three reference implementations. Code can signal completion at any point, not just via return value. - llm_query_batched(prompts): parallel recursive sub-calls via tokio::spawn, matching fast-rlm's asyncio.gather pattern and Prime Intellect's llm_batch. - Output truncation increased to 8000 chars (from 120), matching Prime Intellect's 8192 default. Shows [TRUNCATED: last N chars] or [FULL OUTPUT]. - Step 0 orientation preamble: auto-injects context metadata (message count, total chars, goal, last user message preview) before first code step, matching fast-rlm's auto-print pattern. - Error-to-LLM flow: Python parse errors, runtime errors, NameErrors, OS errors, and async errors now flow back as stdout content instead of terminating the step, enabling LLM self-correction on next iteration. Only VM panics (catch_unwind) terminate as EngineError. 74 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): update architecture plan with RLM cross-reference learnings Comprehensive update after cross-referencing against official RLM (alexzhang13/rlm), fast-rlm (avbiswas/fast-rlm), Prime Intellect (verifiers/RLMEnv), rlm-rs (zircote/rlm-rs), and Google ADK RLM. Changes: - Mark Phases 1-3 as DONE with commit refs and test counts - Add "Key Influences" section documenting all reference implementations - Phase 3: full table of implemented RLM features with sources - Phase 3: "Remaining gaps" table with which phase addresses each - Phase 4: expanded with compaction (85% context), rlm_query() (full recursive sub-agent), dual model routing, budget controls (USD, timeout, tokens, consecutive errors), lazy loading, pass-by-reference - Add "RLM Execution Model" cross-cutting section - Add "Implementation Progress" tracking table - Remove stale "TO IMPLEMENT" markers (all Phase 3 work is done) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 4 — budget controls, compaction, reflection pipeline Budget enforcement in ExecutionLoop: - max_tokens_total: cumulative token limit, checked before each iteration - max_duration: wall-clock timeout for entire thread - max_consecutive_errors: consecutive error steps threshold (resets on success, matching official RLM behavior) - All produce ThreadOutcome::Failed with descriptive messages Context compaction (from RLM paper, 85% threshold): - estimate_tokens(): char-based estimation (chars/4, matching RLM) - should_compact(): triggers when tokens >= threshold_pct * context_limit - compact_messages(): asks LLM to summarize progress, replaces history with [system, summary, continuation_note], preserves intermediate results - Configurable via ThreadConfig: model_context_limit, compaction_threshold Dual model routing: - LlmCallConfig gains depth field (0=root, 1+=sub-call) - Implementations can route to cheaper models for sub-calls - ExecutionLoop passes thread depth to every LLM call Reflection pipeline (reflection/pipeline.rs): - reflect(thread, llm): analyzes completed thread via LLM - Produces Summary doc (always), Lesson doc (if errors), Issue doc (if failed) - Builds transcript from thread messages + error events - Returns ReflectionResult with docs + token usage ThreadConfig extended with: max_tokens_total, max_consecutive_errors, model_context_limit, enable_compaction, compaction_threshold, depth, max_depth. 78 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 5 — conversation surface separated from execution Conversation is now a UI layer, not an execution boundary. Multiple threads can run concurrently within one conversation; threads can outlive their originating conversation. New types (types/conversation.rs): - ConversationSurface: channel + user + entries + active_threads - ConversationEntry: sender (User/Agent/System) + content + origin_thread_id - ConversationId, EntryId (UUID newtypes) - EntrySender enum (User, Agent{thread_id}, System) ConversationManager (runtime/conversation.rs): - get_or_create_conversation(channel, user) — indexed by (channel, user) - handle_user_message() — injects into active foreground thread or spawns new - record_thread_outcome() — adds agent/system entries, untracks completed threads - get_conversation(), list_conversations() This enables the key architectural insight: a user can ask "what's the weather?" while a deployment thread is still running. Both produce entries in the same conversation. 85 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): simplify execution tiers — Monty-only for CodeAct/RLM Restructure phases 6-8 to clarify execution model: - Monty is the sole Python executor for CodeAct/RLM. No WASM or Docker Python runtimes for LLM-generated code. - WASM sandbox is for third-party tool isolation (existing infra, Phase 8) - Docker containers are for thread-level isolation of high-risk work (Phase 8) - Two-phase commit moves to Phase 6 (integration) at the adapter boundary Phase renumbering: - Old Phase 6 (Tier 2-3) → removed as separate phase - Old Phase 7 (integration) → Phase 6 - Old Phase 8 (cleanup) → Phase 7 - New Phase 8: WASM tools + Docker thread isolation (infra integration) Updated progress table: Phases 1-5 marked DONE with test counts and commits. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 6 — bridge adapters for main crate integration Strategy C parallel deployment: when ENGINE_V2=true env var is set, user messages route through the engine instead of the existing agentic loop. All existing behavior is unchanged when the flag is off. Bridge module (src/bridge/): - LlmBridgeAdapter: wraps LlmProvider as engine LlmBackend, converts ThreadMessage↔ChatMessage, ActionDef↔ToolDefinition, depth-based model routing (primary vs cheap_llm) - EffectBridgeAdapter: wraps ToolRegistry+SafetyLayer as EffectExecutor, routes tool calls through existing execute_tool_with_safety pipeline - InMemoryStore: HashMap-backed Store impl (no DB tables needed yet) - EngineRouter: is_engine_v2_enabled() + handle_with_engine() that builds engine from Agent deps and processes messages end-to-end Integration touchpoint (4 lines in agent_loop.rs): After hook processing, before session resolution, check ENGINE_V2 flag and route UserInput through the engine path. Accessor visibility widened: llm(), cheap_llm(), safety(), tools() changed from pub(super) to pub(crate) for bridge access. 85 engine tests + main crate clippy clean. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): add user message and system prompt to thread before execution The ExecutionLoop was sending empty messages to the LLM because the thread was spawned with the user's input as the goal but no messages. Fixes: - ThreadManager.spawn_thread() now adds the goal as an initial user message before starting the execution loop - ExecutionLoop.run() injects a default system prompt if none exists Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): match existing LLM request format to prevent 400 errors The LLM bridge was missing several defaults that the existing Reasoning.respond_with_tools() sets: - tool_choice: "auto" when tools are present (required by some providers) - max_tokens: 4096 (default) - temperature: 0.7 (default) - When no tools (force_text): use plain complete() instead of complete_with_tools() with empty tools array — matches existing no-tools fallback path Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): persist conversation context across messages The engine was creating a fresh ThreadManager and InMemoryStore per message, losing all context between turns. A follow-up question like "what are the latest 10 issues?" had no memory of the prior "how many issues" response. Fixes: - EngineState (ThreadManager, ConversationManager, InMemoryStore) now persists across messages via OnceLock, initialized on first use - ConversationManager builds message history from prior conversation entries (user messages + agent responses) and passes it to new threads - ThreadManager.spawn_thread_with_history() accepts initial_messages that are prepended before the current user message - System notifications (thread started/completed) are filtered out of the history (not useful as LLM context) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): enable CodeAct/RLM mode with code block detection The engine now operates in CodeAct/RLM mode: System prompt (executor/prompt.rs): - Instructs LLM to write Python in ```repl fenced blocks - Documents available tools as callable Python functions - Documents llm_query(), llm_query_batched(), FINAL() - Documents context variables (context, goal, step_number, previous_results) - Strategy guidance: examine context, break into steps, use tools, call FINAL() Code block detection (bridge/llm_adapter.rs): - extract_code_block() scans LLM text responses for ```repl or ```python blocks - When detected, returns LlmResponse::Code instead of LlmResponse::Text - The ExecutionLoop routes Code responses through Monty for execution No structured tool definitions sent to LLM: - Tools are described in the system prompt as Python functions - The LLM call sends empty actions array, forcing text-mode responses - This ensures the LLM writes code blocks (CodeAct) instead of structured tool calls (which would bypass the REPL) 85 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(engine): add 8 CodeAct/RLM E2E tests with mock LLM Comprehensive test coverage for the Monty Python execution path: - codeact_simple_final: Python code calls FINAL('answer') → thread completes - codeact_tool_call_then_final: code calls test_tool() → FunctionCall suspends VM → MockEffects returns result → code resumes → FINAL() - codeact_pure_python_computation: sum([1,2,3,4,5]) → FINAL('Sum is 15') with no tool calls — pure Python in Monty - codeact_multi_step: first step prints output (no FINAL), second step sees output metadata and calls FINAL — tests iterative REPL flow - codeact_error_recovery: first step has NameError → error flows to LLM as stdout → second step recovers with FINAL — tests error transparency - codeact_context_variables_available: code accesses `goal` and `context` variables injected by the RLM context builder - codeact_multiple_tool_calls_in_loop: for loop calls test_tool() 3 times → 3 FunctionCall suspensions → all results collected → FINAL - codeact_llm_query_recursive: code calls llm_query('prompt') → VM suspends → MockLlm provides sub-agent response → result returned as Python string variable 93 tests passing (85 prior + 8 new), zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): detect code blocks in plain completion path + multi-block support Two bugs fixed: 1. The no-tools completion path (used by CodeAct since we send empty actions) returned LlmResponse::Text without checking for code blocks. Code blocks were rendered as markdown text instead of being executed. 2. extract_code_block now: - Handles bare ``` fences (skips non-Python languages) - Collects ALL code blocks in the response and concatenates them (models often split code across multiple blocks with explanation) - Tries markers in order: ```repl, ```python, ```py, then bare ``` Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(bridge): add 11 regression tests for code block extraction Covers the exact failure modes discovered during live testing: - extract_repl_block: standard ```repl fenced block - extract_python_block: ```python marker - extract_py_block: ```py shorthand - extract_bare_backtick_block: bare ``` with Python content - skip_non_python_language: ```json should NOT be extracted - no_code_blocks_returns_none: plain text, no fences - multiple_code_blocks_concatenated: two ```repl blocks with explanation between them → concatenated with \n\n - mixed_thinking_and_code: model outputs explanation + two ```python blocks (the Hyperliquid case) → both extracted - repl_preferred_over_bare: ```repl takes priority over bare ``` - empty_code_block_skipped: empty fenced block returns None - unclosed_block_returns_none: no closing ``` returns None Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): detect FINAL() in text responses + regression tests Models sometimes write FINAL() outside code blocks — as plain text after an explanation. The Hyperliquid case: model outputs a long analysis then FINAL("""...""") at the end, not inside ```repl fences. Fixes: - extract_final_from_text(): regex-based FINAL detection in text responses, matching the official RLM's find_final_answer() fallback - Handles: double-quoted, single-quoted, triple-quoted, unquoted, nested parens - Checked in LlmResponse::Text handler BEFORE tool intent nudge (FINAL takes priority) 9 new tests: - codeact_final_in_text_response: FINAL("answer") in plain text - codeact_final_triple_quoted_in_text: FINAL("""multi\nline""") in text - final_double_quoted, final_single_quoted, final_triple_quoted, final_unquoted, final_with_nested_parens, final_after_long_text, no_final_returns_none 102 tests passing (93 + 9 new), zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add crate extraction & cleanup roadmap Documents architectural recommendations from the engine v2 design process for future reference: - Root directory consolidation (channels-src + tools-src → extensions/) - Crate extraction tiers: zero-coupling (estimation, observability, tunnel), trivial-coupling (document_extraction, pairing, hooks), medium-coupling (secrets, MCP, db, workspace, llm, skills), heavy-coupling (web gateway, agent, extensions) - src/ module reorganization into logical groups (core, persistence, infra, media, support) - main.rs/app.rs slimming targets (100/500 lines after migration) - WASM module candidates (document_extraction) and non-candidates (REPL, web gateway → separate crates instead) - Priority ordering for extraction work - Tracks completed items (ironclaw_safety, ironclaw_engine, transcription move) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): live progress status updates via event broadcast Engine v2 now shows live progress in the CLI (and any channel): - "Thinking..." when a step starts - Tool name + success/error when actions execute - "Processing results..." when a step completes Implementation: - ThreadManager holds a broadcast::Sender<ThreadEvent> (capacity 256) - ExecutionLoop.emit_event() writes to thread.events AND broadcasts - ThreadManager.subscribe_events() returns a receiver - Router uses tokio::select! to listen for events while waiting for thread completion, forwarding them as StatusUpdate to the channel This replaces the polling approach with zero-latency event streaming. Agent.channels visibility widened to pub(crate) for bridge access. 102 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): include tool results in code step output for LLM context The LLM was ignoring tool results and answering from training data because the compact output metadata didn't include what tools returned. Tool results lived only as ActionResult messages (role: Tool) which some providers flatten or the model ignores. Now the code step output includes: - stdout from Python print() statements - [tool_name result] with the actual output (truncated to 4K per tool) - [tool_name error] for failed tools - [return] for the code's return value - Total output truncated to 8K chars to prevent context bloat This ensures the model sees web_search results, API responses, etc. in the next iteration and can reason about them instead of hallucinating. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): add debug/trace logging for CodeAct execution Three verbosity levels for debugging the engine: RUST_LOG=ironclaw_engine=debug: - LLM call: message count, iteration, force_text - LLM response: type (text/code/action_calls), token usage - Code execution: code length, action count, had_error, final_answer - Text response: length, FINAL() detection RUST_LOG=ironclaw_engine=trace: - Full message list sent to LLM (role, length, first 200 chars each) - Full code block being executed - stdout preview (first 500 chars) - Per-tool results (name, success, first 300 chars of output) - Text response preview (first 500 chars) Usage: ENGINE_V2=true RUST_LOG=ironclaw_engine=debug cargo run ENGINE_V2=true RUST_LOG=ironclaw_engine=trace cargo run Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): execution trace recording + retrospective analysis Enable with ENGINE_V2_TRACE=1 to get full execution traces and automatic issue detection after each thread completes. Trace recording (executor/trace.rs): - build_trace(): captures full thread state — messages (with full content), events, step count, token usage, detected issues - write_trace(): writes JSON to engine_trace_{timestamp}.json - log_trace_summary(): logs summary + issues at info/warn level Retrospective analyzer detects 8 issue categories: - thread_failure: thread ended in Failed state - no_response: no assistant message generated - tool_error: specific tool failures with error details - code_error: Python errors (NameError, SyntaxError, etc.) in output - missing_tool_output: tool results exist but not in system messages - excessive_steps: >10 steps (may be stuck in loop) - no_tools_used: single-step answer without tools (hallucination risk) - mixed_mode: text responses without code blocks (prompt not followed) Thread state now saved to store after execution completes (for trace access after join_thread). Usage: ENGINE_V2=true ENGINE_V2_TRACE=1 cargo run # After each message: trace JSON + issue log in terminal Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): wire reflection pipeline + trace analysis into thread lifecycle After every thread completes, ThreadManager now automatically runs: 1. Retrospective trace analysis (non-LLM, always): - Detects 8 issue categories (tool errors, code errors, missing outputs, excessive steps, hallucination risk, etc.) - Logs issues at warn level when found 2. Trace file recording (when ENGINE_V2_TRACE=1): - Writes full JSON trace to engine_trace_{timestamp}.json 3. LLM reflection (when enable_reflection=true): - Calls reflection pipeline to produce Summary, Lesson, Issue docs - Saves docs to store for future context retrieval - Enabled by default in the bridge router All three run inside the spawned tokio task after exec.run() completes, before saving the final thread state. No external wiring needed. Removed duplicate trace recording from the router — it's now handled by ThreadManager automatically. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): convert tool name hyphens to underscores for Python compatibility Root cause from trace analysis: the LLM writes `web_search()` (valid Python identifier) but the tool registry has `web-search` (with hyphen). The EffectBridgeAdapter couldn't find the tool → "Tool not found" error → model fabricated fake data instead. Fixes: - available_actions(): converts tool names from hyphens to underscores (web-search → web_search) so the system prompt lists valid Python names - execute_action(): tries the original name first, then falls back to hyphenated form (web_search → web-search) for tool registry lookup - Same conversion in router's capability registry builder Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): parse JSON tool output to prevent double-serialization From trace analysis: web_search returned a JSON string, which was wrapped as serde_json::json!(string) creating a Value::String containing JSON. When Monty got this as MontyObject::String, the Python code couldn't index it with result['title'] → TypeError. Fix: try parsing the tool output string as JSON first. If valid, use the parsed Value (becomes a Python dict/list). If not valid JSON, keep as string. This means web_search results are directly indexable in Python: results = web_search(query="...") print(results["results"][0]["title"]) # works now Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): persist variables across code steps via `state` dict Monty creates a fresh runtime per code step, so variables are lost between steps. This caused the model to re-paste tool results from system messages, wasting tokens. Fix: maintain a `persisted_state` JSON dict in the ExecutionLoop that accumulates across steps: - Tool results stored by tool name: state["web_search"] = {results...} - Return values stored: state["last_return"], state["step_0_return"] - Injected as a `state` Python variable in each new MontyRun Now the model can do: Step 1: results = web_search(query="...") # tool result saved in state Step 2: data = state["web_search"] # access previous result summary = llm_query("summarize", str(data)) FINAL(summary) System prompt updated to document the `state` variable. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): add state hint on code errors + retrieval engine integration When code fails with NameError/UnboundLocalError (model trying to access variables from a previous step), the error output now includes: [HINT] Variables don't persist between code blocks. Use the `state` dict to access data from previous steps. Available keys: ["web_search", "last_return"] This teaches the model to use `state["web_search"]` instead of `result` after a NameError, reducing wasted steps from 3-4 to 1. Also integrates RetrievalEngine into context building and ThreadManager: - build_step_context() now accepts optional RetrievalEngine to inject relevant memory docs (Lessons, Specs, Playbooks) into LLM context - RetrievalEngine uses keyword matching with doc-type priority scoring - Memory docs from reflection (Phase 4) now feed back into future threads Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: remove trace files and add to .gitignore Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): replace web_fetch example with web_search in CodeAct prompt The system prompt example used web_fetch(url="...") which doesn't exist as a tool. The model learned from the example and tried web_fetch, getting "Tool not found". Changed to web_search(query="...") which is an actual registered tool. Found via trace analysis — reflection pipeline correctly identified this as a "Tool Name Correction" spec doc. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(engine): extract prompt templates to markdown files Prompt templates moved from inline Rust strings to plain markdown files at crates/ironclaw_engine/prompts/ for easy inspection and iteration: - prompts/codeact_preamble.md — main instructions, special functions, context variables, rules - prompts/codeact_postamble.md — strategy section Loaded at compile time via include_str!(), so no runtime file I/O. Edit the .md files and rebuild to iterate on prompts. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): replace byte-index slicing with char-safe truncation Panic: 'byte index 80 is not a char boundary; it is inside ''' when tool output contained multi-byte UTF-8 characters (smart quotes from web search results). Fixed 4 unsafe byte-index slices: - thread.rs:281: message preview &content[..80] → chars().take(80) - loop_engine.rs:556: tool output &str[..4000] → chars().take(4000) - loop_engine.rs:579: output tail &str[len-8000..] → chars().skip() - scripting.rs:82: stdout tail &str[len-N..] → chars().skip() All now use .chars().take() or .chars().skip() which respect character boundaries. Follows CLAUDE.md rule: "Never use byte-index slicing on user-supplied or external strings." Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): fix false positive missing_tool_output warning in trace analyzer The check was looking for "[" + "result]" in System-role messages only, but tool output metadata is added with patterns like "[shell result]" and may appear in messages with any role. Changed to scan all messages for " result]" or " error]" patterns regardless of role. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): update architecture plan with Phase 6 status and approval flow design Phase 6 updated to reflect what was actually built: - Bridge adapters (LLM, Effect, InMemoryStore, Router) — all done - Integration touchpoint (4 lines in handle_message) — done - Live progress via broadcast events — done - Conversation persistence across messages — done - Trace recording + retrospective analysis — done - 8 bugs found and fixed via trace analysis — documented Phase 6 remaining work documented: - Approval flow: detailed 5-step design (send to channel, pause thread, route response, resume execution, always handling) with v1 reference - Database persistence (InMemoryStore → real DB tables) - Acceptance testing (TestRig + TraceLlm fixtures) - Two-phase commit for high-stakes effects Progress table updated: Phase 6 marked as DONE (partial), 134 tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add self-improving engine design plan Designs a system where the engine debugs and improves itself, based on the pattern observed in the last session: 5 consecutive bug fixes all followed trace → read → identify → edit → test, using tools the engine already has access to. Three levels of self-improvement: - Level 1 (Prompt): edit prompts/*.md to prevent LLM mistakes. Auto-apply. - Level 2 (Config): adjust defaults/mappings. Branch + test + PR. - Level 3 (Code): Rust patches for engine bugs. Branch + test + clippy + PR. Architecture: Self-improvement Mission spawns a Reflection thread that reads traces, reads source, proposes fixes, validates via cargo test, and either auto-applies (Level 1) or creates a PR (Level 2-3). Includes: fix pattern database (seeded from our 8 debugging session fixes), feedback loop diagram, safety model, implementation phases (A through D), and what exists vs what's new. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add engine v2 security model and audit Comprehensive security analysis of engine v2 covering: Threat model: 4 attacker profiles (malicious input, prompt injection via tools, poisoned memory, supply chain). Current state audit: 9 controls working (Monty sandbox, safety layer, policy engine, leases, provenance, events) and 9 gaps identified. Critical finding: ALL tools granted by default — CodeAct code can call shell, write_file, apply_patch without approval. Proposed fix: 3-tier tool classification (auto/approve-once/always-approve). CodeAct-specific threats: tool call amplification, prompt injection via search results, data exfiltration via tool chains, Monty escape. Self-improvement security: poisoned trace attacks, memory poisoning via reflection. Mitigations: edit validation, frequency caps, audit trail, auto-rollback, reflection output scanning. 6-layer security architecture proposed: input validation, capability gating, output sanitization, execution sandboxing, self-improvement controls, observability. Prioritized implementation plan with severity/effort ratings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(security): cross-reference v1 controls — use, don't reinvent Updated security plan with detailed audit of ALL existing v1 security controls and how they map to engine v2 bridge gaps: Key finding: v1 already has solutions for every security gap identified. The bridge just needs to wire them in: - Tool::requires_approval() exists but bridge doesn't call it - safety.wrap_for_llm() exists but tool results enter context unwrapped - RateLimiter exists but bridge doesn't check rate limits - BeforeToolCall hooks exist but bridge doesn't run them - redact_params() exists but bridge doesn't redact sensitive params - Shell risk classification (Low/Medium/High) is inherited but ignored Revised priority: most fixes are small wiring tasks in EffectBridgeAdapter, not new security infrastructure. The bridge is the security boundary. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): add missions, reliability tracker, reflection executor, and provenance-aware policy - Add Mission type and MissionManager for recurring thread scheduling - Add ReliabilityTracker for per-capability success/failure/latency tracking - Add reflection executor that spawns CodeAct threads for post-completion reflection - Extend PolicyEngine with provenance-aware taint checking (LLM-generated data requires approval for financial/external-write effects) - Extend Store trait with mission CRUD methods - Add conversation surface tracking, compaction token fix, context memory injection - Wire new modules through lib.rs re-exports and bridge adapters Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): wire v1 security controls into engine v2 adapter Zero engine crate changes. All security controls enforced at the bridge boundary in EffectBridgeAdapter: 1. Tool approval (v1: Tool::requires_approval): - Checks each tool's approval requirement with actual params - Always → returns EngineError::LeaseDenied (blocks execution) - UnlessAutoApproved → checks auto_approved set, blocks if not approved - Never → proceeds - Per-session auto_approved HashSet (for future "always" handling) 2. Hook interception (v1: BeforeToolCall): - Runs HookEvent::ToolCall before every execution - HookOutcome::Reject → blocks with reason - HookError::Rejected → blocks with reason - Hook errors → fail-open (logged, execution continues) 3. Output sanitization (v1: sanitize_tool_output + wrap_for_llm): - Leak detection: API keys in tool output are redacted - Policy enforcement: content policy rules applied - Length truncation: output capped at 100KB - XML boundary protection: prevents injection via tool output 4. Sensitive param redaction (v1: redact_params): - Tool's sensitive_params() consulted before hooks see parameters - Redacted params sent to hooks, original params used for execution 5. available_actions() now sets requires_approval based on each tool's default approval requirement, so the engine's PolicyEngine can gate tools it hasn't seen before. 6. Actual execution timing measured via Instant::now() (replaces placeholder Duration::from_millis(1)). Accessor visibility: hooks() widened to pub(crate). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): implement tool approval flow for engine v2 Adds a complete approval flow that mirrors v1 behavior, using the existing v1 security controls (Tool::requires_approval, auto-approve sets, StatusUpdate::ApprovalNeeded). ## How it works ### Step 1: Tool blocked at execution When the LLM's code calls a tool (e.g., `shell("ls")`): 1. EffectBridgeAdapter.execute_action() looks up the Tool object 2. Calls tool.requires_approval(¶ms) — returns ApprovalRequirement 3. If Always → EngineError::LeaseDenied (always blocks) 4. If UnlessAutoApproved → checks auto_approved HashSet → if not in set, returns EngineError::LeaseDenied 5. If Never → proceeds to execution ### Step 2: Engine returns NeedApproval The LeaseDenied error propagates through: - CodeAct path: becomes Python RuntimeError, code halts, thread returns NeedApproval with action_name + parameters - Structured path: same via ActionResult.is_error ### Step 3: Router stores pending approval - PendingApproval { action_name, original_content } stored on EngineState - StatusUpdate::ApprovalNeeded sent to channel (shows approval card in CLI/web with tool name, parameters, yes/always/no buttons) - Returns text: "Tool 'shell' requires approval. Reply yes/always/no." ### Step 4: User responds handle_message() intercepts Submission::ApprovalResponse when ENGINE_V2: - 'yes' → auto_approve_tool(name) on EffectBridgeAdapter, re-processes original message (tool now passes the approval check on second run) - 'always' → same + logs for session persistence - 'no' → returns "Denied: tool was not executed." ### Key design choice Instead of pausing/resuming mid-execution (which needs engine changes to freeze/restore the Monty VM state), we auto-approve the tool and re-run the full message. The EffectBridgeAdapter's auto_approved set persists across runs, so the second execution passes immediately. This trades one extra LLM call for zero engine modifications. ## Files changed - src/bridge/router.rs: PendingApproval struct, handle_approval(), NeedApproval → StatusUpdate::ApprovalNeeded conversion - src/bridge/mod.rs: export handle_approval - src/agent/agent_loop.rs: intercept ApprovalResponse for engine v2 - src/bridge/effect_adapter.rs: fmt fixes 151 tests passing, clippy + fmt clean. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): demote trace/reflection logging from info to debug INFO-level log output from background tasks (trace analysis, reflection) corrupts the REPL terminal UI. The trace summary, issue warnings, and reflection doc previews were printing mid-approval-card, breaking the interactive display. Fix: all logging in trace.rs changed from info!/warn! to debug!/warn!. Trace analysis and reflection results now only show when RUST_LOG=ironclaw_engine=debug is set. Also added logging discipline rule to global CLAUDE.md: - info! → user-facing status the REPL intentionally renders - debug! → internal diagnostics (traces, reflection, engine internals) - Background tasks must NEVER use info! — it breaks the TUI Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): demote all router info! logging to debug! "engine v2: initializing" and "engine v2: handling message" were printing at INFO level, corrupting the REPL UI. All router logging now uses debug! — only visible with RUST_LOG=ironclaw=debug. Zero info! calls remain in crates/ironclaw_engine/ or src/bridge/. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(safety): demote leak detector warn-action logs from warn! to debug! The leak detector's Warn-action matches (high_entropy_hex pattern on web search results containing commit SHAs, CSS colors, URL hashes) were logging at warn! level, corrupting the REPL UI with lines like: WARN Potential secret leak detected pattern=high_entropy_hex preview=a96f********cee5 These are informational false positives — real leaks use LeakAction::Redact which silently modifies the content. Warn-action matches only log for debugging purposes and should not appear in production output. Changed to debug! level — visible with RUST_LOG=ironclaw_safety=debug. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): strengthen CodeAct prompt to prevent shallow text answers The model was answering "Suggested 45 improvements" as a brief text summary from training data without actually searching or listing them. The trace showed: no code block, no tool calls, no FINAL(). Prompt changes: - Rule 1: "ALWAYS respond with a ```repl code block. NEVER answer with plain text only." (was: "Always write code... plain text for brief explanations") - Rule 2 (NEW): "NEVER answer from memory or training data alone. Always use tools to get real, current information before answering." - Rule 3: FINAL answer "should be detailed and complete — not just a summary like 'found 45 items'" - Rule 8 (NEW): "Include the actual content in your FINAL() answer, not just a count or summary. Users want to see the details." Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): persist reflection docs to workspace for cross-session learning Replaces InMemoryStore with HybridStore: - Ephemeral data (threads, steps, events, leases) stays in-memory - MemoryDocs (lessons, specs, playbooks from reflection) persist to the workspace at engine/docs/{type}/{id}.json On engine init, load_docs_from_workspace() reads existing docs back into the in-memory cache. This means: - Lessons learned in session 1 are available in session 2 - The RetrievalEngine injects relevant past lessons into new threads - The engine genuinely improves over time as reflection accumulates Workspace paths: engine/docs/lessons/{uuid}.json engine/docs/specs/{uuid}.json engine/docs/playbooks/{uuid}.json engine/docs/summaries/{uuid}.json engine/docs/issues/{uuid}.json No new database tables. Uses existing workspace write/read/list. workspace() accessor widened to pub(crate). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): adapt to execute_tool_with_safety params-by-value change Staging merge changed execute_tool_with_safety to take params by value instead of by reference (perf optimization from PR #926). Updated bridge adapter to clone params before passing. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): add web gateway integration plan to Phase 6 Documents three gaps between engine v2 and the web gateway: 1. No SSE streaming (engine emits ThreadEvent, gateway expects SseEvent) 2. No conversation persistence (engine uses HybridStore, gateway reads v1 DB) 3. No cross-channel visibility (REPL ↔ web messages invisible to each other) Implementation plan: bridge ThreadEvent→AppEvent, write messages to v1 conversation tables after thread completion. Prerequisite: AppEvent extraction PR (in progress separately). Also updated DB persistence status: HybridStore with workspace-backed MemoryDocs is now implemented (partial persistence). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): document routine/job gap and SIGKILL crash scenario Routines are entirely v1 — not hooked up to engine v2. When a user asks "create a routine" as natural language, engine v2 tries to call routine_create via CodeAct, but the tool needs RoutineEngine + Database refs that the bridge's minimal JobContext doesn't provide. This caused a SIGKILL crash during testing. Options documented: block routine tools in v2 (short term), pass refs through context (medium), replace with Mission system (long term). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: extract AppEvent to crates/ironclaw_common SseEvent was defined in src/channels/web/types.rs but imported by 12+ modules across agent, orchestrator, worker, tools, and extensions — it had become the application-wide event protocol, not a web transport concern. Create crates/ironclaw_common as a shared workspace crate and move the enum there as AppEvent. Also move the truncate_preview utility which was similarly leaked from the web gateway into agent modules. - New crate: crates/ironclaw_common (AppEvent, truncate_preview) - Rename SseEvent → AppEvent, from_sse_event → from_app_event - web/types.rs re-exports AppEvent for internal gateway use - web/util.rs re-exports truncate_preview - Wire format unchanged (serde renames are on variants, not the enum) Aligned with the event bus direction on refactor/architectural-hardening where DomainEvent (≡ AppEvent) is wrapped in a SystemEvent envelope. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): integrate with web gateway via AppEvent + v1 conversation DB Three changes to make engine v2 visible in the web gateway: 1. SSE event streaming (AppEvent broadcast): - ThreadEvent → AppEvent conversion via thread_event_to_app_event() - Events broadcast to SseManager during the poll loop - Covers: Thinking, ToolCompleted (success/error), Status, Response - Web gateway receives real-time progress without any gateway changes 2. Conversation persistence to v1 database: - After thread completes, writes user message + agent response to v1 ConversationStore via add_conversation_message() - Uses get_or_create_assistant_conversation() for per-user per-channel - Web gateway reads from DB as usual — chat history appears 3. Final response broadcast: - AppEvent::Response with full text + thread_id sent via SSE - Web gateway renders the response in the chat UI New EngineState fields: sse (Option<Arc<SseManager>>), db (Option<Arc<dyn Database>>). Both populated from Agent.deps. Agent.deps visibility widened to pub(crate). Depends on: ironclaw_common crate with AppEvent type (PR #1615). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): complete Phase 6 — v1-only tool blocking, rate limiting, call limits Three security/stability improvements in EffectBridgeAdapter: 1. V1-only tool blocking: - routine_create, create_job, build_software (and hyphenated variants) return helpful error: "use the slash command instead" - Filtered out of available_actions() so system prompt doesn't list them - Prevents crash from tools needing RoutineEngine/Scheduler refs 2. Per-step tool call limit: - Max 50 tool calls per code block (AtomicU32 counter) - Prevents amplification: `for i in range(10000): shell(...)` - Returns "call limit reached, break into multiple steps" 3. Rate limiting: - Per-user per-tool sliding window via RateLimiter - Checks tool.rate_limit_config() before every execution - Returns "rate limited, try again in Ns" Architecture plan updated: - Gateway integration: DONE - Routines: BLOCKED (gracefully, with slash command fallback) - Rate limiting: DONE - Call limit: DONE - Phase 6 status: DONE (remaining: acceptance tests, two-phase commit) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add Mission system design — goal-oriented autonomous threads Missions replace routines with evolving, knowledge-accumulating autonomous agents. Unlike routines (fixed prompt, stateless), Missions: - Generate prompts from accumulated Project knowledge (lessons, playbooks, issues from prior threads) - Adapt approach when something fails repeatedly - Track progress toward a goal with success criteria - Self-manage: pause when stuck, complete when goal achieved Architecture: MissionManager with cron ticker spawns threads via ThreadManager. Meta-prompt built from mission goal + Project MemoryDocs via RetrievalEngine. Reflection feeds back automatically. 6-step implementation plan: cron trigger, meta-prompt builder, bridge wiring, CodeAct tools, progress tracking, persistence. Includes two worked examples: daily tech news briefing (ongoing) and test coverage improvement (goal-driven, self-completing). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): extend Mission types with webhook/event triggers + evolving strategy Mission types updated to support external activation sources: MissionCadence expanded: - Cron { expression, timezone } — timezone-aware scheduling - OnEvent { event_pattern } — channel message pattern matching - OnSystemEvent { source, event_type } — structured events from tools - Webhook { path, secret } — external HTTP triggers (GitHub, email, etc.) - Manual — explicit triggering only The engine defines trigger TYPES. The bridge implements infrastructure (cron ticker, webhook endpoints, event matchers). GitHub issues, PRs, email, Slack events all use the generic Webhook cadence — no special-casing in the engine. Webhook payload injected as state["trigger_payload"] in the thread's Python context. Mission struct extended: - current_focus: what the next thread should work on (evolving) - approach_history: what we've tried (for adaptation) - max_threads_per_day / threads_today: daily budget - last_trigger_payload: webhook/event data for thread context Plan updated with trigger type table and webhook integration design. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): implement MissionManager execution with meta-prompts The MissionManager now builds evolving meta-prompts and processes thread outcomes for continuous learning: fire_mission() upgraded: - Loads Project MemoryDocs via RetrievalEngine for context - Builds meta-prompt from: goal, current_focus, approach_history, project knowledge docs, trigger payload, thread count - Spawns thread with meta-prompt as user message - Background task waits for completion and processes outcome - Daily thread budget enforcement (max_threads_per_day) Meta-prompt structure: # Mission: {name} Goal: {goal} ## Current Focus (evolves between threads) ## Previous Approaches (what we've tried) ## Knowledge from Prior Threads (lessons, playbooks, issues) ## Trigger Payload (webhook/event data if applicable) ## Instructions (accomplish step, report next focus, check goal) Outcome processing: - Extracts "next focus:" from FINAL() response → updates current_focus - Detects "goal achieved: yes" → completes mission - Records accomplishment in approach_history - Failed threads recorded as "FAILED: {error}" Cron ticker: - start_cron_ticker() spawns tokio task, ticks every 60s - Checks active Cron missions, fires those past next_fire_at 151 tests passing. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): wire MissionManager into engine v2 for CodeAct access Missions are now callable from CodeAct Python code: ```python # Create a daily briefing mission result = mission_create( name="Tech News", goal="Daily AI/crypto/software news briefing", cadence="0 9 * * *" ) # List all missions missions = mission_list() # Manually fire a mission mission_fire(id="...") # Pause/resume mission_pause(id="...") mission_resume(id="...") ``` Implementation: - MissionManager created on engine init, cron ticker started - EffectBridgeAdapter intercepts mission_* function calls before tool lookup and routes to MissionManager - parse_cadence() handles: "manual", cron expressions, "event:pattern", "webhook:path" - Mission functions documented in CodeAct system prompt - MissionManager set on adapter via set_mission_manager() after init (avoids circular dependency) System prompt updated with mission_create, mission_list, mission_fire, mission_pause, mission_resume documentation. 151 tests passing. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): map routine_* calls to mission operations in v2 When the model calls routine_create, routine_list, routine_fire, routine_pause, routine_resume, or routine_delete, the bridge now routes them to the MissionManager instead of blocking with an error. Mapping: routine_create → mission_create (with cadence parsing) routine_list → mission_list routine_fire → mission_fire routine_pause → mission_pause routine_resume → mission_resume routine_update → mission_pause/resume (based on params) routine_delete → mission_complete (marks as done) Routine tools removed from v1-only blocklist and restored in available_actions(). The model can use either "routine" or "mission" vocabulary — both work. Still blocked: create_job, cancel_job, build_software (need v1 Scheduler/ContainerJobManager refs). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(engine): add E2E mission flow tests — 7 new tests Comprehensive mission lifecycle tests: - fire_mission_builds_meta_prompt_with_goal: verifies thread spawned with project context and recorded in history - outcome_processing_extracts_next_focus: "Next focus: X" in FINAL() response → mission.current_focus updated - outcome_processing_detects_goal_achieved: "Goal achieved: yes" → mission status transitions to Completed - mission_evolves_via_direct_outcome_processing: 3-step evolution: step 1 sets focus to "db module", step 2 evolves to "tools module", step 3 detects goal achieved → mission completes. Tests the full learning loop without background task timing dependencies. - fire_with_trigger_payload: webhook payload stored on mission and threads_today counter incremented - daily_budget_enforced: max_threads_per_day=1 → first fire succeeds, second returns None 157 tests passing (151 prior + 6 new mission E2E). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): self-improving engine via Mission system Wire the self-improvement loop as a Mission with OnSystemEvent cadence, inspired by karpathy/autoresearch's program.md approach. The mission fires when threads complete with issues, receives trace data as trigger payload, and uses tools directly to diagnose and fix problems. Key changes: Engine self-improvement (Phase A+B from design doc): - Add fire_on_system_event() to MissionManager for OnSystemEvent cadence - Add start_event_listener() that subscribes to thread events and fires matching missions when non-Mission threads complete with trace issues - Add ensure_self_improvement_mission() with autoresearch-style goal prompt (concrete loop steps, not vague instructions) - Add process_self_improvement_output() for structured JSON fallback - Seed fix pattern database with 8 known patterns from debugging - Runtime prompt overlay via MemoryDoc (build_codeact_system_prompt now async + Store-aware, appends learned rules from prompt_overlay docs) - Pass Store to ExecutionLoop for overlay loading Bridge review fixes (P1/P2): - Scope engine v2 SSE events to requesting user (broadcast_for_user) - Per-user pending approvals via HashMap instead of global Option - Reset tool-call limit counter before each thread execution - Only persist auto-approval when user chose "always", not one-off "yes" - Remove dead store/mission_manager fields from EngineState Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Add checkpoint-based engine thread recovery * feat(engine): add Python orchestrator module and host functions Add the orchestrator infrastructure for replacing the Rust execution loop with versioned Python code. This commit adds the module and host functions without switching over — the existing Rust loop is unchanged. New files: - orchestrator/default.py: v0 Python orchestrator (run_loop + helpers) - executor/orchestrator.rs: host function dispatch, orchestrator loading from Store with version selection, OrchestratorResult parsing Host functions exposed to orchestrator Python via Monty suspension: __llm_complete__, __execute_code_step__ (nested Monty VM), __execute_action__, __check_signals__, __emit_event__, __add_message__, __save_checkpoint__, __transition_to__, __retrieve_docs__, __check_budget__, __get_actions__ Also makes json_to_monty, monty_to_json, monty_to_string pub(crate) in scripting.rs for cross-module use. Design doc: docs/plans/2026-03-25-python-orchestrator.md Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): switch ExecutionLoop::run() to Python orchestrator Replace the 900-line Rust execution loop with a ~80-line bootstrap that loads and runs the versioned Python orchestrator via Monty VM. The orchestrator Python code (orchestrator/default.py) is the v0 compiled-in version. Runtime versions can override it via MemoryDoc storage (orchestrator:main with tag orchestrator_code). Key fixes during switchover: - Use ExtFunctionResult::NotFound for unknown functions so Monty falls through to Python-defined functions (extract_final, etc.) - Move helper function definitions above run_loop for Monty scoping - Use FINAL result value (not VM return value) in Complete handler - Rename 'final' variable to 'final_answer' to avoid Python keyword Status: 171/177 tests pass. 6 remaining failures are step_count and token tracking bookkeeping — the orchestrator manages these internally but doesn't yet update the thread's counters via host functions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): all 177 tests pass with Python orchestrator - Increment step_count and track tokens in __emit_event__("step_completed") so thread bookkeeping matches the old Rust loop behavior - Remove double-counting of tokens in bootstrap (orchestrator handles it) - Match nudge text to existing TOOL_INTENT_NUDGE constant - Fix FINAL result propagation (use stored final_result, not VM return) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): orchestrator versioning, auto-rollback, and tests Add version lifecycle for the Python orchestrator: - Failure tracking via MemoryDoc (orchestrator:failures) - Auto-rollback: after 3 consecutive failures, skip the latest version and fall back to previous (or compiled-in v0) - Success resets the failure counter - OrchestratorRollback event for observability Update self-improvement Mission goal with Level 1.5 instructions for orchestrator patches — the agent can now modify the execution loop itself via memory_write with versioned orchestrator docs. 12 new tests: version selection (highest wins), rollback after failures, rollback to default, failure counting/resetting, outcome parsing for all 5 ThreadOutcome variants. 189 tests pass, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add engine v2 architecture, self-improvement, and dev history Three new docs for contributors: - engine-v2-architecture.md: Two-layer architecture (Rust kernel + Python orchestrator), five primitives, execution model with nested Monty VMs, bridge layer, memory/reflection, missions, capabilities - self-improvement.md: Three improvement levels (prompt/orchestrator/ config/code), autoresearch-inspired Mission loop, versioned orchestrator with auto-rollback, fix pattern database, safety model - development-history.md: Summary of 6 Claude Code sessions that built the system, key design decisions and debugging moments, architecture evolution from 900-line Rust loop to Python orchestrator Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): complete v2 side-by-side integration with gateway API Wire engine v2 into the full submission pipeline and expose threads, projects, and missions through the web gateway REST API. Bridge routing — route ExecApproval, Interrupt, NewThread, and Clear submissions to engine v2 when ENGINE_V2=true. Previously only UserInput and ApprovalResponse were handled; all other control commands fell through to disconnected v1 sessions. Bridge query layer — add 11 read-only query functions and 6 DTO types so gateway handlers can inspect engine state (threads, steps, events, projects, missions) without direct access to the EngineState singleton. Gateway endpoints — new /api/engine/* routes: GET /threads, /threads/{id}, /threads/{id}/steps, /threads/{id}/events GET /projects, /projects/{id} GET /missions, /missions/{id} POST /missions/{id}/fire, /missions/{id}/pause, /missions/{id}/resume SSE events — add ThreadStateChanged, ChildThreadSpawned, and MissionThreadSpawned AppEvent variants. Expand the bridge event mapper to forward StateChanged and ChildSpawned engine events to the browser. Engine crate — add ConversationManager::clear_conversation() for /new and /clear commands. Code quality — replace 10 .expect() calls with proper error returns, remove dead AgentConfig.engine_v2 field, log silent init errors, fix duplicate doc comment, improve fallthrough documentation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): empty call_id on ActionResult and trace analyzer false positives Fix structured executor not stamping call_id onto ActionResult — the EffectExecutor trait doesn't receive call_id, so the structured executor must copy it from the original ActionCall after execution. Empty call_id caused OpenAI-compatible providers to reject the next LLM request with "Invalid 'input[2].call_id': empty string". Fix trace analyzer false positives: - code_error check now only scans User-role code output messages (prefixed with [stdout]/[stderr]/[code ]/Traceback), not System prompt which contains example error text - missing_tool_output check now recognizes ActionResult messages as valid tool output (Tier 0 structured path) - Add NotImplementedError to detected code error patterns New trace checks: - empty_call_id: detect ActionResult messages with missing/empty call_id before they reach the LLM API (severity: Error) - llm_error: extract LLM provider errors from Failed state reason - orchestrator_error: extract orchestrator errors from Failed state Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): add Missions tab to gateway UI Add a full Missions page to the web gateway with list view, detail view, and action buttons (Fire, Pause, Resume). Backend: add /api/engine/missions/summary endpoint returning counts by status (active/paused/completed/failed). Frontend: - New "Missions" tab between Jobs and Routines - Summary cards showing mission counts by status - Table with name, goal, cadence type, thread count, status, actions - Detail view with goal, cadence, current focus, success criteria, approach history, spawned thread list, and action buttons - Fire/Pause/Resume actions with toast notifications - i18n support (English + Chinese) - CSS following the existing routines/jobs patterns Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): eagerly initialize engine v2 at startup The gateway API endpoints (/api/engine/missions, etc.) call bridge query functions that return empty results when the engine state hasn't been initialized yet. Previously, initialization only happened lazily on the first chat message via handle_with_engine(). Now when ENGINE_V2=true, the engine is initialized in Agent::run() before channels start, so the self-improvement mission and other engine state is available to gateway API endpoints immediately. Also rename get_or_init_engine → init_engine and make it public so it can be called from agent_loop.rs at startup. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): improve mission detail with markdown goal and thread table - Goal rendered as full-width markdown block instead of plain-text meta item (uses existing renderMarkdown/marked) - Current focus and success criteria also rendered as markdown - Spawned threads shown as a clickable table with goal, type, state, steps, tokens, and created date instead of a UUID list - Clicking a thread row opens an inline thread detail view showing metadata grid and full message history with markdown rendering - Back button returns to the mission detail view - Backend: mission detail now returns full thread summaries (goal, state, step_count, tokens) instead of just thread IDs Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(web): close SSE connections on page unload to prevent connection starvation The browser limits concurrent HTTP/1.1 connections per origin to 6. Without cleanup, SSE connections from prior page loads linger after refresh/navigation, eating into the pool. After 2-3 refreshes, all 6 slots are consumed by stale SSE streams and new API fetch calls queue indefinitely — the UI shows "connected" (SSE works) but data never loads. Add a beforeunload handler that closes both eventSource (chat events) and logEventSource (log stream) so the browser can reuse connections immediately on page reload. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(web): support multiple gateway tabs by reducing SSE connections Each browser tab opened 2 SSE connections (chat events + log events). With the HTTP/1.1 per-origin limit of 6, the 3rd tab exhausted the pool and couldn't load any data. Three changes: 1. Lazy log SSE — only connect when the logs tab is active, disconnect when switching away. Most users rarely view logs, so this saves a connection slot per tab. 2. Visibility API — close SSE when the browser tab goes to background (user switches to another tab), reconnect when it becomes visible. Background tabs don't need real-time events. 3. Combined with the existing beforeunload cleanup, this means: - Active foreground tab: 1 connection (chat SSE only, +1 if logs tab) - Background tabs: 0 connections - Closed/refreshed tabs: 0 connections (beforeunload cleanup) This allows many gateway tabs to coexist within the 6-connection limit. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): route messages to correct conversation by thread scope Messages sent from a new conversation in the gateway always appeared in the default assistant conversation because handle_with_engine ignored the thread_id from the frontend. Two fixes: 1. Engine conversation scoping — when the message carries a thread_id (from the frontend's conversation picker), use it as part of the engine conversation key: "gateway:<thread_id>" instead of just "gateway". This creates a distinct engine conversation per v1 thread, so messages don't cross-contaminate. 2. V1 dual-write targeting — write user messages and assistant responses to the v1 conversation matching the thread_id (via ensure_conversation), not the hardcoded assistant conversation. Falls back to the assistant conversation when no thread_id is present (e.g., default chat). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): richer activity indicators for engine v2 execution The gateway UI showed only generic "Thinking..." during engine v2 execution with no visibility into CodeAct code execution, tool calls, or reflection. Now the event mapping produces detailed status updates: Step lifecycle: - "Calling LLM..." when a step starts (was "Thinking...") - "Step complete — N in / M out tokens" when done (was "Processing...") Tool execution: - Emit ToolStarted + ToolCompleted SSE events so the frontend renders proper tool cards with spinner → checkmark/error transitions - Duration shown in parameters field (e.g., "42ms") CodeAct visibility: - "Executing code..." when assistant produces a code block - "Code executed" / "Code executed (no output)" for successful runs - "Code error — retrying..." when Monty raises an exception Reflection: - "Reflecting on execution..." when post-thread analysis starts - "Reflection complete — N insight(s) saved" when done Also refactored thread_event_to_app_event → thread_event_to_app_events (returns Vec<AppEvent>) to support emitting ToolStarted before ToolCompleted in a single event handler pass. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): resolve tool names as callable stubs in CodeAct runtime When LLM-generated code calls `mission_list()` or any tool function, Monty's Python execution model first resolves the name (`mission_list`) as a NameLookup before invoking it as a FunctionCall. The NameLookup handler always returned Undefined, causing NameError before the function call could dispatch to the effect executor. Fix: before starting the Monty VM, collect all known tool names from the effect executor's available_actions(). In the NameLookup handler, if the name matches a known tool, return a MontyObject::Function stub instead of Undefined. Monty then yields FunctionCall for the stub, which dispatches to the normal tool execution pipeline. This enables CodeAct code to call any registered tool as a Python function: mission_list(), mission_create(), routine_list(), web_search(), memory_search(), etc. — all without explicit imports or __execute_action__ boilerplate. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): consolidate action execution, remove reflection, add learning missions Three major changes to the v2 engine: 1. **Consolidated action execution** — `handle_execute_action` in Rust is now the single source of truth for lease lookup, policy check, lease consumption, action execution, event emission, and ActionResult message recording. The Python orchestrator no longer duplicates event/message logic. This fixes the empty call_id bug (OpenAI HTTP 400) and the missing tool_calls on assistant messages (Codex "No tool call found" error). 2. **Removed reflection system** — Deleted the per-thread reflection pipeline (pipeline.rs, executor.rs), ThreadState::Reflecting, ThreadType::Reflection, enable_reflection config, and all 3 reflection event kinds. Learning is now handled entirely by event-driven missions that fire selectively. 3. **Three learning missions** replace reflection: - `self-improvement` — fires on trace issues (error diagnosis, prompt fixes) - `playbook-extraction` — fires on successful 5+ step threads (reusable procedures) - `conversation-insights` — fires every 5 threads per project (user preferences, domain knowledge, workflow patterns) Additional fixes: - llm_query()/llm_query_batched() always include system message (Codex compat) - handle_llm_complete adds assistant message with structured action_calls for Tier 0 responses (prevents "No tool call found" errors) - Gateway broadcasts without thread_id emit as Status events instead of being dropped - Comprehensive tests for call_id propagation and trace analysis (17 new tests) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(skills): extract ironclaw_skills crate and integrate with v2 engine Extract the skills system into a standalone `ironclaw_skills` crate (following the ironclaw_safety pattern) and wire it into the v2 engine for deterministic skill selection, CodeAct code injection, and confidence tracking. **ironclaw_skills crate** (94 tests): - Core types: SkillManifest, ActivationCriteria, LoadedSkill, SkillTrust - V2 types: V2SkillMetadata, CodeSnippet, SkillMetrics, V2SkillSource - Deterministic 4-phase selector (gating→scoring→budget→attenuation) - apply_confidence_factor() for extracted skill scoring - SKILL.md parser, validation/escaping, gating, registry, catalog - Feature-gated: catalog (reqwest), registry (filesystem) **Engine integration** (14 new tests): - DocType::Skill with retrieval weight 0.45 - SkillSelector bridges MemoryDoc→LoadedSkill for shared scoring - SkillTracker for usage/version/rollback confidence tracking - System prompt injection via <skill> XML blocks - CodeAct snippet injection via Monty NameLookup - Skill extraction mission replaces playbook extraction - ThreadManager.set_skill_selector() for runtime wiring **Bridge + migration**: - skill_migration.rs: v1 SKILL.md → v2 MemoryDoc (idempotent) - init_engine() migrates v1 skills, builds SkillSelector - src/skills/mod.rs → re-export shim **E2E test** (tests/engine_v2_skill_codeact.rs): - Full CodeAct loop: skill selected → LLM returns Python code → Monty executes http() → mock returns canned GitHub JSON → FINAL() terminates → thread completes with canned data - GitHub SKILL.md in skills/github/ as reference implementation Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Documenting research around how to extend to more integrations * docs: update engine-v2-architecture for missions and skills - Replace "Reflection Pipeline" with "Learning Missions" (self-improvement, skill-extraction, conversation-insights) - Add "Skills System" section covering ironclaw_skills crate, deterministic selection pipeline, CodeAct integration, confidence tracking, v1 migration - Update MemoryDoc types table (add Skill, remove Playbook as primary) - Update Integration Scaling section: Skills replace Capabilities-as-knowledge as the concrete implementation - Update example from Capability YAML to SKILL.md format with credentials - Fix thread state machine (remove Reflecting state) - Update key files table and test counts - Add self-improvement feedback loop diagram Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: clean up legacy playbook references in engine crate - Rename PLAYBOOK_MIN_STEPS/ACTIONS → SKILL_EXTRACTION_MIN_STEPS/ACTIONS - Fix pattern DB uses DocType::Note instead of DocType::Playbook - Update CLAUDE.md: skill-extraction mission, DocType list, module map Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(skills): credential specs in skill frontmatter, HTTP tool hardening, mission leases Skills can now declare API credentials in YAML frontmatter (SkillCredentialSpec, SkillCredentialLocation, SkillOAuthConfig, ProviderRefreshStrategy). Valid specs are registered into SharedCredentialRegistry at startup; the HttpTool auto-injects credentials for matching hosts — same zero-exposure model as WASM tools. HTTP tool security hardening: - Block LLM-provided auth headers for hosts with registered credentials - Return structured authentication_required error for missing credentials - Strip sensitive response headers (Set-Cookie, WWW-Authenticate, Authorization) - Scan response body through LeakDetector before returning to LLM Mission capability leases: registered mission_create/list/fire/pause/resume/delete as a "missions" capability so threads receive leases. Removed routine_* aliases from effect adapter — descriptions mention "routine" for LLM intent mapping. Includes 10 integration tests (tests/skill_credential_injection.rs) covering the full pipeline: YAML parsing → validation → registry → HttpTool wiring → per-user isolation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore(engine): remove legacy Playbook doc type, superseded by Skill Drop DocType::Playbook variant and all references — playbook extraction mission was already renamed to skill extraction in the previous session. Updates CLAUDE.md, architecture docs, context builder, retrieval weights, mission comments, and store adapter path mapping. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(engine): move skill selection and injection to Python orchestrator Skill selection was in Rust (SkillSelector in loop_engine.rs) — now it's in the Python orchestrator where the self-improvement mission can evolve it. Rust provides data access via two new host functions: - __list_skills__() — loads DocType::Skill MemoryDocs from Store - __record_skill_usage__(doc_id, success) — confidence tracking Python orchestrator handles everything else: - score_skill() — keyword/tag/confidence scoring (~40 lines) - select_skills() — budget-aware top-N selection (~15 lines) - format_skills() — XML block injection into system prompt (~20 lines) - Injection at step 0 with active_skill_ids stored in state Removed from Rust: - SkillSelector field + builder on ExecutionLoop and ThreadManager - format_skills_section() from prompt.rs - Rust-side skill injection block in loop_engine.rs - SkillSelector wiring in bridge/router.rs E2E test updated: skills stored in TestStore, Python orchestrator finds them via __list_skills__() and injects based on goal keywords. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: annotate v1-only code for removal after migration Mark modules and functions that exist solely for the v1 agent with "remove after v1 migration" notes: - src/skills/mod.rs ��� shim, attenuation, credential registration - src/skills/attenuation.rs — trust-based tool filtering (v1 only) - ironclaw_skills: selector, gating, registry, catalog modules - ironclaw_engine: skill_selector.rs (superseded by Python orchestrator) - src/bridge/skill_migration.rs — one-time v1→v2 conversion Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore(engine): remove unused skill_selector.rs Rust-side skill selection was moved to the Python orchestrator in7f87d179. This module had no production callers — only its own tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(skills): compile-time skill bundling infrastructure Add support for embedding skills into the binary at compile time: - build.rs: embed_skills() collects skills/*/SKILL.md into embedded_skills.json - src/skills/bundled.rs: loads embedded skills via include_str! - SkillRegistry: with_bundled_content(), load_from_content(), step 4 in discover_all() - Bundled skills are Trusted (ship with binary), lowest discovery priority - 4 new tests for bundled loading, user override, gating, and removal rejection - Cargo.toml: add serde_json build-dependency [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): non-blocking auth signal, NeedAuthentication flow, timeout safety When the HTTP tool detects a missing credential for a registered host: 1. EffectBridgeAdapter emits SSE AuthRequired event (best-effort, for connected frontends — silently dropped for missions/background threads) 2. Error flows back to LLM as normal ActionResult (non-blocking) 3. LLM tells the user to authenticate This avoids the blocking interruption approach which would hang mission threads and sub-threads that have no channel context. Engine additions: - EngineError::NeedAuthentication variant for structured auth failures - ThreadOutcome::NeedAuthentication for batch interruption when needed - structured.rs handles NeedAuthentication by interrupting the batch (stops subsequent calls, returns outcome to orchestrator) - Auth callback on EffectBridgeAdapter (optional, set by router for SSE) - extract_credential_name parser for HTTP tool error messages - routine_* tools added to is_v1_only_tool blocklist Safety: added 5-minute timeout to await_thread_outcome to prevent infinite hangs (e.g. after denied tool approval where thread fails to resume). Tests: 3 structured executor tests (NeedAuthentication interrupts batch, stops subsequent calls, regular errors don't interrupt) + 7 effect adapter tests (credential extraction, callback firing, v1-only tools). Also adds Linear API skill (skills/linear/SKILL.md). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): platform self-awareness, event pipeline fix, globals() builtin, prompt templates Session 9 changes driven by live trace analysis: - CodeAct event pipeline: handle_execute_code_step now transfers CodeExecutionResult events to thread.events and broadcasts via event_tx (fixes false-positive no_tools_used trace warnings) - Monty globals()/locals() builtins: returns dict of available action names from capability leases, enabling "tool_name" in globals() probing - PlatformInfo injection into system prompts (version, LLM backend, model, database, channels, owner, repo URL) - Mission goal prompts moved to prompts/*.md files (include_str! pattern) - /expected command for triggering self-improvement from user feedback - Session 9 development history Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): auto-approve http calls with registered credentials in v2 The v1 approval flow (interactive yes/no prompt) doesn't exist in v2. When the http tool returned UnlessAutoApproved for credentialed hosts, the effect adapter blocked with LeaseDenied — making all skill-based API calls fail. Fix: credential-backed http calls bypass the v1 approval check. The user authorized by storing the credential; the v1 interactive prompt is redundant in v2's lease-based security model. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(ui): show activated skills in CLI and gateway End-to-end skill activation display: 1. Python orchestrator emits __emit_event__("skill_activated", skill_names=...) after select_skills() picks skills for the conversation 2. Rust host function parses the comma-separated names into EventKind::SkillActivated 3. Router forwards to channels as StatusUpdate::SkillActivated 4. REPL renders: ◈ skills: github, linear (cyan) 5. Web gateway emits AppEvent::SkillActivated SSE event for frontend display Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(cli): show auth prompt in REPL when credential is missing The AuthRequired SSE event was emitted but only reached the web gateway. The REPL never saw it because it receives events through forward_event_to_channel which converts ThreadEvents to StatusUpdates. Fix: when forward_event_to_channel sees an ActionFailed with "authentication_required" in the error, emit StatusUpdate::AuthRequired to the channel. Also add AuthRequired/AuthCompleted rendering to the REPL (was missing — fell through to unmatched arm). CLI now shows: ⚿ Authentication required: github_token Store the credential with: ironclaw secret set <name> <value> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(ui): show tool arguments in CLI and gateway Add params_summary to ActionExecuted/ActionFailed events so the CLI and gateway can display what tools are doing: ● http(https://api.github.com/repos/nearai/ironclaw/issues) ● web_search(latest AI news) ● memory_read(HEARTBEAT.md) The summarize_params() helper extracts the most relevant argument per tool type (URL for http, query for search, path for memory, etc.) and truncates to 80 chars. Sensitive params are not included. Router forwards the summary in both StatusUpdate (CLI/REPL) and AppEvent (web gateway SSE) display names. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: handle Python None params in http tool, add params_summary to CodeAct dispatch Two fixes from live testing: 1. http tool: treat null headers/body as empty (Python's None becomes JSON null via Monty). Previously headers=None errored with "'headers' must be an object or array of {name, value}". 2. scripting.rs: compute params_summary before dispatching actions in the CodeAct path (was always None). Now http calls show their URL in the CLI: ● http(https://api.github.com/repos/.../issues) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: remove glob re-exports, fix clippy warnings, clean up duplicates - Remove `pub use ironclaw_safety::*` from src/safety/mod.rs and migrate all 20+ call sites to import directly from `ironclaw_safety` - Remove `pub use ironclaw_skills::*` from src/skills/mod.rs and migrate all 15+ call sites to import directly from `ironclaw_skills` - Fix 4 clippy warnings: 2 shadow imports, 2 collapsible if-let chains - Add missing SkillActivated arm to WASM channel StatusUpdate match - Remove duplicate AuthRequired/AuthCompleted arms in repl.rs - Update CLAUDE.md extracted crates guidance and prompt template rule - Fix bench imports (safety_check, safety_pipeline) 46 files changed, zero warnings, 3836 tests passing. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(auth): guided credential flow — prompt for token and retry When a thread completes with authentication_required, the router enters "auth mode" for that user: 1. Detects credential_name from the error in the thread response 2. Looks up setup_instructions from the skill's credential spec 3. Emits AuthRequired to CLI/gateway with instructions 4. Stores PendingAuth — next user message is treated as a token 5. Stores the token in SecretsStore 6. Retries the original user request automatically CLI flow: › create an issue in github ⚿ Authentication required: github_token Create a PAT at https://github.com/settings/tokens Paste your token below (or type 'cancel'): › ghp_abc123... ✓ github_token authenticated: Credential stored. Retrying... ● http(https://api.github.com/repos/.../issues) Issue created: https://github.com/... Gateway flow: same but AuthRequired SSE event shows the auth modal. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): skill-based OAuth flow tests 6 E2E tests covering the full skill credential lifecycle via the gateway API: - test_github_skill_loaded: github skill with credential spec loaded - test_no_github_token_initially: no stored secrets before auth - test_http_tool_returns_auth_required: http tool signals missing cred - test_guided_auth_flow: request → auth prompt → paste token → retry - test_auth_required_sse_event: SSE stream includes auth/skill events - test_different_users_isolated: per-user credential scoping Includes mock API server (aiohttp) requiring Bearer auth with token tracking for assertions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: document Monty runtime limitations in CodeAct prompt, fix new-thread read-only - Add "Runtime environment" section to codeact_preamble.md documenting Monty's restrictions: no stdlib imports, single imports only, no classes/ with/match/del/yield, available builtins and modules, workarounds - Add MONTY.md tracking current pin, all limitations, upgrade process, and changelog for future Monty updates - Fix gateway createNewThread() not resetting read-only state — new threads now eagerly enable chat input instead of waiting for async loadThreads() callback Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): transition thread to Waiting on NeedApproval The orchestrator Python returned {"outcome": "need_approval"} without calling __transition_to__("waiting"), leaving the thread in Running state. When the user later approved/denied, resume_thread rejected it with "thread is not resumable from Running". - Add __transition_to__("waiting", "approval needed") in both code-step and action-call approval paths in default.py - Add Rust safety net in loop_engine.rs: if orchestrator returns NeedApproval but thread isn't Waiting, force the transition Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): restructure workspace storage for human readability Rewrite HybridStore (src/bridge/store_adapter.rs) to produce a developer-friendly workspace layout: - Knowledge docs use frontmatter+markdown with slugified filenames instead of UUID.json with wrapped structs - Orchestrator code, prompt overlays, and failure tracker grouped under engine/orchestrator/ - Missions nested under their project in named folders with room for working files alongside mission.json - Runtime state (threads, leases, events) under engine/.runtime/ - Terminal threads archived to compact summaries, dead leases cleaned on startup - Auto-generated engine/README.md with knowledge counts, mission status, and thread stats Also includes: /expected command, approval state fix, platform self-awareness, Monty limitations in preamble, prompt template extraction. See docs/development-history.md Session 10 for details. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(skills): explicit /skill-name activation in messages Users can now write /github or /file-issues anywhere in their message to force-activate a skill. The /skill-name is replaced with the skill's description so the sentence reads naturally for the LLM: "fetch issues from /github" → "fetch issues from GitHub API" "please /file-issues for all bugs" → "please file detailed GitHub issues for all bugs" Implementation: - extract_skill_mentions() in selector.rs scans for /name patterns, matches against available skills, returns matched skills + rewritten message - select_active_skills() returns (skills, rewritten_message) — explicit mentions merged with score-based selection - dispatcher.rs rewrites the last user message in LLM context with expanded text - 8 tests covering: basic mention, description expansion, hyphenated names, multiple mentions, unknown skills, URLs not matched Also includes: seed_orchestrator_v0() for workspace visibility of compiled-in orchestrator code. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): wire NeedAuthentication and NeedApproval through v2 CodeAct path Production traces revealed tool result desync on RequireApproval (no ActionResult message → OpenAI 400), auth flow not triggering in CodeAct (EffectAdapter returned Ok instead of Err(NeedAuthentication)), and HTTP tool blocking unauthenticated requests. Fixes: - Add emit_and_record() to RequireApproval branch in handle_execute_action - Wire NeedAuthentication through scripting.rs DispatchResult, orchestrator host functions, default.py, loop_engine safety net - Add EngineError::NeedApproval variant; effect adapter returns it instead of LeaseDenied for tools needing approval - HTTP tool: inject-if-available (proceed without auth, error only on 401) - HTTP_ALLOW_LOCALHOST env flag for E2E testing with mock servers - host_matches_pattern supports port in pattern (127.0.0.1:8080 matches host_str() output 127.0.0.1) - CodeAct postamble: error recovery guidance - Orchestrator user_id from thread.metadata instead of hardcoded "orchestrator" Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): v1/v2 history, approval routing, cancel cleanup Multiple v2 engine bridge fixes discovered by E2E tests: - Write response to v1 DB for ALL thread outcomes (not just Completed), so history API shows NeedApproval/NeedAuthentication responses - Remove v1 thread_id hint from pending_approval lookup (v1/v2 use different UUID spaces) - Add has_pending_auth() check in agent_loop: route "cancel"/"no" through handle_with_engine when PendingAuth is active (SubmissionParser parsed "cancel" as ApprovalResponse, bypassing auth flow) - Add engine_thread_id to PendingAuth; stop_thread on cancel - Write cancel response to v1 DB - NeedAuthentication handler enters guided auth flow with setup hints Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): comprehensive v2 engine test suite (12 tests, 5 files) E2E tests for the v2 engine covering auth flow, approval lifecycle, error handling, and edge cases. Uses mock API servers with strict token validation, dedicated ironclaw server fixtures per module, and the mock LLM's tool call pattern system. Tests: - Auth flow: skill activation, NeedAuthentication → token → retry, credential persistence across threads, cancel during auth, empty token treated as cancel, special character injection safety - Approval: approve yes (text-based), deny, always (persists across threads), prompt mentions tool name - Error handling: max iterations (30 step limit), tool intent nudge (LLM recovery after "let me search") Infrastructure: - mock_llm.py: runtime-configurable github_api_url, tool call patterns for issues/loop/drive, canned responses for intent nudge - HTTP_ALLOW_LOCALHOST=true + SECRETS_MASTER_KEY in fixtures - Separate server instances for cancel tests (cancel contaminates conversation state) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add Session 11 — E2E test suite + engine hardening Documents 14 bugs found across two production traces and E2E test execution, the test infrastructure design (mock servers, dedicated fixtures, HTTP_ALLOW_LOCALHOST), and the architecture evolution from trace analysis → code fix → test to prevent regression. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(auth): kernel-level pre-flight auth gate for engine v2 Transform authentication from a reactive post-execution error to a proactive pre-flight check. The EffectBridgeAdapter now checks credentials BEFORE executing tool calls, preventing wasted HTTP requests and 401 errors from reaching the LLM. Key changes: - New AuthManager (src/bridge/auth_manager.rs) centralizes credential checking, setup instruction lookup, and tool readiness queries - Pre-flight auth gate in execute_action() checks SharedCredentialRegistry + SecretsStore before tool execution - Post-install auth pipeline: after tool_install, kernel auto-checks readiness and initiates auth flow or appends setup instructions - tool_auth and tool_activate filtered from v2 LLM tool list and blocked in execute_action() — auth is kernel-level in v2 - Text-based auth detection kept as defense-in-depth fallback with tracing when it fires - Setup instruction lookup deduplicated via AuthManager - ExtensionManager gains check_tool_auth_status_pub() for auth queries Also fixes pre-existing DocType::Plan exhaustiveness errors in the engine crate and re-exports PlanStepDto from ironclaw_common. Includes 10 unit tests (AuthManager + is_v1_auth_tool) and 5 E2E tests covering pre-flight blocking, auth-then-retry, credential persistence, v1 auth tools hidden, and auth cancellation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add Session 12 — kernel-level auth rework decisions and rationale Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(plan): autonomous plan mode via v2 primitives (MemoryDoc, Mission, SSE) Add plan mode for autonomous long-running task execution, composing existing v2 engine primitives rather than new engine states. Inspired by OpenAI Codex's update_plan checklist and Claude Code's file-based plan mode — both enforce planning through prompts, not tool removal. Engine: DocType::Plan variant for MemoryDoc (project-scoped, retrievable). Events: PlanUpdate SSE event with PlanStepDto for live checklist rendering. Tool: plan_update tool broadcasts structured plan progress via SSE. Command: /plan (create/approve/status/revise/list) rewrites to UserInput with [PLAN MODE] prefix to activate the plan-mode skill. Skill: skills/plan-mode/SKILL.md defines full plan protocol — creation (memory_write), approval (mission_create + mission_fire), execution (step-by-step with plan_update), and revision flows. UI: Inline chat checklist widget with status badges, step icons (checkmark/spinner/circle), results, and progress summary. Tests: 5 E2E scenarios + mock LLM patterns + helper selectors. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(auth): scope pending approvals by thread to prevent cross-thread leakage The `engine_pending_approval()` handler was ignoring the v1 thread_id parameter, passing `None` to the resolver. This caused two bugs: 1. An approval pending on thread A would appear in thread B's history 2. Multiple concurrent approvals for the same user returned Ambiguous Fix: pass the v1 thread_id as a hint and update `resolve_pending_approval_for_thread()` to match against both engine thread UUIDs (direct match) and v1 session UUIDs embedded in the conversation channel key ("web:{v1_uuid}"). This eliminates the unused `thread_id` variable warning in chat.rs:293. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(auth): fix gateway auth card for skill credentials + fallback token storage Two bugs in the gateway auth flow for v2 engine skill credentials: 1. **Frontend: auth card not shown for skill credentials** When `auth_url` is None but `instructions` ARE present (skill-based credential like `github_token`), the frontend incorrectly called `showConfigureModal()` (extension setup UI) instead of `showAuthCard()` (token paste UI). The configure modal fails for skill credentials (they're not extensions), permanently blocking the chat input. Fix: show `showAuthCard()` when instructions are present, regardless of `auth_url`. The configure modal is now only used when neither `auth_url` nor `instructions` are provided (pure extension setup). 2. **Backend: /api/chat/auth-token doesn't handle skill credentials** The auth-token endpoint calls `ext_mgr.configure_token()` which fails for skill credentials ("extension not installed"). The token is never stored, leaving the user stuck. Fix: when `configure_token()` fails with "not installed"/"not found", fall back to storing the token directly in SecretsStore via the tool registry. This bridges the frontend auth card and the v2 engine's skill credential system. Includes 3 E2E tests (test_v2_kernel_auth_gateway_flow.py) covering the auth-token API path, chat-message token path, and cancel flow. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: resolve pre-existing clippy warnings and E2E test failures Clippy fixes: - Collapse nested `if let` into `&&` chains (effect_adapter.rs, http.rs) - Replace `match` with `if let` for single-pattern destructure (store_adapter.rs) E2E test fixes (test_v2_engine_oauth_google.py): - Add `HTTP_ALLOW_LOCALHOST` and `SECRETS_MASTER_KEY` to test env (mock API runs on localhost — without this the HTTP tool silently blocks the request) - Skip `test_oauth_redirect_flow` when extension returns "not installed" (was only checking for HTTP 404, but the endpoint returns 200 with success:false for missing extensions) - Fix NoneType crash: `turns[-1].get("response", "")` returns None when key exists with None value — use `(... or "")` pattern instead - Skip `test_invalid_token_paste` when credentials already stored from prior test (test ordering dependency) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(auth): add logging to skill credential fallback in auth-token endpoint Add debug/warn logging when the skill credential fallback path fires in chat_auth_token_handler, making it easier to diagnose when the secrets store is unavailable. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(auth): strict pre-flight gate E2E test + unit integration test Add strict assertion: mock API must receive ZERO requests when pre-flight gate blocks (was previously lenient). The test was failing because the E2E conftest built the binary to `target/debug/` while `cargo build` with shared-target outputs to `~/.cargo/shared-target/debug/`. Fixed via symlink. Also adds `preflight_gate_blocks_missing_credential` unit test that exercises `execute_action()` directly with a mocked ToolRegistry containing credential mappings — verifies NeedAuthentication is returned without executing the tool. Diagnostic logging: warn-level log when pre-flight gate is skipped due to missing auth_manager or credential_registry dependency. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(auth): harden auth flow — clear v2 pending state, fix binary path, SSE broadcast Three hardening fixes from fragility audit: 1. **Clear v2 pending_auth from API path** (#5/#6): The /api/chat/auth-token endpoint now calls `clear_engine_pending_auth()` after storing credentials. Without this, the next chat message would be intercepted as a token retry even though auth was completed via the API endpoint. 2. **Fix E2E binary path resolution** (#1): conftest.py now resolves the actual cargo target-dir from ~/.cargo/config.toml instead of hardcoding `target/debug/`. Also adds `crates/` to the mtime check inputs so engine crate changes trigger rebuilds. 3. **Send AuthCompleted SSE from chat message path** (#7): The chat-message token submission path now broadcasts AuthCompleted via SSE (same as the API path), so the frontend dismisses the auth card immediately. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(auth): survive SSE reconnect — include pending_auth in history response When SSE drops during an auth flow and the frontend reconnects, the auth card was lost (DOM cleared by loadHistory) but authFlowPending remained true, permanently blocking the chat input. Fix: include `pending_auth` in the `/api/chat/history` response (same pattern as `pending_approval`). The frontend's `loadHistory()` now re-shows the auth card when `pending_auth` is present, and clears stale auth UI state when it's absent. Backend: - Add `PendingAuthInfo` type to gateway types - Add `get_engine_pending_auth()` to router (queries v2 pending_auth) - Include `pending_auth` in all HistoryResponse constructions Frontend: - `loadHistory()` calls `handleAuthRequired()` when `pending_auth` present - Clears `authFlowPending` when `pending_auth` is absent (cleanup stale state) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(auth): robust secrets store fallback in auth-token endpoint The skill credential fallback in /api/chat/auth-token was silently failing when tool_registry.secrets_store() returned None. Fix: try tool_registry first, then fall back to extension_manager.secrets(). If neither is available, return an explicit error instead of falling through to the "Extension not installed" message. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(auth): add credential fallback to ACTUAL auth-token handler in server.rs Root cause: there are TWO `chat_auth_token_handler` functions — one in handlers/chat.rs (dead code, never called) and one in server.rs (the real one registered on the route). All previous fallback fixes went to the wrong file. Fix: add the skill credential fallback (store directly in SecretsStore when extension manager returns NotInstalled) to the REAL handler in server.rs. Uses extension_manager.secrets() as fallback when tool_registry.secrets_store() is None. Also strengthens the E2E test to assert `success: true` in the response body, not just HTTP 200 status (which masked this bug for weeks). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore(engine): update Monty to v0.0.9 (7a0d4b7) Removes three runtime limitations: multi-module imports now work, datetime and json modules now available as builtins. Updates CodeAct preamble and MONTY.md tracking doc accordingly. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(gateway): remove 504 lines of dead chat handlers from handlers/chat.rs Five handler functions in handlers/chat.rs were dead code — identical copies existed in server.rs where the routes are actually registered: - chat_send_handler - chat_approval_handler - chat_auth_token_handler (the root cause of the auth-token fallback bug) - chat_auth_cancel_handler - chat_history_handler (+ engine_pending_approval/auth helpers) These duplicates caused the auth-token fallback bug: fixes were applied to the dead copy in handlers/chat.rs while the real handler in server.rs remained unchanged. Removing them prevents this class of bug entirely. Kept: clear_auth_mode (shared helper), chat_events_handler, chat_ws_handler, chat_threads_handler, chat_new_thread_handler, and unit tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): strengthen assertions to prevent false confidence Audit found 18 weak assertions across 7 E2E test files that could pass even when features are broken. Key patterns fixed: 1. **Require token in mock API** (auth_flow, oauth_google): Replace `token_received OR "improve" in response OR "http" in response` with `assert token in mock_api_tokens` — must verify the mechanism, not just that "something happened" 2. **Remove passive assertions** (auth_flow): Replace `if tokens: assert True; else: pass` with `assert len(tokens) > 0` — silent passes hide failures 3. **Remove generic keyword matches** (approval_flow): Replace `"tool" in response` (matches anything) with specific `pending_approval is None` (verifies state change) 4. **Verify mock API received requests** (preflight, auth_flow): Add `assert request_count > 0` after credential storage to prove credential injection actually worked 5. **Poll for state change, not just text** (approval_flow): Wait for `pending_approval` to be cleared rather than checking for specific keywords in response text 6. **Add negative auth checks after cancel** (auth_cancel): Verify "paste your token" not in response after cancel flow Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): fix skipped OAuth tests — reorder for isolation, add fake WASM extension Two tests were skipped due to test ordering and missing infrastructure: 1. **test_invalid_token_paste**: Was skipped because credentials stored by prior test_api_key_then_api_call prevented auth prompt from triggering. Fix: reorder to run BEFORE api_key test. Now runs without skip. 2. **test_oauth_redirect_flow**: Was skipped because google_drive isn't a real WASM extension. Added fake WASM extension with OAuth capabilities (empty .wasm + capabilities.json). Still skips because wasmtime can't load the empty binary, but now has clear infrastructure for when a real test binary is available. Also made test_api_key_then_api_call resilient to prior bad-token state from test_invalid_token_paste (graceful fallback if no auth prompt). Result: 24 passed, 1 skipped (down from 2 skipped). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): use real google-drive WASM binary for OAuth redirect test Replace the fake empty WASM binary with the real google-drive tool built from tools-src/google-drive/. The extension manager can now activate it via wasmtime, generate a real OAuth URL, and complete the redirect flow. Result: 25 passed, 0 skipped (was 24 passed, 1 skipped). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): full multi-tenant isolation for v2 engine Add user_id as a first-class field to Thread, Mission, MemoryDoc, and Project types. Update Store trait with user-scoped list methods and admin cross-tenant methods. Enforce ownership validation throughout: - ThreadManager: stop/inject/resume require user_id, validate ownership - MissionManager: fire/pause/resume validate ownership, per-user learning missions (self-improvement, skill-extraction, conversation-insights) - ConversationManager: validate conversation ownership on message/clear - Bridge router: all public functions take user_id, web handlers pass AuthenticatedUser identity through Shared space model for system resources: - list_memory_docs_with_shared / list_missions_with_shared merge user's own docs/missions with system-owned ones (admin-installed skills, shared knowledge) - System missions require admin role to manage (403 for non-admins) - Learning missions are per-user: pause/resume is independent per user Legacy migration: on startup, stamps owner_id onto pre-existing records that deserialized with user_id="legacy" (serde default). Event listener fires learning missions with the completed thread's user_id (not a hardcoded owner_id), ensuring artifacts stay user-scoped. 8 new multi-tenancy tests covering isolation, cross-user denial, shared visibility, admin-only management, and per-user event scoping. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): audit fixes — dead code, tautological check, dedup, formatting - Fix tautological trace check in executor/trace.rs that could never fire: the "missing_tool_output" diagnostic now correctly checks User-role messages instead of re-checking ActionResult role - Remove dead code in loop_engine.rs: save_runtime_checkpoint, check_signals, SignalAction (replaced by Python orchestrator); simplify RuntimeCheckpoint to just persisted_state; move extract_final_from_text to #[cfg(test)] - Deduplicate default_user_id() — single definition in types/mod.rs used by thread, memory, project, and mission types - Add PartialEq derive to Provenance enum - Remove phantom skill_selector.rs from CLAUDE.md module map - Fix cargo fmt violations in test code Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(skills): audit fixes — CRLF parser bug, logging levels, dedup, SkillSource::Installed - Fix parse_skill_md to normalize \r\n internally so callers don't need to pre-normalize (find_closing_delimiter byte offset was wrong on CRLF) - Change tracing::info! to tracing::debug! in registry (3 sites) per CLAUDE.md logging policy — info! corrupts REPL/TUI - Extract shared build_loaded_skill() helper, eliminating ~40 duplicated lines between load_and_validate_skill and load_from_content - Add SkillSource::Installed variant so installed-dir skills have correct provenance metadata (was incorrectly using SkillSource::User) - Fix misleading dedup log labels in discover_all override source strings - Log warning on reqwest::Client builder failure instead of silent fallback - Add regression tests for CRLF and mixed line endings in parser - Auto-fix 155 uninlined_format_args clippy warnings in test code Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style(engine): auto-fix clippy uninlined_format_args warnings Formatting-only changes applied by cargo clippy --fix. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(engine): deduplicate test Store mocks with shared InMemoryStore Expand the shared InMemoryStore in lib.rs to support all entity types (threads, steps, events, projects, docs, leases, missions) with proper CRUD semantics and user_id/project_id filtering. Replace 3 duplicate mock Store implementations (~350 lines removed): - executor/context.rs: DocStore → InMemoryStore - memory/retrieval.rs: DocStore → InMemoryStore - memory/store.rs: InMemoryDocStore → InMemoryStore Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): audit fixes — UTF-8 panic, missing Plan type, dedup, logging - Fix UTF-8 panic in truncate_for_readme: use char-based truncation instead of byte-index slicing on user content (thread goals, messages) - Add missing "Plan" arm to deserialize_knowledge_doc — was silently falling through to Note, losing doc type on workspace reload - Extract shared event display helpers (format_action_display_name, interpret_message_event) to deduplicate logic between forward_event_to_channel and thread_event_to_app_events - Change info! to debug! in skill_migration to avoid corrupting REPL Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(engine): move SkillTracker from capability/ to memory/ SkillTracker does MemoryDoc CRUD (load skill → update metrics → save), not capability/lease/policy operations. It belongs with the memory persistence layer, not the access-control layer. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(engine): add v2 acceptance tests with per-agent ENGINE_V2 toggle Add engine v2 acceptance test infrastructure and 8 initial tests proving the v2 pipeline works end-to-end through the real agent loop. Infrastructure: - Add `engine_v2: bool` to AgentConfig (resolved from ENGINE_V2 env var) - Replace process-global `is_engine_v2_enabled()` checks in agent_loop.rs with per-agent `self.config.engine_v2` — safe for parallel test execution - Add `reset_engine_state()` to clear the OnceLock singleton between tests - Add `.with_engine_v2()` to TestRigBuilder and `run_recorded_trace_v2()` Tests (tests/e2e_engine_v2.rs): - v2_smoke_text_response: basic text routing through engine v2 - v2_single_tool_call: echo tool dispatch via EffectBridgeAdapter - v2_multi_tool_chain: sequential echo + time tool execution - v2_tool_error_recovery: JSON parse error propagation and LLM recovery - v2_multi_turn_conversation: context persistence across ConversationManager - v2_status_events: ToolStarted/ToolCompleted event emission - v2_recorded_telegram_check: v1 parity — replay recorded trace through v2 - v2_recorded_weather_sf: v1 parity — HTTP tool with large response Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style(bridge): auto-format effect_adapter, store_adapter, codeact test Formatting-only changes applied by cargo fmt. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(engine): remove executor/intent.rs — duplicated by Python orchestrator signals_tool_intent() and TOOL_INTENT_NUDGE were from the Rust-native loop path. The Python orchestrator (default.py) has its own signals_tool_intent() implementation. No Rust code referenced the module — safe to delete. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * merge: integrate origin/staging — fix ensure_conversation arity Merge staging to pick up the 5th `source_channel` parameter added to `ensure_conversation()` in V15 migration. Fix the v2 bridge call site at router.rs:1390 to pass `Some(&message.channel)`. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): formatting and no-panics check in merged router tests Fix two CI issues from the staging merge: - Reformat make_expected_test_state signature (single-line args) - Add inline // safety: comment on test-only assert! to suppress the no-panics production code check Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): cargo-deny wildcard and git source errors - Pin monty to rev 7a0d4b7 instead of branch=main (deterministic builds) - Add version = "0.1.0" to ironclaw_engine and ironclaw_skills path deps (fixes wildcard dependency errors) - Allow git sources for pydantic/monty and astral-sh/ruff in deny.toml - Set allow-wildcard-paths = true (monty is git-only, no crates.io version) - Add // safety: comments on unwrap() calls guarded by len()==1 checks - Auto-fix clippy warnings and formatting in merged engine files Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): port V1 tool-intent nudge to Python orchestrator The V2 Python orchestrator's signals_tool_intent() was too aggressive — matching "I can" + "call"/"fetch" anywhere in text caused false positives on news content and past-tense summaries, creating a nudge loop that burned 1.6M tokens on a simple query. Ported V1's approach: strip code blocks and quoted strings, check 15 exclusion phrases, then require a future-tense prefix ("let me", "I'll", "I will", "I'm going to") immediately followed by an action verb. Also fixed the nudge counter to use V1's consecutive semantics — it no longer resets on action/code responses, only on non-intent text. Added 11 Monty-based unit tests covering true positives, true negatives, exclusions, code blocks, quoted strings, and 3 regression tests from the trace that triggered this fix. Also includes: parallel store persistence, pre-fetched system docs for orchestrator loading, and parallel action execution in scripting. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): async-first CodeAct tool dispatch via Monty ResolveFutures Tool calls in CodeAct now use Monty's async suspension model: each tool FunctionCall does preflight (lease/policy) synchronously, then spawns a tokio task and calls resume_pending() to return an ExternalFuture to Python. When Python awaits the future (or gathers multiple via asyncio.gather), Monty yields ResolveFutures and the host resolves all pending tools — which ran concurrently as tokio tasks. This replaces the old synchronous dispatch_action + execute_parallel approach with native Python async/await semantics. The LLM writes natural Python: `await tool()` for sequential, `asyncio.gather()` for parallel — no special API needed. Changes: - scripting.rs: async tool dispatch via resume_pending + ResolveFutures handler, preflight_action for lease/policy, PendingTool tracking - Removed: dispatch_action, DispatchResult, handle_execute_parallel, execute_parallel NameLookup entry - Builtins (FINAL, llm_query, etc.) remain synchronous - 9 new tests: single await, 2/3-way gather, sequential chains, error propagation, denied tools, empty/single gather, globals, FINAL sync Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(security): address PR review findings — 11 fixes across engine and bridge Critical: - C3: Fix UTF-8 byte-slice panic in summarize_params (event.rs) — use truncate() helper instead of raw &u[..77] High: - H1: Stop leaking internal errors to HTTP clients — all 12 engine API handlers now return generic "Internal engine error" instead of e.to_string() - H3: Add MAX_AUTH_RETRY_DEPTH=2 recursion limit to auth retry in router - H9: Change default_trust() from Trusted to Installed (fail-closed) - H6: Demote all warn!() to debug!() in engine crate (~25 locations) to prevent REPL/TUI corruption per CLAUDE.md logging policy - H14: Remove no-op test_approval_prompt_contains_tool_name (was just `pass`) - H2/M5: Add credential name validation (alphanumeric+underscore, max 64 chars) Medium: - M15: Replace raw byte-slicing with .get() in deserialize_knowledge_doc - M14: Prevent pending_auth overwrite — check for existing entry before insert in text-fallback path - M13: Log swallowed store errors in record_orchestrator_failure instead of silent unwrap_or_default - LeaseNotFound: Add distinct EngineError::LeaseNotFound variant instead of reusing LeaseExpired for missing leases Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(docker): add python3-dev build dependency for monty/pyo3 The monty crate (embedded Python interpreter) uses pyo3-build-config with the resolve-config feature, which probes for Python 3 headers at compile time. Without python3-dev in the builder stage, the Docker build fails. Added python3-dev to the chef stage's apt-get install. This is a build-only dependency — the runtime image (debian:bookworm-slim) is unchanged and does not include Python. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): make llm_query and llm_query_batched async via ResolveFutures llm_query() and llm_query_batched() now use the same async dispatch as tool calls: spawn tokio task, resume_pending(), resolve in ResolveFutures handler. This enables: import asyncio summary, results = await asyncio.gather( llm_query("summarize this", context=data), web_search(query="latest news"), ) The LLM call and tool call run concurrently — saving 1-3s per step when both are needed. rlm_query() stays synchronous because it spawns a child Monty VM which isn't Send (can't cross tokio::spawn boundary). Refactored PendingTool → PendingFuture enum with Tool and Llm variants. Extracted resolve_tool_future() and resolve_llm_future() helpers for clean resolution in the ResolveFutures handler. Token usage from async LLM calls is accumulated via the (ExtFunctionResult, TokenUsage) return type. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): auto-approve tool after single approval to prevent infinite loop When a user approves a tool call with "yes" (not "always"), the engine resumes the thread and the LLM issues a NEW tool_install call — which triggers another approval prompt, creating an infinite approve loop. Fix: auto-approve the tool for the session on any "yes" approval, not just on "always". The user already consented to this tool — asking again is a UX bug. "always" still works the same (persistent across threads). Discovered via trace analysis: engine_trace_20260331T222859.json showed the GitHub tool_install stuck in a Waiting→approve→Waiting loop. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(http): move leak detection before credential injection The leak detector was scanning outbound HTTP headers AFTER the credential registry injected Authorization headers, causing false positives — legitimate system-injected GitHub tokens were blocked as "secret leaks." Fix: scan the LLM-controlled headers/URL/body first (catches actual exfiltration attempts), THEN inject system credentials (trusted, not LLM-controlled). This preserves leak detection for LLM-crafted headers while allowing the credential injection system to work. Discovered via trace: engine_trace_20260331T225126.json showed http(api.github.com) blocked with "Secret leak blocked: pattern 'header:Authorization' matched 'github_fine_grained_pat'". Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): 4 integration fixes — auth cancel, history, per-user projects, plan scope P1: Clear engine pending auth on /api/chat/auth-cancel The cancel endpoint only cleared v1 session state, leaving bridge::pending_auth active. Next message was consumed as a token value instead of normal input. P2: Populate pending_auth in chat history response All 4 HistoryResponse code paths hardcoded pending_auth: None. On SSE reconnect/refresh during an auth flow, the UI cleared the auth card while the bridge was still waiting for a token. Now surfaces engine pending-auth state via get_engine_pending_auth(). P1: Per-user default project instead of global owner project Engine init created one project under owner_id and used it for all users. In multi-user gateway deployments, non-owner threads and missions were created inside the owner's project, making /api/engine/projects return nothing for non-owner accounts. Added resolve_user_project() that creates per-user projects. P2: Include thread_id in plan_update SSE events plan_update events had thread_id: None, so plan checklists from background threads rendered in whichever chat was open. Now carries ctx.conversation_id so clients can scope plan rendering. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): address review feedback — lease audit, policy logging, char count From ilblackdragon's review: 1. Lease revocation now stores reason for audit trail — added revoked_reason: Option<String> to CapabilityLease, logged at debug! level on revocation (was silently discarding _reason param) 2. Policy denial decisions now logged at debug! level with action name, capability, and reason — enables incident investigation for privilege escalation attempts 3. Fixed byte/char count mismatch in compact_output_metadata — stdout.len() (bytes) was displayed as "chars" but stdout.chars().count() was used for truncation. Now consistent. 4. Store trait splitting (H3) acknowledged as follow-up — documenting that default impls are stubs, not real behavior. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): address zmanian review — orchestrator gate, sandbox tests, TOCTOU, matching C1: Add ORCHESTRATOR_SELF_MODIFY disable flag (default: off) Runtime orchestrator loading is now disabled by default. Only the compiled-in v0 runs unless explicitly opted in. Prevents unreviewed self-improvement patches from executing with full tool access. C2: Add 3 Monty sandbox security negative tests (224 total) - sandbox_denies_os_operations: os.system() blocked - sandbox_enforces_resource_limits: infinite loop terminated - sandbox_restricts_imports: subprocess import blocked H3: Fix TOCTOU race in lease find+consume Added LeaseManager::find_and_consume() that atomically finds a lease and consumes a use under a single write lock. structured.rs now uses this instead of separate find (read lock) + consume (write lock). M3: Fix ActionCondition::ActionMatches substring → exact match "delete" no longer matches "undelete_restore". Changed contains() to == for exact action name matching. Also: log store errors in loop_engine.rs instead of silent unwrap_or_default(). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(security): block orchestrator/prompt writes when self-modify disabled Defense-in-depth: three layers now enforce ORCHESTRATOR_SELF_MODIFY: 1. memory_write tool: blocks writes to orchestrator:* and prompt:* paths with a clear error message when the flag is off 2. HybridStore adapter: save_memory_doc rejects protected docs (except system-internal v0 seeding and failure tracking) with EngineError::AccessDenied 3. Mission system: process_self_improvement_output skips prompt additions when the flag is off, logging the skip at debug level Previously, ORCHESTRATOR_SELF_MODIFY only controlled loading — the LLM could still write malicious orchestrator/prompt MemoryDocs via memory_write, which would take effect once self-modify was enabled. Now writes are blocked at all layers. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(gate): unified ExecutionGate abstraction for approval + auth flows (#1818) * feat(gate): unified ExecutionGate abstraction for approval + auth flows Introduce a composable gate pipeline that structurally prevents the 6 recurring bug categories found across ~50 approval/auth fixes: TOCTOU races, cross-channel hijacking, privilege escalation via composition, silent error swallowing, state loss on restart, and execution path mismatch between approval and authentication. Engine crate (ironclaw_engine::gate): - ExecutionGate trait with priority-ordered GatePipeline (fail-closed) - GateDecision (Allow/Pause/Deny) — no None variant by construction - ResumeKind (Approval/Authentication/External) — unified pause type - GateResolution with Cancelled variant (fixes cancel-as-approval misrouting) - ToolTier classification (ReadOnly < Stateful < Privileged < Administrative) - LeaseGate — deny if no valid capability lease (priority 10) - ThreadOutcome::GatePaused + EngineError::GatePaused variants Lease system: - LeasePlanner now thread-type-aware (was grant-everything): Foreground=all, Research=read+stateful, Mission=no-admin - derive_child_leases() with intersection semantics for child threads - Children never exceed parent expiry or budget Bridge layer (src/gate + src/bridge): - PendingGateStore: Mutex-based (not RwLock), keyed by (user_id, thread_id) - take_verified(): atomic request_id + channel + expiry check — single lock - GatePersistence trait for restart recovery - TRUSTED_GATE_CHANNELS and RESERVED_CHANNEL_NAMES constants - resolve_gate() public API with auto-approve rollback on resume failure - Concrete gates: ApprovalGate, AuthenticationGate, HookGate, RateLimitGate, RelayChannelGate Python orchestrator: - gate_paused outcome handling for both Tier 0 and Tier 1 paths - Rust-side GatePaused error → {"gate_paused": true} JSON mapping - loop_engine.rs safety net includes GatePaused in Waiting transition 46 new tests across both crates, including regression tests for:74cbe5c2,52d935d7,5d1d504e,427f908e,92138b8c,aa151d9f,e3b66f69,09e1c97a,0e5f1b12, e75fa8c4,49b4c398Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(gate): add integration tests for unified gate lifecycle 17 integration tests covering the full gate abstraction: Engine-level (ThreadManager → EffectExecutor → GatePaused → Waiting): - gate_paused_transitions_thread_to_waiting: Tier 0 tool call → GatePaused outcome → thread state == Waiting (regression:67d5a473) - gate_paused_authentication_carries_credential_name: auth gate carries credential info through the outcome PendingGateStore lifecycle: - pending_gate_full_lifecycle: insert → peek → take_verified → removed - cross_channel_approval_blocked: telegram gate rejected from slack (5d1d504e) - trusted_channel_can_resolve_any_gate: web/gateway bypass channel check - gate_scoped_to_thread_no_leakage: thread A gate invisible to B (e3b66f69) - expired_gate_cannot_be_resolved: TTL enforcement - wrong_request_id_does_not_consume_gate: stale ID doesn't eat gate (74cbe5c2) - concurrent_resolution_exactly_one_succeeds: TOCTOU prevention (52d935d7) - persistence_round_trip_survives_restart: GatePersistence → restore Lease system: - lease_planner_research_excludes_privileged: Research = ReadOnly+Stateful - lease_planner_mission_excludes_denylisted: Mission excludes Administrative - child_lease_inherits_subset_of_parent: intersection semantics - expired_parent_yields_no_child_leases: fail-closed - lease_gate_denies_without_lease / allows_with_valid_lease - pipeline_first_deny_wins: GatePipeline composition Also wires GatePaused through structured.rs and scripting.rs executors so EffectExecutor::execute_action() can return the new error variant. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gate): address review findings — redaction, wildcard lease, TOCTOU, rollback Critical fixes: - C1: Replace .expect() with .ok_or() in take_verified() (production panic) - C2: Redact sensitive params via redact_params() before SSE broadcast and PendingGate storage. Add tools() accessor to EffectBridgeAdapter. - C3: Fix wildcard parent lease (granted_actions=[]) producing wildcard child instead of requested subset. Add regression test wildcard_parent_lease_gives_requested_subset_not_wildcard. Major fixes: - M1: Remove false panic safety documentation from GatePipeline (async catch_unwind impractical with borrowed context). Docs now accurately state gate implementations must not panic. - M2: Batch child lease insertion under single write lock instead of per-iteration locking in derive_child_leases(). - M3: Auto-approve rollback on resume failure now revokes both underscore and hyphenated tool name variants. - M4: Log persistence.remove() failures at debug level instead of silently discarding with let _. - M5: expire_stale() now calls persistence.remove() for each expired gate, preventing indefinite storage accumulation. Minor fixes: - m3: Downgrade gate insert failure from warn! to debug! (AlreadyExists is a normal race condition). - m5: Fix GateContext doc claiming "all fields are borrowed" — ThreadId and ExecutionMode are Copy/inline. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(gate): add InteractiveAutoApprove execution mode Add ExecutionMode::InteractiveAutoApprove for foreground threads with AGENT_AUTO_APPROVE_TOOLS=true. In this mode: - Never tools: allowed (same as all modes) - UnlessAutoApproved tools (shell, file_write, http, etc.): auto-approved without prompting — no approval pause - Always tools (destructive operations): still pause for explicit approval All other safeguards remain active: leases, rate limits, hooks, relay channel checks, authentication gates, parameter redaction. This maps the existing v1 auto_approve_tools config flag into the v2 gate abstraction, providing a "power user" mode where experienced users skip repetitive approval prompts while retaining defense-in-depth. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(cli): add --auto-approve flag for autonomous foreground mode ironclaw run --auto-approve Wires the existing AGENT_AUTO_APPROVE_TOOLS config into a CLI flag via set_runtime_env() (thread-safe override, no unsafe set_var). Activates InteractiveAutoApprove execution mode where: - shell, file_write, http, etc. execute without prompting - Always-gated destructive operations still pause for approval - All other safeguards remain active (leases, rate limits, hooks, auth) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): suppress false positive in check_no_panics for test assert! The CI script's brace-depth tracker loses #[cfg(test)] mod tests {} context when the sanitizer state carries over from earlier string processing. Add // safety: test-only annotation to the affected assert. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Harden engine gate recovery and thread-scoped auth * fix(engine): never delete LLM output data, fix mission thread visibility Mission detail pages showed no threads because cleanup_terminal_state() deleted thread/event/step data from the database. LLM execution data is the most valuable information — it must never be deleted. Changes: - cleanup_terminal_state() now only evicts from in-memory caches, never deletes database rows (threads, events, steps all preserved) - load_thread/load_steps/load_events fall back to database on cache miss - backfill_archived_threads() recovers mission threads on startup from both active DB path and legacy archive summaries - Document "never delete LLM output" principle in CLAUDE.md, engine CLAUDE.md, and database rules - Fix pre-existing compile error in orchestrator (params use-after-move) - Add ApprovalRequested event fields (parameters, description, allow_always, gate_name, params_summary) for richer gate UX Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Refactoring approval gates * Working on improving authentication/approval flows * Updating the implementaton plan * Stabilize engine v2 transcripts and extension tests * fix(engine): address PR #1557 review feedback — security, types, validation - Remove credential-backed HTTP auto-approval bypass (zmanian H2): credential presence no longer skips the approval gate in EffectBridgeAdapter - Add GrantedActions enum (zmanian M1, ilblackdragon #6): replace implicit empty-vec-means-wildcard with explicit All/Specific variants, backward- compatible serde - Validate lease duration/max_uses at grant time (zmanian M2, ilblackdragon #5): reject non-positive durations and zero max_uses - Fix byte/char label mismatch in compact_output_metadata (ilblackdragon #4, zmanian #5): use chars().count() consistently - Add 6 Monty sandbox security negative tests (zmanian C2): OS call denial, file access, socket access, resource limits, lease enforcement, syntax errors - Fix pre-existing clippy warnings in structured.rs and scripting.rs Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: resolve CI failures from staging merge (clippy, no-panics) - router.rs: Box PendingGateResolution::Resolved to fix large_enum_variant, add safety comments for unwrap() calls, simplify Option::map - auth_manager.rs: allow await_holding_lock in tests (env guard must span test) - selector.rs: remove unnecessary double parentheses Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: fix remaining fmt diff in router.rs from staging merge Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: include unstaged merge changes (server.rs 2-arg calls, test updates, snapshots) - server.rs: pass thread_id to clear_engine_pending_auth (2-arg signature) - server.rs: seed workspace on resolve, add test - effect_adapter.rs: update test expectation for approval-before-auth ordering - cli snapshots: add --auto-approve flag Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
262 lines
8.5 KiB
Rust
262 lines
8.5 KiB
Rust
//! Build script: compile Telegram channel WASM from source.
|
|
//!
|
|
//! Do not commit compiled WASM binaries — they are a supply chain risk.
|
|
//! This script builds telegram.wasm from channels-src/telegram before the main crate compiles.
|
|
//!
|
|
//! Reproducible build:
|
|
//! cargo build --release
|
|
//! (build.rs invokes the channel build automatically)
|
|
//!
|
|
//! Prerequisites: rustup target add wasm32-wasip2, cargo install wasm-tools
|
|
|
|
use std::env;
|
|
use std::path::{Path, PathBuf};
|
|
use std::process::Command;
|
|
|
|
fn main() {
|
|
let manifest_dir = env::var("CARGO_MANIFEST_DIR").unwrap();
|
|
let root = PathBuf::from(&manifest_dir);
|
|
|
|
// ── Embed registry manifests ────────────────────────────────────────
|
|
embed_registry_catalog(&root);
|
|
|
|
// ── Embed bundled skills ────────────────────────────────────────────
|
|
embed_skills(&root);
|
|
|
|
// ── Build Telegram channel WASM ─────────────────────────────────────
|
|
let channel_dir = root.join("channels-src/telegram");
|
|
let wasm_out = channel_dir.join("telegram.wasm");
|
|
|
|
// Rerun when channel source or build script changes
|
|
println!("cargo:rerun-if-changed=channels-src/telegram/src");
|
|
println!("cargo:rerun-if-changed=channels-src/telegram/Cargo.toml");
|
|
println!("cargo:rerun-if-changed=wit/channel.wit");
|
|
|
|
if !channel_dir.is_dir() {
|
|
return;
|
|
}
|
|
|
|
// Build WASM module
|
|
let status = match Command::new("cargo")
|
|
.args([
|
|
"build",
|
|
"--release",
|
|
"--target",
|
|
"wasm32-wasip2",
|
|
"--manifest-path",
|
|
channel_dir.join("Cargo.toml").to_str().unwrap(),
|
|
])
|
|
.current_dir(&root)
|
|
.status()
|
|
{
|
|
Ok(s) => s,
|
|
Err(_) => {
|
|
eprintln!(
|
|
"cargo:warning=Telegram channel build failed. Run: ./channels-src/telegram/build.sh"
|
|
);
|
|
return;
|
|
}
|
|
};
|
|
|
|
if !status.success() {
|
|
eprintln!(
|
|
"cargo:warning=Telegram channel build failed. Run: ./channels-src/telegram/build.sh"
|
|
);
|
|
return;
|
|
}
|
|
|
|
let raw_wasm = channel_dir.join("target/wasm32-wasip2/release/telegram_channel.wasm");
|
|
if !raw_wasm.exists() {
|
|
eprintln!(
|
|
"cargo:warning=Telegram WASM output not found at {:?}",
|
|
raw_wasm
|
|
);
|
|
return;
|
|
}
|
|
|
|
// Convert to component and strip (wasm-tools)
|
|
let component_ok = Command::new("wasm-tools")
|
|
.args([
|
|
"component",
|
|
"new",
|
|
raw_wasm.to_str().unwrap(),
|
|
"-o",
|
|
wasm_out.to_str().unwrap(),
|
|
])
|
|
.current_dir(&root)
|
|
.status()
|
|
.map(|s| s.success())
|
|
.unwrap_or(false);
|
|
|
|
if !component_ok {
|
|
// Fallback: copy raw module if wasm-tools unavailable
|
|
if std::fs::copy(&raw_wasm, &wasm_out).is_err() {
|
|
eprintln!("cargo:warning=wasm-tools not found. Run: cargo install wasm-tools");
|
|
}
|
|
} else {
|
|
// Strip debug info (use temp file to avoid clobbering)
|
|
let stripped = wasm_out.with_extension("wasm.stripped");
|
|
let strip_ok = Command::new("wasm-tools")
|
|
.args([
|
|
"strip",
|
|
wasm_out.to_str().unwrap(),
|
|
"-o",
|
|
stripped.to_str().unwrap(),
|
|
])
|
|
.current_dir(&root)
|
|
.status()
|
|
.map(|s| s.success())
|
|
.unwrap_or(false);
|
|
if strip_ok {
|
|
let _ = std::fs::rename(&stripped, &wasm_out);
|
|
}
|
|
}
|
|
}
|
|
|
|
/// Collect all registry manifests into a single JSON blob at compile time.
|
|
///
|
|
/// Output: `$OUT_DIR/embedded_catalog.json` with structure:
|
|
/// ```json
|
|
/// { "tools": [...], "channels": [...], "bundles": {...} }
|
|
/// ```
|
|
fn embed_registry_catalog(root: &Path) {
|
|
use std::fs;
|
|
|
|
let registry_dir = root.join("registry");
|
|
|
|
// Rerun if the bundles file changes (per-file watches for tools/channels
|
|
// are emitted inside collect_json_files to track content changes reliably).
|
|
println!("cargo:rerun-if-changed=registry/_bundles.json");
|
|
|
|
let out_dir = PathBuf::from(env::var("OUT_DIR").unwrap()); // safety: build script
|
|
let out_path = out_dir.join("embedded_catalog.json");
|
|
|
|
if !registry_dir.is_dir() {
|
|
// No registry dir: write empty catalog
|
|
fs::write(
|
|
&out_path,
|
|
r#"{"tools":[],"channels":[],"mcp_servers":[],"bundles":{"bundles":{}}}"#,
|
|
)
|
|
.unwrap();
|
|
return;
|
|
}
|
|
|
|
let mut tools = Vec::new();
|
|
let mut channels = Vec::new();
|
|
let mut mcp_servers = Vec::new();
|
|
|
|
// Collect tool manifests
|
|
let tools_dir = registry_dir.join("tools");
|
|
if tools_dir.is_dir() {
|
|
collect_json_files(&tools_dir, &mut tools);
|
|
}
|
|
|
|
// Collect channel manifests
|
|
let channels_dir = registry_dir.join("channels");
|
|
if channels_dir.is_dir() {
|
|
collect_json_files(&channels_dir, &mut channels);
|
|
}
|
|
|
|
// Collect MCP server manifests
|
|
let mcp_servers_dir = registry_dir.join("mcp-servers");
|
|
if mcp_servers_dir.is_dir() {
|
|
collect_json_files(&mcp_servers_dir, &mut mcp_servers);
|
|
}
|
|
|
|
// Read bundles
|
|
let bundles_path = registry_dir.join("_bundles.json");
|
|
let bundles_raw = if bundles_path.is_file() {
|
|
fs::read_to_string(&bundles_path).unwrap_or_else(|_| r#"{"bundles":{}}"#.to_string())
|
|
} else {
|
|
r#"{"bundles":{}}"#.to_string()
|
|
};
|
|
|
|
// Build the combined JSON
|
|
let catalog = format!(
|
|
r#"{{"tools":[{}],"channels":[{}],"mcp_servers":[{}],"bundles":{}}}"#,
|
|
tools.join(","),
|
|
channels.join(","),
|
|
mcp_servers.join(","),
|
|
bundles_raw,
|
|
);
|
|
|
|
fs::write(&out_path, catalog).unwrap(); // safety: build script
|
|
}
|
|
|
|
/// Collect all `skills/*/SKILL.md` files into an embedded JSON blob.
|
|
///
|
|
/// Output: `$OUT_DIR/embedded_skills.json` — a JSON array of `{"name": "...", "content": "..."}`.
|
|
/// These are loaded at runtime as bundled skills (lowest discovery priority, Trusted trust level).
|
|
fn embed_skills(root: &Path) {
|
|
use std::fs;
|
|
|
|
let skills_dir = root.join("skills");
|
|
|
|
// Rerun when any skill changes
|
|
println!("cargo:rerun-if-changed=skills");
|
|
|
|
let out_dir = PathBuf::from(env::var("OUT_DIR").unwrap()); // safety: build script panics on failure
|
|
let out_path = out_dir.join("embedded_skills.json");
|
|
|
|
if !skills_dir.is_dir() {
|
|
fs::write(&out_path, "[]").unwrap(); // safety: build script
|
|
return;
|
|
}
|
|
|
|
let mut skills: Vec<String> = Vec::new();
|
|
|
|
let mut entries: Vec<_> = fs::read_dir(&skills_dir)
|
|
.unwrap() // safety: build script
|
|
.filter_map(|e| e.ok())
|
|
.filter(|e| e.path().is_dir())
|
|
.collect();
|
|
entries.sort_by_key(|e| e.file_name());
|
|
|
|
for entry in entries {
|
|
let skill_md = entry.path().join("SKILL.md");
|
|
if !skill_md.is_file() {
|
|
continue;
|
|
}
|
|
// Emit per-file watch
|
|
println!("cargo:rerun-if-changed={}", skill_md.display());
|
|
|
|
let name = entry.file_name().to_string_lossy().to_string();
|
|
if let Ok(content) = fs::read_to_string(&skill_md) {
|
|
// Escape for JSON embedding
|
|
let name_json = serde_json::to_string(&name).unwrap(); // safety: build script
|
|
let content_json = serde_json::to_string(&content).unwrap(); // safety: build script
|
|
skills.push(format!(
|
|
r#"{{"name":{},"content":{}}}"#,
|
|
name_json, content_json
|
|
));
|
|
}
|
|
}
|
|
|
|
let catalog = format!("[{}]", skills.join(","));
|
|
fs::write(&out_path, catalog).unwrap(); // safety: build script
|
|
}
|
|
|
|
/// Read all .json files from a directory and push their raw contents into `out`.
|
|
fn collect_json_files(dir: &Path, out: &mut Vec<String>) {
|
|
use std::fs;
|
|
|
|
let mut entries: Vec<_> = fs::read_dir(dir)
|
|
.unwrap()
|
|
.filter_map(|e| e.ok())
|
|
.filter(|e| {
|
|
e.path().is_file() && e.path().extension().and_then(|x| x.to_str()) == Some("json")
|
|
})
|
|
.collect();
|
|
|
|
// Sort for deterministic output
|
|
entries.sort_by_key(|e| e.file_name());
|
|
|
|
for entry in entries {
|
|
// Emit per-file watch so Cargo reruns when file contents change
|
|
println!("cargo:rerun-if-changed={}", entry.path().display());
|
|
if let Ok(content) = fs::read_to_string(entry.path()) {
|
|
out.push(content);
|
|
}
|
|
}
|
|
}
|