Files
ironclaw/src/orchestrator/mod.rs
Zaki Manian 8d052a9eb5 fix(security): scope orchestrator credentials to job creator (#2068) (#2698)
* fix(security): remove cross-tenant credential fallbacks in orchestrator, WASM, and channels (#2068, #2069, #2100)

Three credential isolation fixes that prevent cross-tenant secret leakage:

- Orchestrator: get_credentials_handler now resolves the job creator's
  user_id from job_owner_cache (or DB fallback) instead of using a
  hardcoded global owner_id. Returns 403 when owner cannot be resolved.
  Removes the user_id field from OrchestratorState entirely.

- WASM tools: resolve_host_credentials uses DefaultFallback::Denied
  instead of AdminOnly, preventing any user's WASM tool from falling
  back to "default" scope credentials.

- Channel broadcast metadata: removes legacy migration fallback that
  read broadcast metadata from "default" scope. Channels re-persist
  metadata under the correct owner scope on next incoming message.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(security): extract resolve_job_owner, bound cache, per-job credentials

Address review feedback:
- Extract resolve_job_owner() to DRY up cache-then-DB resolution
- Bound job_owner_cache to 10K entries with batch eviction
- Add register_job_owner() for pre-population at job creation
- get_credentials_handler uses per-job owner instead of global
  state.user_id, preventing cross-tenant credential leakage

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(security): filter empty user_id from cache, unify error codes

Address follow-up review feedback:
- Filter empty user_id before caching to prevent poisoned entries
- Map secret decrypt failures to 403 (not 500) to avoid info leak
  distinguishing "secret missing for user" from "owner unknown"
- register_job_owner is available for callers that have both the
  cache and user_id; DB is required for sandbox credential injection

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: fix rustfmt line-length violation in orchestrator api

Break long method chain in cache eviction across multiple lines to pass
`cargo fmt --check`.

https://claude.ai/code/session_01JRasj3ujmr1uzmfUeLbNFo

* fix(review): address PR feedback — fix test, drop dead helper, bump log level

- Update credentials_uses_job_creator_not_other_user to assert 403 FORBIDDEN.
  The prior assertion of 500 INTERNAL_SERVER_ERROR contradicted the same
  PR's change that mapped all secret-lookup failures to FORBIDDEN, so the
  test failed to even compile-as-regression. Also expand the comment to
  explain why uniform 403 is the correct wire response here.
- Remove register_job_owner: the helper had zero call sites. The cache is
  self-warming because resolve_job_owner inserts on every DB fallback, so
  an explicit registration hook would only save one DB hit on the first
  SSE event of a job. Wiring it into ContainerJobManager::create_job is a
  larger refactor; file a follow-up if eager warming is worth the cost.
- Update job_owner_cache doc comment to describe lazy population — the
  previous "populated when sandbox jobs are created" claim was aspirational.
- Fix MAX_JOB_OWNER_CACHE_SIZE comment: HashMap eviction is not LRU/FIFO.
  Note IndexMap/lru::LruCache as upgrade options if recency matters.
- Bump decrypt-failure log from debug to warn, add env_var for operability.
  Keeps 403 wire response (no existence-leak to the caller) but restores
  operator visibility for real crypto/keychain failures.
- Annotate the job_event_handler unwrap_or_default with a silent-ok comment
  per the error-handling rule — SSE broadcast is best-effort and the empty
  user_id path is already handled below.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
2026-04-22 02:11:13 +09:00

207 lines
8.0 KiB
Rust

//! Orchestrator for managing sandboxed worker containers.
//!
//! The orchestrator runs in the main agent process and provides:
//! - An internal HTTP API for worker communication (LLM proxy, status, secrets)
//! - Per-job bearer token authentication
//! - Container lifecycle management (create, monitor, stop)
//!
//! ```text
//! ┌───────────────────────────────────────────────┐
//! │ Orchestrator │
//! │ │
//! │ Internal API (default :50051, configurable) │
//! │ POST /worker/{id}/llm/complete │
//! │ POST /worker/{id}/llm/complete_with_tools │
//! │ GET /worker/{id}/job │
//! │ GET /worker/{id}/credentials │
//! │ POST /worker/{id}/status │
//! │ POST /worker/{id}/complete │
//! │ │
//! │ ContainerJobManager │
//! │ create_job() -> container + token │
//! │ stop_job() │
//! │ list_jobs() │
//! │ │
//! │ TokenStore │
//! │ per-job bearer tokens (in-memory only) │
//! │ per-job credential grants (in-memory only) │
//! └───────────────────────────────────────────────┘
//! ```
pub mod api;
pub mod auth;
pub mod job_manager;
pub mod reaper;
pub use api::OrchestratorApi;
pub use auth::{CredentialGrant, TokenStore};
pub use job_manager::{
CompletionResult, ContainerHandle, ContainerJobConfig, ContainerJobManager, JobMode,
};
pub use reaper::{ReaperConfig, SandboxReaper};
use std::collections::{HashMap, VecDeque};
use std::sync::Arc;
use tokio::sync::{Mutex, broadcast};
use uuid::Uuid;
use crate::db::Database;
use crate::llm::LlmProvider;
use crate::secrets::SecretsStore;
use ironclaw_common::AppEvent;
/// Resolve the orchestrator port from the `ORCHESTRATOR_PORT` environment
/// variable, falling back to 50051.
fn resolve_orchestrator_port() -> u16 {
std::env::var("ORCHESTRATOR_PORT")
.ok()
.and_then(|v| v.parse().ok())
.unwrap_or(50051)
}
/// Result of orchestrator setup, containing all handles needed by the agent.
pub struct OrchestratorSetup {
pub container_job_manager: Option<Arc<ContainerJobManager>>,
pub job_event_tx: Option<broadcast::Sender<(Uuid, String, AppEvent)>>,
pub prompt_queue: Arc<Mutex<HashMap<Uuid, VecDeque<api::PendingPrompt>>>>,
pub docker_status: crate::sandbox::DockerStatus,
}
/// Detect Docker availability, create the container job manager, and start
/// the orchestrator internal API in the background.
pub async fn setup_orchestrator(
config: &crate::config::Config,
llm: &Arc<dyn LlmProvider>,
db: Option<&Arc<dyn Database>>,
secrets_store: Option<&Arc<dyn SecretsStore + Send + Sync>>,
) -> OrchestratorSetup {
let prompt_queue = Arc::new(Mutex::new(
HashMap::<Uuid, VecDeque<api::PendingPrompt>>::new(),
));
let docker_status = if config.sandbox.enabled {
let detection = crate::sandbox::check_docker().await;
match detection.status {
crate::sandbox::DockerStatus::Available => {
tracing::info!("Docker is available");
}
crate::sandbox::DockerStatus::NotInstalled => {
tracing::warn!(
"Docker is not installed -- sandbox disabled for this session. {}",
detection.platform.install_hint()
);
}
crate::sandbox::DockerStatus::NotRunning => {
tracing::warn!(
"Docker is installed but not running -- sandbox disabled for this session. {}",
detection.platform.start_hint()
);
}
crate::sandbox::DockerStatus::Disabled => {}
}
detection.status
} else {
crate::sandbox::DockerStatus::Disabled
};
let (job_event_tx, container_job_manager) = if config.sandbox.enabled && docker_status.is_ok() {
let (tx, _) = broadcast::channel(256);
let job_event_tx = Some(tx);
let token_store = TokenStore::new();
let orchestrator_port = resolve_orchestrator_port();
let job_config = ContainerJobConfig {
image: config.sandbox.image.clone(),
memory_limit_mb: config.sandbox.memory_limit_mb,
cpu_shares: config.sandbox.cpu_shares,
orchestrator_port,
claude_code_api_key: std::env::var("ANTHROPIC_API_KEY").ok(),
claude_code_oauth_token: crate::config::ClaudeCodeConfig::extract_oauth_token(),
claude_code_model: config.claude_code.model.clone(),
claude_code_max_turns: config.claude_code.max_turns,
claude_code_memory_limit_mb: config.claude_code.memory_limit_mb,
claude_code_allowed_tools: config.claude_code.allowed_tools.clone(),
acp_memory_limit_mb: config.acp.memory_limit_mb,
acp_timeout_secs: config.acp.timeout_secs,
mcp_per_job_enabled: std::env::var("MCP_PER_JOB_ENABLED")
.map(|v| v.eq_ignore_ascii_case("true") || v == "1")
.unwrap_or(false),
claude_code_enabled: config.claude_code.enabled,
acp_enabled: config.acp.enabled,
};
let jm = Arc::new(ContainerJobManager::new(job_config, token_store.clone()));
let orchestrator_state = api::OrchestratorState {
llm: Arc::clone(llm),
job_manager: Arc::clone(&jm),
token_store,
job_event_tx: job_event_tx.clone(),
prompt_queue: Arc::clone(&prompt_queue),
store: db.cloned(),
secrets_store: secrets_store.cloned(),
job_owner_cache: Arc::new(std::sync::RwLock::new(std::collections::HashMap::new())),
};
tokio::spawn(async move {
if let Err(e) = OrchestratorApi::start(orchestrator_state, orchestrator_port).await {
tracing::error!("Orchestrator API failed: {}", e);
}
});
if config.claude_code.enabled {
tracing::info!(
"Claude Code sandbox mode available (model: {}, max_turns: {})",
config.claude_code.model,
config.claude_code.max_turns
);
}
if config.acp.enabled {
tracing::info!("ACP agent sandbox mode available");
}
(job_event_tx, Some(jm))
} else {
(None, None)
};
OrchestratorSetup {
container_job_manager,
job_event_tx,
prompt_queue,
docker_status,
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::config::helpers::lock_env;
#[test]
fn resolve_orchestrator_port_from_env() {
let _guard = lock_env();
// Safety: env-var mutation requires unsafe in edition 2024;
// lock_env() serializes concurrent access from other test threads.
// Absent env var → default 50051
unsafe { std::env::remove_var("ORCHESTRATOR_PORT") };
assert_eq!(resolve_orchestrator_port(), 50051);
// Valid custom port
unsafe { std::env::set_var("ORCHESTRATOR_PORT", "50052") };
assert_eq!(resolve_orchestrator_port(), 50052);
// Non-numeric value → fallback to default
unsafe { std::env::set_var("ORCHESTRATOR_PORT", "not_a_port") };
assert_eq!(resolve_orchestrator_port(), 50051);
// Out of u16 range → fallback to default
unsafe { std::env::set_var("ORCHESTRATOR_PORT", "99999") };
assert_eq!(resolve_orchestrator_port(), 50051);
// Cleanup
unsafe { std::env::remove_var("ORCHESTRATOR_PORT") };
}
}