* feat(fleet): your fleet is the models you added, and it comes first
Design MODEL-ROUTING-CATALOG-20260901 §10, slice F1. A fleet model is a Pod
member: the selected Pod file's operator route plus every member row that
pins an exact provider + model; the roles a model fills are the member rows
that pin it. No second store.
- crate::fleet::members: fleet_models / add_fleet_model / remove_fleet_model
/ toggle_fleet_model + change_receipt; Config::fleet_members(workspace) is
the read seam for the operator-awareness slice (F2).
- /pod models | add <provider> <model> [role…] | remove <provider> <model>
(also via the /fleet alias). A model the configured provider does not
serve is rejected; the first add creates and selects a user-global Pod
named 'My fleet'.
- /model picker: ⇧F adds or removes the row's exact route; fleet models
lead the list labelled 'fleet · <roles>', ahead of ⇧P pins and providers.
- /models prints the fleet before the provider list ('Your fleet is the
session model only' when empty).
- PickerActionFleet message in all 15 locales; docs/FLEET.md 'Your fleet
as models'.
Tests: scripts/dev-test.sh tui fleet::members groups::core::fleet
model_picker format_helpers — Summary 37 tests run: 37 passed, 11834
skipped.
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
* fix(fleet): pass slugify by name (clippy redundant_closure)
cargo clippy -p codewhale-tui --all-targets -- -D warnings -A clippy::too_many_arguments -A clippy::uninlined_format_args -A clippy::unnecessary_map_or: no findings.
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
* fix(tui): review fixes for fleet toggle and /pod add provider validation\n\n- Reject unconfigured provider ids in "/pod add" before writing, reusing\n the existing provider_is_configured_for_active predicate and custom\n provider table checks.\n- Add App.config snapshot so commands can consult the loaded config.\n- Update the stale DEFAULT_FLEET_NAME doc comment to mention ⇧F.\n- Sync crates/tui/CHANGELOG.md.
* style: cargo fmt
* fix(web): align react with react-dom 19.2.8 to unbreak npm ci
Dependabot #5801 bumped react-dom to 19.2.8, whose peer range requires
react 19.2.8; the lockfile still resolved react 19.2.6, so 'npm ci' in
web/ failed ERESOLVE on main and on every branch that merged it
(Lint & Type Check red). Align react to 19.2.8; install verified clean.
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
* brand: trace supplied whale assets
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* brand: align icon ombre and generated tokens
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* brand: use white icon tile
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* brand: deepen ombre light stop
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* brand: wordmark takes the blue ombre
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* tui: recover from image-input rejections by non-vision routes
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* tui: localize image rejection recovery
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* chore: format 0.9.12 mega branch
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* Redesign Fleet role labels and agent cards
* feat(tui): launch hero as wordmark + small surfacing mark
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* design(tui): retune whale palette to codewhale navy / ombre sky
Field, chrome, panel, plate and raised surfaces move onto the brand navy
(#070C1D → #142352 → #1A2C63); interaction blue becomes the ombre sky
#6AA6DC, light-mode action the ombre cobalt #1535B2; ice/cyan/border/tool
tints follow. web/app/tokens.css regenerated via
scripts/export-design-tokens.py.
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* test(tui): re-bless ink goldens for navy palette
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* web: Space Mono wordmark, quiet layout refresh, fleet vocabulary in site + docs
Space Mono (OFL) outlined wordmark rebuilt via scripts/build-wordmark.py,
wired as --font-display through next/font/google; body stays IBM Plex Sans,
code stays JetBrains Mono. Nav loses the issue strip, strapline, Discord badge
and second filled CTA; home loses the ticker, seals and tilt figure; docs
shell hero collapses to a one-line band; footer uses the inverted wordmark.
Public noun is fleet (/fleet, codewhale fleet, /docs/fleet canonical; /pod,
codewhale pod, /docs/pod remain compatibility aliases) across docs/, site
dictionaries, vocabulary contract and public-surface facts.
No-Issue: 0.9.12 website lane
* brand: keep the traced wordmark; drop Space Mono outline build
* web: IBM Plex Sans Condensed as display face
* brand: Plex Sans Condensed wordmark; nav mark; drop fabricated home demos; AA meta text
* feat(tui): bottom dock tabs — clickable panel switch + close
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* Fix Fleet role migration verification
* fix(tui): dock keys yield Tab to mode/permission cycles
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>
* web: Impeccable polish — type floors, heading outline, docs measure; add PRODUCT.md/DESIGN.md
* Resolve canonical Fleet roles to legacy members
* web: flat hero — drop cyan glow/gradients/shadow, muted eyebrows
* design: PRODUCT.md/DESIGN.md at repo root — shell direction, bottom dock, anti-slop rules, 0.9.12 tokens
* Auto-enroll used models into the Fleet
* brand: keep the founder's wordmark rasters as the source of truth
The web lane replaced brand/wordmark0901.png and brand/wordmarkinverted.png
with Plex Sans Condensed renders. The founder-supplied PNGs are the brand
source; the SVGs are re-traced from them in a following commit.
* tui(mark): the launch mark has one rung
The hero now paints the small mark over the wordmark, so the medium and
large rungs and the for_area ladder have no consumer and fail the
dead-code lint. Remove them rather than allow them.
* brand: trace the founder's wordmark to SVG
brand/wordmark.svg and wordmark-inverted.svg were an IBM Plex Sans
Condensed text render; the founder's wordmark is the rounded monoline in
brand/wordmark0901.png. scripts/brand/trace-brand.py now traces that PNG
(magick threshold 60% + trim, potrace -s --flat -t 20 -O 0.4 -a 1.2),
folds potrace's transform into one compact path in a tight 1874x264
viewBox, and writes the navy #142352 and white colourways from the same
geometry. The Plex builder scripts/build-wordmark.py is gone with it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* web: derive icons and OG image from the traced mark
app/icon.svg is now the white whale on the #142352 rounded tile as on the
founder's sheet; favicon.ico (48/32/16), apple-icon.png, icon-192.png and
icon-512.png are rasterised from it by scripts/brand/trace-brand.py, and
the manifest colours are the same navy. The social card keeps the navy
ground, white mark and traced wordmark and restores the identity phrase
the page-meta contract expects.
The nav sits on the dark field on every route, so it renders the inline
Whale (white brand ink) and the inverted wordmark instead of a
prefers-color-scheme picture pair; the wordmark box uses object-fit so
the ~7.1:1 trace scales inside the compact nav instead of squashing.
Exploration rasters web/public/brand/codewhale-mark-*.png and their
web/brand/mark tile sources had no consumers and are removed;
codewhale-mark.png stays (public-auth-routes pins its hash).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* web: map stray hard-coded colours to navy tokens
globals.css carried a handful of literal navy-family greys, ice hairlines,
seafoam borders and a cyan glow beside the generated --whale-* tokens.
Each now reads the token it was approximating (whale-bg/chrome/panel,
whale-ice, whale-accent-secondary, whale-action, whale-cyan,
whale-text-dim), and the docs light sheet inks the mark in the brand
navy via --whale-composer (#142352).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* palette: inventory WHALE_* tokens before the one-name-per-colour collapse
Shell design §2.6 (SHELL-DESIGN-20260901) measured "58 WHALE_* symbols;
one colour under five names; 5 dead tokens". Receipt before touching
anything, generated from crates/tui/src/palette/tokens.rs. "uses" is the
whole-word count across crates/ excluding the const's own definition and
`use`/`pub use` lines (wrapper consts inside tokens.rs count).
name value alias-of dead uses
WHALE_BG_RGB (7, 12, 29) 3
WHALE_CHROME_RGB (12, 21, 49) 1
WHALE_PANEL_RGB (16, 28, 64) 3
WHALE_COMPOSER_RGB (20, 35, 82) 3
WHALE_ELEVATED_RGB (26, 44, 99) 4
WHALE_SELECTION_RGB (30, 60, 143) 3
WHALE_TEXT_BODY_RGB (246, 242, 232) 10
WHALE_TEXT_SOFT_RGB (182, 192, 212) 4
WHALE_TEXT_MUTED_RGB (147, 160, 184) 3
WHALE_TEXT_HINT_RGB (138, 153, 179) 3
WHALE_TEXT_DIM_RGB (105, 119, 145) yes 0
WHALE_ACTION_RGB (106, 166, 220) 6
WHALE_COBALT_RGB (21, 53, 178) yes 0 (web: --whale-cobalt x3)
WHALE_ICE_RGB (221, 238, 249) yes 0 (web: --whale-ice, rustRgb("WHALE_ICE"))
WHALE_CYAN_RGB (120, 188, 232) 2
WHALE_ACCENT_SECONDARY_RGB (79, 209, 197) 11
WHALE_BRAND_ORANGE_RGB (255, 138, 61) 1
WHALE_BRAND_MAGENTA_RGB (240, 78, 184) 1
WHALE_HUMAN_RGB (246, 196, 83) 5
WHALE_ACCENT_PRIMARY_RGB = WHALE_ACTION_RGB WHALE_ACTION_RGB 8
WHALE_WORKING_GREEN_RGB (155, 214, 111) 5
WHALE_ACCENT_ACTION_RGB = WHALE_ACTION_RGB WHALE_ACTION_RGB yes 0
WHALE_ERROR_RGB (255, 134, 178) 9
WHALE_ERROR_HOVER_RGB (255, 156, 194) 3
WHALE_ERROR_SURFACE_RGB (43, 21, 34) 6
WHALE_ERROR_BORDER_RGB = WHALE_ERROR_RGB WHALE_ERROR_RGB 3
WHALE_ERROR_TEXT_RGB (255, 219, 232) 3
WHALE_WARNING_RGB (255, 122, 89) 4
WHALE_SUCCESS_RGB = WHALE_WORKING_GREEN_RGB WHALE_WORKING_GREEN_RGB 9
WHALE_INFO_RGB = WHALE_ACTION_RGB WHALE_ACTION_RGB 17
WHALE_BORDER_RGB (42, 63, 114) 1
WHALE_REASONING_TEXT_RGB (224, 153, 72) 13
WHALE_REASONING_SURFACE_RGB (42, 34, 24) 3
WHALE_REASONING_TINT_RGB (22, 36, 74) 7
WHALE_DIFF_ADDED_RGB (87, 199, 133) 3
WHALE_DIFF_DELETED_RGB = WHALE_ERROR_RGB WHALE_ERROR_RGB yes 0
WHALE_DIFF_ADDED_BG_RGB (18, 42, 34) 3
WHALE_DIFF_DELETED_BG_RGB (52, 24, 39) 3
WHALE_MODE_AGENT_RGB (126, 180, 232) 4 (via MODE_AGENT: 13)
WHALE_MODE_YOLO_RGB (255, 112, 160) 4 (via MODE_YOLO: 13)
WHALE_MODE_PLAN_RGB (185, 220, 236) 4 (via MODE_PLAN: 13)
WHALE_MODE_OPERATE_RGB (173, 136, 255) 4 (via MODE_OPERATE: 14)
WHALE_TOOL_LIVE_RGB = WHALE_ACCENT_SECONDARY_RGB 3 (via ACCENT_TOOL_LIVE: 5)
WHALE_TOOL_ISSUE_RGB = WHALE_ERROR_RGB 3 (via ACCENT_TOOL_ISSUE: 5)
WHALE_TOOL_OUTPUT_RGB = WHALE_TEXT_SOFT_RGB 3 (via TEXT_TOOL_OUTPUT: 13)
WHALE_TOOL_SURFACE_RGB (15, 26, 58) 3 (via SURFACE_TOOL: 5)
WHALE_TOOL_ACTIVE_RGB (24, 44, 94) 3 (via SURFACE_TOOL_ACTIVE: 9)
WHALE_ACCENT_PRIMARY Color(WHALE_ACCENT_PRIMARY_RGB) -> WHALE_ACTION 9
WHALE_ACTION Color(WHALE_ACTION_RGB) 85
WHALE_LIVE Color(WHALE_ACCENT_SECONDARY_RGB) 17
WHALE_HUMAN Color(WHALE_HUMAN_RGB) 41
WHALE_INFO Color(WHALE_INFO_RGB) -> WHALE_ACTION 105
WHALE_BG Color(WHALE_BG_RGB) 96
WHALE_CHROME Color(WHALE_CHROME_RGB) 5
WHALE_PANEL Color(WHALE_PANEL_RGB) 15
WHALE_COMPOSER Color(WHALE_COMPOSER_RGB) 5
WHALE_ERROR Color(WHALE_ERROR_RGB) 35
57 WHALE_* consts. Pure aliases (9): ACCENT_PRIMARY_RGB, ACCENT_ACTION_RGB,
ERROR_BORDER_RGB, SUCCESS_RGB, INFO_RGB, DIFF_DELETED_RGB, TOOL_LIVE_RGB,
TOOL_ISSUE_RGB, TOOL_OUTPUT_RGB. #[expect(dead_code)] (5): TEXT_DIM_RGB,
COBALT_RGB, ICE_RGB, ACCENT_ACTION_RGB, DIFF_DELETED_RGB.
Non-WHALE aliases of the same blue in tokens.rs: STATUS_INFO (8 uses),
ACCENT_PRIMARY (dead, 0). One colour, #6AA6DC, under seven symbols:
WHALE_ACTION(_RGB), WHALE_INFO(_RGB), WHALE_ACCENT_PRIMARY(_RGB),
WHALE_ACCENT_ACTION_RGB, STATUS_INFO, ACCENT_PRIMARY — 225 call sites.
Script: python3 over tokens.rs + grep -rnw crates; kept out of scripts/
(one-off receipt, the numbers live here).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* palette: one name per colour — collapse WHALE_INFO / WHALE_ACCENT_PRIMARY into WHALE_ACTION
Shell design §2.6: "`WHALE_INFO`, `WHALE_ACTION`, `WHALE_ACCENT_PRIMARY`
and their `_RGB` twins are one colour. Collapse to `WHALE_ACTION`."
Mechanical rename across crates/ (word-boundary sed, no value changes):
WHALE_INFO, WHALE_INFO_RGB -> WHALE_ACTION, WHALE_ACTION_RGB
WHALE_ACCENT_PRIMARY(_RGB) -> WHALE_ACTION(_RGB)
palette::STATUS_INFO -> palette::WHALE_ACTION
WHALE_ACCENT_ACTION_RGB, ACCENT_PRIMARY (dead aliases) -> deleted
The `STATUS_INFO` static in commands/groups/config is an unrelated
CommandInfo and is untouched.
Where two names met in one predicate (adapt.rs light/solarized/community
remaps, grayscale text-soft bucket, SemanticForegroundRole::Action) the
duplicate disjuncts are dropped; `use` lists deduped; the
"primary accent aligns with action" test collapses to its one live
assertion (action blue != human gold). The Blue Stage doc comment moves
onto WHALE_ACTION_RGB. `palette::grammar` untouched: it reads UiTheme
slots, never these consts.
Evidence (CARGO_BUILD_BUILD_DIR=.../mega-tokens):
cargo check -p codewhale-tui --all-targets -> Finished, 0 warnings
cargo clippy -p codewhale-tui --all-targets --all-features --locked
-- -D warnings (CI allow-list) -> clean
cargo test -p codewhale-tui --lib palette::tests:: --locked -- --skip command_palette
-> test result: ok. 59 passed; 0 failed
RUST_MIN_STACK=67108864 cargo test -p codewhale-tui --lib startup_ink --locked
-> test result: ok. 1 passed; 0 failed (ink goldens unchanged)
cargo test -p codewhale-tui --lib --locked -- menu_style cursor_accent color_compat
-> test result: ok. 36 passed; 0 failed
Pre-existing, not from this diff (reproduced on the stashed tree):
tui::command_palette tests, feat012_ac1 and the startup_ink golden
overflow the default test-thread stack in a debug build; they pass with
RUST_MIN_STACK=64MiB.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* palette: delete the dead whale tokens
Shell design §2.6: "delete the five dead tokens". Two of the five went
with the alias collapse (WHALE_ACCENT_ACTION_RGB, ACCENT_PRIMARY); this
removes the rest that have no consumer in crates/ and no web consumer:
WHALE_TEXT_DIM_RGB (105,119,145) 0 uses, no --whale-text-dim on the site
WHALE_DIFF_DELETED_RGB = WHALE_ERROR_RGB 0 uses, no --whale-diff-deleted
ACCENT_SECONDARY Color(WHALE_ACCENT_SECONDARY_RGB) 0 uses (TEXT_ACCENT/WHALE_LIVE carry it)
STATUS_NEUTRAL = TEXT_MUTED 0 uses
Kept, with a comment saying why the `#[expect(dead_code)]` is honest:
WHALE_COBALT_RGB and WHALE_ICE_RGB have no TUI consumer but the site
reads them through the token export (`--whale-cobalt` x3, `--whale-ice`
and `rustRgb("WHALE_ICE")` in web/lib/blue-stage-contract.test.ts).
Mode and tool-surface `_RGB` tuples stay: each is consumed through its
Color wrapper (MODE_AGENT/YOLO/PLAN/OPERATE 13-14 call sites each,
themes.rs + color_compat.rs; SURFACE_TOOL 5, SURFACE_TOOL_ACTIVE 9,
ACCENT_TOOL_LIVE 5, ACCENT_TOOL_ISSUE 5, TEXT_TOOL_OUTPUT 13). The §1
"12 tokens with zero consumers" counted the tuples, not their wrappers.
Evidence: cargo check -p codewhale-tui --all-targets -> Finished, 0 warnings;
cargo test -p codewhale-tui --lib palette::tests:: --locked -- --skip command_palette
-> test result: ok. 59 passed; 0 failed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* web: regenerate tokens.css after the whale token collapse
scripts/export-design-tokens.py (never hand-edited). Ten lines gone:
--whale-accent-primary(-rgb), --whale-accent-action(-rgb),
--whale-info(-rgb), --whale-text-dim(-rgb), --whale-diff-deleted(-rgb).
No site stylesheet or component consumed any of them
(grep -rn "\-\-whale-" web/app web/components web/lib); the only
reference was the alias-chain example in web/lib/whale-tokens.ts's doc
comment, now `--whale-success` -> `--whale-working-green` -> `#9bd66f`
(the old example also quoted a hex that stopped being true a retune ago).
Evidence:
cd web && python3 ../scripts/export-design-tokens.py --check (CI: npm run check:tokens)
-> design tokens up to date (1 file(s), 42 tokens)
vitest run lib/blue-stage-contract lib/docs-theme-contract
-> Test Files 2 passed (2) / Tests 6 passed (6)
(vitest ran against the main checkout's node_modules via a temporary
symlink; this worktree has none installed.)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* docs(design): the status-bar grammar names the one blue token
STATUS_BAR_COLOR_GRAMMAR.md never named a collapsed token, so this is the
one sentence it needed: the Identity blue is `WHALE_ACTION`, its former
aliases (`WHALE_INFO`, `WHALE_ACCENT_PRIMARY`, `STATUS_INFO`) are gone, and
the whale theme's `info` / `accent_primary` slots both hold it. No other
document in the repo named them (grep over *.md, *.ts, *.tsx, *.css,
*.py, *.toml, *.yml, *.json outside node_modules); the root DESIGN.md
already speaks in CSS names.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* palette: the field follows the terminal-owned shell; `underwater` aliases deepsea
Shell design §2.0 decision 1 (founder: "We aren't supposed to be using a
blue background anymore"): ground is the terminal's; the navy field is
painted only under the opt-in deepsea column.
What was already true, verified before changing anything:
- Settings::default().theme is "terminal" (settings.rs:103).
- The whale pair (UI_THEME / LIGHT_UI_THEME) ends in
`.with_terminal_native_shell()`: surface, panel, composer, header and
footer are `Color::Reset`, pinned by
`whale_pair_flat_shells_are_terminal_native_without_erasing_semantic_surfaces`.
- The ink goldens' legend reads `a reset on reset`.
- OceanTreatment::Flat is the default; Deepsea repaints Reset cells through
OceanRamp::for_theme, which matches the whale pair by name + Reset shell.
The reviewer's citations (tokens.rs:6/:250/:465) are the token definitions
deepsea and the semantic surfaces still need, not the theme.
What was not true: ~90 direct `bg(palette::WHALE_BG)` paints in pickers,
overlays and full-screen views (provider_picker 14, views/mod.rs 11,
user_input, live_transcript, help, session/file/model pickers ...) bypass
the theme, and `adapt_bg_for_theme` only remapped them for
`theme_remap_active` presets. On the whale theme they laid navy patches
over the terminal ground. Rung 2 fix, one rule in palette/adapt.rs: the
field (`WHALE_BG` / `BACKGROUND_DARK`) always follows `ui.surface_bg` —
Reset on the whale pair, the user's `background_color` override when set,
the preset surface elsewhere. Panels, selection, elevation, error and
diff surfaces are untouched; no widget file changes.
`underwater` is now an accepted alias of `deepsea` in
settings.rs (normalize + `set`), OceanTreatment::parse and the
config_ui serde enum. Tests extended in place; the color_compat light
test now expects the Reset shell it already had for theme consumers.
DESIGN.md "Field" says the TUI ground is the terminal's own background
and the navy field is deepsea-only.
Contrast, all whale text/accent tokens on #000, #1e1e1e (VS Code),
#282c34 (One Dark), #002b36 (Solarized dark), #300a24 (Ubuntu), #0d1117
(GitHub dark), #282a36 (Dracula): body 12.5-18.8:1, soft 7.7-11.5, muted
5.3-8.0, hint 4.9-7.3 (floor 3:1), action 5.4-8.1, human 8.6-12.9, live
7.5-11.3, error 6.2-9.3, warning 5.5-8.2, green 8.2-12.2, reasoning
5.9-8.8. Only `border` (#2A3F72) is low, 1.4-2.1:1: a non-text hairline.
No token value changed.
Seen, tmux 80x24 PTY, TERM=xterm-256color COLORTERM=truecolor, hermetic
HOME, debug build, counting `48;2;R;G;B` background sequences:
default (Terminal theme): startup, /theme picker, Help — no painted
background before or after (picker shows only accent swatches).
Blue Stage selected via T/Down/Enter, then F1 Help:
before: 15 x `48;2;7;12;29` (WHALE_BG) + 1 x selection row
after: 1 x `48;2;30;60;143` (selection row only)
Startup stage on Blue Stage: none, before and after.
Evidence (CARGO_BUILD_BUILD_DIR=.../mega-tokens, RUST_MIN_STACK=16 MiB as CI):
cargo check -p codewhale-tui --all-targets -> Finished, 0 warnings
cargo test -p codewhale-tui --lib --locked -- color_compat palette::tests::
ocean:: ocean_treatment live_transcript views::tests startup_ink
--skip command_palette -> test result: ok. 238 passed; 0 failed
cargo test -p codewhale-tui --lib --locked (full) ->
test result: FAILED. 11901 passed; 7 failed; 13 ignored
1 was this change (color_compat light test, updated above); the other 6
are role-name / slash-list assertions from other lanes on this branch
(scout<->explore, worker<->general, slash.impeccable) and untouched.
Ink goldens unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* tests(palette_audit): re-pin whale roles to the retuned palette
`whale_roles_are_pinned_and_non_colliding` still pinned the pre-navy
values (WHALE_BG (3,7,13), ACTION (106,174,242), ...) and failed on this
branch before the token slice touched anything. Pins now match tokens.rs;
no colour value changes.
cargo test -p codewhale-tui --test integration --locked palette_audit
-> test result: ok. 3 passed; 0 failed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* tui/cli/web: fleet is the public product term; /pod, codewhale pod stay aliases
Founder decision 2026-09-01: "fleet" is the customer-facing name for the
assembled model team and "Pod" is retired from product copy. `/fleet` is
the canonical slash command and `codewhale fleet` the canonical CLI verb;
`/pod`, `codewhale pod`, `loadout`, and `party` remain parser aliases.
Storage keys, the ledger file name, config tables, protocol identifiers,
and MessageId variant names keep their current spelling.
- CommandInfo name/aliases/usage, help text, and the unknown-verb error
flip to /fleet; `/fleet fleets` (saved/manage) is the saved-fleet picker
with `/fleet pods` kept as an alias.
- All 15 locale packs: localized values say fleet; the settings goldens
follow. `KbCompleteCycleModes` names the modes as Plan → Work → Operate
(Act is only a compatibility alias per docs/MODES.md).
- `scripts/check-tui-product-vocabulary.sh` now rejects `Pod` in en.json
instead of rejecting `fleet` in every pack.
- Hotbar id `slash.fleet` is canonical; persisted `slash.pod` normalizes.
- Fleet store error prose says fleet.
- Docs: PRODUCT.md lists the current role tokens (general, explore,
planner, reviewer, implement, test, advisor, custom) and names the old
spellings as aliases; docs/FLEET.md uses one role vocabulary;
web/lib/content/vocabulary.ts ADVISORY_ROLE is Advisor with consultant/
oracle as the legacy spellings (matches fleet/profile.rs migration).
Evidence:
sh scripts/check-tui-product-vocabulary.sh -> exit 0
cargo test -p codewhale-config -p codewhale-lane --locked
-> 638 passed; 0 failed / 62 passed; 0 failed
cargo test -p codewhale-cli --locked -- fleet pod -> 3 passed; 0 failed
cargo test -p codewhale-tui --lib --locked -- fleet::store fleet::members
fleet::identity -> 24 passed; 0 failed
cargo test -p codewhale-tui --lib --locked -- groups::core::fleet
localization command_palette hotbar fleet_roster settings widgets
fleet::control pod_workers -> 605 passed; 1 failed (the failure is
slash_source_matches_command_palette_command_entries, which reads the
machine's ~/.claude/skills and finds an `impeccable` skill; it fails
identically without this change)
cd web && npm test -- lib/content/vocabulary.test.ts -> 11 passed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* chore(tui): clear the six clippy 1.98 errors the base lanes left
needless_borrow on provider_identity_for_persistence (five sites) and a
collapsible_if in the work-surface mouse path. No behaviour change.
* tui(composer): restore double-tap Enter as the send-now gesture
While a turn is running, the first bare Enter queues the message (as
before) and opens a 500 ms window (`App::DOUBLE_TAP_WINDOW`, the value
the removed code in c5c42b7d91 used). A second bare Enter inside that
window with an empty composer promotes the just-queued message to a
Steer through `attempt_steer_with_queue_fallback` — the same path
Ctrl+Enter takes, so there is one steering path. A second Enter with
new text is an ordinary queue; Ctrl+Enter still steers immediately;
outside a turn Enter is unchanged. `enter_with_double_tap` is the one
decision point again (`take_queued_for_double_tap_steer` routes through
it), and `submit_disposition_does_not_mutate_the_queue` stays true.
The posture bar advertises the gesture while the window is open
(`PostureHintEnterAgain`, next commit).
Tests (cargo test -p codewhale-tui --lib <filter> --locked):
double_tap: test result: ok. 3 passed; 0 failed
enter_with: test result: ok. 5 passed; 0 failed
submit_disposition: test result: ok. 6 passed; 0 failed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* tui(shell): one owner per fact — posture bar, metrics line, no dead hints
Design: SHELL-DESIGN-20260901 §2.0 item 3, §2.2, §2.3, §2.3b, §2.11 and
the founder's 2026-09-02 redirect (Claude Code's grammar, less always-on
information). Under the composer there are exactly two chrome rows in
the default state, then the work surface only when it has content:
▶▶ ask (Shift+Tab) · work (Tab) · 2 agents · Esc to interrupt /rc …
deepseek-v4 · ctx 61% · $0.42 · ttft 400ms · 40 tok/s · ↓ 1.2K Ctrl+/ help
Fact → owner, before → after (composed 80x24 / 120x32 frames, working
turn with two sub-agents; "strip" = the work-surface bottom view):
fact before after
context % info line metrics line
cost footer (+ info line when priced) metrics line
model info line metrics line
provider info line (wide) metrics line (wide)
ttft / tok/s / ↓ behind /cost only metrics line
repo slug info line (+ idle empty state) launch header / git view (not chrome)
branch info line (+ idle empty state) launch header / git view (not chrome)
mode footer posture bar
permission footer posture bar
phase word footer ("sub-agents underway") transcript active row (not chrome)
elapsed footer ("1m 15s") roster rows (per agent)
agent count indicator row + info "pod 2/2" + "whales 2/4" posture bar ("2 agents")
+ dock tab + strip header (+ dock tab, strip header — other slice)
task/shell count indicator row above the composer posture bar
help hint footer keys legend (F1) + info line (Ctrl+/) metrics line (Ctrl+/), from the binding
other key hints footer "⌥V:output", compact "? help" none; cycle keys next to the chip they cycle
live hint footer "Esc to interrupt" posture bar hint slot
≥80 % microcopy footer right slot posture bar hint slot (outranks the hint)
notice / rc footer right slot / — posture bar right slot
Dead key hints removed: `F1:keys` / `fn+F1:keys` (Help binding's
`footer_chord` is now `Ctrl+/`; `info_help_hint` derives from the
binding), compact `? help`, and the `footer_action_hints` family. The
mode/permission cycle keys print only when the binding table admits
them at the current focus (no `(Tab)` on the launch stage).
Row order: composer → posture bar → metrics line → roster/to-do. The
#5286 background-work chip above the composer is gone (it repeated the
posture bar's counts); `PendingWork` stays as the counts' source.
Goldens re-blessed and read: footer_* (posture bar), infoline_startup_*,
infoline_work_* (metrics line), settings_* (the settings preview's
bottom row); infoline_settings_* deleted with the settings-path segment.
Commands run (CARGO_BUILD_BUILD_DIR=…/mega-frame, RUST_MIN_STACK=16777216):
cargo check -p codewhale-tui --all-targets clean
cargo test -p codewhale-tui --lib infoline --locked test result: ok. 11 passed; 0 failed
cargo test -p codewhale-tui --lib tideline_tests test result: ok. 64 passed; 0 failed
cargo test -p codewhale-tui --lib one_owner_tests test result: ok. 4 passed; 0 failed
cargo test -p codewhale-tui --lib shell_key_routing test result: ok. 13 passed; 0 failed
cargo test -p codewhale-tui --lib localization::tests test result: ok. 49 passed; 0 failed
cargo test -p codewhale-tui --lib --locked test result: FAILED. 11893 passed; 8 failed
(config_panel golden re-blessed after; the other 7:
4 fail on HEAD without this change (fleet rename
in flight), tmux clipboard passes alone, none in
files this change touches)
cargo clippy … -D warnings 6 pre-existing errors, none in this change's hunks
(config.rs:2106/2796, apply.rs:759, event_loop.rs:464,
session_state.rs:1004, work_surface/input.rs:401)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* wip(launch): checkpoint for overnight takeover — braille mark + kitty tier + Claude-Code launch header compile and pass lib tests; PTY cucumber tests still wait on the old heading
Compiles (cargo check --all-targets clean). Lib tests for mark::, underwater::,
mouse_ui::, localization:: pass: test result: ok. 120 passed; 0 failed
(needs RUST_MIN_STACK=16777216 like scripts/dev-test.sh; the rust_i18n static
overflows a 2 MiB test thread with or without this change). Startup goldens
re-blessed and read. Clippy is red only in files outside this slice
(config.rs, apply.rs, session_state.rs, work_surface/input.rs, and a
pre-existing event_loop.rs borrow).
Not done: crates/tui/tests/cucumber/{screen_mode_inline_pty,
active_composer_pointer_pty,plugin_e2e_acceptance}.rs still wait for
"What are we working on?" and press 'w'; they need the new marker
("Codewhale v") and a typed message + Enter to begin the session.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* wip(rail): checkpoint for overnight takeover — dock views compile, 136/137 work_surface tests pass, files/notepad/git views are stubs
Foundation for the one bottom dock (founder redirect 2026-09-02):
RailPanel is now the eight-view cycle (agents, tasks, background, files,
notepad, context, git, price; Pinned folded into tasks), an auto rule opens
agents/tasks/background while they have content, explicit picks stick until
Esc, and Ctrl+Tab / Ctrl+] (fwd) + Ctrl+Shift+Tab (back) cycle. Context and
price views render as rows; files, notepad, git are stubs in views.rs. The
classic sidebar line panels and their dead consumers are deleted.
Known: agent_rows_show_role_assignment_and_open_the_agent_transcript fails
(role_label 'worker' vs 'general'); role derivation is untouched here and
the failure is believed to predate this work — unverified.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* wip(operate): checkpoint for overnight takeover — auto-goal + contract land; one Operate approval test needs a goal-complete mock
Operate now turns a non-trivial prompt into the goal through the same
GoalState::create path as explicit_goal_directive, appends the Operate
contract once as a user-role runtime message (append-only history, not
the prefix), shows the Operate goal receipt, and carries the new mode copy
in all 15 locale packs plus docs/MODES.md.
Compiles. Passing: goal (134), prompts (133, incl.
every_mode_shares_one_prompt_per_host), localization (49),
runtime_handoff (14), session_peek (15), history_cells (2), both new
engine tests. Known failing:
core::engine::tests::operate_model_shell_uses_normal_approval_and_workspace_sandbox
— its mocked model never reports the auto-set goal complete, so the turn
re-prompts to max_steps (wiremock expect(1) sees 199). Six clippy
needless_borrow/collapsible_if hits pre-exist on the branch base.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* wip(fleet): checkpoint for overnight takeover — #5815 review findings 1-9 fixed, compiles, fleet tests green
Findings against the fleet-as-models work (verified against the current
tree, then fixed):
1. `toggle_fleet_model` decides presence by member rows, not the projected
role list (a role-less row projects to no role, so the old
`all(== "operator")` was vacuously true). Regression test
`toggle_removes_a_role_less_member_row` fails on the previous commit
("got Unchanged { … operator route … }") and passes now.
2. `selected_or_default` loads an existing personal `My fleet` instead of
overwriting it and never writes or selects before the add succeeds;
`FleetModelChange::Added` carries `created_fleet` + `selected_fleet`.
3. `fleet_models` returns `Result<Vec<_>, FleetStoreError>`: a broken
explicit selection is surfaced in `/fleet models`, `/models`, and the
picker's ⇧F instead of reading as "session model only".
4. `add_fleet_model` dedupes roles (case-insensitive) and returns
`Unchanged` without touching the file when every role is present
(test compares bytes and mtime).
5. `App.config` startup snapshot removed. `/fleet add|remove` now return
`AppAction::FleetAddModel|FleetRemoveModel`; the UI arm validates the
provider against the live `Config` (`fleet_provider_rejection`,
`fleet_catalog_rejection`, re-exported from `commands`).
6. ⇧F applies the same provider gate as `/fleet add`.
7. One roster path: `sync_fleet_roster` (extracted from the
FleetStoreChanged arm) plus `App::fleet_roster_stale`, flushed once per
event-loop iteration; `/fleet add|remove`, ⇧F, and every UI-side
auto-enroll site set it (`auto_enroll_fleet_model` now returns bool).
8. ⇧F receipts go through `push_status_toast` (Success/Info, 6 s) and
`set_sticky_status` (Error); no new `status_message` writes.
9. All new fleet prose is `tr(locale, MessageId::Fleet…)` (27 keys,
translated in all 15 packs); `FleetModelError` is typed with a
localized `message(locale)`.
10. No stale "`a` in /models" doc comment exists in the current tree.
Also re-blessed `config_panel_{80x24,120x32}` goldens (the Config tab
label says Fleet); the diff is that one label.
Evidence (this tree):
cargo check -p codewhale-tui --all-targets --locked -> Finished
cargo test -p codewhale-tui --lib --locked -- fleet::members
groups::core::fleet model_picker format_helpers fleet_roster
localization golden hotbar command_palette fleet::store
-> 257 passed; 1 failed (slash_source_matches_command_palette_
command_entries: reads ~/.claude/skills and finds `impeccable`;
fails identically on main in this environment)
cargo clippy … -D warnings (CI flags) -> the only remaining error is
crates/tui/src/tui/work_surface/input.rs:401 collapsible_if, which
belongs to the work_surface lane and predates this commit
cargo fmt --all -- --check -> clean
sh scripts/check-tui-product-vocabulary.sh -> exit 0
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH
* feat(tui): launch card, canonical role vocabulary, DashScope descriptor, test fixes
- Launch is now our own card take (founder, 2026-09-02): thin top line
⑂ branch path; centred bordered card with the whale mark, Codewhale +
version, one true announcement (no-model warning / MCP news), and the
menu New worktree / Resume session / Changelog / Quit with real chords
right-aligned; Enter runs the highlighted entry, Up/Down move it, and
typing goes straight to the composer. The card dissolves on the first
keystroke or command (≤240ms, instant under reduced motion); the
working screen then shows ⑂ branch path + ⋮ MCP n/m, the
session_start receipt, and the posture bar + metrics line (hidden
while the card is up). The composer's bottom rule carries
model (effort) · permission — the route's one launch reading.
- Role vocabulary: fixtures and the stopship fleet/workflow now use the
canonical tokens (explore/implement/reviewer/test); the workflow JS
wire accepts canonical spellings with the pre-rename ones as aliases
(AgentType serde rename+alias, serialized form is canonical).
- Alibaba Model Studio (DashScope) joins the data-driven descriptor
table: international compatible-mode endpoint, DASHSCOPE_API_KEY,
live /v1/models as the Qwen model authority (never a compiled id).
- Tests: role-keyed gate fixtures moved to canonical tokens; the operate
model-shell test now seals the goal through the deferred update_goal
tool (deferral retry included) instead of pausing blindly; the
slash-source hotbar test isolates HOME; ⚠ and ⋮ gained ASCII
fallbacks; launch goldens re-blessed for the card.
* feat(tui): retire Pod from copy; canonical workflow fixture; gate clean-up
- Pod literal sweep across fleet views, worker runtime, sub-agent tool,
managed API, and command groups: user-facing copy now says Fleet
(founder vocabulary decision; /fleet canonical, /pod and
'codewhale pod' stay as compatibility aliases). Roster tests that
encoded the retired Pod-public/Fleet-internal split now assert the
public Fleet vocabulary.
- workflows/stopship + fleets/stopship use canonical role names
(explore/implement/reviewer/test); the workflow crate's own stopship
tests and required-roles list follow.
- Operate mode-picker hint shortened to fit 80 columns in every locale.
- Cucumber PTY launch flows: wait for the launch card, type the first
prompt and press Enter; the live shell is proven by the launch stage's
top line disappearing and the metrics line's ctx reading (the
interrupt hint needs a live turn, which an offline route never
starts). The stopship acceptance feature expects the canonical /fleet
help copy.
- CHANGELOG receipts synced; DESIGN.md shell direction records the card,
posture bar + metrics line, and the bottom view cycle.
* test(cucumber): launch-card PTY contract fixes
- The launch-card wait uses the menu's New worktree entry — unique to the
card; the bare wordmark also matches onboarding copy.
- The live-shell proof is the launch stage's top line disappearing plus
the metrics line's ctx reading; the interrupt hint needs a live turn,
which an offline route never starts, and the help hint sheds first at
the 40-column floor by design (SHELL-DESIGN-20260901 §2.2).
- The pointer-submit queue proof takes the offline onboarding seed into
account: the receipt toast proves the gesture, and the queue count
grows by exactly one.
* docs(readme): restore the canonical product screenshot the web contract pins
The brand header redesign dropped the assets/screenshot.webp embed; the
web public-surface contract pins the README and the website to one
canonical optimized screenshot (byte-identical, 1562x1256 lossless
VP8L). Re-embed it.
* test: platform-robust pointer-submit wait and cap-warning diagnostics
- The pointer queue proof accepts either the transient receipt toast or
the queue-count increment: toast timing differs across runners, and a
20 s wait missed a toast the queue dump proved had fired.
- The context-cap posture test dumps the drawn rows when the warning
count misses, instead of a bare 0 != 1, so a platform-specific shed
(the hint sheds first when the left run exceeds its budget) is
visible in CI.
* test: fix the linux-only context-cap shed; bounded pointer-click retry
- The context-cap posture test drew at 100 columns, where a backend-less
platform (linux CI paints 'files: workspace (unenforced)') sheds the
cap hint first, so the warning count read 0. Draw at 140 columns,
where the hint survives with the notice present; verified locally.
- The pointer queue proof retries the [↑] click once, re-finding the
affordance first: under runner load a redraw can shift cells between
the find and the click, so the first SGR gesture lands nowhere.
* fix(gates): tool-catalog budget covers the fleet rename; readme stamps; pointer baseline
- The Pod->Fleet sweep grew every mode's tool-schema surface by 58 bytes
(+14 tokens). The receipts are re-measured and the one-way ceilings in
scripts/runtime-contract-budget.json are raised to them as the
explicit maintainer decision the gate asks for (the rename is the
founder's 2026-09-01 vocabulary call).
- The README screenshot embed changed README.md; the 18 translated
READMEs re-stamp (the embed is language-neutral HTML - no prose
changed, so no retranslation was needed).
- The pointer queue baseline is captured while the composer is empty:
the pending preview row hides while a draft sits in the composer, so
the pre-click depth read None and the growth proof could not fire.
* test: re-click then keep polling until the deadline
The qa_harness Instant wrapper does not implement Div, and the retry's
single read raced the app processing the second gesture: poll to the
full deadline, re-click once at the half-way point.
* test: pointer queue diagnostics (baseline/expected/last-seen) in the failure output
* test: pointer poll keeps per-iteration state only (unused-assignment gate)
* fix: Copilot review findings — planner wire spelling and Advisor copy
- workflow::AgentType::Plan serializes as the canonical 'planner'
('plan'/'awaiter' stay accepted aliases), matching the FleetRole
vocabulary the mega PR declares.
- Web: the vocabulary docs metadata, the vocabulary module header, and
the docs-map topic description say Advisor (the public advisory term)
instead of the retired Consultant spelling.
- Polish home dictionary: restore 'Podwodna powłoka terminala' — the
fleet-vocabulary sweep had merged 'Fleet' into the compound word
'Podwodna' (underwater), producing the non-word 'fleetwodna'.
* test: pointer proof accepts preview-appears when no baseline count is painted
* test: the tolerant preview-appears proof (the arm the last commit missed)
---------
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: CodeWhale Bot <bot@codewhale.net>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
38 KiB
Agent fleet
Baca terjemahan bahasa Indonesia: id/FLEET.md
Agent fleet is the local-first roster and member-selection layer for durable
multi-worker runs. It does not execute or authorize work. After fleet resolves
who should participate, the delegated coordinator launches a headless
codewhale exec run and the Runtime tracks it durably. See
AGENT_RUNTIME.md for how sub-agents, exec, and
fleet-backed workers converge on one runtime. In product language, a user may
still "open a sub-agent"; in architecture language, durable nested work uses a
fleet member identity with delegated runtime execution.
Naming and compatibility boundary
Fleet is the public product noun. The durable ledger, saved rosters, config
tables, and --fleet flag share that name:
| Surface | Canonical | Compatibility alias |
|---|---|---|
| CLI | codewhale fleet … |
codewhale pod … |
| Slash command | /fleet … |
/pod … |
/pod and codewhale pod remain accepted as compatibility aliases.
These shared names are load-bearing wherever changing them would break existing workspaces, receipts, or scripts:
- the durable ledger
.codewhale/fleet.jsonland the log directories.codewhale/fleet/and.codewhale/fleet-host/; - saved rosters
fleets/<name>.tomland theirschema = "fleet"header; - the
[fleet]and[fleets.*]config tables; - the
codewhale workflow run --fleet <name>flag; - wire, receipt, and control-plane operation ids such as
fleet.status.
The rest of this document uses fleet as the public product noun and retains these literal paths, keys, flags, and ids.
Use a fleet roster rather than anonymous short-lived agent fanout whenever a
delegated run needs stable member identities across retries, sleep/restart,
remote execution, receipts, or a ledgered audit trail. The initial CLI surface
is:
For a guided start-to-monitor walkthrough that combines fleet task specs with Workflow authoring, see fleet + Workflow Tutorial.
codewhale fleet init
codewhale fleet run tasks.json --max-workers 4
codewhale fleet status
codewhale fleet inspect <worker-id>
codewhale fleet logs <worker-id>
codewhale fleet artifacts <worker-id>
codewhale fleet interrupt <worker-id>
codewhale fleet restart <worker-id>
codewhale fleet resume <run-id>
codewhale fleet stop --all
codewhale fleet resume <run-id> is the restart-recovery verb: it replays the
ledger, reconciles any in-flight lease whose worker stopped heartbeating
(retrying within the task's budget, else failing and escalating per the alert
policy), and prints the post-resume status. It launches no new work and is
idempotent, so it is safe to run after a manager exit, laptop sleep, or runtime
restart.
Coordinator state for fleet-backed runs is stored under the workspace in
.codewhale/fleet.jsonl. Worker logs and adapter logs are stored under
.codewhale/fleet/ and .codewhale/fleet-host/.
Public contract: identity, membership, and selection
Fleet = the user's model inventory: who is in the roster and which member is selected.
A public fleet identity consists only of:
- a stable member id and an optional user-facing name;
- a semantic role, such as
explore,implement, orreviewer(the legacy spellingsworker,scout,builder,verifier,consultant, andoracleare still accepted on input and map togeneral,explore,implement,test, andadvisor); - an exact provider/model identity, or an explicit inherited route;
- visible roster state or origin.
Project or workspace trust, filesystem and network reach, secret access, approval mode, sandboxing, tool authorization, and every other form of runtime authority are separate delegated-coordination and Runtime policy inputs. They are never fleet identity fields, and they never select or reroute a fleet member. The Runtime applies and clamps those policies only after member selection; if the selected member cannot run inside the effective envelope, launch fails closed instead of choosing somebody else.
Natural-language member selection is deterministic. A caller may name:
- an exact member id, optionally as
member:<id>orid:<id>; - a unique user-facing member name, optionally as
name:<name>; - a unique semantic role, for example
exploreorrole:explore; - an exact pinned model id, for example
deepseek-v4-flash, or its offline display name, for exampleDeepSeek V4 Flash; or - an exact
route:<provider>/<model>.
An unqualified exact member id wins. Every other match succeeds only when it
identifies one distinct roster member. Multiple matches produce an ambiguity
error that names the candidates and asks for member:<id>; Codewhale never
picks whichever match happened to be listed first. Users do not need to know
an internal role label such as explore: a unique member name, display model, or
exact model id is equally valid. Saved v2 fleets store that optional human name
as display_name (the input alias name is also accepted); it must be one
trimmed printable line of at most 80 characters.
Your fleet as models
The same fleet file answers a third question: which models has this person
put in their fleet? Every exact provider + model pin in the selected
fleet — the operator route and each pinned member — is a fleet model, and the
member rows that pin it are the roles it fills. There is no second list.
/fleet modelsprints the fleet:provider/model · roles · price · context · tools, facts read from the model catalog. With no selected fleet the line reads "Your fleet is the session model only"./fleet add <provider> <model> [role…]adds a model (one member row per role; none for a role-less add). The provider must be one you configured and, when the catalog knows the provider, must serve that exact id. With no fleet selected, a user-global fleet namedMy fleetis created and selected first./fleet remove <provider> <model>drops every row that pins the route; the operator route is changed with/fleet save, not removed.- In
/model,⇧Fon a row adds or removes that exact route the same way; fleet models are listed first, labelledfleet · <roles>, ahead of your own⇧Ppins and the provider lists./modelsprints the fleet before the provider's list.
The operator model reads this list when it assigns sub-agents (design
MODEL-ROUTING-CATALOG-20260901.md §10, slice F2).
Interactive and persistent status
/fleet status and codewhale fleet status are the same command on two
surfaces. Both read the durable .codewhale/fleet.jsonl ledger for the
workspace, through one shared control-plane contract, and both report the same
verb id (fleet.status), read-vs-write authority, persistence scope, and
receipt. When the workspace has no ledger they say so with a typed reason
(no_fleet_ledger) instead of rendering an empty-looking "all clear" — and
neither creates the ledger as a side effect of reading it.
The current interactive session's sub-agents are a different set, and now have their own name:
/fleet workers(or/subagents, orn) shows sub-agents attached to the current TUI session. It does not read the persistent ledger./fleet list|status|interrupt|resumeandcodewhale fleet list|status|interrupt|resumeact on the durable ledger.codewhale fleet restart <worker-id>is CLI-only: it re-leases the task and then drives the manager loop to completion./fleet restartdoes not silently do a smaller thing — it reportssurface_not_supportedand names the CLI command.
Before v0.9.2, /fleet status showed session sub-agents. That reading is gone;
/fleet workers replaces it.
The contract behind this — descriptors, availability reasons, exact-identity
targets, receipts, typed unknowns, and bounds — is documented in
docs/COMMAND_CONTROL_PLANE.md.
Authoring agent profiles (/fleet setup)
/fleet setup (also /fleet setup edit / new) opens an in-TUI wizard for
authoring a reusable agent-team profile. Bare /fleet and the
roster/roles/profiles/party aliases open the selected fleet's member
roster. /fleet saved opens the named saved-fleet picker. /fleet workers opens the
current-session worker view; /subagents is a
compatibility shortcut for that view. For durable run history, use
/fleet status or the shell command codewhale fleet status described above —
they are the same command.
The wizard is progressive: you make one focused choice at a time — a role,
then a model (inherit, or a concrete model from any configured
provider, not only the one the parent session is currently using), then
where the profile lives, and finally a review of the member identity
and route. When the review also previews thinking, tools, approvals, or another
execution control, those rows summarize separate Runtime policy; they do not
become fleet identity or member selectors. The header shows "Saves to: …" on
every step — the choice you still have to make, or the exact resolved file
once you have made it. Nothing is written until you activate the save control
on the review step.
The Destination step is a focused two-option list (arrows move, Enter or Space chooses; Tab never changes the destination):
- This project writes
<workspace>/.codewhale/agents/<role>.toml. It applies to this project only and takes precedence over a Personal profile with the same id. When project profiles are disabled for the session (--no-project-config) or the workspace folder is unavailable, the option is shown disabled with that reason; the wizard never falls back to Personal on its own. - Personal writes
$CODEWHALE_HOME/agents/<role>.tomland is available in every project on this machine, except where a project has its own profile with the same id.
For the highlighted option the step shows the exact file, whether saving would
create a new file or replace an existing one, and the precedence
consequence for the roster. The review step repeats those facts under
"Saves to" and names the final action by its effect — Save to this
project, Save as Personal profile, or Replace …. Replacing an
existing file needs a second Enter on the save control. Tab / Shift+Tab (or
←/→) move focus between the save control, Change destination, and
Back; s is a secondary shortcut back to the Destination step. Reopening a
saved member from /fleet starts from what is on disk: its member identity,
route, and save scope. Thinking (inherit, off, low, medium, high,
max, or auto) is adjusted on the review step with t, but remains a route
execution setting rather than part of the member's fleet identity.
Profile scope controls where a role definition is reusable; it does not widen
the authority of a running operation and is not a project-trust setting. To
coordinate several nearby repositories, start Codewhale from their shared
parent directory so that parent is the workspace. Project/workspace trust,
external paths, filesystem and network reach, secrets, approvals, sandboxing,
and tool authorization come from delegated-coordination and Runtime policy.
For nested delegation, Runtime intersects the requested child posture with the
live parent. For standalone codewhale fleet execution, Runtime instead uses
the bounded tool-authority envelope minted from the task's explicit write
scope together with live config, sandbox, and platform enforcement. Neither
path reads authority from the profile's storage scope or identity selector.
Picking a concrete model pins its provider explicitly: the saved profile records both
model and provider fields, so the route it names doesn't depend on
whichever provider happens to be active when the profile is later loaded.
Pressing Enter ("start") on the review step previews the exact starter
profile TOML inline on that same screen; nothing is written until you save it.
The provider field may be a built-in provider id such as openrouter or a
user-named OpenAI-compatible provider configured under [providers.<name>]
such as lm-studio; the launch path preserves that id and fails closed if the
provider is not configured.
Profiles are also how the model-facing agent tool selects a route since the
v0.9.9 schema slim (#5324, #5123): the advertised surface no longer carries
model or thinking — a child either runs as a profile (whose saved route
and thinking tier it uses exactly) or inherits the operator's model. Removed
fields stay parse-accepted for saved transcripts, ACP/MCP clients and fleet
configs; see docs/SUBAGENTS.md for the advertised 12-field list and the
compat list.
When a provider is configured, the review step also offers model-assisted drafting behind an explicit preview-before-save gate:
- Press
mto have your first configured model draft the profile. The draft arrives sanitized and bounded. Separately, the Runtime keeps its conservative execution floor (no shell or trust escalation and approval required) regardless of what the model proposes. - Drafting is not saving. The exact rendered TOML preview renders
inline on the review step (not in a separate scrollable viewer), so nothing
is saved until you press
gor Enter to save (or pressmagain to redraft). Saving writes the profile to the project or personal scope shown in the preview.
Naming: Modes, Workflow, and fleet
These names describe different layers, not competing systems. Plan and Act are the everyday work modes. Operate accepts ordinary messages and keeps the parent's normal tool surface under the same approval, sandbox, shell, ask-rule, and repository protections as Act. It prefers background fleet workers for independent, parallel, isolated, or long-running work, but does not require a worker for every executable step. Workflow is an optional orchestration overlay for work that needs ordering, gates, shared budgets, replay, or deterministic fan-in.
The short public vocabulary is:
-
Fleet is the durable roster and deterministic member-selection surface. It records member ids and names, semantic roles, provider/model identities, and roster state. Fleet is also the name used by storage and wire formats.
-
Workflow = what order the work follows: phases, gates, budgets, replay, and fan-in.
-
Lane = one running Workflow instance and its live progress.
-
Runtime = where, how, and with what authority selected work executes. Runtime owns the local or remote process, provider route, project/workspace trust, filesystem, network, secrets, approvals, sandbox, tools, and API boundary.
-
Workflow is the repeatable plan and user-facing orchestration overlay: a script/IR that decides which phases and agents run next, keeps intermediate results out of the main conversation, and can be inspected or rerun. A Workflow run should have a visible progress view and a clear active header state instead of feeling like a hidden background task.
-
Fleet is the durable roster and deterministic member-selection surface: member ids and names, semantic roles, pinned or inherited provider/model identities, and roster state. The delegated coordinator and Runtime own launch concurrency, leases, heartbeats, logs, receipts, tools, sandboxing, approvals, and authority.
-
High fan-out is a behavior of a Workflow run, not a separate system: when a phase needs many workers at once, Workflow dispatches them as a fleet-backed run (durable workers, receipts, goal re-dispatch) rather than reviving prompt-only sub-agent fanout.
-
Fan-in is explicit: when the user needs one combined result, an owner aggregates, verifies, and synthesizes the worker receipts. Independent tasks may finish separately; dispatch is never presented as completion.
UI guidance: keep the main transcript calm. A Workflow run should appear as a compact progress card plus work-bar rows (the strip above the transcript, or a side rail) with phase names, worker counts, receipts, and nested indentation for child workers. Use the whale mark sparingly as an active header/status signal; avoid repeating emoji-heavy rows for every worker.
Saved fleets and the Reasoning Router
A selected v2 fleet freezes each selected member's id, semantic role, provider,
and model identity into the durable run before a Workflow starts. Save the
fleet as fleets/<name>.toml in the workspace or under $CODEWHALE_HOME.
Models cannot replace those identity or route assignments at runtime:
schema = "fleet"
schema_revision = 2
name = "release"
[operator]
provider = "deepseek"
model = "deepseek-v4-pro"
[[members]]
id = "implementer"
display_name = "Release Builder"
role = "implement"
provider = "zai"
model = "glm-5.2"
[[members]]
id = "advice"
role = "advisor"
provider = "openai"
model = "gpt-5.6"
The workflow crate's older schema = "exact", revision 1 files are migration
input only. Do not author them for v0.9.11; the selected roster and setup UI
read and write only schema = "fleet", revision 2.
Reasoning is a separate route-execution decision, not fleet identity. The
optional Reasoning Router is a reusable Runtime service, not a fleet member.
Save one profile at routers/<name>.toml in either search root and reference it
from any number of fleets:
name = "luna-low"
schema = "reasoning_router"
schema_revision = 1
provider = "openai"
model = "gpt-5.6-luna"
call_reasoning = "low"
At runtime it may choose only the reasoning tier for an already-frozen worker
route. It cannot change the member, provider, model, or semantic role. The
Router call itself is capped at off or low; more expensive values are
rejected. A manually selected worker reasoning tier makes no Router call. Route
and reasoning receipts name the worker model and, when used, the Router's exact
provider/model so the operator can see which model did which job. If the same
bare Router or fleet name exists in both roots, qualify it as
workspace/<name> or codewhale_home/<name> instead of relying on shadowing.
Compatibility schemas may serialize reasoning, permissions, tool hints, or
other execution settings beside a member. Those values are not fleet identity,
member selectors, or active authority. A valid legacy schema = "exact" roster snapshot
retains its old permissions bytes only while verifying and replaying that
snapshot's recorded content hash; a fresh capture emits the authority-free
member shape. New-run validation rejects legacy roster
security_policy and worker trust_level fields; configure execution authority
through Runtime policy. The delegated coordinator resolves and durably freezes
the member first. Runtime then applies either the delegating parent's effective
ceiling or, for standalone fleet CLI work, Runtime execution configuration plus
live sandbox/platform enforcement. That boundary may
reduce or refuse the selected worker's execution surface, but it must never
choose a different member or route. See
docs/MODES.md, docs/SUBAGENTS.md, and
docs/AGENT_RUNTIME.md for the enforcement contract.
Reasoning receipts record the requested tier and the tier the provider was
actually asked for. Those differ whenever a route cannot express the requested
one — Codewhale's route normalizer sends high for a requested low on most
routes, and Z.AI's GLM routes express only thinking on/off — so the receipt
reports the real request rather than the label that was selected. The value a
call actually carries is spelled by that route's own normalizer, not by the tier
label: an OpenAI Codex route is asked for xhigh, not max, and cannot be
asked for off at all.
A v0.9.11 durable fleet CLI receipt keeps the selected profile id in
effective_permissions.profile_id, the resolved semantic role in
resolved_route.role, and the effective Runtime surface in the permission,
shell, and tool-scope fields. An exact Workflow launch receipt records
member_role separately from an optional Runtime posture_role, plus the
fingerprint of the effective authority envelope checked at the spawn boundary.
A member named auditor can therefore retain that identity while Runtime
reports a custom posture and independently proves the narrower surface it
enforced.
A Workflow start fails closed on anything decidable locally: an unresolvable
provider or model, a missing credential, a client that cannot be built for a
member's route, or an auto member with no usable Reasoning Router. Per-task
validation that the spawn boundary would refuse anyway — notably a write-capable
member with no declared write_roots/exact_files/coordination_contracts —
is checked before the Router is called, so an invalid task never spends a
routing request. If a spawn fails after a Router decision, the receipt is
still recorded: the tokens were spent, and any cross-provider disclosure already
happened.
Manager-owned Workflow fan-in
When parallel work must return one combined answer, use a manager-owned
Workflow instead of a flat agent fan-out:
- Cast one manager (operator or workflow orchestrator).
- Fan out child tasks through
workflow(task(),parallel(),pipeline(),phase()) or a single manager session that owns the children. - Wait for child receipts or completion events.
- Aggregate and verify load-bearing claims before treating them as facts.
- Synthesize one result the operator can depend on.
Raw agent fan-out is appropriate only for independent, fire-and-forget work
where no single fan-in result is required. If results must be merged, compared,
or verified, route through workflow so the manager owns fan-in.
Workflow on fleet
The intended high-capability path is agent-authored. When the main agent decides a task needs more durable coordination than turn-by-turn sub-agent calls, it drafts a Workflow script/IR, presents the run plan according to the active permission mode, and the runtime compiles it into typed fleet work.
fleet remains the sub-agent roster and member-selection surface. It owns member
identity, membership, semantic roles, saved provider/model pins or inheritance,
and roster state. Workflow owns the orchestration plan:
branch, sequence, loop, expand, review, and reduce decisions. The delegated
coordinator and Runtime own slot admission, launch concurrency, the execution
ledger, and every authority decision. A workflow script receives no direct
shell, filesystem, network, provider-secret, cancellation, or TUI authority;
workers perform real work as codewhale exec processes under the effective
Runtime policy.
Default Workflow-to-fleet validation is intentionally bounded:
- 1,000 total worker agents per Workflow run;
- 16 live worker agents at once; larger populations queue (block) on the host's per-run concurrency gate until a live slot frees, then route through fleet;
- Workflow IR structural nesting no deeper than 5;
- Runtime child delegation defaults to 3 levels and has an opt-in hard ceiling of 8, independently of the Workflow document's structural depth;
- bounded loops only (
max_iterationsrequired); - bounded dynamic expansion only (
max_childrenplus a template required).
These are delegated-coordination population limits, not fleet identity and not
a demand to launch everything at once. A 1,000-agent Workflow should still
drain through the configured Runtime worker pool. They are also not model-step
budgets: omitted or zero max_steps remains unbounded. An explicit positive
max_steps may cap that task, while wall-clock timeouts, cancellation,
provider safeguards, heartbeats, and admission controls remain independent.
Recommended model layouts, such as a DeepSeek Pro orchestrator with Flash
workers in the first ring and cheaper workers farther out, are presets only.
Every slot can inherit the active model or carry an explicit model override.
Inheritance is literal: the model you select in /model is the operator
(the pinned first row in /fleet roster), and any worker whose task spec and
roster profile pin no model runs on that session model. Once a selected member
has an exact provider/model pin, the Runtime does not silently reroute that
identity because a policy input differs; it either runs that route inside the
effective envelope or fails closed. Route receipts record the requested and
resolved identity.
The setup UI should render this as an expanding grid: an orchestrator plus a small number of visible sub-agent slots, with Right/Enter drilling into a slot's next recursive ring rather than trying to show the whole tree at once.
Task Spec
codewhale fleet run accepts JSON or TOML. A minimal JSON spec:
{
"name": "local smoke",
"tasks": [
{
"id": "lint",
"name": "Lint",
"instructions": "Run the lint check and report failures.",
"expected_artifacts": ["log"]
}
]
}
Workers are optional. If omitted, Codewhale creates local worker slots up to
--max-workers.
Task specs are typed in Rust and keep verification data separate from worker
transcripts. Only the worker member/role reference participates in fleet
identity selection. The remaining execution fields are delegated-coordination
or Runtime inputs applied after the member is resolved. A task can declare:
id,name,description,objective, andinstructionsworkerrole, tool profile, tools, and required capabilitiesworkspaceroot, required files, writable paths, and environment allowlistinput_files, extracontext,budget,timeout_seconds, andretry_policyexpected_artifacts,scorer,tags, and free-formmetadata
None of those execution-policy fields becomes part of a fleet identity or an
alternate member selector. Omitted or zero max_steps means no model-step
ceiling; Codewhale must not synthesize a default step budget. Explicit positive
step limits, timeouts, cancellation, provider safeguards, heartbeats, and
admission control are enforced independently by the delegated coordinator and
Runtime.
Workers write bounded artifact files under .codewhale/fleet/ and ledger only
the artifact refs: kind, path, checksum, MIME type, and size. Receipts record
pass, fail, partial, skip, or timeout; failed receipts may also mark
the source as transport, task, or verifier. codewhale fleet status
surfaces those failure-source counts separately.
Deterministic built-in scorers are exit_code, file_exists, regex_match,
and json_path. Specs may also declare command,
code_whale_verifier_prompt, or manual; those record a partial receipt until
an explicit verifier pass completes.
Using Role Presets
Tasks can reference a semantic role name to select one unique roster member.
Built-in role names (smoke-runner, reviewer, builder, read-only) remain
available for compatibility, and custom roles may be defined in
[fleet.roles].
{
"name": "smoke check",
"tasks": [
{
"id": "lint",
"name": "Lint check",
"instructions": "Run lint and report failures.",
"worker": { "role": "smoke-runner" },
"expected_artifacts": ["log"]
}
]
}
After identity resolution, compatibility role presets may provide tool, timeout, or retry defaults to the delegated coordinator. Those defaults do not grant authority, do not change which member was selected, and remain subject to Runtime clamping. A task spec may request its execution settings explicitly:
{
"id": "deep-review",
"name": "Deep review",
"instructions": "Review the entire crate for soundness issues.",
"worker": {
"role": "reviewer",
"tools": ["cargo", "rg", "git"],
"capabilities": ["rust"]
},
"input_files": ["crates/**/*.rs"],
"budget": { "max_tokens": 32000 },
"expected_artifacts": ["log", "report"],
"scorer": { "kind": "regex_match", "path": ".codewhale/fleet/report.md", "pattern": "finding|all clear" }
}
Multi-Task Run Example
A single fleet run can dispatch several independent tasks in parallel:
{
"name": "CI gate",
"tasks": [
{
"id": "check",
"name": "Compile check",
"instructions": "Run cargo check --workspace and report errors.",
"worker": { "role": "builder" },
"expected_artifacts": ["log"],
"scorer": { "kind": "exit_code" }
},
{
"id": "clippy",
"name": "Clippy lint",
"instructions": "Run cargo clippy --workspace and report warnings.",
"worker": { "role": "reviewer", "tools": ["cargo", "cargo-clippy"] },
"expected_artifacts": ["log"],
"scorer": { "kind": "exit_code" }
},
{
"id": "security",
"name": "Secret audit",
"instructions": "Search for plaintext secrets and report any matches.",
"worker": { "role": "read-only", "tools": ["rg"] },
"input_files": ["crates/**/*.rs"],
"expected_artifacts": ["log", "report"],
"retry_policy": { "max_attempts": 1 }
}
]
}
Alerts
fleet alerting is disabled by default. A caller must supply an enabled alert config before anything is sent. Routes match typed fleet event classes, not log strings:
stalerestart_exhaustedneeds_humanbudget_exceededverifier_failedrun_completed
Adapter config stores environment variable names, not secret values. Send-time
code resolves those names from the environment or a future secrets provider.
Ledger records store only audit labels such as slack, webhook, or
pagerduty; task specs persisted in the ledger redact webhook URLs and routing
keys.
Example alert config shape:
{
"enabled": true,
"dry_run": true,
"routes": [
{
"events": ["stale", "restart_exhausted", "verifier_failed"],
"adapter": "ops-slack"
},
{
"events": ["restart_exhausted"],
"adapter": "pager"
}
],
"adapters": {
"ops-slack": {
"kind": "slack",
"webhook_env": "CODEWHALE_FLEET_SLACK_WEBHOOK",
"channel": "#codewhale-fleet"
},
"pager": {
"kind": "pager_duty",
"routing_key_env": "CODEWHALE_FLEET_PAGERDUTY_ROUTING_KEY",
"severity": "critical"
}
}
}
Use dry-run to inspect a redacted adapter payload without sending:
codewhale fleet alert-dry-run \
--event stale \
--run-id fleet-demo \
--worker-id fleet-demo-local-1 \
--task-id release-triage \
--reason "worker heartbeat stale since 2026-06-13T02:00:00Z" \
--adapter slack
The payload includes the run id, worker id, task id, status, short reason, and
safe inspection commands such as codewhale fleet status and
codewhale fleet inspect <worker-id>. Endpoints, webhook secrets, and
PagerDuty routing keys are shown as <redacted:env:...>.
Status Surfaces
codewhale fleet status shows compact counts for queued, running, completed,
partial, failed, restarted, escalated, cancelled, stale, and verifier/transport
failure sources. inspect shows the worker state plus the current task
objective, role, host, heartbeat, latest event, artifact refs, latest error, and
alert state. logs prints bounded log artifact contents, and artifacts lists
artifact refs without embedding large payloads.
The Runtime API exposes the same ledger-backed projection behind the existing runtime auth middleware:
GET /v1/fleet/runs
GET /v1/fleet/runs/{run_id}
GET /v1/fleet/runs/{run_id}/workers
GET /v1/fleet/workers/{worker_id}
POST /v1/fleet/workers/{worker_id}/interrupt
POST /v1/fleet/workers/{worker_id}/restart
POST /v1/fleet/runs/{run_id}/stop
Action endpoints call the same manager controls as the CLI and record their decisions in the fleet ledger.
Manager-Agent Runbook
Manager agents should treat fleet operations as typed, ledgered control-plane
work. Start with codewhale fleet status, then inspect one run or worker with
codewhale fleet inspect <worker-id>, logs, and artifacts. Use direct
reads of .codewhale/fleet.jsonl, host logs, or remote files only when the
typed CLI/API surface cannot provide the required evidence.
Classify the worker before taking action:
transient failure: stale heartbeat, host timeout, interrupted transport, retryable provider/network error, or an adapter status that can plausibly recover without changing the task.task failure: the worker completed but produced an incorrect result, domain failure, missing required artifact, or explicit task-level error.verifier failure: the worker result exists, but the scorer/verifier failed, timed out, or disagrees with the receipt.needs-human: missing authority, secret request, destructive operation, repeated restart exhaustion, ambiguous product decision, or conflicting evidence that the manager cannot resolve from typed artifacts.
Choose one typed action:
- Restart a worker only when the failure is transient, retry budget remains,
the task is idempotent or retry-safe, and no permission or secret boundary is
involved:
codewhale fleet restart <worker-id>. - Interrupt or stop only when the current task is unsafe to continue or the
operator explicitly asks for cancellation:
codewhale fleet interrupt <worker-id>orcodewhale fleet stop --all. - Do not restart pure task failures by default; preserve artifacts and hand the receipt to the task owner unless the task spec says retrying can produce new evidence.
- For verifier failures, inspect scorer inputs and artifact refs first. If the verifier cannot be corrected through typed fleet actions, escalate for human review.
- For
needs-human, draft an escalation instead of sending it unless alert config explicitly authorizes sending.
Safe Slack or PagerDuty draft:
Codewhale fleet needs attention
Run: <run-id>
Worker: <worker-id>
Task: <task-id or unknown>
Classification: <transient failure | task failure | verifier failure | needs-human>
Reason: <one sentence, no secrets>
Latest typed evidence: codewhale fleet inspect <worker-id>; codewhale fleet artifacts <worker-id>
Safe log excerpt: <3 lines max or "see artifact <ref>">
Requested decision: <restart approval | verifier review | task owner review | permission decision>
Post-run summaries should include the run id, workers checked, classification, typed action taken or drafted, expected ledger effect, artifact refs reviewed, and next owner. Keep summaries bounded; link artifact refs instead of copying full logs or transcripts.
The bundled fleet-manager skill mirrors this runbook for manager agents. It
is a first-party system skill and should be discoverable through the normal
skill registry after system skills are installed or refreshed.
Host Adapters
The Runtime host-adapter boundary supports local child processes and explicit SSH workers. Host choice is Runtime placement on a worker spec, not fleet member identity or a member selector. It does not authenticate the host or grant access. Adapters expose the same operations: start, read status, read bounded logs, interrupt, restart, stop, and cleanup.
Local workers run as child processes with stdin closed and stdout/stderr written
to bounded host-adapter logs. They inherit only a small safe base environment
such as PATH and explicitly allowlisted variables.
SSH workers run through the system ssh client with BatchMode=yes and a
bounded connect timeout. Remote environment variables are sent with OpenSSH
SendEnv; values are not embedded in the local ssh argv or fleet logs.
Example SSH worker spec:
{
"id": "builder-1",
"name": "Builder 1",
"host": {
"kind": "ssh",
"host": "builder.example.com",
"user": "codewhale",
"port": 22,
"identity": "~/.ssh/codewhale_fleet",
"working_directory": "/srv/codewhale/work",
"env_allowlist": ["CODEWHALE_PROFILE"],
"codewhale_binary": "/usr/local/bin/codewhale"
},
"capabilities": ["local", "linux", "tests"],
"max_concurrent_tasks": 1
}
Defaults are intentionally conservative:
- no hosted control plane or cloud provisioning is enabled;
- SSH requires an explicit host, working directory, and Codewhale binary path;
- secret-like environment names such as
TOKEN,SECRET,PASSWORD,API_KEY, andPRIVATE_KEYare rejected from adapter allowlists; - secrets should remain in Codewhale config providers or remote host config, not in task instructions, argv, or fleet logs.
Runtime policy and authority are not fleet identity
fleet does not define a project/workspace trust level, filesystem or network reach, secret access, approval mode, sandbox, tool set, or execution authority. Those belong to delegated-coordination and Runtime policy. This separation is load-bearing:
- member resolution considers only member id/name, semantic role, provider/model identity, and roster state;
- the selected identity is frozen before any authority policy is evaluated;
- the Runtime applies the live parent ceiling when one exists; standalone fleet CLI launches instead carry an explicit bounded authority envelope, and both paths remain subject to live sandbox and platform enforcement;
- no trust, permission, capability, secret, sandbox, approval, or tool-policy value may select another member or silently change its provider/model route; and
- receipts report requested and effective Runtime posture separately from the fleet member identity.
Older persisted configuration and protocol shapes may still contain fields such as
security_policy, trust_level, permissions, capability_grants, secret
references, host authentication, environment allowlists, or tool profiles.
They remain deserializable for ledger replay, but new fleet run creation rejects
security_policy and worker trust_level rather than pretending they grant
authority. Their presence in old data does not make them fleet variables or
grants. The active Runtime remains the final authority and fails closed when a
requested operation cannot be enforced.
For current enforcement behavior, use Modes, Sub-agents, Agent Runtime, and the Command Control Plane. Keep secret values out of task instructions, arguments, logs, and receipts; adapter and Runtime layers must continue to redact or reject them independently of fleet selection.