Files
DeepSeek-TUI/docs/MODES.md
Hunter Bown 329960fcbf feat: Codewhale 0.9.12 shell, brand, fleet, and Operate (mega) (#5826)
* feat(fleet): your fleet is the models you added, and it comes first

Design MODEL-ROUTING-CATALOG-20260901 §10, slice F1. A fleet model is a Pod
member: the selected Pod file's operator route plus every member row that
pins an exact provider + model; the roles a model fills are the member rows
that pin it. No second store.

- crate::fleet::members: fleet_models / add_fleet_model / remove_fleet_model
  / toggle_fleet_model + change_receipt; Config::fleet_members(workspace) is
  the read seam for the operator-awareness slice (F2).
- /pod models | add <provider> <model> [role…] | remove <provider> <model>
  (also via the /fleet alias). A model the configured provider does not
  serve is rejected; the first add creates and selects a user-global Pod
  named 'My fleet'.
- /model picker: ⇧F adds or removes the row's exact route; fleet models
  lead the list labelled 'fleet · <roles>', ahead of ⇧P pins and providers.
- /models prints the fleet before the provider list ('Your fleet is the
  session model only' when empty).
- PickerActionFleet message in all 15 locales; docs/FLEET.md 'Your fleet
  as models'.

Tests: scripts/dev-test.sh tui fleet::members groups::core::fleet
model_picker format_helpers — Summary 37 tests run: 37 passed, 11834
skipped.

Signed-off-by: CodeWhale Bot <bot@codewhale.net>

* fix(fleet): pass slugify by name (clippy redundant_closure)

cargo clippy -p codewhale-tui --all-targets -- -D warnings -A clippy::too_many_arguments -A clippy::uninlined_format_args -A clippy::unnecessary_map_or: no findings.

Signed-off-by: CodeWhale Bot <bot@codewhale.net>

* fix(tui): review fixes for fleet toggle and /pod add provider validation\n\n- Reject unconfigured provider ids in "/pod add" before writing, reusing\n  the existing provider_is_configured_for_active predicate and custom\n  provider table checks.\n- Add App.config snapshot so commands can consult the loaded config.\n- Update the stale DEFAULT_FLEET_NAME doc comment to mention ⇧F.\n- Sync crates/tui/CHANGELOG.md.

* style: cargo fmt

* fix(web): align react with react-dom 19.2.8 to unbreak npm ci

Dependabot #5801 bumped react-dom to 19.2.8, whose peer range requires
react 19.2.8; the lockfile still resolved react 19.2.6, so 'npm ci' in
web/ failed ERESOLVE on main and on every branch that merged it
(Lint & Type Check red). Align react to 19.2.8; install verified clean.

Signed-off-by: CodeWhale Bot <bot@codewhale.net>

* brand: trace supplied whale assets

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* brand: align icon ombre and generated tokens

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* brand: use white icon tile

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* brand: deepen ombre light stop

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* brand: wordmark takes the blue ombre

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* tui: recover from image-input rejections by non-vision routes

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* tui: localize image rejection recovery

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* chore: format 0.9.12 mega branch

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* Redesign Fleet role labels and agent cards

* feat(tui): launch hero as wordmark + small surfacing mark

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* design(tui): retune whale palette to codewhale navy / ombre sky

Field, chrome, panel, plate and raised surfaces move onto the brand navy
(#070C1D → #142352 → #1A2C63); interaction blue becomes the ombre sky
#6AA6DC, light-mode action the ombre cobalt #1535B2; ice/cyan/border/tool
tints follow. web/app/tokens.css regenerated via
scripts/export-design-tokens.py.

Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* test(tui): re-bless ink goldens for navy palette

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* web: Space Mono wordmark, quiet layout refresh, fleet vocabulary in site + docs

Space Mono (OFL) outlined wordmark rebuilt via scripts/build-wordmark.py,
wired as --font-display through next/font/google; body stays IBM Plex Sans,
code stays JetBrains Mono. Nav loses the issue strip, strapline, Discord badge
and second filled CTA; home loses the ticker, seals and tilt figure; docs
shell hero collapses to a one-line band; footer uses the inverted wordmark.

Public noun is fleet (/fleet, codewhale fleet, /docs/fleet canonical; /pod,
codewhale pod, /docs/pod remain compatibility aliases) across docs/, site
dictionaries, vocabulary contract and public-surface facts.

No-Issue: 0.9.12 website lane

* brand: keep the traced wordmark; drop Space Mono outline build

* web: IBM Plex Sans Condensed as display face

* brand: Plex Sans Condensed wordmark; nav mark; drop fabricated home demos; AA meta text

* feat(tui): bottom dock tabs — clickable panel switch + close

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* Fix Fleet role migration verification

* fix(tui): dock keys yield Tab to mode/permission cycles

Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Hunter Bown <hmbown@gmail.com>

* web: Impeccable polish — type floors, heading outline, docs measure; add PRODUCT.md/DESIGN.md

* Resolve canonical Fleet roles to legacy members

* web: flat hero — drop cyan glow/gradients/shadow, muted eyebrows

* design: PRODUCT.md/DESIGN.md at repo root — shell direction, bottom dock, anti-slop rules, 0.9.12 tokens

* Auto-enroll used models into the Fleet

* brand: keep the founder's wordmark rasters as the source of truth

The web lane replaced brand/wordmark0901.png and brand/wordmarkinverted.png
with Plex Sans Condensed renders. The founder-supplied PNGs are the brand
source; the SVGs are re-traced from them in a following commit.

* tui(mark): the launch mark has one rung

The hero now paints the small mark over the wordmark, so the medium and
large rungs and the for_area ladder have no consumer and fail the
dead-code lint. Remove them rather than allow them.

* brand: trace the founder's wordmark to SVG

brand/wordmark.svg and wordmark-inverted.svg were an IBM Plex Sans
Condensed text render; the founder's wordmark is the rounded monoline in
brand/wordmark0901.png. scripts/brand/trace-brand.py now traces that PNG
(magick threshold 60% + trim, potrace -s --flat -t 20 -O 0.4 -a 1.2),
folds potrace's transform into one compact path in a tight 1874x264
viewBox, and writes the navy #142352 and white colourways from the same
geometry. The Plex builder scripts/build-wordmark.py is gone with it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* web: derive icons and OG image from the traced mark

app/icon.svg is now the white whale on the #142352 rounded tile as on the
founder's sheet; favicon.ico (48/32/16), apple-icon.png, icon-192.png and
icon-512.png are rasterised from it by scripts/brand/trace-brand.py, and
the manifest colours are the same navy. The social card keeps the navy
ground, white mark and traced wordmark and restores the identity phrase
the page-meta contract expects.

The nav sits on the dark field on every route, so it renders the inline
Whale (white brand ink) and the inverted wordmark instead of a
prefers-color-scheme picture pair; the wordmark box uses object-fit so
the ~7.1:1 trace scales inside the compact nav instead of squashing.

Exploration rasters web/public/brand/codewhale-mark-*.png and their
web/brand/mark tile sources had no consumers and are removed;
codewhale-mark.png stays (public-auth-routes pins its hash).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* web: map stray hard-coded colours to navy tokens

globals.css carried a handful of literal navy-family greys, ice hairlines,
seafoam borders and a cyan glow beside the generated --whale-* tokens.
Each now reads the token it was approximating (whale-bg/chrome/panel,
whale-ice, whale-accent-secondary, whale-action, whale-cyan,
whale-text-dim), and the docs light sheet inks the mark in the brand
navy via --whale-composer (#142352).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* palette: inventory WHALE_* tokens before the one-name-per-colour collapse

Shell design §2.6 (SHELL-DESIGN-20260901) measured "58 WHALE_* symbols;
one colour under five names; 5 dead tokens". Receipt before touching
anything, generated from crates/tui/src/palette/tokens.rs. "uses" is the
whole-word count across crates/ excluding the const's own definition and
`use`/`pub use` lines (wrapper consts inside tokens.rs count).

  name                        value                       alias-of                    dead uses
  WHALE_BG_RGB                (7, 12, 29)                                                    3
  WHALE_CHROME_RGB            (12, 21, 49)                                                   1
  WHALE_PANEL_RGB             (16, 28, 64)                                                   3
  WHALE_COMPOSER_RGB          (20, 35, 82)                                                   3
  WHALE_ELEVATED_RGB          (26, 44, 99)                                                   4
  WHALE_SELECTION_RGB         (30, 60, 143)                                                  3
  WHALE_TEXT_BODY_RGB         (246, 242, 232)                                               10
  WHALE_TEXT_SOFT_RGB         (182, 192, 212)                                                4
  WHALE_TEXT_MUTED_RGB        (147, 160, 184)                                                3
  WHALE_TEXT_HINT_RGB         (138, 153, 179)                                                3
  WHALE_TEXT_DIM_RGB          (105, 119, 145)                                         yes    0
  WHALE_ACTION_RGB            (106, 166, 220)                                                6
  WHALE_COBALT_RGB            (21, 53, 178)                                           yes    0  (web: --whale-cobalt x3)
  WHALE_ICE_RGB               (221, 238, 249)                                         yes    0  (web: --whale-ice, rustRgb("WHALE_ICE"))
  WHALE_CYAN_RGB              (120, 188, 232)                                                2
  WHALE_ACCENT_SECONDARY_RGB  (79, 209, 197)                                                11
  WHALE_BRAND_ORANGE_RGB      (255, 138, 61)                                                 1
  WHALE_BRAND_MAGENTA_RGB     (240, 78, 184)                                                 1
  WHALE_HUMAN_RGB             (246, 196, 83)                                                 5
  WHALE_ACCENT_PRIMARY_RGB    = WHALE_ACTION_RGB          WHALE_ACTION_RGB                   8
  WHALE_WORKING_GREEN_RGB     (155, 214, 111)                                                5
  WHALE_ACCENT_ACTION_RGB     = WHALE_ACTION_RGB          WHALE_ACTION_RGB            yes    0
  WHALE_ERROR_RGB             (255, 134, 178)                                                9
  WHALE_ERROR_HOVER_RGB       (255, 156, 194)                                                3
  WHALE_ERROR_SURFACE_RGB     (43, 21, 34)                                                   6
  WHALE_ERROR_BORDER_RGB      = WHALE_ERROR_RGB           WHALE_ERROR_RGB                    3
  WHALE_ERROR_TEXT_RGB        (255, 219, 232)                                                3
  WHALE_WARNING_RGB           (255, 122, 89)                                                 4
  WHALE_SUCCESS_RGB           = WHALE_WORKING_GREEN_RGB   WHALE_WORKING_GREEN_RGB            9
  WHALE_INFO_RGB              = WHALE_ACTION_RGB          WHALE_ACTION_RGB                  17
  WHALE_BORDER_RGB            (42, 63, 114)                                                  1
  WHALE_REASONING_TEXT_RGB    (224, 153, 72)                                                13
  WHALE_REASONING_SURFACE_RGB (42, 34, 24)                                                   3
  WHALE_REASONING_TINT_RGB    (22, 36, 74)                                                   7
  WHALE_DIFF_ADDED_RGB        (87, 199, 133)                                                 3
  WHALE_DIFF_DELETED_RGB      = WHALE_ERROR_RGB           WHALE_ERROR_RGB             yes    0
  WHALE_DIFF_ADDED_BG_RGB     (18, 42, 34)                                                   3
  WHALE_DIFF_DELETED_BG_RGB   (52, 24, 39)                                                   3
  WHALE_MODE_AGENT_RGB        (126, 180, 232)                                                4  (via MODE_AGENT: 13)
  WHALE_MODE_YOLO_RGB         (255, 112, 160)                                                4  (via MODE_YOLO: 13)
  WHALE_MODE_PLAN_RGB         (185, 220, 236)                                                4  (via MODE_PLAN: 13)
  WHALE_MODE_OPERATE_RGB      (173, 136, 255)                                                4  (via MODE_OPERATE: 14)
  WHALE_TOOL_LIVE_RGB         = WHALE_ACCENT_SECONDARY_RGB                                   3  (via ACCENT_TOOL_LIVE: 5)
  WHALE_TOOL_ISSUE_RGB        = WHALE_ERROR_RGB                                              3  (via ACCENT_TOOL_ISSUE: 5)
  WHALE_TOOL_OUTPUT_RGB       = WHALE_TEXT_SOFT_RGB                                          3  (via TEXT_TOOL_OUTPUT: 13)
  WHALE_TOOL_SURFACE_RGB      (15, 26, 58)                                                   3  (via SURFACE_TOOL: 5)
  WHALE_TOOL_ACTIVE_RGB       (24, 44, 94)                                                   3  (via SURFACE_TOOL_ACTIVE: 9)
  WHALE_ACCENT_PRIMARY        Color(WHALE_ACCENT_PRIMARY_RGB)  -> WHALE_ACTION               9
  WHALE_ACTION                Color(WHALE_ACTION_RGB)                                       85
  WHALE_LIVE                  Color(WHALE_ACCENT_SECONDARY_RGB)                             17
  WHALE_HUMAN                 Color(WHALE_HUMAN_RGB)                                        41
  WHALE_INFO                  Color(WHALE_INFO_RGB)       -> WHALE_ACTION                  105
  WHALE_BG                    Color(WHALE_BG_RGB)                                           96
  WHALE_CHROME                Color(WHALE_CHROME_RGB)                                        5
  WHALE_PANEL                 Color(WHALE_PANEL_RGB)                                        15
  WHALE_COMPOSER              Color(WHALE_COMPOSER_RGB)                                      5
  WHALE_ERROR                 Color(WHALE_ERROR_RGB)                                        35

  57 WHALE_* consts. Pure aliases (9): ACCENT_PRIMARY_RGB, ACCENT_ACTION_RGB,
  ERROR_BORDER_RGB, SUCCESS_RGB, INFO_RGB, DIFF_DELETED_RGB, TOOL_LIVE_RGB,
  TOOL_ISSUE_RGB, TOOL_OUTPUT_RGB. #[expect(dead_code)] (5): TEXT_DIM_RGB,
  COBALT_RGB, ICE_RGB, ACCENT_ACTION_RGB, DIFF_DELETED_RGB.
  Non-WHALE aliases of the same blue in tokens.rs: STATUS_INFO (8 uses),
  ACCENT_PRIMARY (dead, 0). One colour, #6AA6DC, under seven symbols:
  WHALE_ACTION(_RGB), WHALE_INFO(_RGB), WHALE_ACCENT_PRIMARY(_RGB),
  WHALE_ACCENT_ACTION_RGB, STATUS_INFO, ACCENT_PRIMARY — 225 call sites.

Script: python3 over tokens.rs + grep -rnw crates; kept out of scripts/
(one-off receipt, the numbers live here).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* palette: one name per colour — collapse WHALE_INFO / WHALE_ACCENT_PRIMARY into WHALE_ACTION

Shell design §2.6: "`WHALE_INFO`, `WHALE_ACTION`, `WHALE_ACCENT_PRIMARY`
and their `_RGB` twins are one colour. Collapse to `WHALE_ACTION`."

Mechanical rename across crates/ (word-boundary sed, no value changes):
  WHALE_INFO, WHALE_INFO_RGB           -> WHALE_ACTION, WHALE_ACTION_RGB
  WHALE_ACCENT_PRIMARY(_RGB)           -> WHALE_ACTION(_RGB)
  palette::STATUS_INFO                 -> palette::WHALE_ACTION
  WHALE_ACCENT_ACTION_RGB, ACCENT_PRIMARY (dead aliases) -> deleted
The `STATUS_INFO` static in commands/groups/config is an unrelated
CommandInfo and is untouched.

Where two names met in one predicate (adapt.rs light/solarized/community
remaps, grayscale text-soft bucket, SemanticForegroundRole::Action) the
duplicate disjuncts are dropped; `use` lists deduped; the
"primary accent aligns with action" test collapses to its one live
assertion (action blue != human gold). The Blue Stage doc comment moves
onto WHALE_ACTION_RGB. `palette::grammar` untouched: it reads UiTheme
slots, never these consts.

Evidence (CARGO_BUILD_BUILD_DIR=.../mega-tokens):
  cargo check -p codewhale-tui --all-targets      -> Finished, 0 warnings
  cargo clippy -p codewhale-tui --all-targets --all-features --locked
    -- -D warnings (CI allow-list)                -> clean
  cargo test -p codewhale-tui --lib palette::tests:: --locked -- --skip command_palette
    -> test result: ok. 59 passed; 0 failed
  RUST_MIN_STACK=67108864 cargo test -p codewhale-tui --lib startup_ink --locked
    -> test result: ok. 1 passed; 0 failed   (ink goldens unchanged)
  cargo test -p codewhale-tui --lib --locked -- menu_style cursor_accent color_compat
    -> test result: ok. 36 passed; 0 failed
Pre-existing, not from this diff (reproduced on the stashed tree):
tui::command_palette tests, feat012_ac1 and the startup_ink golden
overflow the default test-thread stack in a debug build; they pass with
RUST_MIN_STACK=64MiB.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* palette: delete the dead whale tokens

Shell design §2.6: "delete the five dead tokens". Two of the five went
with the alias collapse (WHALE_ACCENT_ACTION_RGB, ACCENT_PRIMARY); this
removes the rest that have no consumer in crates/ and no web consumer:

  WHALE_TEXT_DIM_RGB     (105,119,145)  0 uses, no --whale-text-dim on the site
  WHALE_DIFF_DELETED_RGB = WHALE_ERROR_RGB  0 uses, no --whale-diff-deleted
  ACCENT_SECONDARY       Color(WHALE_ACCENT_SECONDARY_RGB)  0 uses (TEXT_ACCENT/WHALE_LIVE carry it)
  STATUS_NEUTRAL         = TEXT_MUTED  0 uses

Kept, with a comment saying why the `#[expect(dead_code)]` is honest:
WHALE_COBALT_RGB and WHALE_ICE_RGB have no TUI consumer but the site
reads them through the token export (`--whale-cobalt` x3, `--whale-ice`
and `rustRgb("WHALE_ICE")` in web/lib/blue-stage-contract.test.ts).

Mode and tool-surface `_RGB` tuples stay: each is consumed through its
Color wrapper (MODE_AGENT/YOLO/PLAN/OPERATE 13-14 call sites each,
themes.rs + color_compat.rs; SURFACE_TOOL 5, SURFACE_TOOL_ACTIVE 9,
ACCENT_TOOL_LIVE 5, ACCENT_TOOL_ISSUE 5, TEXT_TOOL_OUTPUT 13). The §1
"12 tokens with zero consumers" counted the tuples, not their wrappers.

Evidence: cargo check -p codewhale-tui --all-targets -> Finished, 0 warnings;
cargo test -p codewhale-tui --lib palette::tests:: --locked -- --skip command_palette
-> test result: ok. 59 passed; 0 failed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* web: regenerate tokens.css after the whale token collapse

scripts/export-design-tokens.py (never hand-edited). Ten lines gone:
--whale-accent-primary(-rgb), --whale-accent-action(-rgb),
--whale-info(-rgb), --whale-text-dim(-rgb), --whale-diff-deleted(-rgb).
No site stylesheet or component consumed any of them
(grep -rn "\-\-whale-" web/app web/components web/lib); the only
reference was the alias-chain example in web/lib/whale-tokens.ts's doc
comment, now `--whale-success` -> `--whale-working-green` -> `#9bd66f`
(the old example also quoted a hex that stopped being true a retune ago).

Evidence:
  cd web && python3 ../scripts/export-design-tokens.py --check   (CI: npm run check:tokens)
    -> design tokens up to date (1 file(s), 42 tokens)
  vitest run lib/blue-stage-contract lib/docs-theme-contract
    -> Test Files 2 passed (2) / Tests 6 passed (6)
  (vitest ran against the main checkout's node_modules via a temporary
  symlink; this worktree has none installed.)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* docs(design): the status-bar grammar names the one blue token

STATUS_BAR_COLOR_GRAMMAR.md never named a collapsed token, so this is the
one sentence it needed: the Identity blue is `WHALE_ACTION`, its former
aliases (`WHALE_INFO`, `WHALE_ACCENT_PRIMARY`, `STATUS_INFO`) are gone, and
the whale theme's `info` / `accent_primary` slots both hold it. No other
document in the repo named them (grep over *.md, *.ts, *.tsx, *.css,
*.py, *.toml, *.yml, *.json outside node_modules); the root DESIGN.md
already speaks in CSS names.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* palette: the field follows the terminal-owned shell; `underwater` aliases deepsea

Shell design §2.0 decision 1 (founder: "We aren't supposed to be using a
blue background anymore"): ground is the terminal's; the navy field is
painted only under the opt-in deepsea column.

What was already true, verified before changing anything:
- Settings::default().theme is "terminal" (settings.rs:103).
- The whale pair (UI_THEME / LIGHT_UI_THEME) ends in
  `.with_terminal_native_shell()`: surface, panel, composer, header and
  footer are `Color::Reset`, pinned by
  `whale_pair_flat_shells_are_terminal_native_without_erasing_semantic_surfaces`.
- The ink goldens' legend reads `a reset on reset`.
- OceanTreatment::Flat is the default; Deepsea repaints Reset cells through
  OceanRamp::for_theme, which matches the whale pair by name + Reset shell.
The reviewer's citations (tokens.rs:6/:250/:465) are the token definitions
deepsea and the semantic surfaces still need, not the theme.

What was not true: ~90 direct `bg(palette::WHALE_BG)` paints in pickers,
overlays and full-screen views (provider_picker 14, views/mod.rs 11,
user_input, live_transcript, help, session/file/model pickers ...) bypass
the theme, and `adapt_bg_for_theme` only remapped them for
`theme_remap_active` presets. On the whale theme they laid navy patches
over the terminal ground. Rung 2 fix, one rule in palette/adapt.rs: the
field (`WHALE_BG` / `BACKGROUND_DARK`) always follows `ui.surface_bg` —
Reset on the whale pair, the user's `background_color` override when set,
the preset surface elsewhere. Panels, selection, elevation, error and
diff surfaces are untouched; no widget file changes.

`underwater` is now an accepted alias of `deepsea` in
settings.rs (normalize + `set`), OceanTreatment::parse and the
config_ui serde enum. Tests extended in place; the color_compat light
test now expects the Reset shell it already had for theme consumers.

DESIGN.md "Field" says the TUI ground is the terminal's own background
and the navy field is deepsea-only.

Contrast, all whale text/accent tokens on #000, #1e1e1e (VS Code),
#282c34 (One Dark), #002b36 (Solarized dark), #300a24 (Ubuntu), #0d1117
(GitHub dark), #282a36 (Dracula): body 12.5-18.8:1, soft 7.7-11.5, muted
5.3-8.0, hint 4.9-7.3 (floor 3:1), action 5.4-8.1, human 8.6-12.9, live
7.5-11.3, error 6.2-9.3, warning 5.5-8.2, green 8.2-12.2, reasoning
5.9-8.8. Only `border` (#2A3F72) is low, 1.4-2.1:1: a non-text hairline.
No token value changed.

Seen, tmux 80x24 PTY, TERM=xterm-256color COLORTERM=truecolor, hermetic
HOME, debug build, counting `48;2;R;G;B` background sequences:
  default (Terminal theme): startup, /theme picker, Help — no painted
    background before or after (picker shows only accent swatches).
  Blue Stage selected via T/Down/Enter, then F1 Help:
    before: 15 x `48;2;7;12;29` (WHALE_BG) + 1 x selection row
    after:  1 x `48;2;30;60;143` (selection row only)
  Startup stage on Blue Stage: none, before and after.

Evidence (CARGO_BUILD_BUILD_DIR=.../mega-tokens, RUST_MIN_STACK=16 MiB as CI):
  cargo check -p codewhale-tui --all-targets -> Finished, 0 warnings
  cargo test -p codewhale-tui --lib --locked -- color_compat palette::tests::
    ocean:: ocean_treatment live_transcript views::tests startup_ink
    --skip command_palette -> test result: ok. 238 passed; 0 failed
  cargo test -p codewhale-tui --lib --locked (full) ->
    test result: FAILED. 11901 passed; 7 failed; 13 ignored
    1 was this change (color_compat light test, updated above); the other 6
    are role-name / slash-list assertions from other lanes on this branch
    (scout<->explore, worker<->general, slash.impeccable) and untouched.
Ink goldens unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* tests(palette_audit): re-pin whale roles to the retuned palette

`whale_roles_are_pinned_and_non_colliding` still pinned the pre-navy
values (WHALE_BG (3,7,13), ACTION (106,174,242), ...) and failed on this
branch before the token slice touched anything. Pins now match tokens.rs;
no colour value changes.

cargo test -p codewhale-tui --test integration --locked palette_audit
  -> test result: ok. 3 passed; 0 failed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* tui/cli/web: fleet is the public product term; /pod, codewhale pod stay aliases

Founder decision 2026-09-01: "fleet" is the customer-facing name for the
assembled model team and "Pod" is retired from product copy. `/fleet` is
the canonical slash command and `codewhale fleet` the canonical CLI verb;
`/pod`, `codewhale pod`, `loadout`, and `party` remain parser aliases.
Storage keys, the ledger file name, config tables, protocol identifiers,
and MessageId variant names keep their current spelling.

- CommandInfo name/aliases/usage, help text, and the unknown-verb error
  flip to /fleet; `/fleet fleets` (saved/manage) is the saved-fleet picker
  with `/fleet pods` kept as an alias.
- All 15 locale packs: localized values say fleet; the settings goldens
  follow. `KbCompleteCycleModes` names the modes as Plan → Work → Operate
  (Act is only a compatibility alias per docs/MODES.md).
- `scripts/check-tui-product-vocabulary.sh` now rejects `Pod` in en.json
  instead of rejecting `fleet` in every pack.
- Hotbar id `slash.fleet` is canonical; persisted `slash.pod` normalizes.
- Fleet store error prose says fleet.
- Docs: PRODUCT.md lists the current role tokens (general, explore,
  planner, reviewer, implement, test, advisor, custom) and names the old
  spellings as aliases; docs/FLEET.md uses one role vocabulary;
  web/lib/content/vocabulary.ts ADVISORY_ROLE is Advisor with consultant/
  oracle as the legacy spellings (matches fleet/profile.rs migration).

Evidence:
  sh scripts/check-tui-product-vocabulary.sh -> exit 0
  cargo test -p codewhale-config -p codewhale-lane --locked
    -> 638 passed; 0 failed / 62 passed; 0 failed
  cargo test -p codewhale-cli --locked -- fleet pod -> 3 passed; 0 failed
  cargo test -p codewhale-tui --lib --locked -- fleet::store fleet::members
    fleet::identity -> 24 passed; 0 failed
  cargo test -p codewhale-tui --lib --locked -- groups::core::fleet
    localization command_palette hotbar fleet_roster settings widgets
    fleet::control pod_workers -> 605 passed; 1 failed (the failure is
    slash_source_matches_command_palette_command_entries, which reads the
    machine's ~/.claude/skills and finds an `impeccable` skill; it fails
    identically without this change)
  cd web && npm test -- lib/content/vocabulary.test.ts -> 11 passed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* chore(tui): clear the six clippy 1.98 errors the base lanes left

needless_borrow on provider_identity_for_persistence (five sites) and a
collapsible_if in the work-surface mouse path. No behaviour change.

* tui(composer): restore double-tap Enter as the send-now gesture

While a turn is running, the first bare Enter queues the message (as
before) and opens a 500 ms window (`App::DOUBLE_TAP_WINDOW`, the value
the removed code in c5c42b7d91 used). A second bare Enter inside that
window with an empty composer promotes the just-queued message to a
Steer through `attempt_steer_with_queue_fallback` — the same path
Ctrl+Enter takes, so there is one steering path. A second Enter with
new text is an ordinary queue; Ctrl+Enter still steers immediately;
outside a turn Enter is unchanged. `enter_with_double_tap` is the one
decision point again (`take_queued_for_double_tap_steer` routes through
it), and `submit_disposition_does_not_mutate_the_queue` stays true.

The posture bar advertises the gesture while the window is open
(`PostureHintEnterAgain`, next commit).

Tests (cargo test -p codewhale-tui --lib <filter> --locked):
  double_tap:          test result: ok. 3 passed; 0 failed
  enter_with:          test result: ok. 5 passed; 0 failed
  submit_disposition:  test result: ok. 6 passed; 0 failed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* tui(shell): one owner per fact — posture bar, metrics line, no dead hints

Design: SHELL-DESIGN-20260901 §2.0 item 3, §2.2, §2.3, §2.3b, §2.11 and
the founder's 2026-09-02 redirect (Claude Code's grammar, less always-on
information). Under the composer there are exactly two chrome rows in
the default state, then the work surface only when it has content:

  ▶▶ ask (Shift+Tab) · work (Tab) · 2 agents · Esc to interrupt   /rc …
  deepseek-v4 · ctx 61% · $0.42 · ttft 400ms · 40 tok/s · ↓ 1.2K  Ctrl+/ help

Fact → owner, before → after (composed 80x24 / 120x32 frames, working
turn with two sub-agents; "strip" = the work-surface bottom view):

  fact              before                                        after
  context %         info line                                     metrics line
  cost              footer (+ info line when priced)              metrics line
  model             info line                                     metrics line
  provider          info line (wide)                              metrics line (wide)
  ttft / tok/s / ↓  behind /cost only                             metrics line
  repo slug         info line (+ idle empty state)                launch header / git view (not chrome)
  branch            info line (+ idle empty state)                launch header / git view (not chrome)
  mode              footer                                        posture bar
  permission        footer                                        posture bar
  phase word        footer ("sub-agents underway")                transcript active row (not chrome)
  elapsed           footer ("1m 15s")                             roster rows (per agent)
  agent count       indicator row + info "pod 2/2" + "whales 2/4" posture bar ("2 agents")
                    + dock tab + strip header                     (+ dock tab, strip header — other slice)
  task/shell count  indicator row above the composer              posture bar
  help hint         footer keys legend (F1) + info line (Ctrl+/)  metrics line (Ctrl+/), from the binding
  other key hints   footer "⌥V:output", compact "? help"          none; cycle keys next to the chip they cycle
  live hint         footer "Esc to interrupt"                     posture bar hint slot
  ≥80 % microcopy   footer right slot                             posture bar hint slot (outranks the hint)
  notice / rc       footer right slot / —                         posture bar right slot

Dead key hints removed: `F1:keys` / `fn+F1:keys` (Help binding's
`footer_chord` is now `Ctrl+/`; `info_help_hint` derives from the
binding), compact `? help`, and the `footer_action_hints` family. The
mode/permission cycle keys print only when the binding table admits
them at the current focus (no `(Tab)` on the launch stage).

Row order: composer → posture bar → metrics line → roster/to-do. The
#5286 background-work chip above the composer is gone (it repeated the
posture bar's counts); `PendingWork` stays as the counts' source.

Goldens re-blessed and read: footer_* (posture bar), infoline_startup_*,
infoline_work_* (metrics line), settings_* (the settings preview's
bottom row); infoline_settings_* deleted with the settings-path segment.

Commands run (CARGO_BUILD_BUILD_DIR=…/mega-frame, RUST_MIN_STACK=16777216):
  cargo check -p codewhale-tui --all-targets            clean
  cargo test -p codewhale-tui --lib infoline --locked   test result: ok. 11 passed; 0 failed
  cargo test -p codewhale-tui --lib tideline_tests      test result: ok. 64 passed; 0 failed
  cargo test -p codewhale-tui --lib one_owner_tests     test result: ok. 4 passed; 0 failed
  cargo test -p codewhale-tui --lib shell_key_routing   test result: ok. 13 passed; 0 failed
  cargo test -p codewhale-tui --lib localization::tests test result: ok. 49 passed; 0 failed
  cargo test -p codewhale-tui --lib --locked            test result: FAILED. 11893 passed; 8 failed
                                                        (config_panel golden re-blessed after; the other 7:
                                                        4 fail on HEAD without this change (fleet rename
                                                        in flight), tmux clipboard passes alone, none in
                                                        files this change touches)
  cargo clippy … -D warnings                            6 pre-existing errors, none in this change's hunks
                                                        (config.rs:2106/2796, apply.rs:759, event_loop.rs:464,
                                                        session_state.rs:1004, work_surface/input.rs:401)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* wip(launch): checkpoint for overnight takeover — braille mark + kitty tier + Claude-Code launch header compile and pass lib tests; PTY cucumber tests still wait on the old heading

Compiles (cargo check --all-targets clean). Lib tests for mark::, underwater::,
mouse_ui::, localization:: pass: test result: ok. 120 passed; 0 failed
(needs RUST_MIN_STACK=16777216 like scripts/dev-test.sh; the rust_i18n static
overflows a 2 MiB test thread with or without this change). Startup goldens
re-blessed and read. Clippy is red only in files outside this slice
(config.rs, apply.rs, session_state.rs, work_surface/input.rs, and a
pre-existing event_loop.rs borrow).

Not done: crates/tui/tests/cucumber/{screen_mode_inline_pty,
active_composer_pointer_pty,plugin_e2e_acceptance}.rs still wait for
"What are we working on?" and press 'w'; they need the new marker
("Codewhale v") and a typed message + Enter to begin the session.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* wip(rail): checkpoint for overnight takeover — dock views compile, 136/137 work_surface tests pass, files/notepad/git views are stubs

Foundation for the one bottom dock (founder redirect 2026-09-02):
RailPanel is now the eight-view cycle (agents, tasks, background, files,
notepad, context, git, price; Pinned folded into tasks), an auto rule opens
agents/tasks/background while they have content, explicit picks stick until
Esc, and Ctrl+Tab / Ctrl+] (fwd) + Ctrl+Shift+Tab (back) cycle. Context and
price views render as rows; files, notepad, git are stubs in views.rs. The
classic sidebar line panels and their dead consumers are deleted.

Known: agent_rows_show_role_assignment_and_open_the_agent_transcript fails
(role_label 'worker' vs 'general'); role derivation is untouched here and
the failure is believed to predate this work — unverified.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* wip(operate): checkpoint for overnight takeover — auto-goal + contract land; one Operate approval test needs a goal-complete mock

Operate now turns a non-trivial prompt into the goal through the same
GoalState::create path as explicit_goal_directive, appends the Operate
contract once as a user-role runtime message (append-only history, not
the prefix), shows the Operate goal receipt, and carries the new mode copy
in all 15 locale packs plus docs/MODES.md.

Compiles. Passing: goal (134), prompts (133, incl.
every_mode_shares_one_prompt_per_host), localization (49),
runtime_handoff (14), session_peek (15), history_cells (2), both new
engine tests. Known failing:
core::engine::tests::operate_model_shell_uses_normal_approval_and_workspace_sandbox
— its mocked model never reports the auto-set goal complete, so the turn
re-prompts to max_steps (wiremock expect(1) sees 199). Six clippy
needless_borrow/collapsible_if hits pre-exist on the branch base.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* wip(fleet): checkpoint for overnight takeover — #5815 review findings 1-9 fixed, compiles, fleet tests green

Findings against the fleet-as-models work (verified against the current
tree, then fixed):

1. `toggle_fleet_model` decides presence by member rows, not the projected
   role list (a role-less row projects to no role, so the old
   `all(== "operator")` was vacuously true). Regression test
   `toggle_removes_a_role_less_member_row` fails on the previous commit
   ("got Unchanged { … operator route … }") and passes now.
2. `selected_or_default` loads an existing personal `My fleet` instead of
   overwriting it and never writes or selects before the add succeeds;
   `FleetModelChange::Added` carries `created_fleet` + `selected_fleet`.
3. `fleet_models` returns `Result<Vec<_>, FleetStoreError>`: a broken
   explicit selection is surfaced in `/fleet models`, `/models`, and the
   picker's ⇧F instead of reading as "session model only".
4. `add_fleet_model` dedupes roles (case-insensitive) and returns
   `Unchanged` without touching the file when every role is present
   (test compares bytes and mtime).
5. `App.config` startup snapshot removed. `/fleet add|remove` now return
   `AppAction::FleetAddModel|FleetRemoveModel`; the UI arm validates the
   provider against the live `Config` (`fleet_provider_rejection`,
   `fleet_catalog_rejection`, re-exported from `commands`).
6. ⇧F applies the same provider gate as `/fleet add`.
7. One roster path: `sync_fleet_roster` (extracted from the
   FleetStoreChanged arm) plus `App::fleet_roster_stale`, flushed once per
   event-loop iteration; `/fleet add|remove`, ⇧F, and every UI-side
   auto-enroll site set it (`auto_enroll_fleet_model` now returns bool).
8. ⇧F receipts go through `push_status_toast` (Success/Info, 6 s) and
   `set_sticky_status` (Error); no new `status_message` writes.
9. All new fleet prose is `tr(locale, MessageId::Fleet…)` (27 keys,
   translated in all 15 packs); `FleetModelError` is typed with a
   localized `message(locale)`.
10. No stale "`a` in /models" doc comment exists in the current tree.

Also re-blessed `config_panel_{80x24,120x32}` goldens (the Config tab
label says Fleet); the diff is that one label.

Evidence (this tree):
  cargo check -p codewhale-tui --all-targets --locked -> Finished
  cargo test -p codewhale-tui --lib --locked -- fleet::members
    groups::core::fleet model_picker format_helpers fleet_roster
    localization golden hotbar command_palette fleet::store
    -> 257 passed; 1 failed (slash_source_matches_command_palette_
    command_entries: reads ~/.claude/skills and finds `impeccable`;
    fails identically on main in this environment)
  cargo clippy … -D warnings (CI flags) -> the only remaining error is
    crates/tui/src/tui/work_surface/input.rs:401 collapsible_if, which
    belongs to the work_surface lane and predates this commit
  cargo fmt --all -- --check -> clean
  sh scripts/check-tui-product-vocabulary.sh -> exit 0

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HSVsAXZJnKGZmqkwH1CeKH

* feat(tui): launch card, canonical role vocabulary, DashScope descriptor, test fixes

- Launch is now our own card take (founder, 2026-09-02): thin top line
  ⑂ branch  path; centred bordered card with the whale mark, Codewhale +
  version, one true announcement (no-model warning / MCP news), and the
  menu New worktree / Resume session / Changelog / Quit with real chords
  right-aligned; Enter runs the highlighted entry, Up/Down move it, and
  typing goes straight to the composer. The card dissolves on the first
  keystroke or command (≤240ms, instant under reduced motion); the
  working screen then shows ⑂ branch  path + ⋮ MCP n/m, the
  session_start receipt, and the posture bar + metrics line (hidden
  while the card is up). The composer's bottom rule carries
  model (effort) · permission — the route's one launch reading.
- Role vocabulary: fixtures and the stopship fleet/workflow now use the
  canonical tokens (explore/implement/reviewer/test); the workflow JS
  wire accepts canonical spellings with the pre-rename ones as aliases
  (AgentType serde rename+alias, serialized form is canonical).
- Alibaba Model Studio (DashScope) joins the data-driven descriptor
  table: international compatible-mode endpoint, DASHSCOPE_API_KEY,
  live /v1/models as the Qwen model authority (never a compiled id).
- Tests: role-keyed gate fixtures moved to canonical tokens; the operate
  model-shell test now seals the goal through the deferred update_goal
  tool (deferral retry included) instead of pausing blindly; the
  slash-source hotbar test isolates HOME; ⚠ and ⋮ gained ASCII
  fallbacks; launch goldens re-blessed for the card.

* feat(tui): retire Pod from copy; canonical workflow fixture; gate clean-up

- Pod literal sweep across fleet views, worker runtime, sub-agent tool,
  managed API, and command groups: user-facing copy now says Fleet
  (founder vocabulary decision; /fleet canonical, /pod and
  'codewhale pod' stay as compatibility aliases). Roster tests that
  encoded the retired Pod-public/Fleet-internal split now assert the
  public Fleet vocabulary.
- workflows/stopship + fleets/stopship use canonical role names
  (explore/implement/reviewer/test); the workflow crate's own stopship
  tests and required-roles list follow.
- Operate mode-picker hint shortened to fit 80 columns in every locale.
- Cucumber PTY launch flows: wait for the launch card, type the first
  prompt and press Enter; the live shell is proven by the launch stage's
  top line disappearing and the metrics line's ctx reading (the
  interrupt hint needs a live turn, which an offline route never
  starts). The stopship acceptance feature expects the canonical /fleet
  help copy.
- CHANGELOG receipts synced; DESIGN.md shell direction records the card,
  posture bar + metrics line, and the bottom view cycle.

* test(cucumber): launch-card PTY contract fixes

- The launch-card wait uses the menu's New worktree entry — unique to the
  card; the bare wordmark also matches onboarding copy.
- The live-shell proof is the launch stage's top line disappearing plus
  the metrics line's ctx reading; the interrupt hint needs a live turn,
  which an offline route never starts, and the help hint sheds first at
  the 40-column floor by design (SHELL-DESIGN-20260901 §2.2).
- The pointer-submit queue proof takes the offline onboarding seed into
  account: the receipt toast proves the gesture, and the queue count
  grows by exactly one.

* docs(readme): restore the canonical product screenshot the web contract pins

The brand header redesign dropped the assets/screenshot.webp embed; the
web public-surface contract pins the README and the website to one
canonical optimized screenshot (byte-identical, 1562x1256 lossless
VP8L). Re-embed it.

* test: platform-robust pointer-submit wait and cap-warning diagnostics

- The pointer queue proof accepts either the transient receipt toast or
  the queue-count increment: toast timing differs across runners, and a
  20 s wait missed a toast the queue dump proved had fired.
- The context-cap posture test dumps the drawn rows when the warning
  count misses, instead of a bare 0 != 1, so a platform-specific shed
  (the hint sheds first when the left run exceeds its budget) is
  visible in CI.

* test: fix the linux-only context-cap shed; bounded pointer-click retry

- The context-cap posture test drew at 100 columns, where a backend-less
  platform (linux CI paints 'files: workspace (unenforced)') sheds the
  cap hint first, so the warning count read 0. Draw at 140 columns,
  where the hint survives with the notice present; verified locally.
- The pointer queue proof retries the [↑] click once, re-finding the
  affordance first: under runner load a redraw can shift cells between
  the find and the click, so the first SGR gesture lands nowhere.

* fix(gates): tool-catalog budget covers the fleet rename; readme stamps; pointer baseline

- The Pod->Fleet sweep grew every mode's tool-schema surface by 58 bytes
  (+14 tokens). The receipts are re-measured and the one-way ceilings in
  scripts/runtime-contract-budget.json are raised to them as the
  explicit maintainer decision the gate asks for (the rename is the
  founder's 2026-09-01 vocabulary call).
- The README screenshot embed changed README.md; the 18 translated
  READMEs re-stamp (the embed is language-neutral HTML - no prose
  changed, so no retranslation was needed).
- The pointer queue baseline is captured while the composer is empty:
  the pending preview row hides while a draft sits in the composer, so
  the pre-click depth read None and the growth proof could not fire.

* test: re-click then keep polling until the deadline

The qa_harness Instant wrapper does not implement Div, and the retry's
single read raced the app processing the second gesture: poll to the
full deadline, re-click once at the half-way point.

* test: pointer queue diagnostics (baseline/expected/last-seen) in the failure output

* test: pointer poll keeps per-iteration state only (unused-assignment gate)

* fix: Copilot review findings — planner wire spelling and Advisor copy

- workflow::AgentType::Plan serializes as the canonical 'planner'
  ('plan'/'awaiter' stay accepted aliases), matching the FleetRole
  vocabulary the mega PR declares.
- Web: the vocabulary docs metadata, the vocabulary module header, and
  the docs-map topic description say Advisor (the public advisory term)
  instead of the retired Consultant spelling.
- Polish home dictionary: restore 'Podwodna powłoka terminala' — the
  fleet-vocabulary sweep had merged 'Fleet' into the compound word
  'Podwodna' (underwater), producing the non-word 'fleetwodna'.

* test: pointer proof accepts preview-appears when no baseline count is painted

* test: the tolerant preview-appears proof (the arm the last commit missed)

---------

Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Signed-off-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: CodeWhale Bot <bot@codewhale.net>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-02 09:32:14 -07:00

22 KiB

Modes and Permission Postures

阅读简体中文版:zh_hans/MODES.md

Codewhale has three related concepts:

  • TUI mode: what kind of visible interaction you're in (Plan/Work/Operate).
  • Permission posture: how aggressively the UI asks before executing tools.
  • Workflow overlay: optional long-running orchestration that can run on top of any TUI mode when a task needs many coordinated workers.

Model selection is separate. --model auto and /model auto route each turn to a concrete model and thinking level; they are not TUI modes and are not part of the Tab cycle.

Workflow is also separate from the mode itself. It is the visible ordered orchestration layer for repeatable workflows and fleet workers. High fan-out routes through durable fleet-backed workers instead of prompt-only sub-agent fanout. The active mode still controls permissions; Workflow controls whether a large task is planned into a resumable workflow with its own progress view.

TUI Modes

Press Tab to complete composer menus or cycle through the visible modes when the composer is empty: Plan → Work → Operate → Plan. Tab never sends or queues composer text; use Enter to send or queue it. Press Shift+Tab to cycle permission posture (Ask → Auto-Review → Full Access). Press Ctrl+T to cycle reasoning effort. Run /mode to open the mode picker, or switch directly with /mode work, /mode plan, or /mode operate.

  • Plan: design-first prompting. The stable primitive names remain familiar, but the runtime centrally refuses file mutation and shell execution. Read-only inspection and policy-allowed research, including deferred Web search/fetch, remain available.
  • Work (internally agent): ordinary multi-step execution. The small first-turn toolbox is read, write, edit, bash, agent, and todo_write; approval, sandbox, repository law, and managed policy still decide what may execute.
  • Operate: multitask conductor posture. Operate turns your prompt into a goal and works it in parallel: background workers for separable streams, verified before it stops. It has the same primitive identities and execution authority as Work. When no unfinished goal exists, a submitted prompt that is real work (not a greeting, acknowledgement, or short question) becomes the session goal automatically, with continuation on; the transcript shows ◆ goal set · Operate keeps working until it is verified · /goal to edit. An explicit /goal declaration still wins, /goal still edits it, and an existing goal is never replaced. The parent session is the operator: dispatching background workers is the default for independent or parallel work. Handle small or tightly coupled tasks in the parent; use background agent workers for separable streams, and use Workflow when order, phases, gates, shared budgets, or deterministic fan-in matter. Dispatch is not completion — write-capable children must return real verification evidence. The first Operate turn of a session appends this contract once as a user-role runtime message (append-only history, never the pinned system prompt), so Plan, Work, and Operate keep one shared prompt prefix.

Act and /mode act remain compatibility aliases for Work. Saved settings still normalize to the internal value agent.

Tool availability by mode

Tool family Plan Work Operate
read and policy-allowed deferred research tools yes yes yes
write and edit visible names; execution denied approval- and policy-gated same as Work
bash visible name; execution denied approval- and policy-gated same as Work; delegation is preferred when parallelism or isolation helps
agent yes, subject to child-depth authority yes, subject to child-depth authority yes, subject to child-depth authority
Deferred native, MCP, and plugin tools discoverable through tool_search when policy permits same same
Paid or external-service tools follows permission posture follows permission posture follows permission posture
Access outside the workspace root explicit trusted paths only only through trusted paths or trust mode same trusted-path/trust policy as Work; fleet profiles never widen it

Operate changes scheduling emphasis, not authority. It neither adds a mode-specific tool denial nor bypasses the active approval, sandbox, shell, ask-rule, repository-law, or managed-policy boundary. Plan remains the mode-specific execution boundary for shell and write-capable tools; that authority difference does not require a different primitive vocabulary.

Operate loop (one screen)

User message
  → small / chat / one-file?  → parent does it (Work-equivalent tools)
  → real / multi-stream work? → goal (set from the prompt) → dispatch background workers
       → each write child: implement → VERDICT PASS/FAIL with evidence
       → ordered / gated fan-in? → Workflow (operate_* starters)
       → high-stakes ambiguous? → best-of-n (N worktrees + reviewer; apply on PASS)
  → parent synthesizes receipts; stays free for the next ask

Lifecycle claims stay exact: dispatched ≠ settled ≠ verified.

allow_shell controls whether bash can execute; it does not rename the tool or make mode the approval authority. Durable tasks and automation keep conservative omitted-field defaults and receive shell authority only when their settings explicitly grant it. Stateful terminal/background controls are specialized deferred tools rather than fields on the small foreground bash schema. Full Access changes the permission posture while hard safety and repository-policy holds remain authoritative.

Action-capable modes can discover the deferred rlm family through tool_search; its open, eval, configure, and close actions own persistent RLM sessions. The legacy split rlm_* spellings remain replay-only aliases. Inside an RLM Python REPL, sub_query_batch fans out 1-16 cheap parallel child calls pinned to deepseek-v4-flash.

The fast deepseek-v4-flash / thinking-off path is called Fin in the product language. Fin is a seam for routing, summaries, cheap child calls, and coordination work; it does not change approval behavior.

The orchestration controls remain available without taking over the starting screen: /auto turns on Auto-Review so the agent just works, /goal keeps one objective across turns, and /workflow prepares a repeatable ordered or fan-out workflow. They are directly callable and searchable through the full command palette, but are not pinned to the starter slash menu, idle welcome, footer, or default Hotbar. A bare / instead opens the small task-oriented starter set; use /help or the command palette for the complete inventory.

/goal <objective> sets a session objective with an optional token budget and keeps active objectives visible as Work context. The agent may also create the goal itself when a direct request describes a verifiable end state that will take more than one turn ("until the tests pass", "make X work end to end"); it then shows one receipt line and you can /goal pause or /goal clear it. Bare /goal shows progress (state, elapsed, continuations, and how to continue when no turn is running); with no goal and no conversation yet it prints usage. /goal pause stops goal continuation without changing the objective, /goal resume resumes and sends the objective back into the turn, /goal complete marks it done, /goal blocked marks it blocked, and /goal clear removes it. Goal state does not change the active TUI mode, permission posture, or model route. This remains distinct from --model auto, which only controls model and thinking selection.

Workflow builds on the same separation: a goal can ask the agent to keep working, while Workflow supplies the repeatable workflow/progress surface for large fanout. In the UI, a Workflow run should be shown as an overlay on the main screen, not as another mode beside Plan, Work, and Operate.

App-server clients can persist a thread-scoped goal with thread/goal/set, read it with thread/goal/get, and clear it with thread/goal/clear. That persisted record carries active, paused, blocked, usage_limited, budget_limited, or complete status plus token/time accounting fields for clients that need thread resume semantics.

Mode Persistence

Choosing a mode interactively also sets the mode a fresh session starts in. Tab/Shift+Tab cycling, the Alt+A / Alt+P / Alt+Y shortcuts, the hotbar's Plan/Work/Operate actions, and /mode all write default_mode to ~/.codewhale/settings.toml, so switching to Operate survives a restart. The write happens off the event loop; if it fails, the TUI says so in a warning toast rather than reverting silently on the next launch.

Mode, thinking level, and the model picker share one serialized writer, so the selection you made last is the one on disk — a burst of Tab presses cannot end up persisting whichever write happened to finish last — and a mode write never rolls back an unrelated key such as default_model.

Two paths deliberately do not rewrite the startup default: restoring a saved session (which re-installs the mode that session was in) and a mode change refused because a turn is in flight. The legacy yolo entry point installs Work plus Full Access, and agent is what it persists — yolo is a permission alias, never a startup mode.

Re-selecting the mode you are already in is not a no-op. After a restored session the live mode and default_mode routinely disagree, so choosing the live mode again is how you make it durable; Codewhale confirms with a "saved as startup default" receipt rather than reporting "already in that mode".

While a turn is running, every change to the live route is refused — mode, model, thinking level, and provider — no matter which surface you use. That now includes the slash surfaces (/mode, /model, /config <key> <value>, /config preset), which are reachable mid-turn. Press Esc to interrupt first. The restart-only default_mode key is exempt, because it does not touch the running turn.

Codewhale writes settings.toml under a lock that spans processes, and replaces the file atomically, so a second Codewhale instance on the same home directory cannot lose your selection or read a half-written file. At exit, queued writes are flushed before the terminal is restored; anything that failed is printed on the way out instead of disappearing with the alternate screen.

Compatibility Notes

  • Older settings files with default_mode = "normal" still load as agent; saving rewrites the normalized value.

Escape Key Behavior

Esc is a cancel stack, not a mode switch.

  • Close slash menus or transient UI first.
  • Cancel the active request if a turn is running.
  • Discard a queued draft if the composer is empty.
  • Clear the current input if text is present.
  • Otherwise it is a no-op.

Permission Posture

Permission posture controls tool approval and whether a turn may pause for a missing user decision. It is one layer of the full authorization order, not a bypass for tool admission, repository law, or sandbox enforcement. Cycle it with Shift+Tab, or edit it at runtime:

/config
# edit the approval_mode row to: suggest | auto | never

Legacy note: /set approval_mode ... was retired in favor of /config.

  • suggest (Ask, default): tool approvals may interrupt, and Codewhale asks when an unresolved user choice materially changes authority, cost, scope, or outcome.
  • auto (Auto-Review): the fully autonomous posture. It never opens a user question; the model resolves ambiguity from context, chooses a safe reversible interpretation, or reports that it cannot proceed safely. Tool safety holds remain separate from user questions. Two layers decide approvals. The deterministic floor (configured block rules plus the built-in safety floor) allows proven-safe calls and hard-blocks publish-like actions and destructive background/headless work; it is never model-reviewed. Fallback holds — calls the deterministic engine could not prove safe — escalate to a one-shot model guardian (v0.9.8) that returns risk, allow/deny, and a rationale. The guardian sees the exact held call and deterministic observations in separate JSON fields; conversation history, skill instructions, attachments, and expanded model context are excluded. It does not infer user intent or compute a generic user-intent score. High or critical risk cannot auto-run even if the model says allow. It has no tools, remembers no rules, and denies rather than truncates an oversized exact call. Exactly one reviewer request is made; incomplete or malformed output, timeout, cancellation, or provider failure fails closed. Headless adapters use the deterministic-only tier. Repo-law holds that explicitly require a person block in Auto-Review rather than opening a hidden approval modal.

The LLM reviewer is closest to OpenAI Codex's experimental Auto-Review at commit 6fc6b9d6d2580d62622fc9884b5f5707f6505a5e. Codex's guardian entry point reconstructs conversation context and runs a dedicated review session. Codewhale deliberately adopts only the exact-action structured decision, 90-second deadline, and fail-closed result. It does not copy Codex's transcript reconstruction, user-authorization score, reviewer tools, retries, persistent review session, or denial ledger.

Kimi Code at commit 1414d4602898f406e540b23342cb18db23ff9efc also has no LLM reviewer. Its ordered permission policy applies explicit deny rules and then its Auto policy returns approve directly. Codewhale borrows Kimi's no-question autonomous UX, not that blanket approval rule.

The sandbox and escalation baseline is grounded in DeepSeek Harness 0.1.0-rc.5 at commit 47f943859bef60e4160492346772ded9b24f765a: its sandbox contract defines per-call read-only, workspace-write, and danger-full-access boundaries and forbids silent unconfined fallback; its approval contract grants only allowed-once and fails closed on rejection, cancellation, or an unavailable answerer; and its sandbox result contract tells the model to retry a denied command exactly once with the narrowest wider mode plus a justification. DeepSeek Harness does not add an LLM reviewer to that path. Codewhale's autonomous posture adds only the single stateless guardian request described above; deterministic hard blocks remain non-bypassable.

  • bypass (Full Access): ordinary tool calls do not show approval prompts, while deliberate user questions remain available. Non-bypassable registered holds auto-approve instead of opening a contradictory modal. Repository-law and managed-policy holds fail closed as hard blocks instead of contradicting Full Access with an approval modal.
  • never: blocks any tool that is not considered safe/read-only; deliberate user questions remain available.

The effective posture and its question discipline are projected into every turn from the same runtime authority that gates tools. A mode/posture change is therefore visible to the next turn. Untrusted runtime-generated input is narrowed before metadata is built and cannot invent approval authority. An explicit Full Access sub-agent handoff preserves the parent's standing posture so ordinary child work does not begin prompting again.

Children (sub-agents and fleet workers)

Children inherit the session posture faithfully rather than a bare auto-approve bit:

  • Auto-Review: a worker's held call goes through the same deterministic policy (proven-safe calls run; publish-like and destructive background work is hard-blocked) and, for holds it cannot prove safe, the same one-shot model guardian using the child's own session client. No prompt is ever opened for a child; an unavailable guardian denies, fail closed.
  • Ask: a call the role may delegate runs. A held call is raised as an approval prompt in the parent's UI (agent:<id>:approval:<n>) when the host is an interactive TUI; the worker waits visibly (waiting for user) and the person's answer is routed back to it, whether the parent turn is idle or itself awaiting an approval. Hosts that cannot prompt deny with the reason.
  • Full Access: ordinary calls run; destructive detached work still fails closed, because children are background workers.

Role posture and the execution envelope are checked before and after this gate and never widen. Every decision a person did not make at a prompt is written to the audit log and to the child's transcript as a one-line note (Auto-Review allowed 'bash' (low risk, model guardian): …), visible when the worker is focused.

Small-Screen Status Behavior

When terminal height is constrained, the status area compacts first so header/chat/composer/footer remain visible:

  • Loading and queued status rows are budgeted by available height.
  • Queued previews collapse to compact summaries when full previews do not fit.
  • /queue workflows remain available; compact status only affects rendering density.

Workspace Boundary and Trust Mode

By default, file tools are restricted to the --workspace directory. Enable trust mode to allow file access outside the workspace:

/trust on

Bare /trust (like /trust status) only reports the current setting — it does not enable anything. Use /trust off to restrict access again.

Full Access enables trust mode automatically.

MCP Behavior

MCP tools are exposed as mcp_<server>_<tool> and use the same approval flow as built-in tools. Read-only MCP helpers may auto-run in Ask and Auto-Review when policy permits; MCP tools with possible side effects require approval. Full Access does not bypass hard policy holds.

See MCP.md.

Run codewhale --help for the canonical list. Common flags:

  • -p, --prompt <TEXT>: one-shot prompt mode (prints and exits)
  • codewhale exec --auto --output-format stream-json <PROMPT>: run the tool-backed non-interactive agent and emit one JSON object per line for harnesses and backend wrappers. Exit codes: 0 on success, 1 for genuine task/agent failures, 75 (EX_TEMPFAIL) when the turn ended on a retryable infrastructure failure (provider/transport network/timeout after all in-session retries) so harnesses can tell a retryable infra exit apart from a task failure; the terminal stream metadata event's error_category carries the same classification
  • codewhale exec --resume <ID|PREFIX> <PROMPT> / --session-id <ID|PREFIX>: continue a saved session non-interactively
  • codewhale exec --continue <PROMPT>: continue the most recent saved session for this workspace non-interactively
  • codewhale fork <ID|PREFIX> / codewhale fork --last: copy a saved session into a new sibling session; forked sessions retain additive parent-session metadata and show that lineage in session listings
  • --model <MODEL>: when using the codewhale facade, forward a DeepSeek model override to the TUI
  • --workspace <DIR>: workspace root for file tools
  • -r, --resume <ID|PREFIX|latest>: resume a saved session
  • -c, --continue: resume the most recent session in this workspace
  • --max-subagents <N>: clamp to 1..=128
  • --mouse-capture / --no-mouse-capture: opt in or out of internal mouse scrolling, transcript selection, right-click context actions, and transcript scrollbar dragging. Mouse capture is enabled by default on non-Windows terminals and on Windows Terminal/ConEmu/Cmder so drag selection copies only transcript text, removes visual wrap-column line breaks from paragraphs, and stays scoped to the transcript pane; hold Shift while dragging or use --no-mouse-capture for raw terminal selection. It defaults off on legacy Windows console (CMD without WT_SESSION / ConEmuPID) and inside JetBrains JediTerm — PyCharm/IDEA/CLion/etc. — where the terminal advertises mouse support but forwards SGR mouse events as raw text (#878, #898). Use --mouse-capture to opt in anywhere it's defaulted off. Raw terminal selection may cross the right sidebar and include visual wraps because the terminal, not the TUI, owns the selection.
  • --profile <NAME>: select config profile
  • --config <PATH>: config file path
  • -v, --verbose: verbose logging

Branching and Rollback

Codewhale has three related but intentionally separate recovery paths:

  • codewhale fork <ID> creates a new saved session from an existing saved conversation and records the source session id. This is the safe way to explore a different answer path without overwriting the original session.
  • Esc-Esc backtrack rewinds the live transcript to a previous user prompt and restores that prompt into the composer for editing.
  • /restore and the revert_turn tool restore workspace files from side-git snapshots. /restore list [N] lists more snapshot options before choosing a rollback point. They do not rewrite conversation history.

A Pi-style in-file tree browser is a larger UI/data-model project. v0.8.40 ships the bounded fork/backtrack primitives and explicit lineage metadata.