Commit Graph

161 Commits

Author SHA1 Message Date
Henry Park
5b4a86690c ci: parallelize affected crate tests with nextest (#8013)
* ci: parallelize affected crate tests with nextest

* ci: preserve process-local test contracts under nextest

- serialize composition runtime and CLI trace-policy tests\n- preserve zero-test package buckets\n- restore fail-closed metadata routing

* fix(ci): validate cargo runner bucket subsets

* Address PR review feedback (#8013)

- Skip nextest installation for exact-target and Cargo-only buckets.
- Validate serialized nextest groups against compiled test inventory.

* fix(ci): parse nextest inventory specifications

Replace the YAML-sensitive heredoc with a Bash array and syntax-check the rendered crate runner in the workflow contract suite.

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-09-01 21:54:43 +00:00
hylin
2b52d1f7a5 perf(github): compact repository list responses (#7996)
* perf(github): compact repository list responses

* fix(github): validate compact repository responses

---------

Co-authored-by: linhongyu510 <linhongyu510@users.noreply.github.com>
2026-09-01 04:23:54 +00:00
Henry Park
24ff93f435 ci: unify bounded integration execution (#7992)
* docs(ci): plan unified bounded integration execution

* ci: unify selected integration tests under nextest

* ci: retire duplicate integration runners

* docs(ci): remove implementation plan from PR

* Address PR review feedback (#7992)

- preserve every unique Cargo integration target in execution inventory
- cover shared-source aliases through the inventory and lane-runner contracts

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-31 19:01:16 +00:00
jinxin
ea111c67b9 fix(ci): stabilize main branch coverage checks (#7995)
* fix(ci): stabilize main branch coverage checks

* fix(ci): complete main CI stabilization

* fix(notifications): backfill terminal approval cleanup

* test(ci): harden hooks parity timeout contract

* test(composition): cover approval backfill replay wiring

* test(ci): keep timeout decoys lint clean

* test(ci): address final review feedback

* ci: restore main-only workflow triggers

* fix(security): update Wasmtime to 47.0.4
2026-08-31 12:21:42 +00:00
Henry Park
1c373882f7 ci: validate integration group topology (#7980)
* ci: validate integration group topology

* Address PR review feedback (#7980)

* validate group registrations before integration-path filtering
* enforce live group topology in the always-running CI contract suite

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-28 23:09:44 +00:00
jinxin
ec4c62df70 fix(notifications): persist auth gates before enrichment (#7901)
* fix(notifications): persist auth gates before enrichment

* fix(notifications): reject gate-less auth inbox records

* fix(notifications): release auth delivery capacity

* fix(notifications): publish webui auth gates

* fix(notifications): require durable auth publication

* fix(notifications): reconcile OAuth auth lifecycle

* fix(notifications): close resumed auth notifications

* fix(notifications): isolate auth rollout backfill

* fix(notifications): preserve current auth gate

* fix(notifications): harden auth recovery cleanup

* fix(notifications): reopen recurring auth gates

* ci: re-seed composition budget for auth notifications

* fix(notifications): preserve recurring auth gate lifecycle

* fix(notifications): scope auth recovery cleanup

* test(composition): complete notification source fixtures

* fix(notifications): harden auth fan-out reconciliation
2026-08-28 16:56:34 +00:00
jinxin
0fb92368ff feat(notifications): publish durable resource blocks (#7900)
* feat(notifications): publish durable resource blocks

* fix(notifications): preserve blocked-run delivery lifecycle

* fix(notifications): reconcile resource blocks from journal

* fix(notifications): harden resource block rollout

* ci(composition): recapture blocked-run wiring budget

* fix(notifications): reconcile blocked delivery evidence

* fix(notifications): preserve unconfirmed delivery evidence

* fix(notifications): recheck replacement gate state

* fix(notifications): stabilize backfill reconciliation

* fix(notifications): reopen authoritative resource gates
2026-08-28 15:10:30 +00:00
Henry Park
3d2bcbf11f ci: centralize integration test inventory (#7967)
* ci: centralize integration test inventory

* Address PR review feedback (#7967)

- reject unsupported planner test names with a controlled error

- require integer schema version and partition count fields

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-28 10:58:54 +00:00
Henry Park
c8bdbfa5f7 feat(github): decode repository file content (#7963)
* feat(github): decode repository file content

* Address PR review feedback (#7963)

- stabilize unsupported-encoding failures and cover malformed provider shapes

- verify binary and malformed results through the caller-level runtime path

* Fix provider output preview assertion (#7963)

Assert the projected text contract instead of reparsing it as JSON.

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-28 04:44:17 +00:00
Henry Park
859c028dc5 ci: compile integration batches once (#7943)
* ci: compile integration batches once

* Address PR review feedback (#7943)

- preserve the 120-minute push coverage backstop

- budget 300 minutes for combined non-push integration batches

- pin the event-specific timeout in workflow contract tests

* Fix CI after dependency registry drift (#7943)

- align the yanked chacha20 lock entry with current main

* Address PR review feedback (#7943)

- remove dead runner override from the batch test

- cover CI fail-closed and local Cargo fallback paths

* Address PR review feedback (#7943)

- fail fast when the missing-nextest fixture exposes a host binary

* Fix merge-queue group test toolchain compatibility (#7943)

- preserve the nightly integration lane while bypassing manifest version checks already covered by the flat batch

- pin the canonical group runner command in the hermetic workflow contract

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-28 03:39:40 +00:00
Henry Park
b63333f9a7 ci: scope merge queue to affected areas (#7898)
* ci: scope merge queue to affected areas

* fix(ci): close merge-group scope gaps

* Address PR review feedback (#7898)

- make workspace manifests exhaustive in merge-group scope

- resolve renamed reverse-dependency edges

- shallow-fetch merge-group diff endpoints

- restore the established Windows Clippy target scope

* Address PR review feedback (#7898)

- preserve uncapped canonical crate buckets for merge-group runs
- execute PR and merge-group shell diff branches in regression coverage
- pin the affected-package selector CLI JSON contract

* Address PR review feedback (#7898)

- keep CLI subprocess fixtures portable across Windows and POSIX

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-27 16:50:10 +00:00
Henry Park
ee3b42a47d feat(tools): add bounded selectable JSON result views (#7928)
* feat(tools): add selectable JSON result views

* Address PR review feedback (#7928)

- harden structured JSON paging redaction, validation, and continuation progress
- preserve unified durable result reads across live-shell and stress harness paths
- correct first-look sizing and E2E result-read selection
- split and extend integration coverage for structured and legacy result views

* Simplify first-look observation sizing (#7928)

Reuse the canonical observation constructor as the fit check instead of duplicating serialized envelope accounting. This also restores the enforced composition mass budget.

* Preserve typed result preview provenance (#7928)

- keep LoopResultRef typed through first-look rendering
- carry structured-page provenance out of band on host tool-result messages
- reject provider attempts to forge structured observation provenance
- retain the composition mass budget and trace recording behavior

* Fix result-view CI regressions (#7928)

- preserve safe-summary fallback for invalid stored observations
- carry structured-view provenance through the shared trace harness
- correlate Responses API tool results by provider call IDs

* Update result-view integration contracts (#7928)

* Tighten result-view integration assertions (#7928)

* test: follow durable extension result continuations

* style: avoid implicit path concatenation

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-27 16:20:41 +00:00
Henry Park
7b41f0da18 fix(docker): repair home ownership so a root-written file cannot brick boot (#7924)
* fix(docker): repair home ownership so a root-written file cannot brick boot

A hosted instance crash-looped on every restart with:

    Error: could not resolve LLM environment fallback: Provider
    provider_registry request failed: failed to read provider registry
    overlay `/data/ironclaw-reborn/providers.json`: Permission denied (os
    error 13)

while `config.toml` in the same directory loaded fine — so this was one
file's ownership, not directory traversal. The provider-registry overlay is
fail-closed by design (an unreadable explicit overlay must not silently fall
back to compiled-in defaults), so a single root-owned file ends the boot
rather than degrading it.

Root cause is the interaction between two deliberate choices. #7723 changed
the runtime image from `USER ironclaw` to `USER root`, because sshd must
start as root before the entrypoint drops via `gosu` — so for the first time
a root process can write to the persistent volume. And the root pass's
`chown` is non-recursive, because start-sshd.sh refuses to run unless it owns
`$IRONCLAW_REBORN_HOME/ssh`. Together they mean the entrypoint fixes the home
DIRECTORY and leaves its CONTENTS untouched, and a volume outlives the image
that wrote it.

Repair ownership of the home's contents in the root pass, excluding `ssh`.
Recursing into `ssh` would trade this crash loop for a broken SSH listener,
which is why the exclusion is explicit rather than incidental.

Reproduced before fixing, in debian:bookworm-slim, seeding a volume the way
the failing instance looks:

    BEFORE  -rw-r--r-- ironclaw ironclaw  config.toml
            -rw------- root     root      providers.json
    AFTER entrypoint (root pass ran):  unchanged
    uid 1000 reading providers.json:   Permission denied

With the fix, providers.json and nested files become ironclaw-owned and
readable, while `ssh/` and its host key stay root:root.

Regression test `home_contents_ownership` asserts both halves — that every
top-level home entry is handed to chown, and that `ssh` is NOT. It fails
without the fix and passes with it. The chown stub records argv rather than
executing, so the test pins what the root pass asks for.

Scope note: this is a robustness fix, not a diagnosis of one box. I could not
inspect the instance, so I have not proven how its providers.json became
root-owned — only that this mechanism produces exactly the logged error and
that `USER root` is what makes root-owned files possible at all. A persistent
volume shared across image versions with differing runtime users should not be
able to brick boot regardless of which write created the file.

Also applies to the 1.3 line, which has carried `USER root` since #7723.

Verified: entrypoint self-tests pass on debian:bookworm-slim as a non-root
user (as CI runs them); the real-sshd regression harness still reports 16/16
against the release/1.3.1 baseline; a full entrypoint -> start-sshd -> real
SSH login still succeeds as agent:1000 with the ssh state dir root-owned;
shellcheck clean apart from two pre-existing SC2016 notices.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(docker): follow a symlinked home and pin the repair's real properties

Review of the first cut found three defects, two of which meant the fix or its
test did not do what they claimed.

**The repair silently did nothing when the home is a symlink.** `find` defaults
to a physical walk, so `find "$IRONCLAW_REBORN_HOME" -mindepth 1` matches
nothing when the start path is a symlink -- no error, no repair, and the boot
loop persists with no diagnostic. Verified in debian:bookworm-slim:
`find /link -mindepth 1 -maxdepth 1` returns 0 entries, `find -H /link` returns
2. Now uses `-H`, which follows the start path only, so symlinks planted
*inside* the home are still never followed; combined with `chown -h`, a
planted symlink cannot redirect ownership outside the home. Confirmed: a
`planted -> /etc` symlink leaves /etc and /etc/passwd root-owned.

**The test did not pin recursion.** It asserted the top-level `nested` entry
appeared in the recorded argv, which stays true even with a non-recursive
chown -- while `nested/deep.db` would keep its old ownership and still brick
startup. It now asserts the nested file itself.

**The test swallowed the entrypoint's exit status** with `) || true`, so a
run that failed after the ownership commands could still pass on argv alone.
The run is now required to succeed.

Also dropped a `! -user ironclaw` filter added in the first cut as a
steady-state optimization. Its own tests caught why it was wrong: the
predicate resolves the name at find time and errors with
`find: 'ironclaw' is not the name of a known user` wherever it does not
resolve, which under `set -e` aborts the entire boot. That trades a rare
ownership repair for a guaranteed crash, and no review finding asked for it.

Each property is now proven discriminating by mutation:
- drop `-H`      -> FAIL symlinked_home_ownership
- top-level only -> FAIL home_contents_ownership (nested/deep.db)
- stop pruning ssh -> FAIL home_contents_ownership (ssh chowned)

Verified: entrypoint self-tests pass on debian:bookworm-slim as a non-root
user; the real-sshd harness still reports 16/16 against the release/1.3.1
baseline; shellcheck clean apart from two pre-existing SC2016 notices.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 03:18:13 +00:00
Henry Park
053c83f513 fix(docker): forward-port the 1.3 in-worker SSH and workspace-root fixes to main (#7723, #7804) (#7915)
* fix(docker): restore Reborn in-worker SSH (#7723)

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
(cherry picked from commit b0ba342268)

* fix(workspace): honor IRONCLAW_REBORN_WORKSPACE_ROOT on 1.3 (#7804)

* fix(workspace): honor IRONCLAW_REBORN_WORKSPACE_ROOT on 1.3

The durable workspace-root override landed on release/2026-08-11 in
e12bbae4d2 and was never forward-ported: neither origin/main nor
release/2026-08-17 has it. Both CLI boot paths resolve the workspace
root from `std::env::current_dir()`, so a container's project files and
landed attachments are written under the container workdir instead of
the mounted volume, and do not survive a redeploy.

Port the workspace-root slice only. The two release-branch commits that
carry it (e12bbae4d2, bdf8aef022) also carry the whole rc1->1.1/1.2
startup-migration program (~11.7k lines); that stack is out of scope
here and is already staged on
origin/port/firat-workspace-artifacts-1.2-to-1.3.

- `local_runtime_workspace_root` resolves the override and falls back to
  cwd, used by both the standalone and hosted single-tenant builders.
  `optional_path_env` rejects an empty value rather than silently
  treating it as unset.
- The Docker entrypoint defaults the root to `$IRONCLAW_REBORN_HOME/workspace`
  and fails closed on Railway when it resolves outside
  RAILWAY_VOLUME_MOUNT_PATH, matching the existing IRONCLAW_REBORN_HOME guard.
- The root pass also creates and chowns the workspace root before the
  gosu privilege drop. This has no counterpart on release/2026-08-11,
  which has no root pass; without it a workspace root outside
  IRONCLAW_REBORN_HOME fails the later mkdir as the unprivileged user.

Test: `scripts/ci/test-reborn-docker-entrypoint.sh` gains a default-root
assertion and an explicit-override case, and the existing ssh_root chown
assertion now pins the workspace root. Both new assertions were verified
to fail when the entrypoint default is broken.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(workspace): resolve workspace root before the Railway guard

Addresses review on #7804.

Drop the workspace root from the root pass's mkdir/chown (P2). The
containment check runs after the gosu privilege drop, so the root pass
reached chown before validation: `IRONCLAW_REBORN_WORKSPACE_ROOT=/etc`
would have had its ownership changed to `ironclaw` before startup failed
closed. Removing it also deletes this port's only deviation from the
release/2026-08-11 original. The default root still works unprivileged —
IRONCLAW_REBORN_HOME is chowned in the root pass, so the later profile-
gated mkdir creates the subdirectory as `ironclaw`. A root outside
IRONCLAW_REBORN_HOME now fails loudly at that mkdir instead of silently
widening ownership.

Compare canonicalized paths in the Railway containment guard (P1). A
symlink beneath the mount whose target is outside it, or a `..` segment,
passed the lexical prefix test while the runtime resolved the real
ephemeral target — booting a deployment whose project files silently do
not persist, the exact failure the guard exists to prevent. Both sides
are resolved, so a mount path that itself traverses a symlink does not
reject every root. The diagnostic reports resolved and original spellings.

Test: the entrypoint self-test gains a rejected escaping-symlink case, a
contained positive control, and a trailing-slash normalization case. The
new checks use `assert_eq`, which the regression-test gate recognizes as
a meaningful shell assertion — `[ ... != ... ]` is not in its vocabulary,
which is why the required check rejected the previous commit. Each
assertion was verified to fail under a targeted mutation (lexical
containment restored, default root broken, normalization removed).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit a5b8345d21)

* fix(docker): harden the forward-ported SSH and workspace-root paths

Review of the two forward-ported commits surfaced four defects. All of them
are latent in the shipped 1.3.x code, not introduced by the port -- `git blame`
puts every one on d4cf2411b0 (#7723) or ab7ac3d87e (#7804). They are fixed here
rather than ported forward knowingly.

**Workspace root is now prepared before the privilege drop.** The root branch
chowned only $IRONCLAW_REBORN_HOME and /workspace, then exec'd gosu; the
resolved $IRONCLAW_REBORN_WORKSPACE_ROOT was created afterwards, unprivileged.
An explicit override onto a fresh root-owned volume therefore hit EACCES and,
under `set -eu`, aborted the boot. The chown stays non-recursive so
start-sshd.sh keeps root ownership of $IRONCLAW_REBORN_HOME/ssh.

**The Railway containment guard no longer accepts an unresolved path.**
`readlink -f` exits non-zero when any non-final component is missing, so the
`|| printf` fallback handed the raw spelling to a glob comparison, and
`$RAILWAY_VOLUME_MOUNT_PATH/missing/../../tmp` passed the very `..`-escape
check the guard exists to enforce -- silently placing project files on
ephemeral storage. Verified empirically in debian:bookworm-slim (coreutils
9.1). Now `readlink -m`, which canonicalizes missing components, with no
raw-spelling fallback: an unresolvable path fails closed.

**sshd forwarding is disabled explicitly.** OpenSSH defaults
AllowTcpForwarding and AllowAgentForwarding to yes; the generated config now
sets those and PermitTunnel to no.

**A pasted private key is refused before it reaches disk.** `ssh-keygen -l -f`
exits 0 on a private key, so one supplied by mistake was written verbatim into
authorized_keys (mode 644, persisted). It never granted a login -- PEM lines do
not match authorized_keys grammar -- so the harm was secret persistence plus a
silently broken setup. The value is now rejected on a PEM header or on carrying
more than one key, without ever echoing the key material.

That last check needs care: an operator almost always supplies the key via
`cat id_ed25519.pub`, whose value carries a trailing newline. A naive
line-count rejection would refuse the most common paste, and because the
entrypoint runs start-sshd under `set -e` that would abort the whole container
boot rather than merely disabling SSH. Surrounding whitespace and blank lines
are therefore stripped before the one-key check, and the PEM check runs first
so a real private key still gets the actionable "supply the matching .pub"
message.

Verified against a real sshd in debian:bookworm-slim, comparing behavior with
the shipped release/1.3.1 script as the baseline:

- 16/16 regression checks, against a 12-pass/1-fail baseline. Every baseline
  pass survives; real logins still succeed for ed25519, RSA and ECDSA.
- 10/10 operator-paste cases: plain, trailing newline, leading newline,
  trailing spaces, CRLF and surrounding blank lines all accept and complete a
  real SSH login; private key, two keys, garbage and whitespace-only are
  refused with accurate messages.
- Unset key still exits 0 with no listener, preserving the opt-in contract.
- Generated config still passes `sshd -t` with every original auth directive.

The trailing-newline case is pinned in CI by a second container in the existing
`Verify in-worker SSH` step, reusing the image already built there -- it adds
about ten seconds to a job that takes roughly twenty-two minutes, and no new
job. The other three fixes carry shell-suite regression tests, each proven to
fail before its fix. Fixes 3 and 4 cannot be reached from that suite
(start-sshd.sh needs real root and a real sshd binary), which is precisely why
the CI container check exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(docker): document the in-worker SSH ingress and its privilege model

`IRONCLAW_REBORN_SSH_PUBLIC_KEY` was documented nowhere, despite being the
single switch that stands up an SSH listener in the runtime image. AGENTS.md
requires environment variables to be documented in `.env.example`.

- `.env.example` gains the variable next to IRONCLAW_REBORN_WORKSPACE_ROOT:
  off unless set, enables sshd on container port 2222, public-key-only login
  as `agent`, and the port must be published by the orchestrator to be
  reachable.
- The operator Docker guide gains an SSH Access section covering the same
  three facts plus the auth directives the generated sshd_config actually
  sets, and states plainly that `agent` is a uid-1000 alias of `ironclaw` --
  an SSH session holds the full runtime identity, so the private key deserves
  the same care as shell access to the service.
- Two Dockerfile comments (no build-logic change) record why `agent` is a
  deliberate UID alias, so a later reader does not mistake it for a
  lower-privileged account and widen the SSH surface on that assumption, and
  that the entrypoint performs the only privilege drop -- anything bypassing
  it (`docker run --entrypoint`, `docker exec` without `--user`, a platform
  custom start command) runs as root.

Every claim was checked against docker/reborn/start-sshd.sh, entrypoint.sh and
the Dockerfile runtime stage rather than restated from the review comments.

Verified: check-guidance.py OK, docs_publication_boundary.py OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(docker): never chown an unvalidated workspace root as root

The previous commit's workspace-root fix introduced a privilege escalation.
It added $IRONCLAW_REBORN_WORKSPACE_ROOT to the root branch's `chown`, but the
Railway containment check that validates an operator-supplied override runs
*after* the privilege drop -- it needs $effective_profile, which is resolved
from the config file further down, past `exec gosu`. The raw value was
therefore chowned to the runtime uid before anything could reject it.

Reproduced in debian:bookworm-slim: with
`IRONCLAW_REBORN_WORKSPACE_ROOT=/etc` on a Railway-profile boot, /etc went
from root:root to ironclaw:ironclaw, and only then did startup fail. A
non-recursive chown of /etc still lets uid 1000 create entries there
(/etc/ld.so.preload being the obvious one), so this is a real escalation
primitive, and the container filesystem keeps it across restarts.

The root pass now pre-creates the workspace root only when it provably lives
inside a directory this entrypoint already manages -- under
$IRONCLAW_REBORN_HOME, or under the Railway volume mount -- comparing
canonicalized paths so `$RAILWAY_VOLUME_MOUNT_PATH/../etc` cannot spell its
way in. Anything else is left untouched: on Railway the containment check
rejects it moments later, and elsewhere the later unprivileged `mkdir -p`
reports the failure exactly as it did before. That still covers the case the
fix exists for, an override onto a fresh root-owned volume mount, which is
what the deployment guide documents.

Regression tests (`workspace_root_outside_managed`) assert the root pass never
passes an out-of-tree path to chown, in both a plain and a `..`-spelled form;
both fail without the guard and pass with it. Two existing cases were
reconciled with the new shape: the chown stub now appends, because the root
pass legitimately issues two chown calls rather than one, and an overwriting
stub silently dropped the first; and `workspace_root_privdrop` now sets
RAILWAY_VOLUME_MOUNT_PATH, making explicit the volume-mount scenario its own
comment already described.

Not addressed here: $IRONCLAW_REBORN_HOME is chowned on the same path with no
validation either. That predates this PR, and unlike the workspace root it has
no containment contract to respect -- any path is a legitimate home by design
-- so tightening it is a behavior change to shipped configuration handling
rather than a fix to this regression. Tracked with the other uid-alias
findings.

Verified: entrypoint self-tests pass on debian:bookworm-slim as a non-root
user (as CI runs them); the real-sshd regression harness still reports 16/16
against the release/1.3.1 baseline; ws12 workflow contracts and check-guidance
pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(docker): refuse a filesystem-root home or workspace root

`${VAR%/}` turns the single-character input `/` into an empty string, and
neither root path handled that. `IRONCLAW_REBORN_HOME=/` reached
`mkdir -p "" /workspace` and died with a confusing "cannot create directory
''". `IRONCLAW_REBORN_WORKSPACE_ROOT=/` reached `readlink -m ""` in the
containment guard, which exits non-zero and printed nothing at all -- under
`set -eu` the boot died silently, with no diagnostic for an operator to act
on. That silent mode was introduced by the containment guard two commits ago.

The obvious repair -- restoring `/` the way the file already does for
RAILWAY_VOLUME_MOUNT_PATH eight lines above -- is wrong here, and testing it
is what showed why. Restoring the value makes the root pass's
`chown ironclaw:ironclaw "$IRONCLAW_REBORN_HOME"` newly *reachable* as
`chown ironclaw:ironclaw /`, handing uid 1000 ownership of the container's
entire root filesystem. Verified in a container: the chown succeeds. That
trades a loud crash for a total loss of the root boundary, which is strictly
worse than the bug being fixed.

Both values are therefore refused with a diagnostic naming the variable.
Nothing legitimately runs with either path set to the filesystem root, so the
safe reading of `/` is operator error -- stop and say so.

Regression tests (`slash_root`) drive both variables through the root pass and
assert three things: the run is refused, the diagnostic names the variable,
and *nothing is chowned on the way to that refusal*. That last assertion is
the one that matters -- it fails against the restore-to-`/` shape as well as
against the original bug, so it pins the escalation rather than just the
crash. All three assertions fail without the guards and pass with them.

Verified in debian:bookworm-slim: `IRONCLAW_REBORN_HOME=/` leaves `/` owned by
root:root and exits with the diagnostic; the entrypoint suite passes as a
non-root user (as CI runs it); the real-sshd regression harness still reports
16/16 against the release/1.3.1 baseline, including a full
entrypoint -> start-sshd -> real SSH login as agent:1000 with the ssh state
directory still root-owned.

Also adds the missing Rust coverage for the same variable:
`local_runtime_workspace_root` now has tests for an explicit override, the
unset fallback to the current directory, and the set-but-empty failure
asserting the error names IRONCLAW_REBORN_WORKSPACE_ROOT. Each was proven
discriminating by breaking the production function and confirming only the
matching test failed. Scoped to the resolution function deliberately: proving
the value reaches the composed RebornHostBindings would require a new public
accessor in ironclaw_composition, which is a production API change and does
not belong in a release-blocking forward-port.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 00:14:28 +00:00
firat.sertgoz
d112612840 feat(hooks): AfterTurn lifecycle point + memory curation as its first consumer (#7770 phase 1) (#7765)
* feat(memory): periodic memory-curation pass ("dreaming"), first slice (#7276)

Memory only ever grew. Writes accumulate, nothing prunes, and the standing
document has a byte budget, so redundancy crowds out what matters. No human
reads the file, so the decay is invisible.

This adds the Hermes-shaped answer: every N completed user turns, the agent
runs with no user present, re-reads its standing memory, and tidies it —
merging duplicates, resolving superseded facts, tightening wording. Its
output is the edits plus a structured report; nothing is sent to anyone.

Buildable now because unbound turns landed (#7562/#7634): a run with no
conversation and no reply target. The pass is submitted through the same
`UnboundTurnService` door OpenAI-compat and subagent spawn already use.

Shape. The loop tier owns only the observation ("an ordinary user turn
completed, under this scope") and reports it through a port; every policy
decision lives in the product tier. The port vocabulary sits in
`ironclaw_loop_contracts` rather than the runner because WS1.7 deliberately
removed `ironclaw_turn_runner` as a production dependency of
`ironclaw_assistant`, and this must not reverse that.

The load-bearing guard: an unbound run NEVER triggers curation. A pass is
itself unbound, so triggering on unbound completion would let each pass
schedule its successor — an unbounded background loop running the model
against a user's memory forever. Pinned by test, both unbound profiles.

Also fixed along the way: `UnboundTurnSubmission` had no way to declare
limits, so it always inherited the profile's 1024-iteration budget and no
wall clock. Fine for a user waiting on a panel, wrong for an unwatched
background chore — an unconverged pass would burn tokens against a user's
memory until that ceiling, and nobody would notice. Added narrowing-only
limits (existing callers unchanged, explicitly defaulted) and the pass
declares 6 model calls / 12 capability calls / 90s.

Safety properties pinned by tests: the pass acts as the owner and never as
an operator-config caller; it gets the three memory capabilities and nothing
else; its id doubles as the idempotency key so a crash-retry converges on
the same pass; a failed submission is swallowed at debug (post-terminal
background path — info!/warn! would corrupt the REPL).

Concurrency is safe without batch-atomic memory ops: memory writes are
compare-and-swap, so a pass racing a live conversation loses the write
rather than clobbering it. The failure mode is a lost curation pass, never
a lost memory.

Not wired into composition yet — no deployment runs this. Wiring, the
gate-behavior decision (unbound runs abort on approval gates, so users with
auto-approve off need skip-not-abort), and an integration scenario follow.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(memory): avoid an extension name in curation comments

The extension-specificity gate scans generic code for concrete extension
names; "with slack for one retry" tripped it on the English word. Reworded
rather than allowlisted — the allowlist is for pre-existing debt, not for
new code that can simply say something else.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(memory): move the curation contract into ironclaw_memory

Memory vocabulary belongs with the memory contract. "Curation" means
nothing outside memory, and the signal exists only to decide whether a
user's memory needs tidying — putting it in ironclaw_loop_contracts made
the loop-contracts crate carry a memory concept it has no stake in.

Both tiers already depend on ironclaw_memory (the runner for after-turn
recording, the product tier for the memory service), so this pulls in no
new edge; it only puts the type where its domain lives.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(hooks): add privileged AfterTurn lifecycle point

Adds `HookPointSpec::AfterTurn`, a privileged-only hook point that fires
once after a turn's run reaches a terminal state — the seam for work about
the turn as a whole rather than about one model call, capability
invocation, or checkpoint.

- `AfterTurnHookContext` (`points/turn.rs`) carries tenant/user/agent/
  project plus a `completed` flag. `user_id` is non-optional and there is
  deliberately no `unbound` field: the dispatch call site never fires this
  point for unbound runs, because hook-started background work runs
  unbound and firing on unbound completion would let each background pass
  schedule its own successor forever. Observing background runs stays with
  `EventTriggered` + `LoopCompleted`, which is observer-only.
- `PrivilegedAfterTurnHook` takes no sink: an AfterTurn hook may hold its
  own collaborators and start follow-on work as a side effect. The
  sealed-return-type law stays scoped to points untrusted tiers can reach.
- `install_after_turn` rejects `Installed` and `SelfAuthored` at install
  time; `install_observer` rejects the point outright.
- `dispatch_after_turn` mirrors the observer dispatch shape (ordered
  snapshot, poison handling, failure policy, telemetry) with a 5s per-hook
  timeout, and never propagates a hook failure to the caller.
- New `DecisionKind::Lifecycle` (three in-crate consumers, all updated):
  act-capable but fails isolated, since the run it observes is already
  terminal.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(memory): curation rides the AfterTurn hook point, bespoke port deleted

#7765 landed memory curation on a bespoke `AfterTurnCurationPort` because no
general lifecycle seam existed yet. The `AfterTurn` hook point now exists, so
the port is deleted and curation becomes one privileged hook among others.

- `ironclaw_memory` sheds `src/curation.rs` entirely: memory carries no
  hook-framework vocabulary and no bespoke port.
- `ironclaw_turn_runner` gains `after_turn_hooks::after_turn_hook_context`,
  which keeps the two guards centrally so no hook has to remember them: an
  unbound run never fires the point (hook-started background work runs
  unbound, so firing on unbound completion would let each pass schedule its
  own successor forever), and an actorless run never fires it (nothing to
  attribute follow-on work to).
- The executor's `after_turn_curation` field becomes
  `after_turn_hooks: Option<Arc<HookDispatcher>>` with `with_after_turn_hooks`.
  The 5s bound survives as an OUTER backstop around the whole dispatch; the
  dispatcher already bounds each hook.
- Semantic widening: the point fires for ANY terminal state of an ordinary
  actor-bearing run, not just `Completed`. Hooks that only want successes read
  `ctx.completed` — which `MemoryCurationService` does, first thing, because a
  failed turn says nothing about whether memory needs tidying and counting it
  would drift the interval.
- `MemoryCurationService` implements `PrivilegedAfterTurnHook`; every policy
  decision (interval, per-owner counters, pass building, idempotency key)
  is unchanged. `ironclaw_assistant` takes a normal `ironclaw_hooks`
  dependency — products→loops, the edge it already has via `ironclaw_loop_host`.
- `AfterTurnHookContext::new` added: the struct is `#[non_exhaustive]` and the
  call site is outside `ironclaw_hooks`, so a struct literal is unavailable.

The dispatcher is un-wired (`None`) after this commit; composition follows.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(memory): wire curation through composition behind [memory] config

Phase 1 of #7770 ends where it should: the `AfterTurn` point has a live
consumer. Composition registers the memory-curation hook, so after every Nth
completed turn the agent goes off on its own and tidies the user's standing
memory document (#7276).

- `[memory].curation_interval_turns` (`ironclaw_config`): opt-in, serde-default
  absent. Absent means the hook is NEVER REGISTERED — disabled is expressed by
  not wiring, never by a sentinel, so a written `0` is rejected at parse time
  rather than clamped downstream into "after every turn". Config-only, no env
  override: that matches `provider`/`admin_overrides`, and only the mem0
  connection fields carry an env convention.
- `ironclaw_assistant::memory_curation::after_turn_curation_dispatcher` owns the
  assembly — which hook, at which phase (`Telemetry`: the run is already
  terminal, so it enforces nothing), under which trust class (`Builtin`), behind
  the stable `HookId::for_builtin` path. Composition calls it; per AGENTS.md the
  wiring root does not own module policy. Its own small dispatcher, not the
  per-run middleware one: `after_turn` fires once per run from a
  process-lifetime `Arc`.
- `DefaultPlannedRuntimeParts::after_turn_hook_dispatcher_factory` is a factory,
  not a ready dispatcher, because the `UnboundTurnService` the hook submits
  through is built from the coordinator the same function builds. Handed
  `AfterTurnHookDeps` once, after those exist; may still decline.
- Two conditions gate registration in composition: an operator asked for an
  interval AND a memory provider resolved. A pass over a document no provider
  backs would submit a run whose only three tools do not exist.

Gate posture (#7770's skip-and-note) is deliberately NOT implemented; a
`DECISION #7770:` comment at the submission site records why. No read-only
"would this capability gate for this scope" query exists: the answer needs the
descriptor's effects and origin-gate matrix, the run's `ApprovalPolicy`, the
`TrustDecision`, grants, and leases composed inside
`authorize_dispatch_with_trust` at dispatch time, with an origin that does not
exist until the run is executing. Approximating it from
`ApprovalSettingsProvider::global_auto_approve` alone would duplicate gate
composition in a product service. The seam that is actually missing is at the
gate strategy: a `GateOutcome` that skips the capability for the model instead
of aborting the unbound run.

Tests: two group scenarios drive the wired path end to end — the pass's thread
id is its idempotency key and therefore deterministic, which is what lets the
harness script the background pass's model at all. The positive scenario runs N
ordinary turns and asserts the tidied text reaches a LATER conversation's prompt
under the same user's own memory lane; the negative asserts an empty pass script
below the interval and then corroborates it by crossing the interval one turn
later, so "empty" cannot be latency. Both falsified by moving the interval.
`with_memory_curation_interval()` on the group builder mirrors production's
opt-in exactly; the wiring-parity tripwire and composition mass gate move with
the new field.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(memory): review round — per-trigger pass identity, conversation-only triggers, fail-closed install

Six review findings on #7765 (epic #7770 phase 1).

- **Pass identity was the number of OWNERS, not passes.** The curation pass id
  was `…-{counters.len()}`, which for one user is forever `1`: every interval
  after the first reused the same public id and idempotency key, so the unbound
  accept door REPLAYED the first pass instead of running a new one — the
  document would be curated exactly once, ever, with nothing surfacing it.
  `AfterTurnHookContext` now carries `run_id` (the terminal run that fired the
  point), the runner threads it through, and the pass id is
  `memory-curation-{tenant}-{user}-{run_id}`: distinct per trigger, and
  replayed as-is by a crash-retry of the same trigger, with no durable counter.
- **Scheduled-trigger fires and subagent children no longer count.** A trusted
  fire keeps its creator as `TurnActor` and runs a non-unbound profile, so it
  passed both original guards and could launch a write-capable pass with no
  user present. The derivation is now an ALLOWLIST of conversation profiles
  (`reborn-planned-default`, `interactive_default`, `default`); the
  denylist shape failed open for every profile added later.
- **Curation install fails closed.** `AfterTurnHookDispatcherFactory` returns
  `Result` and the runtime build propagates it as
  `DefaultPlannedRuntimeBuildError::AfterTurnHooks`. Declining is expressed by
  supplying no factory, never by a swallowed error that leaves a deployment
  believing memory is being tidied.
- **A zero interval is unrepresentable downstream.** Config already rejected
  `curation_interval_turns = 0`; `NonZeroU32` now carries through the input
  builder into `MemoryCurationService`, and the clamp is gone.
- **Typed error and typed counter key.** `CurationPassSubmitter::submit_pass`
  returns `UnboundTurnError`; counters key on a `(TenantId, UserId)` struct.
- `// arch-exempt:` on the executor's hook field uses the enforced
  `plan #NNNN` form.

Tests: distinct-vs-converging pass ids; scheduled-trigger and subagent profiles
yield no context, planned-default does; the executor actually dispatches at the
seam (recording hook over a completed bound run, and never for an unbound one);
`accept_and_submit` journals the declared `TurnLimits`. The two curation
scenarios script the pass by owner-scoped thread PREFIX — a new test-support
`register_scope_script_prefix_for_test` — because a per-run pass id is not
knowable before the triggering turn runs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): panic-free interval const + QA harness field the sweep missed

Two breaks, one class: struct call sites in test bins the local
verification set never compiled.

- The production panic baseline scans syntactically, so the compile-time
  `match … unreachable!()` NonZeroU32 constructor counted as a new panic.
  Replaced with `NonZeroU32::MIN.saturating_add(9)` — const, panic-free,
  and the comment says why the odd spelling exists.
- `reborn_parity_qa/binary_e2e.rs` initializes DefaultPlannedRuntimeParts
  and needed the new `after_turn_hook_dispatcher_factory` field (None: QA
  replay drives no lifecycle hooks).

Verified with `cargo check --workspace --tests` — the command that
covers every bin, which the per-crate verification lists did not.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(filesystem): satisfy the Rust 1.98 chunks_exact_to_as_chunks lint

Rust stable 1.98 rolled through CI today and its new clippy lint fails
every branch on decode_embedding_blob's chunks_exact. as_chunks is the
better code anyway: const chunk size yields [u8; 4] directly, so the
per-element indexing disappears. Behavior pinned by the existing vector
tests.

Not this branch's code — the same fix goes to main in its own PR so
every other open branch stops failing too.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(lints): complete the Rust 1.98 clippy migration

Full-workspace sweep under 1.98 (the toolchain CI now runs), on top of the
vector.rs fix already on this branch:

- result_large_err: GoogleCredentialError boxes its Recovery projection
  (one variant, nine sites' worth of warnings); agent_loop's batch error
  boxes its host error; turn_runner boxes only HostFinalizationFailed's
  payload — DriverError stays unboxed because five match sites destructure
  it by pattern, and it is not the oversized member.
- chunks_exact_to_as_chunks: the two UTF-16 decoders in coding/text.rs.
- useless_format in a trace_commons test.

All private types or contained call sites; no public API changes beyond
the boxed variant payloads inside their own crates.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(lints): last two 1.98 sites — map_or_identity, test-support large errors

The tracing-syntax architecture test's map_or(len, |end| end) becomes
unwrap_or; db_write_measurement's error enum boxes its DbProbeError
payloads (test-support only, ~5 construction sites).

Full-workspace clippy --tests under 1.98: clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(hooks): Lifecycle-vs-Effect rationale + amend the side-effect invariant

Approach-audit disposition on #7770 (accepted findings ST3/SP3):

- trust.rs documents why Lifecycle is not a duplicate of Effect: Effect is
  permitted for Installed/SelfAuthored by default — the third-party class
  for post-durable-fact event hooks — while turn completion must not carry
  that default. Folding them would silently widen who may react to a
  finished turn.
- The hooks contract's side-effect invariant now names mediated
  prepared-context turn submission as a sanctioned route for Lifecycle
  hooks, instead of the code silently diverging from a list written before
  unbound turns existed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(memory): audit round 2 — fail-closed curation gate, per-run hook dispatcher

Second approach audit on #7765 (#7770 phase 1). Six accepted findings plus
the documentation gaps they exposed.

Fail closed on a provider that cannot curate. Composition registered curation
whenever an interval was configured and any memory provider resolved, but a
pass REPLACES the standing document and a bound third-party provider may
reject that write outright — a deployment would spawn passes forever that all
fail, with nothing surfacing it. `curation_interval_for_binding` now gates on
the resolved binding and turns a configured-but-unservable curation into a
startup error naming the provider and how to disable it. Nothing in a manifest
declares "supports standing-document replacement" (`[memory].lifecycle` is
about read/record hooks), so the gate is the native binding, with the missing
declaration named in the comment as the seam for #7664.

Hook poison is run-scoped by contract, so the executor now holds a per-run
dispatcher FACTORY instead of one process-lifetime dispatcher: a panic or
timeout is barred for the run it happened in and retried on the next, instead
of disabling curation until restart. The curation SERVICE stays one long-lived
instance — its per-owner counters must accumulate across runs — and each fresh
dispatcher installs a binding over that same service.

Blocked states no longer dispatch. `after_turn_hook_context` requires
`TurnStatus::is_terminal()`: a gated-then-resumed turn fired the point twice,
once while still running.

Also: tier-specific `install_builtin_after_turn` / `install_trusted_after_turn`
replace the trust-class-parameterized installer (an invalid tier is now
unrepresentable, not rejected at runtime); the executor's outer dispatch bound
moves 5s -> 30s so it can never preempt the dispatcher's own per-hook timeout
classification; the unused default-interval constant is deleted and its "ten
matches Hermes" rationale moved to the config field a deployer reads; the
hooks consumer inventory gains `ironclaw_assistant`; and `points/turn.rs` now
states plainly that the point fires only for exits the executor applies —
scheduler failure terminalization does not dispatch it, tracked as a follow-up
on #7770.

Composition budget 42198 -> 42316 (both records, dated): +7 wiring, +109 for
the fail-closed gate and its tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(hooks): enforce the after_turn tier gate in the registry, and close the review gaps

CodeRabbit round three on #7765.

`HookRegistry::insert` now refuses `Installed` / `SelfAuthored` bindings at
`HookPointSpec::AfterTurn`, alongside the phase-vs-trust gate it already
carries. The tier-split installers encoded the restriction, but raw bindings
reach the registry through `from_bindings` and the public builder's
`insert_binding`, which bypass them — the point is act-capable, so an
untrusted binding there would surface as a malformed binding mid-dispatch
instead of an install-time refusal.

The dispatcher's per-hook `after_turn` budget becomes injectable
(`HookDispatcherBuilder::with_after_turn_timeout`, defaulting to
`AFTER_TURN_HOOK_TIMEOUT`), which is what makes the timeout-race regression
affordable: the executor-seam test wedges one hook against a millisecond
budget and proves the hook ordered after it still runs, that the wedged one is
recorded as a Timeout failure, and that the already-terminal run is unaffected.
That asymmetry — outer backstop strictly larger than per-hook budget times hook
count — was fixed earlier but never pinned.

Executor-seam coverage also gains the two non-success terminal states: a FAILED
and a CANCELLED conversation run each dispatch exactly once with
`ctx.completed == false`.

The below-threshold curation scenario no longer rests on a single empty
reading, which a queued-but-unstarted pass would also produce. After crossing
the interval it now requires EXACTLY ONE pass — one pass's worth of model calls
and no more — which is what makes the earlier zero real rather than latency.

The group harness mirrors production's two-part activation gate: curation wires
only when an interval AND a bound memory provider are present, not from the
interval alone.

Version claims in two comments are reworded to name the lint rather than a
toolchain release nobody can verify offline.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(composition): re-measure the mass budget after the main merge

The merge combined main's composition growth (notification inbox #7697,
subagent slice #7788) with this branch's curation wiring; the two
ceilings merged textually without a git conflict while the sum exceeded
both — the gate caught exactly the case it exists for. Ceiling and the
mirrored COMPOSITION_ABSOLUTE_SRC_LOC move together to the measured
42479, dated rationale in the toml. No composition code changes here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(memory): give the curation pass report headroom — live-test finding

The 2026-08-21 live test (DeepSeek-V4-Flash, isolated home, interval 2)
proved the machinery end to end — the pass fired exactly once, acted as
the user, consolidated two wordings of one fact into a correct merged
line, and read its own write back to verify — and then terminated
`Failed { model_call_limit }` before emitting its structured report. A
real model spends calls a scripted one does not: three writes where the
prompt asks for one, plus a fumbled read.

Two changes, both evidence-backed:
- MEMORY_CURATION_MAX_ITERATIONS 6 -> 10. The ceiling still hard-stops
  an unconverged pass; it now leaves room for the report after ordinary
  real-model imperfection.
- The prompt's Finishing section states the budget and the exact
  sequence (read -> at most one write -> result tool), and says plainly
  that a pass dying unreported is worse than a pass changing nothing.

The scripted integration scenario hands the model exactly three replies
and structurally cannot see this failure mode; the constants comment
records the live evidence so the next tuner knows where 10 came from.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(extension-contracts): declare [[memory.scheduled_ops]] — pass ops, trust-gated, cost-floored

A memory provider can now declare its own recurring upkeep in its manifest
instead of the host hardcoding which provider gets which background work.
The provider names the work and the cadence; the host keeps the clock, the
invocation envelope, and the authority.

    [[memory.scheduled_ops]]
    trigger = "after_turn"
    interval_turns = 10
    pass = { prompt = "prompts/memory_curation.md", tools = ["ironclaw.memory.read", "ironclaw.memory.write"], max_model_calls = 10 }

Contracts tier only — nothing dispatches or invokes these yet.

`MemoryScheduledTrigger` is a closed host-owned vocabulary with exactly one
v0 entry; an unrecognized token fails the parse rather than being dropped,
because a silently ignored trigger presents as a provider whose declared
upkeep simply never runs. `MemoryScheduledOpKind` is tagged by which key the
entry declares, and `tool = "..."` is RECOGNIZED and REJECTED with its own
message rather than falling through to an unknown-field error, so a manifest
written against the eventual schema fails with intent. Both keys or neither
are errors too. The wire shape and the parsed shape are separate types, so
`MemoryScheduledOp` cannot be built from a manifest without clearing every
per-op rule.

Three bounds, each with its reason in a doc comment and a test:

- `interval_turns >= 2` (`MIN_SCHEDULED_OP_INTERVAL_TURNS`) — a manifest
  declares work that runs on someone else's deployment at their expense, so
  it must not be able to demand per-turn invocation. `NonZeroU32` makes
  "every 0 turns" unrepresentable before the floor even applies.
- `pass.max_model_calls <= 16` (`MAX_SCHEDULED_PASS_MODEL_CALLS`) — a pass is
  unwatched background spend with nobody reading the transcript. The #7770
  live test put the realistic curation need at 10.
- At most one op per trigger — the host holds one interval counter per
  trigger per owner, so a second op has no well-defined cadence.

Two rules need the whole manifest and land in
`ironclaw_extension_registry::v3::validate_memory_scheduled_ops`, beside the
existing `[admin_configuration]` cross-check and for the same reason — only
that layer sees `[[tools]]` and the requested trust class next to `[memory]`:

- A pass's `tools` must be ids the SAME manifest declares. Declaration is
  selection, never authority: a memory provider must not schedule passes
  wielding another extension's tools.
- Only a first-party/system manifest may declare a pass op at all. A pass is
  a manifest-authored prompt running with write tools, as every user, on a
  schedule — a strictly larger grant than a model-chosen tool call, so it
  gets the same default-deny wall as the after-turn hook tiers. Host-bundled
  alone is not enough, pinned by a test that refuses a third-party-trust
  manifest from a host-bundled source.

`scheduled_ops` is serde-defaulted and empty when absent, so every manifest
written before it existed parses unchanged and schedules nothing
(`memory_manifest_without_scheduled_ops_still_parses`,
`scheduled_ops_absent_in_an_older_manifest_means_none`). `pass.prompt` reuses
`guidance_doc`'s validated bundled-asset ref type; asset RESOLUTION stays
host-side and fail-closed.

The §11.2.3 contracts size ceiling moves 10_841 -> 11_451 for the declaration
family and its inline tests, count read from the ratchet's own failure
message.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(memory): scheduled ops drive curation — native declares its pass, opt-in stays

The declaration replaces the hardwired layer (#7664 addendum v2):

- memory-native's manifest declares its curation as `[[memory.scheduled_ops]]`
  (after_turn, recommended cadence 10, pass over its own three memory tools,
  max_model_calls 10 — the live-test calibration). The curation prompt moves
  into the package beside the guidance doc, exported through the same asset
  table, resolved host-side fail-closed.
- `MemoryCurationService` dies; `MemoryScheduledOpRunner` is built FROM the
  resolved declaration (prompt text, tool ids, model-call ceiling), keeping
  the policy that was already pinned: per-owner counters, completed-only
  counting, the `memory-curation-` pass-id prefix as contract, the submitter
  seam, debug-only failure swallowing. The tool-op arm is
  unreachable-by-construction (leg A parse-rejects it) and says so explicitly.
- Composition's native-only gate arm dies: the gate is now "did the bound
  provider declare an op" — a configured interval against a provider that
  declares nothing stays a startup error naming the provider.

OPT-IN preserved (owner decision, 2026-08-22): the declaration ARMS upkeep —
validated shape, resolved prompt, recommended cadence — and
`[memory].curation_interval_turns` ENABLES it. Omitted = nothing runs,
exactly as before this change; a manifest cannot switch on background token
spend for a deployment that never asked. The config floor (>= 2) is now
enforced at parse, where the operator can read why.

Leg B built by a subagent (session-limited mid-flight), completed and
re-verified from the worktree; opt-in flip + config validation + marker
resolution by the orchestrator.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(memory): the curation prompt demands an explicit append:false — live-test v2 finding

The declared-op live re-test (2026-08-23, DeepSeek-V4-Flash, fresh isolated
home): the pass reached its structured report — the model_call_limit death
from the first live test is fixed — but the model's FIRST write omitted
append:false, transiently duplicating the document before it self-corrected
with a proper replace two calls later. The prompt asked for one write; it
never said which KIND. Now it does, with the consequence spelled out.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:53:19 +00:00
Henry Park
9102166541 ci: reduce scope checkout transfer (#7894)
Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-25 20:44:26 +00:00
Henry Park
453dcf554c ci: single setup-rust composite — toolchain pin, mold, centralized build profiles (T1) (#7821)
* perf(build): codify CI's debug=0 policy as the workspace dev/test/release profile

A handful of lanes (reborn-tests.yml's 10 jobs via workflow env,
reborn-e2e.yml's rust-reborn job, platform-and-compat.yml's
hooks-parity-tests/windows-build jobs, coverage.yml's coverage job, one
live-canary.yml step) already build dev/test at CARGO_PROFILE_DEV_DEBUG=0/
CARGO_PROFILE_TEST_DEBUG=0 via duplicated env decls; this is a true no-op
for them. Every other Rust-building job in the repo (code_style.yml's
clippy/test jobs, reborn-playwright.yml, sccache-dist-smoke.yml,
release-plz.yml, nightly-deep-ci.yml, and most of ironclaw-stress.yml/
live-canary.yml) has never set this env and builds at cargo's full-debug
dev default today, so this task changes their build fingerprint and costs
each a one-time Swatinem/rust-cache cold miss on its next run. Chosen
deliberately: this pays that cost once, up front, for the whole repo,
rather than smearing it across the 13 file-swap commits that follow.
ironclaw-stress.yml also separately restated CARGO_PROFILE_RELEASE_DEBUG=0
three times (already cargo's built-in release default, and now explicit
here too) — release build output is unaffected either way. Measured on
ironclaw_common tests: debug=0 175 MiB vs line-tables-only 257 MiB vs full
272 MiB.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: add the setup-rust composite action (unused by any workflow yet)

One component for what toolchain, linker, and profile env a Rust CI job
gets: installs via the pinned dtolnay/rust-toolchain SHA, then exports
RUSTUP_TOOLCHAIN from the action's own resolved-toolchain output so a job's
cargo invocations can never drift from what this step actually installed —
no separate mechanism needed for nightly lanes. Optional mold: true absorbs
the install/verify/RUSTFLAGS-export steps currently copy-pasted per job.
No workflow calls this yet (following tasks swap each file one at a time),
so this cannot change CI behavior. New ws12 structural check on the action
file itself; self-tested.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(sccache-dist-smoke): install Rust via the setup-rust composite

No behavior change: the composite defaults to the same 'stable' toolchain
dtolnay/rust-toolchain already defaulted to; this job now also gets
RUSTUP_TOOLCHAIN protection ahead of rust-toolchain.toml landing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(reborn-playwright): install Rust via the setup-rust composite

No behavior change: same default toolchain, now with RUSTUP_TOOLCHAIN
protection ahead of rust-toolchain.toml landing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(release-plz): install Rust via the setup-rust composite

No behavior change: same default toolchain, now with RUSTUP_TOOLCHAIN
protection ahead of rust-toolchain.toml landing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(nightly-deep-ci): install Rust via the setup-rust composite

No behavior change: same default toolchain, now with RUSTUP_TOOLCHAIN
protection ahead of rust-toolchain.toml landing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(reborn-release-compile): install Rust via the setup-rust composite

No behavior change: same default toolchain and the same matrix.target
input, now with RUSTUP_TOOLCHAIN protection ahead of rust-toolchain.toml.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(ironclaw-stress): install Rust via the setup-rust composite

No toolchain behavior change: same default toolchain, now with
RUSTUP_TOOLCHAIN protection ahead of rust-toolchain.toml landing. Also
deletes the three CARGO_PROFILE_RELEASE_DEBUG=0 env decls, already no-ops
restating cargo's own release-profile default and now Cargo.toml's explicit
[profile.release] debug = 0 (landed in a prior commit).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(code_style): install Rust via the setup-rust composite

No behavior change: same default toolchain and the same clippy/rustfmt
component inputs, now with RUSTUP_TOOLCHAIN protection ahead of
rust-toolchain.toml landing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(platform-and-compat): install Rust via the setup-rust composite

No toolchain behavior change: same default toolchain and the same
wasm32-wasip2 target input, now with RUSTUP_TOOLCHAIN protection ahead of
rust-toolchain.toml landing. Deletes two now-redundant
CARGO_PROFILE_DEV_DEBUG/CARGO_PROFILE_TEST_DEBUG=0 pairs, no-ops since
Cargo.toml's [profile.dev] debug = 0 landed. Also removes the wasm-wit-compat
job's now-empty `env:` key left over once both its lines were deleted (a
dangling null mapping is valid YAML but dead config; no functional change).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(reborn-e2e): install Rust and mold via the setup-rust composite

No behavior change: rust-reborn keeps mold (now via mold: true — same apt
install, same verify commands, same canonical RUSTFLAGS prefix, moved into
the composite); webui-v2-smoke keeps the same default toolchain. Both now
get RUSTUP_TOOLCHAIN protection ahead of rust-toolchain.toml. Deletes the
now-redundant literal mold RUSTFLAGS string and CARGO_PROFILE_*_DEBUG pair.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(coverage): install the nightly toolchain via the setup-rust composite

No behavior change: both jobs pass the same explicit
toolchain: nightly-2025-11-01 / components: llvm-tools-preview /
targets: wasm32-wasip2 as before, now through the composite, which pins
RUSTUP_TOOLCHAIN to that exact spec — the same protection the prior design
gave these two lanes via a hand-paired job env var, now automatic for every
composite call. Deletes the now-redundant CARGO_PROFILE_*_DEBUG pair in the
coverage job. check-reborn-branch-coverage-flags.py and
reborn_coverage_lane_stack_headroom.rs both still pass (their pinned
strings are untouched).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(reborn-tests): install Rust and mold via the setup-rust composite

No behavior change to toolchain resolution (same defaults/nightly pins as
before, now with RUSTUP_TOOLCHAIN protection ahead of rust-toolchain.toml)
or to the mold linker's effective RUSTFLAGS (the composite's mold: true
prepends the canonical prefix onto each job's existing RUSTFLAGS, so the
two nightly jobs' -Zcrate-attr suffix is preserved byte-for-byte). Absorbs
7 duplicated mold-install steps and 2 duplicated 'Verify mold linker'
blocks into the composite. Two disclosed minor deltas in
qa-recorded-fixtures: mold install now includes clang (already present via
the runner image; matches the other 6 sites) and now runs after, not
before, Install Rust. Deletes the workflow-level mold RUSTFLAGS/
CARGO_PROFILE_*_DEBUG trio, now no-ops. CARGO_INCREMENTAL lines and
RUST_MIN_STACK lines are untouched. check-reborn-branch-coverage-flags.py
and reborn_coverage_lane_stack_headroom.rs both still pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(live-canary): install Rust via the setup-rust composite

No toolchain behavior change: same default toolchain and the same 7
wasm32-wasip2 target inputs, now with RUSTUP_TOOLCHAIN protection ahead of
rust-toolchain.toml landing. Deletes the one now-redundant
CARGO_PROFILE_DEV_DEBUG='0' step env, a no-op since Cargo.toml's
[profile.dev] debug = 0 landed. CARGO_INCREMENTAL is untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(qa): drop the now-redundant CARGO_PROFILE_DEV_DEBUG setdefault

Cargo.toml's [profile.dev] debug = 0 is now the workspace default, so this
QA harness's env setdefault restates it for no reason. CARGO_INCREMENTAL is
untouched — incremental compilation stays a per-caller decision, not a
Cargo.toml profile setting (same reasoning applies to the sibling decl in
scripts/reborn_qa_matrix/run_hermetic_qa.py:101, left alone).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* build: pin the local Rust toolchain via rust-toolchain.toml

Pins exactly what CI's composite-installed toolchain already resolves today
(1.98.0 + clippy/rustfmt), so nothing rebuilds and rust-cache keys hold.
Safe now that every CI job installs Rust through .github/actions/setup-rust
(landed across the prior 13 commits), which exports RUSTUP_TOOLCHAIN and so
cannot be overridden by this file landing. Local clippy/rustfmt now match
CI's.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(agents): document the toolchain pin and bump process

AGENTS.md is scanned by check-guidance.py; .github/workflows/README.md is
not (zero hits under ROOT_GUIDANCE/CRATE_GUIDANCE_BASENAMES), so this is
the enforced location for the bump instructions Task 19's ws12 guard
assumes contributors can find.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(gate): forbid direct dtolnay/rust-toolchain use; enforce the toolchain-pin sync

New contracts: no workflow may call dtolnay/rust-toolchain directly (must
go through .github/actions/setup-rust) or write out the canonical mold
RUSTFLAGS prefix by hand (must pass mold: true); rust-toolchain.toml's
channel and the composite's default toolchain input must name the same
version. Both are simple negative/equality checks, not per-site window
scans — collapsed from the abandoned per-input design specifically because
every job now routes through one component. Sabotage tests included.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): let the composite own RUSTFLAGS so job env cannot shadow mold

A job-level `env: RUSTFLAGS:` is re-applied to every step of that job on
top of whatever earlier steps wrote to $GITHUB_ENV, so it shadows the
setup-rust composite's export for the rest of the job — silently dropping
the mold linker flags. That hit the two heaviest lanes: crate-tests (the
required PR lane and the measured critical path) and
reborn-integration-coverage, both of which declared RUSTFLAGS for their
nightly `-Zcrate-attr` features. The pre-composite code manually pasted
the mold prefix into exactly those two job envs despite the
workflow-level env already carrying it, which is the historical symptom
of the same shadowing.

The composite now takes an `extra_rustflags` input and composes the whole
value (mold flags, then the job's flags, then anything already exported),
so no job-level declaration is left to shadow it. Both reborn-tests jobs
and both coverage.yml jobs pass their crate attributes through the input
instead; zero job-level RUSTFLAGS keys remain in the tree. Install Rust
precedes every cargo step in all four jobs, so the export covers them.

`validate_no_job_env_rustflags_with_setup_rust` pins the invariant: a job
installing Rust through the composite may not declare its own RUSTFLAGS.
It found the two coverage.yml jobs that this fix would otherwise have
missed. Three sabotage tests plus a live-tree assertion cover it.

Note for review: GitHub's documentation does not state the
$GITHUB_ENV-vs-job-`env` precedence explicitly. This change is correct
either way — with no job-level key there is nothing to shadow under any
precedence rule — so the fix does not depend on resolving that question.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): map .github/actions/setup-rust/ into the Reborn PR test planner

Detect Reborn test scope failed on this PR's own head commit: the planner's
fail-closed arm for .github/actions/** had no rule for the new setup-rust
composite, so it raised "unmapped test or CI path" and red the whole
Tests (Reborn) roll-up. Every Tests (Reborn) job now installs Rust through
this composite, exactly like the existing setup-sccache-dist entry in
SHARED_REBORN_ACTION_PREFIXES, so it belongs in the same fail-safe-to-full
bucket: no narrow lane can exercise a change to it, and a change here means
run the exhaustive plan.

Generalized the reason string and its test assertion from "shared sccache
action changed" to "shared reborn action changed" since the bucket now
covers two actions, not one; this is a text-accuracy fix, not a weakened
check (still asserts full mode / all partitions / all lanes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): let has_code see the toolchain pin and its own guard

validate_toolchain_pin_sync() only runs inside fast-checks, gated on
has_code || has_guidance. The has_code regex enumerated crates/, tests/,
Cargo.toml, and friends but not rust-toolchain.toml or
.github/actions/setup-rust/ — so a PR touching only the two files the
sync guard exists to police got has_code=false, has_guidance=false, and
the guard never ran. Add both paths to the has_code regex, and pin them
into CRATE_SCOPE_FILTERS' has_code in_scope probes so a future narrowing
of the grep fails validate_crate_scope_filters loudly instead of quietly
dropping them again.

Regression: added a dedicated test asserting rust-toolchain.toml and
.github/actions/setup-rust/action.yml are in has_code.in_scope, and
confirmed test_checked_in_scope_filters_pass failed against the old
regex before fixing .github/workflows/code_style.yml.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): pin the mold verify step's Linux guard in the setup-rust contract

validate_setup_rust_action() checked the Linux `if:` guard on the mold
install and export steps, but the "Verify mold linker is active" step
in between had no constant and no check at all — an unguarded verify
step would run the mold link check on any runner OS (where mold and
clang are never installed) and nothing would catch it. Add
MOLD_VERIFY_STEP to the guarded set validate_setup_rust_action checks.

Regression: added test_missing_mold_verify_linux_guard_fails and
test_missing_mold_verify_step_fails, confirmed both failed with the old
two-step tuple before adding MOLD_VERIFY_STEP to it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): anchor the toolchain-default regex to the toolchain input block

validate_toolchain_pin_sync() used `re.search(r'default:\s*"([^"]+)"',
action_text)` over the WHOLE setup-rust action.yml, resolving to the
first non-empty `default:` in the file. That is only ever the
`toolchain` input's default by accident — every other input's default is
an empty string today and sits after it. A reordered or added input with
its own non-empty default before `toolchain:` silently redirects the
guard onto the wrong value, hiding real drift between rust-toolchain.toml
and the composite.

Add input_body(), a step_body-style helper that bounds one action.yml
`inputs:` entry by the next input heading, and scope the default-value
search to the `toolchain:` input's own block.

Regression: added
test_a_reordered_earlier_input_with_a_non_empty_default_cannot_hide_drift,
confirmed it failed against the old whole-file regex (a decoy default
matching the pinned channel hid an actual drift in toolchain's own
default) before scoping the search to input_body().

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): catch workflow-level env RUSTFLAGS shadowing setup-rust too

JOB_ENV_RUSTFLAGS matched only six-space job-level `env:` indentation
(jobs.<job>.env.RUSTFLAGS). A workflow-level top `env:` block shadows the
composite's mold export identically — GitHub re-applies it to every step
of every job the same way it re-applies job-level env — but sits at
two-space indentation before any job heading, so it was invisible both
to the indentation-bound regex and to the per-job block slicing (which
only starts scanning at each job heading). Add a WORKFLOW_ENV_RUSTFLAGS
check over the file's preamble alongside the existing per-job scan.

Also add the missing PER-JOB skip coverage: the only existing test for
`SETUP_RUST_USES not in block` exercised the FILE-level skip
(`SETUP_RUST_USES not in text`) via a single-job file. Added a same-file
two-job test where one job uses the composite and a sibling does not, to
prove the sibling's own RUSTFLAGS stays allowed.

Regression: added test_workflow_level_rustflags_alongside_setup_rust_fails,
confirmed it failed before adding the WORKFLOW_ENV_RUSTFLAGS preamble
check, then restored it green; also added
test_sibling_job_without_setup_rust_is_allowed_in_a_multi_job_file to
pin the previously-untested per-job skip branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(ci): assert every exhaustive-plan field for the setup-rust widen

test_shared_setup_rust_action_widens_to_exhaustive_plan only checked 4
of the 11 fields _full_plan() returns (mode, root_partitions,
integration_lanes, and a substring of reasons[0]) — the other 7,
including run_sandbox_docker and run_group_tests, were unpinned. This
is ws12's own guardrail framework, so a regression narrowing this exact
plan's blast radius deserved full coverage, not a partial one.

Regression: widened the assertion to full dict equality against every
field _full_plan() returns; confirmed it fails when run_sandbox_docker
is flipped to False (a change the old partial assertion would have let
through silently), then restored it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): glob both .yml and .yaml when loading workflows

load_workflows() globbed only `*.yml` under .github/workflows/. GitHub
Actions accepts either extension for a workflow file, so a `.yaml`
workflow would silently escape every ws12 contract this loader feeds —
latent today since no `.yaml` workflow exists, but a landmine for the
next one added. Glob both extensions.

Regression: added test_discovers_both_yml_and_yaml_extensions,
confirmed it failed against the *.yml-only glob (the .yaml fixture was
silently dropped) before widening the glob.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): guard every Rust bootstrap, not just the vendor action

An approach audit found the migration's "single owner" claim was false and,
worse, unenforceable: .github/workflows/ironclaw-release.yml:133 still
bootstraps Rust with `curl https://sh.rustup.rs | sh`, so the lane that
builds shipped artifacts got no RUSTUP_TOOLCHAIN pin, no mold wiring and no
check against rust-toolchain.toml. validate_no_direct_dtolnay_usage greps
only the literal `dtolnay/rust-toolchain@`, so that path passed the gate
forever, and no test covered it. The file is live, not inert boilerplate —
commit 2a7d49c60 hand-edited that exact line five days before this branch.

That file is regenerated wholesale by cargo-dist (see
[workspace.metadata.dist], cargo-dist-version 0.31.0), so migrating it onto
the composite would be clobbered on the next regeneration. It is therefore
an ACCEPTED exception rather than a silent one:

- validate_no_unmanaged_rust_bootstrap now fails on `sh.rustup.rs`,
  `rustup-init` and `rustup toolchain install` in any workflow, with
  ACCEPTED_RUST_BOOTSTRAPS pinning ironclaw-release.yml to exactly one
  occurrence. A second bootstrap there, or any in a hand-written workflow,
  fails the gate; and if the generator stops emitting it, the count check
  says so instead of leaving a stale exemption behind.
- AGENTS.md and rust-toolchain.toml said "every CI job" and "exclusively".
  Both were false as written. They now say hand-written jobs, name the
  exception, and point at the checker that enforces it.

Proven by sabotage: adding a curl bootstrap to code_style.yml fails the gate;
adding a SECOND one to ironclaw-release.yml fails it with "2 ... 1 accepted
here"; both restored, gate green. Five unit tests cover the hand-written
case, rustup-init/toolchain-install variants, the accepted single, the
accepted-file second, and the live tree.

Also drops INPUT_HEADING/input_body, a byte-identical duplicate of the
existing JOB_HEADING/job_body with one call site — an action's `inputs:`
entry has the same two-space `name:` shape a workflow job does, so
job_body(action_text, "toolchain") does the job.

Verified: ws12 self-tests 120 OK, ws12 live gate passed, planner suite 88 OK,
check-guidance OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(toolchain): name the accepted release-workflow bootstrap in the pin header

The companion commit corrected AGENTS.md but this file's header still said
"CI workflows install Rust exclusively through .github/actions/setup-rust" —
false while ironclaw-release.yml bootstraps Rust itself, and the same
overstatement the audit flagged. (The earlier edit aborted on an unrelated
assertion before reaching this file.)

It now says hand-written workflows, and names the one accepted exception
plus the checker that pins it to a single occurrence, so the header matches
what the gate actually enforces.

Verified: check-guidance OK, ws12 live gate passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): migrate the release lane onto the composite and unbreak two guards

An approach re-audit and a multi-agent code review both rejected the previous
round's work. Three real defects, all in code added by this PR.

1. The "accepted exception" was unjustified. The claim was that cargo-dist
   regenerates ironclaw-release.yml wholesale, so its `curl sh.rustup.rs`
   bootstrap could not be migrated. The repo contradicts that four ways:
   `allow-dirty = ["ci"]` exists precisely so hand-edits survive regeneration
   (.github/workflows/README.md documents it for the permission hardening);
   two other hand-added steps already live in that job; cargo-dist exposes
   `github-build-setup` (already used for Node/pnpm in
   .github/dist-build-setup.yml) as the durable seam for exactly this; and
   `actions/checkout` and `setup-python` already run via `uses:` in that same
   container job, so a composite call is not structurally precluded.

   The bootstrap now lives in .github/dist-build-setup.yml as a
   `uses: ./.github/actions/setup-rust` call, re-included on every
   regeneration, keeping the previous step's `if: ${{ matrix.container }}`
   condition so release behaviour is unchanged. ACCEPTED_RUST_BOOTSTRAPS is
   empty: no lane is exempt, and the prose in AGENTS.md and
   rust-toolchain.toml states that without carve-outs because it is now true.

2. The workflow-level RUSTFLAGS check was dead code. JOB_HEADING matches any
   two-space `key:` line, so `on:`'s children (`push:`, `workflow_call:`)
   matched and headings[0] truncated the preamble at the first TRIGGER —
   before the real top-level `env:`. Verified on the live reborn-tests.yml:
   first match was `workflow_call` at offset 26, `jobs:` at 3177. Headings are
   now bounded to the `jobs:` block. The old unit test passed only because its
   fixture omitted `on:`; the new one carries a realistic trigger block.

3. The per-job check was blind to YAML aliases. release-plz.yml's
   `release-plz-pr` reaches the composite via `- *install-rust` and contains
   no literal `uses:` line, so the scan skipped it — a job-level RUSTFLAGS
   there would have shadowed mold silently. Anchors carrying the composite are
   now resolved and aliased jobs are checked.

Also broadens the bootstrap guard beyond three rustup literals to known
third-party toolchain actions (actions-rs, actions-rust-lang, hecrj, raftario),
after the coverage lane showed the enumeration was trivially evadable. The one
case text cannot see — a `container:` image shipping Rust preinstalled — is
named in a comment as residual risk rather than papered over.

Proven against the real files that defeated the old guards: a workflow-level
RUSTFLAGS injected into reborn-tests.yml is now caught; a job-level RUSTFLAGS
in release-plz.yml's alias-reached job is now caught; both were silent before.
Also removes an orphaned comment left describing the deleted INPUT_HEADING.

Verified: ws12 self-tests 123 OK, ws12 live gate passed, planner suite 88 OK,
check-guidance OK, both changed workflows parse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): catch step-level RUSTFLAGS, not just job- and workflow-level

Third gap the review found in the same guard. JOB_ENV_RUSTFLAGS matched the
exact 6-space job-env depth, but GitHub also lets an individual step declare
`env:` (10 spaces in this repo's workflows, e.g. reborn-tests.yml:255), and a
step-level RUSTFLAGS shadows the composite's $GITHUB_ENV write for that step
exactly like a job-level one — dropping the mold linker flags with no failing
check, which is the precise regression the guard exists to prevent.

Widened to `^ {6,}RUSTFLAGS:` so job- and step-level depths are both caught,
and corrected the error message, which called a step-level key "job-level".

Proven against the live file: injecting a step-level RUSTFLAGS into
reborn-tests.yml's crate-tests job is now reported; the clean tree still
passes, so the wider pattern adds no false positives.

Worth stating plainly: this is the third hole found in this one
substring-matching guard (dead workflow-level check, alias blindness, now
step-level depth). Enforcement by grepping YAML text is structurally weaker
than the composite's own RUSTUP_TOOLCHAIN export, which makes drift
unrepresentable rather than policed. The stdlib-only constraint on
scripts/ci rules out a real YAML parse here, so the residual risk is the
shapes no pattern anticipated — noted rather than claimed away.

Verified: ws12 self-tests 124 OK, ws12 live gate passed, planner suite 88 OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): restore the release lane's Rust install and assert it is there

The previous commit broke the release lane. It deleted the container-only
`curl | sh` bootstrap from .github/workflows/ironclaw-release.yml and added
the composite call only to .github/dist-build-setup.yml — but that fragment
reaches the generated workflow solely through `dist generate`, which was
never run. The checked-in workflow ended up with NO Rust install path at
all: `grep -c 'setup-rust|Install Rust|rustup' ironclaw-release.yml` → 0.
The next tagged release would have failed on `cargo: command not found` in
every container matrix entry — the exact case the deleted step existed for.

Two independent code-review lanes caught it (Critical/confidence 100 and
High/confidence 92), both noting the same root cause: every check in this
file asserts an ABSENCE — no dtolnay, no raw bootstrap, no shadowing
RUSTFLAGS — so deleting a step and adding nothing read as "clean" and the
whole suite stayed green over a broken release lane.

The step is now in both places, deliberately: the fragment so `dist
generate` keeps emitting it, and the checked-in workflow because that is
the file GitHub actually runs.

validate_release_workflow_installs_rust closes the class by asserting
PRESENCE in both files. Proven red-first in both directions: removing the
step from the generated workflow fails with the cargo-not-found rationale;
removing it from the fragment fails with the regeneration rationale;
restoring either passes. Four unit tests plus a live-tree assertion.

This is the second time in this PR that an absence-only guard reported
success over a real defect. Guards that only forbid shapes cannot notice
that the thing they were protecting is gone.

Verified: ws12 self-tests 128 OK, ws12 live gate passed, planner suite 88
OK, ironclaw-release.yml parses.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(ci): split the Rust toolchain contracts out of ws12_workflow_contracts

Pure move, no behavior change. ws12_workflow_contracts.py had grown to 2,171
lines — past the repo's ~1k ceiling — mixing stress-suite parity, crate scope
filters, WebUI site checks, and this PR's six toolchain validators in one file.
Three separate reviews flagged it (two approach audits, one design lane).

scripts/ci/lib/rust_toolchain_contracts.py now owns one question: what
toolchain, linker, and build flags does a Rust job get? Six validators, their
constants, and a module docstring recording the absence-vs-presence asymmetry
that let a release workflow with no Rust install pass this suite last week.

scripts/ci/lib/workflow_text.py owns the four generic YAML text helpers both
modules need. This absorbed a latent defect: JOB_HEADING was defined TWICE in
ws12 (lines 426 and 988), the second shadowing the first for every validator
below it. The two patterns were `[A-Za-z0-9_-]` and `[a-zA-Z0-9_-]` —
semantically identical, so nothing was wrong today, but only by luck. One
definition now.

The test file imports each symbol from its owning module rather than letting
ws12 act as an implicit re-export, so the ownership is checkable.

classify-test-scope.sh: both new paths join the reborn-scoped list. Without
this the split would have silently changed CI behavior — editing a validator
in its new home would classify differently than editing it in ws12 did — which
is exactly the kind of quiet regression a "pure move" is supposed to not have.
test-classify-test-scope.sh pins it; proven red-first by reverting the list
entry and watching the new assertion fail.

Verified after the move, both validators proven to still fire from their new
home by sabotage: breaking the RUSTUP_TOOLCHAIN export fails the gate, and
deleting the release workflow's composite step fails it with the
cargo-not-found rationale. ws12 self-tests 128 OK, live gate passed,
classify-test-scope 69 PASS / 0 FAIL, check-guidance OK.

ws12_workflow_contracts.py: 2171 -> 1833 lines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(agents): point the toolchain pin at its owner instead of restating it

The AGENTS.md paragraph and rust-toolchain.toml's header had grown into two
near-verbatim copies of the same rationale — rustup precedence, the nightly
coverage-lane carve-out, the cargo-dist fragment, why Docker is unaffected.
Two copies of a rationale drift; the one nobody edits goes stale silently.

AGENTS.md keeps what an agent needs to act: the pin is the source of truth,
never install Rust in a workflow directly, and bumping is a two-place edit in
one PR. The why lives in rust-toolchain.toml's header, where it sits next to
the value it explains. 17 lines -> 11.

Also refreshes the enforcement pointer to the module that now owns those
checks after the split, and mentions the two guards added since this paragraph
was written (release-lane presence, RUSTFLAGS shadowing).

Verified: check-guidance OK (2603 path references, including the new
scripts/ci/lib/rust_toolchain_contracts.py), ws12 gate passed, 128 self-tests OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): require every cargo job to reach the composite, not just the release lane

A review lane proved the release-lane guard was the right fix written too
narrowly. It verified the escape by mutation: delete the `uses:
./.github/actions/setup-rust` step from code_style.yml's fast-checks job — no
dtolnay call, no raw bootstrap, no RUSTFLAGS key — and all 128 tests stayed
green. That is the identical bug the release-lane guard was added for last
commit, reproduced in a different file, because "no dtolnay, no bootstrap, no
shadowing RUSTFLAGS" is trivially true of a job that installs nothing at all.

validate_rust_jobs_reach_the_composite generalizes the rule: any job whose
steps invoke cargo/rustc/rustup must reach the composite, directly or through
a YAML alias. Scoped to jobs that actually run the compiler, so docs and
frontend jobs need no exemption, and `Cargo.toml` in a `paths:` filter is not
an invocation.

It ships with NO allowlist because it needs none — all 31 cargo-running jobs
in the tree satisfy it today. An entry here would mean a lane building Rust on
whatever toolchain the runner image happens to ship.

Proven red-first on the reviewer's exact mutation: the release-only guard
reports 0 errors on it, the new one names all three affected jobs
(fast-checks, clippy, clippy-windows). Six unit tests including the alias path
and a live-tree assertion.

Also in this commit, both from the same review:

- workflow_text.py still defined JOB_HEADING twice. My split moved the
  duplicate instead of removing it, so the commit message claiming
  consolidation was wrong. One definition now, and a test asserts each pattern
  is bound exactly once — the shadowing itself is now what fails, not just
  today's instance of it.
- The module docstring claimed five validators assert an absence. Three do.
  The other four assert presence or equality and are deletion-safe for their
  own subject. An inaccurate tally in the very docstring warning about this
  asymmetry is worth more than a typo.

The job/anchor scan the RUSTFLAGS validator had inline is now shared with the
new check (_composite_anchors / _job_blocks / _reaches_composite) rather than
written twice.

Verified: 137 self-tests OK, live gate passed, check-guidance OK,
classify-test-scope 70 PASS / 0 FAIL.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): classify the cargo-dist fragment so the plan step stops failing closed

This PR's own CI was red on `Build affected-area test plan`, and the cause was
this PR: it edits `.github/dist-build-setup.yml`, the planner had no rule for
that path, and the planner fails closed on anything unclassified. The step
died before scheduling a single lane — so a PR touching only CI plumbing could
not report at all.

Reproduced locally with the exact CI invocation over the PR's real changed-file
list: "Reborn PR test planner failed: unclassified pull-request path:
.github/dist-build-setup.yml". Green after the fix, plan mode `full`.

The fragment is workflow source that lives outside `.github/workflows/`:
cargo-dist re-inlines it into `ironclaw-release.yml` on every `dist generate`.
So it is the same static control as the file it becomes, and no Reborn lane
reads it. That is exactly the `pull_request_template.md` gap the suite already
pins, one directory over.

A sweep test over the surfaces that make CI run then found a second, older gap:
`.github/actions/install-cargo-component/action.yml` has been unmapped since it
was introduced, so any PR editing it hits the same dead plan step. It is
consumed by coverage.yml and platform-and-compat.yml, so it takes the
exhaustive plan like its two siblings — the deliberate mapping that arm's own
comment asks for.

The sweep is scoped to `.github/actions/**` plus the fragment, NOT all of
`.github/`. Fail-closed on an undecided path is this planner's intended
behaviour and an existing test pins `.github/labeler.yml` refusing on purpose;
a wider sweep contradicted it. The line is between config someone may leave
undecided and source a job runs.

Reported, not silently mapped: `.github/labeler.yml` and six
`.github/scripts/*` helpers still refuse. `ci-job-result-ok.sh` is invoked by
workflows, so it is a latent break of the same class — but mapping it is a
policy call for the planner's owner, not a drive-by in this PR.

Both fixes proven red-first: removing either mapping fails the suite, restoring
either passes.

Verified: planner 90 OK, ws12 137 OK, ws12 gate passed, staged-paths 4 OK,
classify-test-scope exit 0, check-guidance OK, and the previously-red CI
command now exits 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(ci): close the three open audit findings

All three were NORMAL, none blocking; closing them so review sees the design
rather than a list of known nits.

**One owner for the debug-info policy** (converged finding, reported
independently by both system-audit lanes). The migration deleted these env
pairs from five workflow job envs on the strength of Cargo.toml's
`[profile.dev] debug = 0` owning the value — but left the identical `:-0`
defaults standing in scripts/ci/quality_gate.sh and
reborn-local-coverage-ratchet.sh. The PR's own claim of a single owner was
not true of the tree it shipped. Both removed; no behaviour change, the values
already agreed, and a developer's `CARGO_PROFILE_DEV_DEBUG=2` still reaches
cargo because `env` inherits what it does not override.

Deleting the lines alone would just let them come back, so
`validate_single_debug_policy_owner` makes it structural: no script under
scripts/ may ASSIGN a CARGO_PROFILE_*_DEBUG value. Assignments only —
run-hermetic-test-process.sh names the same variables in a passthrough
allowlist (a `case` pattern, no `=`), which is exactly how that override
survives the hermetic barrier, and matching it would break the documented
escape hatch. Test files are skipped: they carry the string on purpose as
fixtures, and scanning them would make the contract untestable.

**No speculative escape hatch.** `ACCEPTED_RUST_BOOTSTRAPS` was a per-path
override with symmetric over-use AND under-use validation, plus a test pinning
it empty — built in full for a case the module's own comment says does not
exist. Removed; the rule is now unconditional. If an unavoidable bootstrap
ever appears, the hatch gets added then, with that lane as its first entry.
The test that asserted the dict stays empty now asserts the mechanism is gone
rather than merely empty.

**One job-boundary walk.** `_job_blocks` reimplemented the slice-between-
consecutive-headings logic that `extract_job_block` already did in the same
subsystem. `job_blocks()` in workflow_text.py is now the single primitive:
`extract_job_block` filters it to one named job and keeps its exactly-one
refusal, and the toolchain contracts enumerate it from the `jobs:` offset.
That offset is load-bearing and stays — JOB_HEADING matches any two-space
key, so an unbounded walk treats `on:`'s children as jobs.

Verified: ws12 141 tests OK, live gate passed, planner 90 OK, staged-paths
4 OK, classify-test-scope exit 0, check-guidance OK, both edited shell
scripts pass `bash -n`. The debug-policy guard proven red-first by restoring
one deleted line and watching the gate name it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address PR review feedback (#7821)

Four threads triaged as valid and fixed. Two are real guard bypasses in the
module this PR introduces — a reviewer reproduced both, and both are the same
species the module's own docstring warns about: a check reading text that only
LOOKS like an executable step, or reading too narrow a slice of the file.

- **Release check was file-wide, not job-scoped.** Deleting the composite step
  from `build-local-artifacts` while a decoy `# uses: ...` comment sat in an
  unrelated job returned no error. That matters more than it looks: nothing in
  ironclaw-release.yml matches a literal `cargo` (cargo-dist shells out to
  `dist build`), so `validate_rust_jobs_reach_the_composite` never covers that
  file and this was its only guard. Now scoped to the job, and every `uses:`
  match is comment-stripped per line.

- **Anchor scope ran past its own YAML node.** `_composite_anchors` scanned
  from an anchor to the next anchor or job heading, so a comment mentioning the
  composite anywhere in that span marked the anchor as installing Rust — and a
  job that merely aliased it passed while installing nothing. Bounded by
  indentation to the anchor's own node, and comment-stripped. The legitimate
  release-plz alias shape keeps working (pinned by its own test).

- **Root-level `env:` after `jobs:` was invisible.** A top-level mapping key
  need not precede `jobs:`; such a block applies to every job identically, but
  slicing the preamble at `jobs:` hid it while its two-space indent also dodged
  the six-space per-job pattern. It fell through both checks. Now scans the
  whole file.

- **AGENTS.md overstated the contract.** My earlier commit 463b02974 made this
  paragraph MORE absolute than the one it replaced — "single source of truth"
  and "every job" — when there are two synchronized pins checked for equality,
  neither derived from the other, plus a nightly-lane carve-out and Docker
  outside the contract entirely. That is the exact universal-claim pattern
  `.claude/rules/guidance-maintenance.md` forbids. Reworded to say what is
  actually enforced.

Test-file changes and why:
- `ReleaseWorkflowInstallsRustTests` gained a `release()` helper and its
  fixtures now name `build-local-artifacts`. The old fixtures were a bare
  `jobs:\n  build:` which no longer exercises a job-scoped contract. This
  makes the fixture match the real workflow shape; it does not relax an
  assertion.
- `GuardBypassRegressionTests` adds four tests: one per bypass, plus one
  pinning that a real anchor alias still passes so the anchor fix cannot be
  "fixed" by rejecting legitimate aliases.

All three code fixes proven red-first: reverting each one individually fails
its own regression, and restoring it passes.

Verified: ws12 145 tests OK, live gate passed, planner 90 OK, staged-paths
4 OK, classify-test-scope exit 0, check-guidance OK.

Not addressed here, reported on the threads: the `workflow_dispatch` +
local-action-resolution regression needs a design call (20+ sites, 4
workflows), and one thread's claim was already fixed earlier in this branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): install Rust in the one hermetic lane that acquired it lazily

Root cause of the E2E failures that have been red on every run of this branch
while T2/T3/T4 from the same base are green.

`webui-v2-test-lanes` runs a prebuilt binary and compiles nothing, so it never
installed Rust. But `run-hermetic-test-process.sh` probes `rustc --print
sysroot` to build the child PATH and exits 1 if it cannot resolve one. Inside
the repo that probe hits the `rust-toolchain.toml` this PR adds, so rustup
installed the pinned toolchain LAZILY — mid-lane, once per shard, racing its
own component downloads:

    provider operation shard 2/4 failed with status 1
    error: component download failed for rustfmt-x86_64-unknown-linux-gnu:
           could not rename 'downloaded' file ... .partial

Confirmed by controlled experiment (#7852): reverting only `[profile.dev]
debug = 0` and keeping everything else left E2E red, with a DIFFERENT subset of
shards failing — a race signature, and it exonerates the profile block.

A survey of every lane invoking the hermetic runners found exactly one in this
state; the other seven already install through the composite. So this is one
missing step, not a design fault — but it was invisible because the lane
compiles nothing, and the whole point of this PR is that toolchain acquisition
should never be implicit.

`validate_rust_jobs_reach_the_composite` now treats a hermetic-runner
invocation as needing a toolchain, the same as a literal `cargo` line. It could
not see this before: the workflow text only names the script. Proven red-first
— removing the step again fails the gate by name.

I called this failure an unrelated rustup flake twice. It was neither: it was
this PR's own side effect, visible in the logs from the first run.

Verified: ws12 147 tests OK, live gate passed, planner 90 OK,
classify-test-scope exit 0, check-guidance OK, reborn-e2e.yml parses.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): realign the assertion I broke by rewording the guard message

The previous commit widened `validate_rust_jobs_reach_the_composite` to cover
hermetic lanes and reworded its message from "runs cargo but never reaches" to
"needs a Rust toolchain but never reaches" — but an existing test still
asserted the old string. Fast-checks went red, and Code Style followed it
("fast-checks failed: failure"), so both new failures had this one cause.

Worth recording HOW it shipped, because the mistake was in my verification and
not in the edit. The combined-validation command was:

    python3 scripts/ci/test_ws12_workflow_contracts.py 2>&1 | tail -3 | head -2

`tail -3` yields ["Ran N tests in Xs", "", "OK|FAILED"]; `head -2` then drops
the verdict line. The output read "Ran 147 tests in 12.705s" and I took that
as a pass. The suite had already been failing locally at that point — the
command was constructed so it could not tell me.

Validation now keys off exit codes rather than parsed tail text, so a failing
suite cannot render as a passing one.

Verified (exit codes): ws12 self-tests 0, ws12 gate 0, planner 0,
staged-paths 0, suite-shards 0, changed-packages 0, classify-test-scope 0,
check-guidance 0, docs-boundary 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): make the debug-policy guard see where the value was actually written

Two review lanes independently reported that `validate_single_debug_policy_owner`
cannot detect the regression it exists to prevent, and they are right.

The guard scanned `scripts/**` for `KEY=value`. This PR deleted 14
`CARGO_PROFILE_*_DEBUG` lines from `.github/workflows/**`, all in YAML
`KEY: value` form. So the guard covered the 3 script-side deletions and was
blind to the 14 workflow-side ones — the majority of what it was written to
keep deleted. A future PR re-adding a job-env pair would have passed clean.

Reproduced before fixing: a scratch tree with a workflow job env containing
`CARGO_PROFILE_DEV_DEBUG: 0` returned zero errors, and the regex did not match
the YAML form at all.

Now scans `scripts/**` and `.github/**` across .sh/.py/.yml/.yaml, and matches
both `=` and `:`. Verified the widening is safe: nothing in the live tree
matches, and the `run-hermetic-test-process.sh` passthrough allowlist still
does not — that entry is a `case` pattern with no assignment character after
the name, and it is how a developer's `CARGO_PROFILE_DEV_DEBUG=2` override
reaches the child process. A test now pins that it keeps working.

Proven red-first in both directions independently: narrowing the scope back to
`scripts` fails the two workflow tests; narrowing the syntax back to `=` fails
the same two.

This is the third time in this PR that a guard I wrote asserted something
narrower than the thing it was protecting. The pattern is consistent enough to
name: I write the check against the case that prompted it rather than against
the invariant, and the surrounding cases go uncovered.

Also adds the widening test for `.github/actions/install-cargo-component/`
(coverage lane, low severity). Its two sibling shared actions each had one; this
path had only the `.github/actions/**` sweep, which asserts the planner does not
RAISE but never checks the mode it returns — a mis-ordering could drop it to a
`none` plan and the sweep would still pass. Red-first: removing the prefix entry
fails the new test.

Not changed, with measurements: a performance lane flagged the `Install Rust`
step added to the 4 binary-only E2E shards as avoidable cost. Measured on the
green run, it is 9-11s per shard on parallel runners, and it replaces a
toolchain install that was already happening implicitly inside the test step
while racing itself. Removing it entirely means changing the sysroot probe in
run-hermetic-test-process.sh, which 8 lanes share. Not worth that risk for ~10s;
recorded as a possible follow-up.

Verified (exit codes): ws12 150 tests 0, ws12 gate 0, planner 91 tests 0,
staged-paths 0, suite-shards 0, changed-packages 0, classify-test-scope 0,
check-guidance 0, docs-boundary 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(agents): replace a universal claim with a verified, re-checkable count

`.claude/rules/guidance-maintenance.md` rule 3 forbids universal claims without
a count, and rule 2 asks for a one-line command a maintainer can re-run. The
paragraph said "Every Rust-installing CI job goes through the composite" — an
uncounted absolute.

Pointedly, this sentence was introduced by MY fix for the previous
universal-claim finding on this same paragraph. I removed "single source of
truth" and "every job" from one clause and wrote a fresh absolute into the next.

Now: "35 CI jobs across 11 workflow files need a toolchain and all 35 reach the
composite today, with no allowlist", plus the command that re-derives it and the
validator that enforces it. Counted at this head by walking every job block and
testing for a cargo/rustc/rustup invocation or a hermetic-runner call; 35 need
one, 35 reach it.

Verified (exit codes): check-guidance 0, ws12 gate 0, ws12 self-tests 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): pin the release step's condition, and scope the coverage claim to it

From your review (#7821 review 5012242062), Medium/86 on the release matrix.

The valid core: the contract asserted the composite step's TEXT existed in
`build-local-artifacts` and said nothing about WHEN it runs. `if: ${{
matrix.container }}` decides which matrix entries reach the composite at all,
so it could be narrowed, widened, or dropped and no check would notice — while
this PR claimed every Rust job reaches the composite.

Now pinned in both the generated workflow and `.github/dist-build-setup.yml`,
since a mismatch is what `dist generate` would silently apply. Proven red-first
in both directions: changing the condition names the old and new values;
dropping it reports `<unconditional>`. Three regressions, plus the existing
fixtures updated to carry the condition (they previously wrote a bare step, so
they could not have exercised this).

One part of the finding I checked and do not think holds. It says a
non-container release build "bypasses the new pin/export path and relies on
whatever Rust the hosted runner provides". The pin is not bypassed: with
`rust-toolchain.toml` in the repo, rustup resolves it when `dist build` invokes
cargo on the hosted runner. What those entries miss is mold and the explicit
`RUSTUP_TOOLCHAIN` export, not the toolchain version. It is also not a
regression — the pre-composite step carried the identical `if: ${{
matrix.container }}`, so release behaviour is unchanged by this PR.

Where the finding lands hardest is the docs, and that is my error. AGENTS.md
said "35 CI jobs ... all 35 reach the composite" — true for what it counts, but
the count walks `.github/workflows` only and the release lane is outside it
entirely (no job there names `cargo`; it shells out to `dist build`). A reader
would have taken it as universal. That is the second time in two commits I
replaced a universal claim with another one. Now scoped explicitly, naming the
release lane's actual shape and what the non-container entries do and do not
get.

Open question I am not deciding here: whether non-container release builds
should install explicitly rather than resolving the pin lazily. That is a
release-path behaviour change on a Track C surface, and the lazy resolution is
the same mechanism that raced in the E2E lanes — lower risk here (no parallel
shards within the job) but still implicit.

Verified (exit codes): ws12 153 tests 0, ws12 gate 0, planner 0,
classify-test-scope 0, check-guidance 0, docs-boundary 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): close four review findings on the condition and debug-policy guards

Two Major, both real, both in code I added in the previous two commits.

**The condition check had a false positive on the release path.**
`_composite_step_condition` scanned only the lines ABOVE the composite `uses:`
line. YAML mapping keys are unordered, so a step written `uses:` then `if:` is
valid and equivalent — and returned None, reporting `<unconditional>` and
rejecting a correct release workflow. Reproduced directly: the reordered step
returned None. Now bounded by the step's own `- ` marker and the next line at
or left of it, so the whole step body is scanned. Regression both ways: the
reordered step passes, and a neighbouring step's `if:` is still not borrowed.

**Quoted YAML keys bypassed the debug-policy guard.** `"CARGO_PROFILE_DEV_DEBUG": 0`
and `'CARGO_PROFILE_DEV_DEBUG': 0` are the same mapping key as the bare form and
matched nothing. Fixed with an optional quote on either side; the
`run-hermetic-test-process.sh` passthrough allowlist still does not match, which
a test pins.

Two smaller ones on the test I added last commit, both correct:
- the fail-closed case still probed `.github/actions/setup-rust/` — a leftover
  from mirroring the sibling test, so it re-tested the sibling's invariant
  instead of this one. Now probes install-cargo-component.
- ISC004 implicit string concatenation in the expected `reasons` list.

All three code fixes proven red-first: reverting the whole-step scan fails the
reordered-step test; reverting the quoted-key regex fails both quoted forms;
removing the prefix entry fails the mirrored test.

Pattern worth recording, since it is now consistent. Both Major findings are
the same shape as the three before them: I wrote each check against the exact
shape in front of me — `if:` before `uses:`, an unquoted YAML key — rather than
against the invariant, and the neighbouring valid shapes went unhandled. The
guards keep being narrower than the thing they protect. External review has
caught every instance; my own tests passed over all of them, because I wrote
the tests from the same narrow mental model as the code.

Verified (exit codes): ws12 156 tests 0, ws12 gate 0, planner 91 tests 0,
check-guidance 0, classify-test-scope 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 21:56:29 +00:00
Henry Park
6f47395323 feat(suggestions): generate over the user's no-approval, read-only tools (#7812) (#7833)
* refactor(host-api): add the no-approval declaration seam (no behavior change)

Structural only, and a no-op by construction: the field is added, threaded
through the unbound submission path, and every caller passes `false`.

- `PreparedTurnDeclarations::require_no_approval`: a narrowing-only bool beside
  `TurnLimits`. `false` (the default, omitted from the wire form) changes
  nothing, so already-persisted declarations and the OpenAI-compatible chat lane
  are unaffected. A backward-compatibility test loads a record written before
  the field existed and asserts it comes back permissive.
- `CapabilitySurfacePolicy::without_approval_gated()`: a consuming builder
  matching the sibling `deny_capability_ids`/`narrow_to_capability_ids` idiom.
  It sets a field the type already owns; it computes nothing. No call site yet.
- Corrects the `tools` doc comment, which claimed "Empty means no tools" while
  the runner has always treated empty as "keep the profile's surface". Code is
  the shipped behavior; the comment was wrong.

The contracts size ceiling is re-pinned 20_516 -> 20_632 per the procedure in
its own doc comment: a defaulted DTO field plus one builder, no logic. The
behavior that reads them lives in `ironclaw_turn_runner`, so there is no lower
crate to hold it. Count read from the test's own failure message.

Verified standalone: `cargo check -p ironclaw_assistant` clean, and
`reborn_integration_unbound_turns` 20 passed at this commit.

* feat(suggestions): scope the autonomous surface to the user's approvals (#7812)

Suggestion generation ran against a hardcoded four-capability allowlist, so
cards were never grounded in the user's connected accounts — it could see that
you *have* Gmail but never read a message. It now takes the run profile's
surface narrowed by the user's own approval settings, and
`suggestion_tool_allowlist()` is deleted.

The approval decision is not re-derived. `authorize_visible_capability` already
runs the real gate per candidate — per-tool overrides, persistent always-allow
grants, and the global auto-approve toggle whose default-on IS the "system
default is don't ask" the issue asks for. This sets one input that gate already
consumes: drop capabilities resolving to `RequireApproval`, because a gate with
nobody to answer it parks the run rather than asking.

Declarations rather than a new run profile because the unbound lane has no
alternative: `unbound_turn.rs` sets `product_context: None`, so there is no
`TurnExecutionPolicy` to carry it. A suggestion-scoped surface profile was tried
and abandoned — `accept_and_submit` hardcodes `requested_run_profile: None` and
`coordinator.rs` rejects any non-`unbound_default` hint, so the arm would never
fire; forcing it needs a *non*-unbound profile, which loses
`UNBOUND_DENIED_CAPABILITY_IDS` ("a background run minting more background work
is the runaway class this lane must not open"). #7498 built that shape and was
closed unmerged; #7694 landed the declarations route the next day.

SCOPE OF THE SURFACE — please read before approving. The approval hard floor is
only `Financial | ModifyApproval | ModifyBudget`, and global auto-approve
defaults on, so "needs no approval" is wide. This surface includes
`builtin.shell`, `builtin.write_file`, `builtin.outbound_deliver`,
`builtin.extension_install`/`_remove`, `ironclaw.memory.write`, and — once
connected — `gmail.send_message` and the GitHub write tools. Read/list-only
behaviour is carried by the PROMPT, hardened here accordingly. A deliberate
product decision (#7812); `CapabilitySurfacePolicy` retains the effect-narrowing
seam if it is ever revisited.

Also drops the now-unused `ironclaw_memory` dependency from the assistant crate,
orphaned by deleting the allowlist.

* test(suggestions): pin the autonomous surface to the user's approval settings

Integration coverage for #7812, plus the design record. Every test drives the
real stores the dispatch authorizer reads, and each failed first for a real
reason before passing.

- `connected_extensions_are_reachable_under_the_system_default` — the headline:
  connect Gmail and GitHub, leave every setting alone, and all of
  `gmail__list_messages`, `gmail__get_message`, `gmail__send_message` and a
  `github__*` capability reach the model. `send_message` is deliberately in the
  positive assertion: it declares `external_write` and resolves to Allow under
  the system default, so a future change of heart fails loudly.
- `disabling_global_auto_approve_shrinks_the_autonomous_surface` — the system
  default half.
- `per_tool_ask_each_time_override_removes_only_that_tool` and
  `per_tool_disabled_override_removes_the_tool_even_with_auto_approve_on` — the
  per-tool half.
- `connected_gmail_tool_is_reachable_only_via_the_users_own_grant` — system
  default OFF plus one explicit always-allow grant, isolating the per-capability
  decision from the blanket one.
- The whole surface is pinned, deliberately brittle, so no capability joins an
  unattended run unnoticed.

Two things worth a reviewer's attention, both learned from the tests:

1. What the model receives is `permitted ∩ progressive-disclosure`. Tightening
   permissions can make small tools APPEAR, as the surface drops below the
   disclosure token budget and schemas are advertised instead of `tool_search`
   bridges. Permission assertions therefore run through a new
   `start_runtime_without_disclosure` fixture (`with_tool_disclosure(Off)` — a
   per-runtime builder, no env var, no cross-test raciness); the pinned-surface
   test keeps disclosure on and stays production-shaped.

2. `require_no_approval` cannot gate a SYNTHETIC capability at all.
   `SyntheticCapabilityPort::visible_capabilities` delegates to the inner port —
   where effects and approval filtering happen — then appends every synthetic
   descriptor unconditionally; the outer filter re-checks capability ids only.
   That is why the only survivors when auto-approve is off are synthetics, three
   of which mutate state. Documented, pinned, and left for a follow-up because a
   real fix changes classification for every run.

Hosted MCP is an explicit, documented coverage gap: `notion-mcp` declares no
static tools, so an offline test installs the package and gets nothing. Closing
it needs a stubbed network egress for the discovery handshake.

* fix(suggestions): deny mutating synthetic capabilities to unattended runs

Addresses the review finding raised independently by three reviewers, and
rewrites the generation prompt now that the run has real read access.

WHAT THE REVIEWERS GOT RIGHT AND WRONG. They flagged that `project_create`,
`skill_activate` and `notification_channels_set` reach an unattended run, and
proposed "declare effects on synthetics, or deny-list them". The finding is
real; the diagnosis is not. Declaring effects would change nothing, because
synthetics never reach the check that reads effects:
`SyntheticCapabilityPort::visible_capabilities` delegates to the inner port —
where effect and approval filtering happen — then appends every synthetic
descriptor unconditionally. The outer filter re-checks capability ids only. So
an id deny-list is not a fallback, it is the only lever, and it happens to be
dispatch-enforced rather than listing-only.

The set was also incomplete. Measuring rather than reading the review comments
turned up five, not three: `memory.profile_set` and `trace_commons.onboard`
mutate too — the latter would enrol a user in trace contribution with nobody
present.

IMPLEMENTATION. No new branch: the `require_no_approval` branch already existed,
and this chains `deny_capability_ids` — the same builder the global, scheduled
trigger, and unbound deny lists already use — onto the existing call. The list
lives in composition because that is the only layer able to name product and
domain ids; `ironclaw_turn_runner` cannot depend on `ironclaw_assistant` or
`ironclaw_memory`, and hand-mirroring the ids as string literals is the mistake
84c661566 already corrected once in this codebase.

The append is CORRECT for loop infrastructure (`result_read`,
`structured_result`, `capability_info`) and accidental for product tools that
happen to be implemented synthetically. That invariant was previously written
down nowhere; it is now stated where the list lives. The real fix — making the
append policy-aware so no list is needed — changes classification for every run
and stays a follow-up.

This bypass long predates the change: it arrived unconditional in 7791a3129
(2026-05-30) and has never been altered. What this PR did was remove its cover —
the old four-id declaration kept `capability_ids` narrow, so the outer id filter
stripped synthetics as a side effect.

PROMPT. Rewritten for a run that can now actually read the user's connected
tools. It leads with "look before you suggest" (the old prompt never mentioned
the tools), gives concrete weak/strong contrasts rather than adjectives, and
states that `suggested_prompt` is sent verbatim as the user's next message so it
must be written in their voice and stand alone. Read-only stays absolute and is
stated once — "availability is not permission" — since under this posture the
prompt is the guardrail. 3,936 bytes against a 64 KiB ceiling; no term that
trips the credential detector.

Also corrects two comments that still promised a read-only surface after the
effects narrowing was dropped (`suggestions.rs`, `unbound_turn.rs`).

TESTS.
- The auto-approve-off survivor set is now pinned exactly, and is the proof this
  fix works: it was eight capabilities including five mutators, and is now three
  benign synthetics (`result_read`, `outbound_delivery_targets_list`,
  `capability_info`).
- `excluded_capability_call_is_refused_not_dispatched` drives the real
  model-to-dispatch path: with auto-approve off it scripts a `builtin.shell`
  call and asserts it is refused, not executed. Red-proof — re-enabling
  auto-approve makes the shell genuinely run
  (`"output":"excluded-capability-should-not-execute\n"`). The refusal happens
  at the model gateway as `InvalidOutputReason::OutsideCapabilitySurface`, a
  third enforcement point not previously mapped.

Test evidence:
  cargo test -p ironclaw_integration_tests --test reborn_integration_suggestions
    29 passed
  --test reborn_integration_unbound_turns                          20 passed
  cargo test -p ironclaw_turn_runner        283 passed, 1 pre-existing failure
    (trace_capture::capture_skips_when_policy_missing_or_disabled — fails on
     pristine main too; shared ~/.ironclaw state, unrelated)
  cargo test -p ironclaw_architecture_tests                        all green
  cargo clippy -p ironclaw_host_api -p ironclaw_assistant
    -p ironclaw_turn_runner -p ironclaw_composition --all-targets    clean

* ci(composition): re-seed the mass ceiling for the unattended deny list

Fixes the two red checks. "Code Style (fmt + clippy)" was not an independent
failure — it gates on the fast-checks result and exited in 5s having run neither
fmt nor clippy. The single root cause is
`scripts/ci/check-composition-budget.sh`:

    ABSOLUTE MASS EXCEEDED: composition holds 42380 production LOC, 32 over the
      effective ceiling of 42348.

Caused by this branch, and proven so rather than assumed: the PR base
(a1e3ca5bf5) measures 42341, seven lines inside the 150-line tolerance, so main
is green and unrelated main-side growth since 08-19 had already spent almost all
the slack. The delta is entirely `unattended_denied_capability_ids()`.

The gate is a one-directional ratchet whose own error text names raising the
ceiling as the sanctioned remedy for justified growth, and the growth is
justified: the function assembles a config value from constants owned by
ironclaw_assistant, ironclaw_loop_host, ironclaw_memory and
ironclaw_host_runtime, and composition is the only layer permitted to depend on
all four at once. The enforcement itself lives in ironclaw_turn_runner. That is
assembly, which is exactly what this gate exists to permit — moving the list
down to fit would break the dependency-direction rule instead.

Not raised to the first measurement, though. Nine of those 42380 lines were a
rationale comment duplicating the one on
`DefaultPlannedRuntimeConfig::unattended_denied_capability_ids`. That
explanation now has one home and composition points at it, so the seed is the
real 42371 — a shared budget should not pay for a copy of prose. No threshold
logic, tolerance, or test assertion was touched.

Also catalogues the six new scenarios in `tests/AGENTS.md`, whose Suggestions
table already names every other test in that file. Not an enforced gate
(`reborn_pr_test_plan.py` is not in the code-style workflow), but it is the
repo's documented coverage map and a reviewer flagged the omission — leaving it
stale is how scenario coverage silently drifts.

Verification:
  bash scripts/ci/check-composition-budget.sh      OK: within mass + dispatch budget
  bash scripts/ci/test-check-composition-budget.sh 76 passed, 0 failed
  cargo test -p ironclaw_architecture_tests        green (incl. the WS0 floor and
    nudge-window cross-checks in reborn_restructure_baselines.rs)
  cargo fmt --all --check / clippy -p ironclaw_composition   clean

* ci(composition): pair the baseline record with the raised mass ceiling

Review catch, and a correct one: bff9ac257 raised `loc_ceiling`/`loc_observed`
to 42371 but left `COMPOSITION_ABSOLUTE_SRC_LOC` at 42_198, so the manifest and
the architecture-test record disagreed.

The two are paired, and the codebase says so in three places I should have read
before splitting them:
- the constant's own doc comment — "this record moves with
  `scripts/ci/composition-budget.toml`'s `loc_ceiling`/`loc_observed` to the
  same figure in this commit";
- the same file's ratchet note — "Paired with `[gate].loc_ceiling` ... this
  ratchet fails when the two disagree";
- the manifest's `loc_ceiling` comment — "RE-RATCHET AT EVERY WAVE CLOSE".

Git precedent confirms it: b6c33d33d moved `COMPOSITION_ABSOLUTE_SRC_LOC`
42_142 -> 42_198 and `loc_ceiling` 42142 -> 42198 in one commit, and four of the
last five commits touching the manifest also touched the baselines file.

The earlier reasoning for leaving it — "42_198 is a fixed WS0 floor" — was a
factual error, and worth recording so it is not repeated: the frozen floor is a
DIFFERENT metric. `ceiling_bp` (share-bp, 658) is the one the manifest says is
"deliberately NOT re-set to the current observation", alongside
`WS0_INTEGRATION_COVERAGE_FLOOR_PERCENT` whose comment says it is "a floor
rather than a ceiling". The absolute-LOC record is a ratchet, not a floor.
Neither frozen value is touched here.

Headroom is now correctly zero rather than a stale 173-line gap, which is the
no-padding pattern every prior raise used.

Also adds the separation rationale the six new scenario rows in `tests/AGENTS.md`
were missing. §8 of that file requires it — "If you added a new scenario rather
than extending an existing one, the row you write below must explain why an
existing scenario couldn't absorb it" — and my rows described behaviour only.
Each now carries one clause in the file's existing em-dash idiom naming what
makes it un-absorbable: the disclosure-disabled fixture, a distinct
settings-store write, `Disabled` outranking the toggle, the combined
extension+toggle+grant fixture, and the model-gateway seam.

Verification:
  bash scripts/ci/check-composition-budget.sh       OK (42371 vs ceiling 42371)
  bash scripts/ci/test-check-composition-budget.sh  76 passed, 0 failed
  cargo test -p ironclaw_architecture_tests         all green, ratchets armed
  python3 scripts/ci/docs_publication_boundary.py   clean

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-24 16:37:39 +00:00
Illia Polosukhin
0157128a5f feat(loop): keep the cached system prefix byte-stable across model calls (#7001)
* feat(loop): keep the cached system prefix byte-stable across model calls

Closes #6985 (P0 #2 of the pi-harness adoption program, docs/research/
pi-agent-deep-dive.md §7.3).

The coalesced leading system block is the provider-cached prompt
prefix, and two sources rewrote it per call: inline loop-control
nudges (repeated-call warnings, admission rejections, model-error
observations) were pushed ahead of the identity section, and the
runtime context carries a minute-precision clock that changes every
run. Any change invalidated the entire cached prefix — measured as
the 82% -> 29% hit-rate collapse in prompt_cache_activity.

- InstructionBundleBuilder now emits runtime context and inline
  messages AFTER the thread messages, so per-call context rides the
  conversation tail.
- The model gateway coalesces only the LEADING run of system messages
  into the system block; any later system message stays at its
  transcript position as a <system-reminder>-framed user message,
  keeping host authority explicit without granting transcript text
  the system block's standing. This also keeps compaction summaries
  at their transcript position instead of hoisting them into the
  prefix.
- Memory snippets were already pinned once-per-run via the OnceCell
  cache in loop_host::memory_context (issue item 3) — no change.

Integration tests (tests/integration/prompt_prefix_stability.rs) pin
the invariant end-to-end: a repeated-call warning renders mid-run and
the runtime clock renders every turn, both reach the model in the
conversation tail, and every captured system prompt — across
iterations of one run and across turns of one conversation — is
byte-identical (new assert_system_prompts_identical helper). The two
llm_gateway tests that pinned the old hoist-everything coalescing now
pin the new in-place reminder contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(loop): align payload pins and mock LLM with tail-positioned context

CI follow-ups for the #6985 prefix-stability change:

- Golden payload snapshots regenerated: the runtime clock and run
  origin now render as a tail <system-reminder> user message instead
  of the head of the system block; the multi-turn structural pin
  counts the ephemeral per-request reminder.
- tests/e2e/mock_llm.py response matching skips <system-reminder>
  user messages so canned responses still key on the user's actual
  ask (the reminder was matching as "last user message").
- agent_loop_host_contract runtime-context ordering updated to the
  tail contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(loop): anchor e2e trace matching on the real user message

Second CI round for #6985: the harvested-trace replay matcher
(mock_llm_trace.py) keyed expected user input and the tool-result
window on the LAST user-role message, which is now the tail
<system-reminder> runtime-context frame — every QA journey replay
409'd with "user input does not match". Skip host reminders in both
_last_user_content and the after_latest_user tool-result anchor,
mirroring the mock_llm.py fix.

Also repoint the profile integration test at the whole model request:
profile facts render in the runtime-context section, which rides the
tail reminder rather than the system prompt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(loop): fix responses-api mock tool anchoring; exempt ?-artifact lines

Third CI round for #6985:

- mock_llm.py's own _find_tool_results anchored the tool-result window
  on the last user-role message — now the tail <system-reminder> — so
  the Responses API scripted scenarios saw no results and re-issued
  their tool batches. Skip host reminders, same as the two matchers
  fixed in the previous commit.
- Exact-line changed-coverage exemptions for the relocated
  push_runtime_context/push_inline_message closing `)?;` lines, where
  LLVM maps only the ? early-return region; success paths run in every
  integration turn.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(llm): make host reminders a distinct role, not user messages

Review of #7001 (serrrfirat, High): lowering every tail system message to
`Role::User` makes the appended runtime-context reminder the last user
message on normal calls. Three production consumers answer "what did the
user last ask?" by scanning back for `Role::User`:

  - `unavailable_requested_capability_guard` (loop_host model gateway)
  - `SmartRoutingProvider::classify` (complexity scoring / model tier)
  - `RecordingLlm`'s trace request hint + UserInput steps

All three would have read the runtime clock instead of the request — the
guard would stop inspecting the real capability ask, every turn would score
identically for routing, and recorded traces would embed a per-run timestamp
as the replay anchor (the one thing that hint is documented to avoid).

Add `Role::HostReminder`. It renders in the user shape on the wire — no
vendor API has a role for it — but the distinct variant means:

  - `role == Role::User` filters exclude reminders for free, so all three
    consumers are correct without touching their logic (comments added at
    each so a later refactor to `renders_as_user()` cannot silently undo it);
  - the compiler enumerates all 9 exhaustive `match msg.role` sites, forcing
    each provider to state its rendering decision rather than inheriting one.

Two adapters needed more than a match arm. The tail reminder lands directly
after a tool result, and `anthropic_oauth`/`rig_adapter` push user messages
unconditionally — `rig_adapter`'s own comment records why that is a bug
("multiple consecutive User messages (which Anthropic rejects)"); it merged
tool-result-into-tool-result only. Both now merge same-role, which is what
`bedrock::push_message` already did. Regression tests drive the real
`convert_messages` for both.

Framing moves into `ChatMessage::host_reminder` so it cannot be built
half-formed, and embedded delimiters are neutralized there (CodeRabbit,
Major): a compaction summary quotes the transcript it replaces, so it can
carry a literal `</system-reminder>` that would otherwise close the block
and leave the rest reading as ordinary conversation. Escaping is surgical —
only the two exact delimiters, case-insensitively — so code, generics, and
markup inside a summary survive, which blanket HTML-escaping would destroy.

Regression tests: 4 in `provider.rs` (escaping, non-destruction of ordinary
angle brackets, degenerate frames, `renders_as_user()` vs `== Role::User`
pinned apart), 2 in `anthropic_oauth.rs`, 1 in `rig_adapter.rs`, and a
compaction-summary case in the loop_host gateway suite covering the
demotion #7001 described as a "side benefit" and left untested.

* test(loop): gate prompt-cache reuse on every mock-LLM test

A prompt-cache regression is invisible to functional assertions: the model
still answers correctly, it just costs several times more (#6985 measured
82% -> 29%). Only a cache-shaped check can see it, and pinning that in one
dedicated scenario protects one scenario.

So every test that drives the mock now asserts it. `mock_llm.py` records
each request's leading system block and tool surface per conversation, and
the autouse `assert_prompt_cache_reuse` fixture fails the test when the
system block changed with no tool-surface change to explain it.

What is gated vs. measured matters here:

  - GATED: system block churned, tool surface identical. No legitimate
    cause — per-call content (a clock, a nudge, a counter) is in the
    cached prefix.
  - NOT GATED: both changed. Installing an extension really does rewrite
    the capability list; that invalidation is the price of a real change,
    and gating it would make the check unusable for the extension
    lifecycle scenarios.
  - MEASURED ONLY: history reuse. Compaction legitimately rewrites
    history, so it is a statistic for a ratchet, not a per-test assertion.

Opt out with `@pytest.mark.allow_prompt_cache_churn` (declared in
pyproject.toml so an exemption stays reviewable).

The gate carries its own regression tests per .claude/rules/review-discipline.md
("Guardrails are code"): `test_prompt_cache_gate.py` proves it catches the
real pre-#6985 shape, permits the post-fix tail reminder, allows an
explained surface change, does not cross-contaminate conversations, and
under-reports rather than false-positives after compaction (a documented
limitation of keying chains on the first user message).

The same rule is available at the Rust integration tier via
`TraceLlm::prompt_cache_prefix_churn` and
`assert_prompt_cache_prefix_stable`, over the capture the other assertions
already read.

Also addresses the remaining review comments on #7001:

  - `assert_system_prompts_identical` now compares per REQUEST rather than
    a flattened message list, and rejects a request carrying no system
    block — flattening let a request with no prefix pass silently
    (CodeRabbit, Major).
  - It reports the first differing BYTE and clamps excerpts to char
    boundaries. The old version collapsed to "[non-boundary]" exactly when
    the prompt held multi-byte characters — which it always does, the
    runtime text has em dashes (Copilot).
  - New `assert_rides_conversation_tail` pins position AND role, not just
    presence plus prefix-exclusion: a value emitted ahead of the thread
    messages, or at the wrong role, satisfied both old assertions and still
    broke caching (CodeRabbit, Major).
  - `agent_loop_host_contract`'s runtime-ordering fixture now contains a
    thread message, so the thread boundary is actually exercised; it was
    empty, proving only ordering against identity/instructions
    (CodeRabbit, Minor).
  - Both e2e mocks share one host-reminder predicate requiring a COMPLETE
    frame. Matching the opening tag alone swallowed a genuine user message
    starting with the literal text (CodeRabbit, Minor).
  - Duplicate `[[exemption]]` entry removed; the entry re-anchored after
    the rebase onto main (CodeRabbit, Major).

Golden snapshots regenerated: the reminder's role renders as
"host_reminder" rather than "user", which is the contract worth pinning.

tests/CLAUDE.md rows added per its binding maintenance rule.

* ci: map shared trace support to prompt tests

* ci: refresh prompt coverage exemption path

* fix: address final review findings

* test: keep trace captures atomically paired

* fix(deps): upgrade lru past unsound advisory

* fix(tests): keep trace support tests after items

* fix: fmt

* fix: keep stable prompt prefix snapshots and clean gateway

- snapshots: restore feature's stable prefix (system block without runtime clock, host_reminder at tail) for 4 conflicted payloads
- model_gateway: drop UnavailableCapabilityGuard (dead code after main removed caller) to satisfy clippy -D warnings
- agent_message: map HostReminder to User for provider neutrality
- anthropic_oauth: restore cache_control handling from main and merge HostReminder tail logic
- instruction_bundle: drop concurrency_hint field (struct has no such field)

* fix(loop): keep dynamic memory out of cached prefix

* test(loop): pin recalled memory after transcript

* test(memory): assert recalled context at request seam

* test(loop): expect structured guidance as host reminder

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: serrrfirat <f@nuff.tech>
2026-08-24 15:52:50 +00:00
firat.sertgoz
71bdd6b94f fix(google-docs): add semantic editing tools (#7728)
* fix(google-docs): add semantic editing tools

* fix(google-docs): address review feedback (#7718)

* test(google-docs): expect tab-aware document reads (#7718)

* fix(google-docs): tolerate legacy table index responses (#7718)

* fix(google-docs): locate provider-normalized tables (#7718)

* test(google-docs): cover unsupported table provider state (#7718)

* fix(google-docs): reject post-insert revision drift (#7718)

* fix(google-docs): guard every table readback revision (#7718)

* fix(google-docs): address coderabbit review — fail closed on missing revisions (#7728)

* chore(deps): clear h2 advisory gate (#7728)

* test(google-docs): prove successful table revision flow
2026-08-24 11:31:54 +00:00
Henry Park
3fd143933a docs(guidance): repo-wide agent-guidance audit — fix drift, prune 21.5k lines, consolidate tests/ onto AGENTS.md convention (#7797)
* docs(guidance): repo-wide agent-guidance audit — fix drift, prune 21.5k lines, consolidate tests/ onto AGENTS.md convention

Full-layer audit of the agent-guidance system (root contracts, .claude rules/
skills/commands, family and crate AGENTS.md, CONTRACT specs, tests guidance,
docs/internal), verified reference-by-reference against HEAD.

- Fix stale/ghost references: UserSandboxProcessPort, ProductSurfaceError,
  LlmError::ContextLengthExceeded, INVERTED_PORT_IMPLEMENTORS, split channel
  traits (ChannelIngress/ChannelReply/ChannelDelivery), memory-native's
  never-implemented EmbeddingProvider seam, wrong layer/crate/module counts.
- Convert unpinned prose numbers to regeneration commands or pinning-test
  citations across root, family, and crate guidance (drift-proofing).
- tests/: rename CLAUDE.md -> AGENTS.md with CLAUDE.md symlinks (crates/
  convention), extend scripts/ci/check-guidance.py discovery to tests/,
  delete stale e2e scenario tables, dedupe tier taxonomy against
  .claude/rules/testing.md.
- Commands/skills: delete six dead v1 commands (add-tool, review-pr,
  review-crate, fix-issue, respond-pr, add-sse-event) and the v1-teaching
  architecture-video skill; convert ironclaw-reborn-skill-maintainer into
  the auto-loading rule .claude/rules/guidance-maintenance.md; fix clippy
  -D warnings and portable date in surviving commands; triggers-only
  frontmatter; add automations section to reborn-feature.
- Rules: rename gateway-events.md -> events.md; revive
  scripts/check-type-duplicates.py (glob matched zero types since the
  family reorg); index all 15 rules in root AGENTS.md for Codex parity.
- docs/internal: delete 70 superseded plans/specs/design docs (~21.5k
  lines, each re-verified unreferenced); fix misleading v1-migration
  status lines; rewrite the contracts index as a recipe; restore two docs
  that proved live-referenced.
- Trim composition CONTRACT.md route-mirror sections (invariants kept).

Verified: check-guidance.py (384 files, 0 grandfathered),
docs_publication_boundary.py, cargo test -p ironclaw_architecture_tests,
scripts/ci test-plan suite (87/87) — all green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(guidance): address PR #7797 review comments

Verified each of the ~40 bot/reviewer findings against the tree; applied the
valid mechanical fixes, rebutted the rest with evidence (see PR comment).

- Test planner: add renamed tests/ guidance aliases to IGNORED_GUIDANCE_PATHS
  (reproduced the fail-closed abort on this PR's own changed-file list) and
  extend the planner test to all six guidance paths.
- check-guidance: add self-tests proving tests/-tree discovery and alias
  enforcement (48 tests, was 46); new self-test file for
  check-type-duplicates.py (4 tests).
- Portability: replace GNU-only date -d in triage commands with a python3
  one-liner (works on macOS BSD and Linux).
- Count/claim accuracy: product_contracts manager-port prose 4 -> 3 (matches
  INVERTED_PORT_IMPLEMENTORS), kernel grep -cF for literal #[test] (was regex
  char class, 2 vs 23), rg -o|wc -l for a true total in architecture.md,
  Rust-scoped LlmProvider count (catches 5 generic impls), measured
  1/73-crate dual-backend claim in pr-shepherd, executable wc -l in
  assistant guidance, AST/pytest recipes for the e2e test-count figures.
- Content: deslop co-author line no longer hardcodes an address; ship.md
  surfaces Postgres-skip counts; risk-label guidance documents the crates/**
  labeler blind spot; e2e authoring recipe leads with reborn_v2_* fixtures;
  stale CLAUDE.md line citation replaced with a stable anchor; ✎ provenance
  notes for two deleted-plan citations; unified-channel-model status text
  reconciled with an explicit ChannelDelivery-only exception note.

Gates: check-guidance (384 files, 0 grandfathered), docs boundary, planner
tests 87/87, check-guidance self-test 48/48, type-dup self-test 4/4 - green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): classify scripts/test-check-type-duplicates.py in the PR test planner

The self-test added in 77bc8f8c5 was never registered in
PR_STATIC_CONTROL_PATHS, so the planner's fail-closed unmapped-path arm
aborted 'Detect Reborn test scope' and cascaded into the whole Reborn
matrix skipping. Classified like its subject (deliberately CI-unwired
local dev tool, per the existing entry's rationale) and pinned in the
static-control planner test alongside it.

Verified: planner tests OK; planner run against this PR's full
changed-file list now returns mode=selected with the path owned by
static checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(docs): drop parity-QA path reference from tests/integration/AGENTS.md

The guidance dedup in this PR added a literal
tests/support/reborn_parity_qa reference to tests/integration guidance,
which scripts/ci/check-test-suite-boundaries.sh correctly flags: the
one-way dependency guard covers docs too, and origin/main's version of
this file carried no such reference. Fix the content, not the check -
the tier comparison is reworded to describe the RebornBinaryE2EHarness
seam difference without naming the parity/QA tree.

Verified: check-test-suite-boundaries.sh OK; check-guidance OK; the
full 'Detect Reborn test scope' job reproduced locally end-to-end.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: run the type-duplicates self-test in Code Style; exercise the production report path

Address the two follow-up review findings on 77bc8f8c5/9ed30caf8:

- Wire scripts/test-check-type-duplicates.py into code_style.yml's
  Static-check self-tests step (next to test-check-guidance.py) so the
  regression test is actually enforced; update the planner's
  PR_STATIC_CONTROL_PATHS comment accordingly (classification unchanged
  - Code Style is the static lane).
- Strengthen the semantic-duplicate self-test to also drive the
  production main() report path and assert on its printed candidate
  output, instead of only re-computing similarity locally. Strictly
  stronger; the other three tests are untouched.

Verified: self-test 4/4, planner tests 87/87, workflow-contract
self-test 94/94, planner simulation over both changed paths classifies
cleanly (no unmapped-path abort).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: trigger fast-checks on type-duplicates script changes; symbol recipe for the tier reference

- code_style.yml's has_code filter covered scripts/ci/ but not the bare
  scripts/ type-duplicates pair, so a diff touching only the new
  self-test skipped the fast-checks job that runs it (the previous
  commit's 'runs unconditionally' claim was wrong - corrected the
  planner comment too). Added the two files to the filter following the
  check_no_panics precedent and pinned them as in_scope probes in the
  ws12 workflow-contracts routing test.
- tests/integration/AGENTS.md tier reference is now re-verifiable via a
  single-hit symbol recipe (rg 'struct RebornBinaryE2EHarness') instead
  of a path citation, which the test-suite boundary guard forbids from
  this subtree.

Verified: ws12 workflow contracts 94/94, planner tests 87/87,
check-test-suite-boundaries OK, check-guidance OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 21:26:50 +00:00
Henry Park
85b11d4515 feat(wasm): typed tool response, guest migration, and dispatch-error cleanup (#7711)
* refactor(runtime): centralize capability outcome processing

Route fresh invocation, approval resume, and auth resume through one
concrete capability response processor in ironclaw_host_runtime. Delete
the CapabilityInvocationResult wrapper, fold the completion/error
translation helpers, and remove the RuntimeCapabilityOutcome::Unknown
variant with a persisted-record compatibility test replacing its role.
Behavior-preserving: no outcome semantics change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(runtime): pin inline invocation modes through the response processor

Caller-level coverage for fresh invoke, approval resume, and auth resume
through DefaultHostRuntime, plus a compile-time pin that
RuntimeCapabilityOutcome never derives Serde.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs+test(runtime): fix dispatch-result attribution, pin approval-lookup error paths

Review follow-up: correct the requests-module attribution for
CapabilityDispatchResult in the capabilities AGENTS.md, and add
caller-level tests for the approval-required-but-not-persisted failure
and approval-store outage propagation in the response processor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(runtime): assert grant consumption in approval-resume contract test

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(runtime): restore relocated rationale comments, drop stale references

Review follow-ups: restore the completion-choke-point and
persistence-outage rationale comments at their new processor locations,
fix a stale 'just above' pointer, stop citing an uncommitted plan doc,
and document (rather than churn) the resume-path resource-scope clone
that a stacked PR will make used again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(qa): await fired routine finalization

* fix(mcp): reject CallToolResult.isError instead of passing it as success

MCP JSON-RPC results with isError: true were treated as successful tool
output. Surface them as typed rejections with bounded diagnostics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(extensions): normalize provider failures for model context

Consolidate provider rejection into DispatchError::{AuthRequired,
Rejected} carrying bounded ProviderDiagnostic and
DispatchAttemptAccounting. Add the protocol-neutral ToolError::Rejected
variant and wire it through registry, resolver, and extension tool
binder; share one provider_diagnostic_model_cause helper across
production, loop capability port, and auth-gate diagnostics; extend MCP
usage accounting across handshake and tool-call rejection.

Legacy lane-specific DispatchError variants remain until the PR 4
cleanup; attempt-accounting settlement is still adapter-local pending
centralization at the dispatcher seam.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): terminalize resumed dispatch failures

A Dispatch-class failure during approval or auth resume returned Failed
to the caller while leaving the durable invocation record blocked
forever. Extend the processor's fail_dispatch_run recording to all
three inline modes; gate arms for repeated approval/auth requirements
are unaffected. Pre-existing defect surfaced on #7686 review, fixed
here where response semantics already change.

Also repairs a rebase-fallout compile break in production.rs's test
module: 8 CapabilityInvocationError::Dispatch test literals predate
the provider_diagnostic field this PR's rebase merged in, so they were
missing the field.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): keep preflight missing credentials routed to the auth gate

Provider-failure normalization taught the no-applicable-binding rule to
fail without a gate; that rule is for provider-observed rejections and
must not swallow the preflight missing-credential path, which still
gates as AuthRequired.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): scrub provider diagnostics before guest-visible failures

Provider 403 bodies can echo credential material; route the diagnostic
through the host secret scrub before it reaches any guest- or
model-visible failure string.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(extensions): cover ToolError::Rejected in remaining exhaustive matches

The new protocol-neutral rejection variant missed exhaustive matches in
crates outside the targeted test gates; cover them and verify with a
workspace-wide check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test+fix(extensions): pin the scrub chokepoint and provider-diagnostic edges

Review follow-ups: leak-pin the sandbox-exit scrub, cover the MCP
rejection fallback, Rejected mapping, bounded-message truncation, and
ToolError equality; dedupe provider-text truncation; redact
required_secrets in RuntimeAuthGate's Debug for consistency with
DispatchError::AuthRequired; fall back to safe_summary when a
diagnostic renders empty.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(contracts): drop the unused dispatch attempt accounting carrier

DispatchAttemptAccounting was populated by a single lane and read by
nothing — no consumer reconciled the usage or receipt it carried.
Remove the type and its field from the rejection errors; the governor's
actual reservation accounting is unchanged. Centralized attempt
settlement can be designed on its own merits later.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(github): carry the provider 401 message in the guest error envelope

The provider-context test asserts the model sees GitHub's rejection
message, but the guest emitted only a stable code, so the assertion
could not pass on this branch. Capture the bounded 401 body message in
the structured envelope the host already decodes, and rebuild the
artifact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(wasm): typed WIT tool response with legacy fallback

Replace the option<string>/option<string> response record with a typed
success/failure variant (near:agent@0.4.0): failure carries a closed
error kind plus bounded optional code/message/retry-after. Hosts keep a
legacy 0.3.0 binding fallback through the existing string decode until
PR 4 removes it. The sandbox-exit secret scrub now applies to the typed
failure's free-text fields.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(extensions): migrate bundled WASM guests to the typed tool response

All six wasm-src packages emit the typed success/failure variant
instead of JSON-in-string error envelopes; plain-string Google error
paths are now classified into the closed error-kind vocabulary with
bounded messages. Artifacts rebuilt from source; digests updated via
the freshness script.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(runtime): move host-runtime WAT harness to the 0.4.0 tool contract

Stage-1 gap: the services-contract harness still emitted 0.3.0
imports. Fixtures now target the typed response layout; one
deliberately-legacy fixture pins the 0.3.0 fallback until PR 4 removes
it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(wasm): carry provider 401 code and message through the typed response

The typed-response migration left the GitHub guest's 401 path hardcoding
`message: None` on the `guest-failure` (crates/extensions/packages/github/wasm-src/src/lib.rs)
while the provider's response body (e.g. "Bad credentials") was discarded
in `github_request` (request.rs). Capture the body's `message` field for
the 401 case specifically and thread it onto the typed failure so both the
bounded provider message and the stable code reach the auth-gate diagnostic
and eventual model context. Scoped to 401 only, not every non-2xx status,
to avoid widening what a guest-authored response body can put in front of
the model. Also keeps lane doc comments vendor-neutral.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(wasm): review follow-ups — bound 401 message, gate legacy retry, dedupe decoders

Bound the github guest's captured 401 provider message to the sibling
512-char convention and rebuild the artifact; retry the legacy world
only on a version/import mismatch and keep both errors on double
failure; share one provider code/message decode helper between the
typed and legacy paths; drop the unconsumed legacy-version re-export;
add sheets/slides unit tests for the new typed classification and
guards.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(wasm): more review follow-ups on the typed tool response

- Dedupe the Google Docs UTF-8 decode-failure builder and route the
  negative-index error through the shared input_failure() helper.
- Add 401->auth-required / non-401->client coverage for Docs and Drive
  api_status_error, and a GitHub-specific regression pinning that a
  credential-shaped token in a 401 provider message is scrubbed before
  it reaches the guest-visible failure.
- Add a real component-model regression that decodes the
  WitToolOutcome::Failure variant (prior fixtures only ever exercised
  Success), and one pinning that a double current+legacy world
  instantiation failure preserves both underlying error messages
  instead of dropping the first.
- Update the wasm crate's AGENTS.md/README and
  docs/internal/reborn/contracts/wasm.md for the near:agent@0.4.0 typed
  response ABI and the legacy 0.3.0 fallback removed in PR 4.

* fix(safety): redact credential-shaped tokens in provider error bodies

The github_token pattern required 36 trailing characters, so a shorter
credential-shaped value passed through unredacted — the gap that led the
GitHub guest to capture provider messages only on 401. Lower the floor
to 16 (fail-closed for a prefix that cannot be anything but a token) and
raise the guest-message bound from 256 to 2048 bytes so real provider
explanations survive to the model instead of being clipped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(github): keep one 401 message-capture mechanism after the typed migration

The stringly and typed capture paths both survived the rebase, so the
guest returned a pre-built envelope that was then wrapped again. Keep
the typed thread-local path, drop the superseded pre-built envelope,
and rebuild the artifact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(contracts): retire lane-named DispatchError variants

Mcp/Script/Wasm/FirstParty carried no distinct semantics — every
consumer folded them into the same generic shape. Provider rejection is
now uniformly DispatchError::Rejected with runtime as metadata;
DispatchErrorLane and its classifier are deleted; an architecture test
pins the retired variant names at zero. Model-visible failure strings
are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(wasm): remove the 0.3.0 legacy tool-response fallback

All bundled artifacts ship the typed 0.4.0 contract; the frozen legacy
bindings, string-decode table, and fallback instantiation path are
deleted. An old-contract component now fails instantiation with an
explicit unsupported-contract error.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(contracts): redact rejection detail in Debug, centralize Rejected construction

Review follow-ups: the folded Rejected Debug arm printed the untrusted
failure detail the retired FirstParty variant deliberately omitted —
redact it in both DispatchError and CapabilityInvocationError and pin
with a leak test; add DispatchError::provider_rejected as the one
construction site for cause-only rejections; assert the code-less
safe_summary branch; fix the specificity ratchet comment and an import
split.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(wasm): add direct coverage for classify_instantiation_error

legacy_retry_gate_tests was deleted with the 0.3.0 legacy tool-response
fallback (299c7721f), taking classify_instantiation_error's only direct
unit test with it. The function and its version/import-mismatch branch
still exist post-removal, so restore direct coverage for both the
hinted (near:agent/import substring) and passthrough branches.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(contracts): stop vendor causes from displacing host failure summaries

A folded rejection dropped the host-authored summary whenever a vendor
cause was present, silently replacing the label users see. Precedence
is now explicit — host summary, then validated provider text, then the
kind sentence — with each tier pinned by tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address PR review feedback (#7711)

- preserve host summaries separately from vendor diagnostics
- stamp runtime provenance from trusted bindings
- retain routing failure kinds across the tool ABI
- correct the centralized processor caller count

* Fix stale WASM contract guidance (#7711)

- document near:agent@0.4.0 as the sole supported tool ABI
- remove references to the deleted 0.3.0 fallback and WIT file
- align the living WASM contract with runtime behavior

* Format PR review regression tests (#7711)

* Converge extension tool failures (#7711)

* Rebuild Slack guest for typed failures (#7711)

* test(wasm): cover every failure discriminant

* docs(wasm): remove stale WIT version references

* fix(runtime): address typed failure boundary review

* Address PR review feedback (#7711)

Bound and redact untrusted diagnostic payloads at their trust boundaries.

Narrow extension tool failures to runtime kinds and add regression coverage.

Preserve nested failure detail and typed retry metadata.

* Fix Reborn integration fixture diagnostics (#7711)

Emit the closed provider diagnostic code instead of an extension-minted host summary so the hardened resolver can map the canonical messaging failure.

* Apply Rust 1.98 formatting (#7711)

* Address remaining PR review gaps

Scope the ToolError architecture ratchet and prove it with a malformed-source sabotage fixture.

Pin Google auth-required extension identity and provider scopes in the runtime contract test.

* Fix merge-queue Clippy failure

Remove the stale DispatchFailureKind import left after ToolError switched to RuntimeDispatchErrorKind.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-21 05:28:25 +00:00
firat.sertgoz
9b36cc54f1 feat(sandbox): persistent per-user container with Docker Exec (#7732 Step 1) (#7764)
* docs(internal): sandbox egress spike results + iron-proxy fixtures (#7732 Step 0)

All 9 spike items evidenced on Docker/OrbStack: internal-net + dual-homed
iron-proxy topology, DNS forwarding, default-deny + audit, placeholder
credential swap with require:true, per-runtime TLS trust matrix, exec
stream/kill/zombie mechanics, dead direct/IPv6 egress. Working proxy.yaml
fixtures committed for Step 1 tests; CA material regenerated per run and
not committed.

* feat(sandbox): persistent per-user container with Docker Exec (#7732 Step 1)

Replace create-container-per-command with one reusable container per
(tenant,user), shared across that user's threads:

- stable RebornSandboxUserKey identity and tenant/user-only labels;
  workspace remains per user and scopes without a thread are valid
- ensure/adopt/start/recycle with a per-user lifecycle gate so concurrent
  thread calls converge on one container and shell commands serialize safely
- Docker Exec through ironclaw-exec; per-command process groups, bounded
  TERM/KILL timeout, exact exit codes, capped output
- tini PID 1 in the worker image prevents zombie accumulation
- active guard plus idle sweeper stops only after the current user command
  completes; next use restarts the same compatible container
- posture or image drift recycles lazily; shutdown leaves containers adoptable
- CI planner classifies docker/sandbox/** into the Docker verification lane

Railway, caller APIs, and network posture remain unchanged. Egress mediation
is #7732 Step 2.

* fix(sandbox): harden persistent user-container lifecycle (#7751 review)

- resolve mutable image refs to immutable Docker IDs before adoption
- distinguish real helper deadlines from every ordinary command exit code
  with an invocation-specific final outcome trailer
- keep per-user serialization alive after caller cancellation by detaching
  the bounded execution task
- prevent registry capacity eviction from orphaning live containers
- fix late watchdog signal handling and queued-marker diagnostics
- pin every contemporaneously recorded spike image identity while preserving
  truthful historical command blocks and documenting Alpine evidence gaps

Adds live regressions for same-tag retarget, ordinary exit 124, and aborted
caller serialization; all sandbox, full-turn, architecture, docs, and planner
gates pass.

* fix(sandbox): close lifecycle and process-isolation review gaps (#7751)

- make sync transport construction side-effect-free and route trait-object
  shutdown to the idle supervisor
- reconcile labeled persistent containers after host restart so idle cleanup
  does not lose process-local registry visibility
- replace the shell watchdog with a Python subreaper that terminates detached
  descendants across process groups before returning an authenticated outcome
- cap Bollard framing buffers at 64 KiB
- prove cross-user parallelism, detached-child cleanup, restart reconciliation,
  and same-user cancellation behavior in live Docker tests
- split the oversized live test helper surface into focused test support
- document the transport-local cleanup authority boundary

All sandbox, full-turn, architecture, docs, and planner gates pass.

* fix(sandbox): keep launch-config unit tests daemon-independent (#7751)

Resolve mutable image references in the real run path, then pass the immutable
identity into pure launch-config construction. Unit tests inject a synthetic
image ID and no longer require ironclaw-worker:latest to exist in the crate
bucket, while production still fails closed before adoption when Docker cannot
resolve the configured image.

* docs(sandbox): clarify idle cleanup authority (#7751 review)

* fix(sandbox): launch resolved immutable worker image (#7751 review)

* test(sandbox): document daemon-free launch config fixtures

* fix(sandbox): preserve request preflight before image resolution

* test(sandbox): assert immutable image identity after recycle

* chore(ci): track transitive h2 advisory until libsql upgrade

* fix: address sandbox review findings

* fix(sandbox): address lifecycle review feedback

* test(sandbox): harden Docker removal polling
2026-08-20 11:41:10 +00:00
Henry Park
b93e92d57d fix(extensions): normalize provider failures and auth diagnostics for model context (#7692)
* fix(mcp): reject CallToolResult.isError instead of passing it as success

MCP JSON-RPC results with isError: true were treated as successful tool
output. Surface them as typed rejections with bounded diagnostics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(extensions): normalize provider failures for model context

Consolidate provider rejection into DispatchError::{AuthRequired,
Rejected} carrying bounded ProviderDiagnostic and
DispatchAttemptAccounting. Add the protocol-neutral ToolError::Rejected
variant and wire it through registry, resolver, and extension tool
binder; share one provider_diagnostic_model_cause helper across
production, loop capability port, and auth-gate diagnostics; extend MCP
usage accounting across handshake and tool-call rejection.

Legacy lane-specific DispatchError variants remain until the PR 4
cleanup; attempt-accounting settlement is still adapter-local pending
centralization at the dispatcher seam.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): terminalize resumed dispatch failures

A Dispatch-class failure during approval or auth resume returned Failed
to the caller while leaving the durable invocation record blocked
forever. Extend the processor's fail_dispatch_run recording to all
three inline modes; gate arms for repeated approval/auth requirements
are unaffected. Pre-existing defect surfaced on #7686 review, fixed
here where response semantics already change.

Also repairs a rebase-fallout compile break in production.rs's test
module: 8 CapabilityInvocationError::Dispatch test literals predate
the provider_diagnostic field this PR's rebase merged in, so they were
missing the field.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): keep preflight missing credentials routed to the auth gate

Provider-failure normalization taught the no-applicable-binding rule to
fail without a gate; that rule is for provider-observed rejections and
must not swallow the preflight missing-credential path, which still
gates as AuthRequired.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): scrub provider diagnostics before guest-visible failures

Provider 403 bodies can echo credential material; route the diagnostic
through the host secret scrub before it reaches any guest- or
model-visible failure string.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(extensions): cover ToolError::Rejected in remaining exhaustive matches

The new protocol-neutral rejection variant missed exhaustive matches in
crates outside the targeted test gates; cover them and verify with a
workspace-wide check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test+fix(extensions): pin the scrub chokepoint and provider-diagnostic edges

Review follow-ups: leak-pin the sandbox-exit scrub, cover the MCP
rejection fallback, Rejected mapping, bounded-message truncation, and
ToolError equality; dedupe provider-text truncation; redact
required_secrets in RuntimeAuthGate's Debug for consistency with
DispatchError::AuthRequired; fall back to safe_summary when a
diagnostic renders empty.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(contracts): drop the unused dispatch attempt accounting carrier

DispatchAttemptAccounting was populated by a single lane and read by
nothing — no consumer reconciled the usage or receipt it carried.
Remove the type and its field from the rejection errors; the governor's
actual reservation accounting is unchanged. Centralized attempt
settlement can be designed on its own merits later.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(github): carry the provider 401 message in the guest error envelope

The provider-context test asserts the model sees GitHub's rejection
message, but the guest emitted only a stable code, so the assertion
could not pass on this branch. Capture the bounded 401 body message in
the structured envelope the host already decodes, and rebuild the
artifact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address PR review feedback (#7692)

- redact diagnostic Debug output at its type boundary and scrub WASM trap text
- propagate durable dispatch terminalization failures on resumed invocations
- add nested redaction and resume transition-failure regressions

* Preserve fresh dispatch failures (#7692)

* Apply rustfmt to dispatch processor (#7692)

* Bound WASM guest execution errors (#7692)

- cap redacted guest errors at the canonical model-diagnostic byte budget
- cover exact UTF-8 boundary truncation through WitToolRuntime::execute

* Preserve bounded structured WASM errors (#7692)

* Format structured WASM error handling (#7692)

* Address WASM error review follow-ups (#7692)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
2026-08-20 04:26:55 +00:00
Henry Park
2a7d49c60e fix(ci): bound every unbounded CI operation — apt hangs, uncapped jobs, external downloads (#7756)
* fix(ci): bound every apt fetch so a dark mirror fails fast instead of hanging

GitHub's Ubuntu image lists azure.archive.ubuntu.com as the priority:1 apt
mirror. Since 2026-08-17 21:04 UTC (actions/runner-images incident #5183) that
host intermittently accepts the TCP connection and then stalls mid-transfer
rather than failing. apt therefore never errors and never exits.

install-ci-apt-packages.sh retried `apt-get update` three times, but only on a
non-zero exit -- a hang never produces one, so the retry loop never fired.
There was no `timeout` wrapper and no Acquire::*::Timeout, and `apt-get install`
had neither retry nor timeout. Each hung job burned its entire cap (30-120 min).
On 08-19 that was 14 cancelled merge-queue runs, 175 jobs killed inside
"Install mold and clang", ~271 runner-hours, and repeated PR dequeues.

Bound every apt invocation twice, since the two bounds catch different failures:
Acquire::*::Timeout turns a silent mirror into an apt-level error so the image's
existing fall-through to archive.ubuntu.com becomes reachable, and `timeout` is
the backstop for DNS that never answers, a wedged dpkg frontend, or sudo itself.
Worst case goes from ~120 min of silence to ~13 min of loud, retried failure;
the healthy path is unchanged at ~9s.

Same bug class, elsewhere:
- playwright install --with-deps shells out to apt as root, so bound it in
  reborn-e2e (20-min job), reborn-playwright, and coverage (uncapped job).
- reborn-release-compile installed musl packages via raw apt, bypassing the
  shared script; route it through the script instead of bounding a second path.
- nightly-watchdog's Slack POST had no --max-time, matching the precedent
  already used in main-ci-slack-alerts.yml.

Adds the companion self-test the other scripts/ci/*.sh have, covering both hang
cases, the happy path, retry-on-transient, and the pre-existing Microsoft-source
stripping. It required one seam: APT_SOURCES_DIR, so the stripping scan can be
pointed at a stand-in tree instead of the real /etc/apt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): correct the stall-location rationale in the apt installer comments

Sampling 9 hung jobs shows the priority:1 mirror azure.archive.ubuntu.com is
NOT where jobs wedge -- it fails fast and cleanly (17-26 `Ign:` lines in
seconds), apt has already fallen through to archive.ubuntu.com, and that
fallback is where the transfer stalls. So there is no further mirror to reach
and retrying, not fall-through, is the only recovery path.

Comment-only. Also records the fix's real ceiling: it guarantees a fast, loud
failure, not a green build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): update the release-compile pin to the shared apt installer

Routing the musl toolchain through scripts/ci/install-ci-apt-packages.sh broke
`release_ci_compiles_reborn_for_all_supported_targets`, which pinned the literal
`sudo apt-get install --yes musl-tools binutils file` line.

Re-point the pin at the new invocation rather than dropping it: the assertion
still names all three packages, so its intent -- release CI installs musl-gcc
for C dependencies without overriding Rust's self-contained musl linker -- is
unchanged, and it now additionally pins that the bounded installer is used.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ci): cap every uncapped job and bound the remaining external downloads

44 jobs across 20 workflows had no `timeout-minutes`, so each fell back to
GitHub's 360-minute default. That is the amplifier behind this incident: an
unbounded network step inside an uncapped job can burn six hours of runner time
before anything reaps it. The worst case was coverage.yml's `coverage` job,
which runs both `pip install -e .` and `playwright install --with-deps`.

Caps are sized generously against observed healthy durations -- the goal is to
replace the 360-minute default, not to trim real work. A cap that fires is a
bug report, not a tuning problem.

Also bound the three remaining external downloads that job caps alone would let
run for the whole cap: the cargo-dist and rustup installers in ironclaw-release
and the graph archive fetch in codebase-graph-refresh. The localhost health-check
curls in ironclaw-stress are left alone -- 127.0.0.1 cannot stall the way a
mirror can.

Verified no pin broke: smoke.rs reads code_style.yml, docker.yml,
ironclaw-release.yml and reborn-release-compile.yml, so all 1,456 of its string
literals were checked for presence and occurrence-count against every one of
those files before and after. None changed. Nothing in the repo asserts on
`timeout-minutes`, and smoke.rs contains no `runs-on` literal, so inserting the
cap after `runs-on:` cannot split a pinned block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Henry Park <16583448+henrypark133@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 23:55:55 +00:00
sergeiest
b6c33d33d1 fix(slack): deliver the unlinked-user connect nudge privately, with a one-click connect link (#7681) (#7682)
* fix(slack): deliver the unlinked-user connect nudge privately, with a one-click link (#7681)

An unlinked user who @-mentions the bot in a shared Slack channel got a
"connect it in the Ironclaw web app, then message me here again" notice
that (a) the whole channel could see and (b) dead-ended into a multi-step
manual process with no context carried between steps.

Two changes, both generic rather than Slack-specific:

Private delivery. `OutboundEnvelope` gains an `OutboundVisibility` axis
(`Public` / `EphemeralTo(actor)`), threaded through the notice-class
delivery path only; the policy-routed path stays `Public` by construction.
It is a hint, not a guarantee — an adapter that cannot honor it must still
deliver publicly rather than drop the message. Slack maps `EphemeralTo` to
`chat.postEphemeral` (same `chat:write` scope, no new Slack app
permission). Direct chats keep requesting `Public`: a DM is already 1:1.

One-click link. When a channel's connection strategy is OAuth and the
deployment sets `IRONCLAW_REBORN_WEBUI_PUBLIC_URL`, the host appends
`<url>/chat?connect=<extension>` to the manifest's `connect_required` copy.
`/chat` is an authenticated route, so the link rides the WebUI's existing
`RequireAuth` -> `redirect_after` -> `login_ticket` round-trip unchanged: a
logged-out click lands through login first, a logged-in one lands directly.
The landing hook strips the param (so a reload cannot replay it) and renders
a "Continue to connect Slack" confirmation button — a real user gesture, so
no popup-blocker risk — which drives the same setup -> oauth-start -> popup
sequence the Extensions page already uses. No new backend route.

Deployments that do not set the env var keep today's static, link-free
notice, so this ships dark until configured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(slack): four defects found by live-workspace verification (#7681)

Manual verification against a real Slack workspace surfaced four problems
none of the automated tiers caught. Each fix ships with the coverage that
would have caught it.

1. Connect link failed with "Connection failed". `setup/oauth/start` fails
   closed for an extension absent from the caller's inventory
   (`require_installed_extension` -> 409), and someone arriving from a
   channel nudge has by definition not installed it. The landing hook now
   installs first; install is idempotent, so an already-installed extension
   costs one no-op rather than a pre-flight inventory read.

2. The shared-channel nudge was never visible. Slack renders an ephemeral
   message in a thread only when that thread is already active, and a
   top-level mention self-roots its own — so `thread_ts` pointed at a thread
   with no replies. Slack answered `ok: true` and displayed nothing; the
   durable attempt recorded `delivered` while the recipient saw no message.
   Ephemeral posts now go at channel level, which always renders and stays
   visible only to the target user.

3. The card spun forever after a successful connect. The OAuth completion
   broadcast is same-origin only, so a callback on a different origin than
   the opener (a tunnelled callback against a 127.0.0.1 app) never reaches
   it. Added the durable flow-status poll the in-chat onboarding watcher
   already uses, with the same bounds, so success, terminal failure, and an
   abandoned popup all end the spinner instead of hanging.

4. Run-lifecycle reactions could not work. The documented Slack app manifest
   omits the `reactions:write` bot scope that `reactions.add`/`.remove`
   require, so every reaction failed `missing_scope` -> `authorization_revoked`
   and vanished silently (reactions are best-effort by design). Added the
   scope to docs/channels/slack.mdx. Pre-existing, not introduced here.

Also consolidates configuration: the connect link now reads the existing
IRONCLAW_REBORN_WEBUI_BASE_URL — already the public base URL for Reborn's
OAuth callbacks — instead of the second variable the first cut introduced.
One setting makes both the OAuth redirect and the connect link resolve, and
they cannot drift apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): record chat.postEphemeral in the Slack egress allowlist pin

`first_party_manifest_v3_parity` asserts the live Slack manifest's POST
egress paths exactly, so adding `/api/chat.postEphemeral` to the manifest
had to be recorded here too. Caught by CI, not locally: I ran the slack,
extension-host, assistant and architecture suites but not the composition
crate's, which is where the first-party manifest pins live.

The exact-match assertion is the point — an egress path may not appear
without a reviewer seeing it — so this records the new path with why it is
safe: `chat.postEphemeral` needs the same `chat:write` bot scope as
`chat.postMessage`, so the allowlist widens without the Slack grant widening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(slack): address multi-agent review findings on #7682 (#7710)

* fix(webui): validate the connect-link param and harden its OAuth watcher

Review findings on #7682's `/chat?connect=<extension>` landing:

- Resolve the `connect` query param against the server's installed +
  registry extension inventory and accept it only as an extension with an
  OAuth channel connection, using the server's display name for the label.
  A crafted link can no longer install an attacker-chosen extension and
  start its OAuth flow.
- Clear the stale `oauthError` when a retry starts, so a second identical
  failure still changes the prop the card watches to exit its spinner.
- Guard the flow-status poll with an in-flight latch and catch transient
  errors, so slow calls cannot stack and rejections cannot fire every tick.
- Move the OAuth watcher timeout/poll interval and terminal-status i18n map
  into `lib/product-auth-oauth-events.ts`, replacing the three
  hand-maintained copies.
- Correct the polling comment: it does not recover the card after a reload.

Tests: unknown/non-OAuth param renders no card and installs nothing,
server display name in the label, two identical failures both surface,
abandoned flow times out and stops polling, blocked popup, failed install
closes the placeholder popup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channels): harden the connect-link notice and cover it with tests

Review findings on #7682's connect-link nudge.

The public-origin env var (IRONCLAW_REBORN_WEBUI_BASE_URL) was read raw in
composition while the OAuth path ran the same variable through trim,
trailing-slash stripping, and a blank filter. A deployment with
`IRONCLAW_REBORN_WEBUI_BASE_URL=` in its .env therefore resolved to
Some(""), and the notice posted "Or connect directly: /chat?connect=slack"
— a relative path — into a customer Slack channel. Composition now
normalizes identically and treats a blank value as unset, so the notice
ships link-free instead. Threading the already-validated URL down from the
serve wiring would need a new deployment-config field and CLI plumbing
through the whole composed-runtime build, so the normalization is mirrored
with a comment naming the other consumer.

The extension id is percent-encoded before it goes into the `connect` query
value: it arrives as a raw string from durable installation records, and an
unescaped `&` or `#` would truncate or re-target the advertised link.

connect_required_notice gains unit coverage for every branch (OAuth + base,
trailing-slash trim, unset base, each non-OAuth strategy, and the encoding),
plus a channel_host e2e scenario that drives the production assembly with a
configured origin and asserts the ephemeral nudge ends with
/chat?connect=slack.

Also in this commit:

- Two stale-comment fixes from #7681's ephemeral-delivery inversion: the
  unpaired-mention test is renamed and re-documented for channel-level
  ephemeral delivery (it had kept "threaded" wording while asserting
  thread_ts absent), and SlackChatPostMessageResponse is renamed to
  SlackChatPostResponse now that it decodes postEphemeral, chat.delete, and
  reactions.add as well.
- run_delivery.rs's arch-exempt annotation is rejoined onto one line and
  says `plan #7681` instead of `issue #7681`. No semantic change: the
  ARCH-SPRAWL check in scripts/pre-commit-safety.sh matches only
  `plan #NNNN` on the single line above the `#[allow]`, so the wrapped
  `issue #7681` form failed the gate and blocked every commit on this
  branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(webui): lazy-load the connect-link flow off the initial /chat chunk

The connect-link validation added in the previous commit pushed the eager
/chat closure to 222.2 KB gzip, over its 222.0 KB budget: resolving the
`connect` param pulled the extensions API client and surface schema into
the initial route.

Move the resolver, the install/oauth-start sequence, and the flow watcher
into `pages/chat/lib/connect-link-flow.ts`, loaded through a dynamic
`import()` only when a `connect` param is actually present. The hook keeps
the param detection, the URL strip, and the card state, all of which are
tiny. The click handler still opens its placeholder popup synchronously
before awaiting the module, so the user activation is not burned.

Measured with `vite build` + `scripts/check-bundle-budgets.ts`: 222.2 KB
over budget -> 221.4 KB, which is also 0.1 KB below the 221.5 KB the same
check measures on the pre-PR tree.

No behavior change; all 14 useConnectLinkLanding tests still pass (the
suite now imports the flow module statically so the dynamic import
resolves on a microtask).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* fix(webui): follow the OAuth poll cadence into the shared flow module

#7710 hoisted OAUTH_SETUP_REFRESH_MS into product-auth-oauth-events.ts as
OAUTH_FLOW_POLL_MS, but the static-asset contract test still pinned the old
literal in useExtensions.ts and failed. Assert the cadence where it now lives
and that the hook consumes it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(deps): patch h2 for RUSTSEC-2026-0258 and ignore the libsql-pinned 0.3

RUSTSEC-2026-0258 (h2 queues empty DATA frames without limit) hits both h2
versions in the lockfile and has been failing cargo-deny on main since the
advisory landed, not just on this branch.

Bump our own h2 0.4.15 -> 0.4.16, the patched release. The remaining hit is
h2 0.3.27, which has no patched 0.3.x and is pinned transitively by libsql
0.9.30 (tonic 0.11 -> axum 0.6 -> hyper 0.14); that h2 is a gRPC client to
the libsql server rather than an inbound surface, so ignore it alongside the
other libsql-pinned advisories until that pin is gone.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(ci): re-measure the composition budget on the merged tree

Main's #7717 landing had already consumed the entire prior loc_ceiling +
tolerance window (42141 = 41991 + 150); this branch's own delta is the
#7681 connect-link notice wiring (~11 lines) in extension_host_assembly.rs,
predating that growth. Re-measured on the merged tree with
`bash scripts/ci/check-composition-budget.sh --print` -> 42197 LOC, and
moved loc_ceiling/loc_observed plus the mirrored
COMPOSITION_ABSOLUTE_SRC_LOC constant to that figure together, per both
files' own re-capture instructions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channels): gate the connect link on declared private delivery, stop on rejected install

Two review findings on #7682, both about a promise the code did not check.

The connect notice's privacy comes from an `OutboundVisibility::EphemeralTo`
request that `ChannelReply` documents as a hint: an adapter that cannot honor
it delivers publicly instead. The notice's one-click setup link was appended on
connection strategy alone, so the first OAuth-strategy channel without an
ephemeral endpoint would have broadcast a setup link to every member of a
shared conversation. `ChannelReply` now carries `supports_private_delivery`
(default `false`), Slack declares it, and `connect_required_notice` withholds
the link unless the channel's own bound reply adapter declares it. Delivery
itself stays fail-open — a notice is never dropped.

On the frontend, `installExtension` reports a rejected install in its response
rather than by throwing, so the connect-link flow walked on to setup and OAuth
start for an extension that was never installed and surfaced the resulting 409
as an OAuth failure. It now reads the backend's verdict first.

Tests: the gate's negative case (adapter that never opted in, and a channel
with no reply half) and the notice-level withhold; a frontend regression for
the `{ success: false }` install response. The existing e2e already proves the
positive path through the production Slack assembly.

Also re-pins the `ironclaw_extension_contracts` size ceiling to the merged
tree's measured 10_672 (main alone was already at 10_639; this branch adds 33
declaration lines), and resolves the composition-budget merge conflict by
re-measuring on the merged tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Sergey <sergey@mac.lan>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: firat.sertgoz <firat.sertgoz@near.ai>
Co-authored-by: serrrfirat <f@nuff.tech>
2026-08-19 12:41:35 +00:00
firat.sertgoz
2746fd48db chore(skills): archive parity-blocked bundles (#7641) 2026-08-19 11:36:34 +00:00
firat.sertgoz
c22d38412f perf(agent-loop): make BeforeModel checkpoint batching opt-in and side-effect-safe (#7712)
* wip(7603): preserve in-progress work after agent session limit

Agent was terminated mid-task by a session limit, not by a failure.
This commit preserves the working tree verbatim; the quality gate has
NOT been run and this is not asserted to compile.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(loop-contracts): satisfy clippy and keep the recovery test inside libtest's stack

- Use `is_multiple_of` for the flush-point check (clippy::manual_is_multiple_of).
- Box the side-effect recovery scenario's future: it drives two tool calls plus
  a resumed run, which overflows libtest's default 2 MiB thread stack in debug
  builds. CI sets no RUST_MIN_STACK, so the naked future aborts the whole test
  binary with SIGABRT rather than failing one case.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): reset the parking gate call count when the gate is re-armed

The suggestions suite re-arms one `ParkingModelGate` to park a second
generation (replacement, and the post-restart run). #7603 gave the gate a
`park_at_call` index so the lease-wedge test can interrupt a chosen
iteration, but `park()` counts every call and `rearm()` never reset that
counter — so after the first parked call the index was already past
`park_at_call` and every later call fell through. `wait_until_parked()` then
blocked until its 10s timeout.

Neither change is wrong alone; they only conflict once merged, which is why
this surfaced in the merge queue rather than on either branch. Reset the
counter in `rearm()` so a re-armed gate parks again. `rearm()` and
`parking_call()` have disjoint callers, so the lease-wedge test is unaffected.

Fixes the four suggestions timeouts: dismissing_a_started_suggestion_persists_across_restart,
generation_in_progress_survives_runtime_restart_and_recovers_via_list_view,
replacement_generation_preserves_reservations_and_replaces_cards,
starting_a_replacement_suggestion_creates_one_thread.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(ci): re-measure the composition budget on the merged tree

Merging main took composition to 42142 LOC, 1 over the effective absolute
ceiling (41991 + 150 tolerance), failing Fast deterministic checks. main alone
already measures 42141 — exactly the effective ceiling, with zero headroom — so
this branch's single mechanical line tipped it over.

That line is the whole of this branch's composition delta: the
`before_model_checkpoint_interval: 1` field in an exhaustive `CheckpointPolicy`
struct literal, i.e. assembly keeping up with a contract type gaining a field,
not behavior accreting into composition.

Re-seed loc_ceiling/loc_observed and the mirrored COMPOSITION_ABSOLUTE_SRC_LOC
baseline to the merged-tree measurement, per this file's own "union re-measured
after merging main — the merged tree is measured, not summed" rule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 09:39:50 +00:00
firat.sertgoz
a7f813d270 fix(resources): stop libSQL write-lane starvation from cascading through the resource governor (#7714) (#7717)
* fix(resources): batch governor deltas and stop congestion poisoning the authority

Three defects from #7714, all in the filesystem resource governor.

Delta batching never engaged. Every caller blocks on its own ack, so the
flusher's greedy `try_recv` drain always found an empty channel and wrote
`batch_size=1` under exactly the load the 256-delta group commit was built
for. The flusher now collects for a bounded 2ms window, so concurrent deltas
coalesce into one durable append.

Transient `BackendBusy` poisoned the whole authority. A contended libSQL
writer exhausted the retry window, the governor treated that like corruption,
and the poison/journal-replacement/durable-reload cascade failed unrelated
reservation releases. The journal now reports whether a failure is congestion
or damage: congestion keeps the writer thread alive and only discards the
diverged in-memory authority so the next call replays the append-only log,
while genuine storage errors still invalidate. In-memory state cannot be
rolled back in place because the per-account commit gate is released before
the durable ack is awaited, so replay is the repair.

Released-failed reservations leaked forever as `Active` holds that replay on
every restart. `ReservationRecord` now carries `reserved_at` (absent on older
records, which are never swept), and `sweep_stale_active_reservations` appends
a normal `Release` delta for holds older than a caller-chosen max age. Nothing
is deleted. The host runtime owns the schedule and is wired separately.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): retry failed reservation releases and report lost leases distinctly

Two independent defects from issue #7714 (libSQL write-lane starvation):

Deferred reservation release. When a first-party dispatch failed to
reconcile and the follow-up release also failed, the reservation id was
only logged and the hold leaked permanently — the durable `Reserve` delta
replays as `Active` after restart and no sweeper reclaims it. Failed
releases now go into a bounded (64-entry) in-memory queue on
`FirstPartyRuntimeAdapter`, retried at the start of the next dispatch
against the same governor. Queue-full drops are logged at `debug!` per the
REPL logging rule.

Distinct LeaseLost error. A capability process whose lease expired and was
recovered surfaced as `UnknownInvocation`, which the host runtime rendered
as "process invocation not found" — a lost lease mislabelled as a missing
record. `ProcessInvocationError::LeaseLost` now covers both distinguishable
cases (a terminal process failed `lease_expired`/`crash_retry_exhausted`,
and a present-but-unclaimable process), and the host runtime reports
"process invocation lease lost or expired". The recovery failure categories
became shared constants so the two sites cannot drift.

`ironclaw_capabilities::helpers::invocation_state_error_kind` matches the
enum exhaustively and gains the one-line `LeaseLost` arm.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(libsql): give the governor and process journals their own write lane

On libSQL every durable writer in the process shares one write connection
(LIBSQL_WRITER_POOL_MAX_CONNECTIONS = 1) with a 10s checkout timeout, so bulk
per-turn traffic (events, messages) queues ahead of the resource-governor delta
journal and the process journal. Under PinchBench load the governor's write
spent ~40s waiting for that slot and never reached SQL, which the governor read
as a failed authority: journal replacement, failed reservation releases, and
capability calls ending in "process invocation not found" (nearai/ironclaw#7714).

PR #7471 built the relief valve for exactly this contention, but only for
Postgres (process_journal_pool). This extends it to libSQL: LibSqlRuntime gains
split_journal_lane(), a second admission runtime over the same database, and
composition routes both latency-sensitive journals through a filesystem handle
on that lane. The lane addresses the same rows over a different connection, the
way the Postgres pool split does.

Why a second connection helps on a single-file SQLite database: the lanes then
contend for SQLite's write lock rather than a pool slot. Migrations put the
database in WAL journaling, so a write lock is held for one short transaction
and busy_timeout retries across it, whereas a pool slot is held for as long as
its holder keeps the lease with no fairness for a starved waiter. The journal's
queue goes from "every writer in the process" to "one transaction". Exactly one
extra process-wide writer is admitted; raising the writer pool cap instead would
let unbounded bulk writers fight over the write lock, which is the contention
the single-slot pool exists to prevent.

Postgres is unchanged: its data plane is already a multi-connection pool, so
only the process journal takes the dedicated pool and the governor stays on the
data plane.

Tests:
- ironclaw_libsql_runtime: journal_lane_writes_while_the_data_plane_writer_is_held
  pins that the lane is admitted while the data-plane writer slot is held and
  that its write lands in the same database (verified failing with Elapsed when
  the lane is made to share the data-plane pools).
- ironclaw_composition: libsql_journal_lane_is_a_separate_write_lane_over_the_same_rows
  extends the existing journal-split test set with the libSQL leg — mount
  parity, admission under a held data-plane writer, and read-back over the data
  plane.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(resources): fence queued deltas to the authority generation they were built from

Addresses review on #7717.

IronLoop (High) — discarding the diverged authority did not stop deltas
already queued behind the failed batch. A delta stacked on top of a failed
one could still be appended, writing a log whose prerequisite is missing and
which `replay_journal` cannot apply, failing every later authority load. The
delta journal now carries a generation: an authority records the generation
it loaded under, `enqueue` refuses a delta stamped with any other, and the
flusher retires the generation *before* acking a failed batch — so queued
deltas are dropped and a caller racing the ack is refused at enqueue rather
than appending into the gap. Regression is discriminating: without the fence
the queued delta persists at SeqNo(1).

CodeRabbit (Major) — every failed prepared-reservation cleanup now defers.
The planner-failure and service-resolution-failure branches released
best-effort and only logged, so a failed release leaked the hold with nothing
to reclaim it. Both now go through `release_or_defer`, joining the deferred
retry queue. Caller-level regressions for both paths, each discriminating.

CodeRabbit (Minor) — caller-level coverage for `crash_retry_exhausted`
through `ProcessInvocationStore::complete`; without the classification it
falls back to a generic `Backend` error, which the test now catches.

CodeRabbit (Minor) — `reserved_at` rollback restriction documented on the
field and pinned by a test: a timestampless record still serializes without
the key (older readers keep loading it), while a stamped record is rejected
by a reader predating the field, so rollback needs a pre-upgrade snapshot.

CodeRabbit (Major, x2) — documented rather than changed: both journals share
one libSQL write lane on purpose (#7714 was queue depth, not two producers; a
third process-wide writer would add another contender for SQLite's write
lock), and `split_journal_lane`'s cross-lane transaction constraint is now
stated on the method.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(libsql): address coderabbit review — bound journal lane creation (#7717)

* chore(composition): trim duplicated journal commentary (#7717)

* fix(composition): isolate direct libsql journal writes (#7717)

* fix(composition): route process journal through libsql lane (#7717)

* fix(host-runtime): move test-only process-runtime accessor out of the production impl

`process_runtime_for_test` was added to `HostRuntimeServices` as a per-method
`#[cfg(any(test, feature = "test-support"))]` inside the production impl, which
took the struct-debt ratchet's count for services.rs from its frozen baseline of
1 to 2 and failed
`reborn_production_struct_test_support_and_dead_code_members_do_not_grow` in the
merge queue with "test-support method in .../services.rs: 1" (the delta above
baseline).

The scanner skips an item whose own attribute is a test cfg, `Item::Impl`
included, so the accessor moves into its own cfg-gated `impl` block. The gate
keeps both halves — the sole caller is
`ironclaw_composition/tests/libsql_substrate.rs`, an integration test in another
crate that a plain `#[cfg(test)]` would not reach. No baseline was changed and
no behavior changed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(resources): only vacate the authority when the generation diverged

`enqueue_delta` vacated the authority slot for every error the delta
journal's `enqueue` returned, though the adjacent comment justifies it only
for generation divergence. `enqueue` has three failure kinds, all flattened
to a `ResourceError::Storage` string: a retired generation, a poisoned sender
lock, and a stopped writer thread. Only the first means the in-memory
authority diverged from the log.

For the other two the delta never reached the queue, so the authority is
still consistent with the log. Vacating there made the caller's
`invalidate_authority` find a non-current slot and skip the journal restart —
the only path that replaces a dead writer. A journal-side failure therefore
left the governor with no writer and no way to get one back: a process-wide,
restart-only outage.

`enqueue` now returns a typed `DeltaEnqueueError` so the two are
discriminated by variant rather than by message. Divergence still vacates;
a journal-side failure returns the error with the slot intact, which mirrors
the busy-vs-infrastructure split already used by `fail_delta`.

Reported by PierreLeGuen on #7717.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(composition): keep the process-journal migration cause typed

Both production substrate paths ran the process journal's startup migration and
mapped its `ProcessJournalStoreError` into `InvalidConfig { reason: format!(..) }`.
That flattened the filesystem cause into a string, so an operator whose startup
failed saw "process journal startup migration failed" with no way to tell an
unreachable database from a rejected index declaration — the lost-cause pattern
`.claude/rules/error-handling.md` bans.

`RebornCompositionError::ProcessJournalMigration` and its `RebornBuildError`
counterpart now carry the store error as a source, and both call sites convert
with plain `?`, which is also four lines shorter than the mapping it replaces
(composition's LOC budget had three lines of headroom, so the smaller shape is
load-bearing, not incidental).

Test: process_journal_startup_migration_failure_keeps_its_cause drives the real
migration against a backend that refuses index declaration, then walks the error
chain at both boundaries — composition's error and the build error it converts
into — asserting the backend's own reason is still reachable. Verified failing
before the fix: with the source dropped the chain renders as bare "process
journal startup migration failed".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-18 12:08:32 +00:00
firat.sertgoz
18ab836f2a fix(release): forward-port 1.2 fixes and thread repair (#7663)
* fix(release): isolate smoke workspace on Windows

* fix(filesystem): publish absent writes atomically on Windows

* fix(release): preserve Windows identity in smoke environment

* fix(windows): keep ACL output out of CLI JSON

* chore(release): forward-port 1.2.0 metadata and healthcheck fix

* fix(release): address review feedback

Use create-only Windows publication, preserve successful writes when temp cleanup fails, surface ACL errors, and make release guard assertions structurally precise.

Remove the migration-shaped thread projection repair from the rebuilt PR branch.

* fix(threads): restore projection repair for upgrades

Keep the one-time thread-index projection repair while moving it off the listing request path. Bound concurrent and pending work, retry incomplete repairs, cover oversized scopes, and add a force-write CAS mode for damaged sidecars.

* fix(storage): address review feedback (#7663)

Bound projection repair with stable keyset directory pages and a per-scope retry budget. Run release smoke commands from the isolated workspace and classify Windows publish conflicts from the original OS error.
2026-08-17 19:53:48 +00:00
Josh Ford
c145f6e522 test(docs): doc-fact contract tests for CLI, manifest, and Responses claims (doc-truth PR 3/5) (#7378)
* docs: fix live drift in extension, responses API, and channel docs

The public tutorial taught the retired manifest v2 authoring format
([[host_api]] / [capability_provider.tools] / runtime_credentials), which
the v3 parser hard-rejects, and never mentioned origin_gate_matrix; the
Responses API page claimed temperature is rejected (accepted 0.0-2.0 and
forwarded), claimed model must be "default" (any well-formed name <= 256
bytes), claimed max_output_tokens is rejected (accepted and ignored by DTO
policy), and omitted the required model field from every request example;
the channel tutorial pointed at two files that no longer exist.

- docs/extensions/building-a-tool.md: rewrite manifest sections to the v3
  [[tools]] / [[tools.credentials]] / [auth.<vendor>] shape, document
  origin_gate_matrix (origins, policies, ratchet), correct the hosted-MCP
  [mcp] section, packaging via ironclaw_extension_support package modules,
  and v3 test references; drop the nonexistent script runtime kind.
- docs/api/responses.mdx: correct model/temperature/tools/tool_choice
  rejection rules, document unknown-field tolerance, add the required
  model field to all 15 request examples.
- docs/channels/building-a-channel.mdx: replace dead
  crates/ironclaw_first_party_extensions + available_extensions.rs
  registration instructions with the current package-directory mechanism.
- docs/reborn/contracts/extensions.md: state that production manifests
  author v3 (lowering into the v2 resolved model described there); label
  the v2 examples as legacy.
- docs/reborn/how-to-port-tool-to-reborn.md: superseded banner pointing at
  the v3 guides.

Part of #7317 (doc-truth pipeline, PR 1 of 5).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): extend the reference gate to the docs/ surface

The public Mintlify tree had no path-reference validation — a published
tutorial told contributors to edit files that no longer exist and nothing
caught it. check-guidance.py already owned the machinery (tracked-tree
resolution, fence exclusion, suppress markers, shrink-only debt, fail-closed
floors), so the docs surface joins the same gate rather than a fork.

- discover_guidance() now collects every tracked docs/**.md|.mdx: published
  pages, the zh/ locale mirror, and the living contract corpus
  docs/reborn/contracts/. Dated archives (docs/internal/, the non-contract
  parts of docs/reborn/) are excluded as classes — measured 2026-08-07,
  705 of 709 dangling docs references sat in those historical corpora, and
  forcing dated plans/ADRs to track today's tree would either rewrite
  history or drown KNOWN_MISSING.
- docs/ files extract backticked inline paths only; Mintlify markdown link
  targets are site routes (extensionless pages, site-absolute /using/cli),
  a different namespace than the tracked tree, so the link extractor is off
  there by design.
- _reference_lines learns MDX comments ({/* ... */}), including
  {/* check-guidance: path-ok */} as the .mdx suppress-marker form, with the
  same one-reference-per-marker and multi-line semantics as HTML comments.
- Floors re-measured and re-dated (364 files / 2276 references; floors
  180/1100), plus a dedicated MIN_DOCS_FILES=60 floor: the aggregate floors
  sit below the guidance-only remainder, so the docs branch of discovery
  silently breaking needs its own refusal. --json now reports docs_files.
- Fixes the four real dangles the new scan found in docs/reborn/contracts/
  (moved nested_dispatch_stream.rs test home, retired event-store migrations
  directory, loop_driver_host tests->src move). KNOWN_MISSING stays empty.
- Self-tests: 8 new cases (dangling docs path fails; Mintlify links are not
  references; MDX marker suppresses exactly one reference; multi-line MDX
  comment hides content; zh discovered; archives excluded but contracts
  scanned; docs fence fails closed; docs floor refuses).
- ws12_workflow_contracts.py: docs/api/responses.mdx and docs/zh/index.mdx
  join the has_guidance in-scope probes so a narrowed trigger regex cannot
  silently skip the gate for public docs.

Part of #7317 (doc-truth pipeline, PR 2 of 5); stacked on #7375.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(docs): pin CLI, manifest, and Responses doc claims to code

Three deterministic doc-fact contract tests, each living in the crate that
owns the truth it checks, so the drift #7317 describes fails CI instead of
shipping:

- crates/app/ironclaw_cli/tests/docs_cli_reference.rs: parses the real
  binary's --help and cross-checks docs/using/cli.mdx table rows both ways
  (every visible subcommand documented, any alias form counting; every
  documented command real), with a fail-closed row floor. Doc gaps this
  surfaced are fixed here: ironhub had no rows at all, completion was
  fence-only, and the Trace Commons table lacked the `ironclaw` prefix the
  rest of the page uses.
- crates/extensions/ironclaw_extension_registry/tests/
  docs_manifest_schema_version.rs: walks the published docs tree (the
  frozen .mintignore fence mirrored as constants) and asserts zero
  occurrences of the retired reborn.extension_manifest.v2 literal, fenced
  code included; asserts building-a-tool.md names
  MANIFEST_SCHEMA_VERSION_V3 verbatim and documents origin_gate_matrix.
- crates/product/ironclaw_openai_compat/tests/docs_responses_contract.rs:
  docs/api/responses.mdx now carries a machine-readable
  {/* doc-fact:responses-request-policy */} marker block (invisible when
  rendered); the test parses it and drives every claim through the same
  route-level seam as the sibling *_contract.rs suites — the marker's
  values parameterize the assertions (temperature accepted at the
  documented max and rejected just above it, model accepted at the byte
  cap and rejected past it, tool_choice always 400, tools 400 without /
  registered with external-tool wiring, empty tools treated as omitted,
  unknown fields like max_output_tokens accepted and ignored, and one
  request carrying every documented field accepted).

Part of #7317 (doc-truth pipeline, PR 3 of 5); stacked on #7376.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: address Copilot and CodeRabbit review on doc-drift PR

- responses.mdx: tool_choice is rejected only without external-tools wiring;
  with external tools enabled it passes validation and is currently ignored
  (validate_responses_supported_fields_with_external_tools never checks it).
- building-a-tool.md: clarify that effect-derived host ports are validation
  vocabulary against the HostPortCatalog allowlist; adapters are built by
  host-runtime services after authorization/obligations, never from manifests.
- how-to-port-tool-to-reborn.md: mark the decision tree's RuntimeKind targets
  historical (v3 accepts only wasm|first_party; MCP is top-level [mcp];
  process/CLI work is the sandbox lane).
- building-a-channel.mdx: document the user install flow — virtual package
  root /system/extensions/<id>/manifest.toml, ironclaw extension search /
  install <extension-id> (ID, not path), WebUI Extensions lifecycle.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(responses): align the limits bullet with the corrected tool_choice claim

The rejection list was corrected in the previous commit (tool_choice is
rejected only without external-tools wiring); the "Limits and quirks"
bullet still said "not supported ... rejected with 400". Same claim, one
wording.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(docs): tool_choice is conditionally rejected, not always

Copilot review on the docs PR caught that
validate_responses_supported_fields_with_external_tools never checks
tool_choice — with external tools wired it is accepted and ignored, not
400'd. The doc-fact marker moves tool_choice into
rejected_without_external_tools, and the dedicated test now proves both
sides: 400 naming the param on the plain router, accepted-and-ignored
(submit succeeds, nothing registers) with external-tool wiring.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: apply verified code-review findings on the drift PR

A full code review of this PR against live code surfaced claims the
original drift pass got wrong or missed; every fix below was re-verified
against the cited source before editing:

- responses.mdx: standard `ironclaw serve` deployments always wire
  external tools (OpenAiCompatRouteMountPorts requires the store/resume
  pair; mount.rs wires them unconditionally), so `tools` is accepted and
  `tool_choice` is accepted-and-ignored on shipped binaries — the
  conditional 400s apply only to custom compositions without the wiring
  (now a Note). temperature is validated and carried in the submitted turn
  payload but not applied as a provider sampling parameter. Non-streaming
  wait timeout is 30 s (DEFAULT_RESPONSES_WAIT_TIMEOUT), not 120. usage on
  retrieval is read best-effort from persisted run state incl. USD cost
  (read_run_usage), not always zero.
- building-a-tool.md: the [auth.example] oauth2_code recipe gains the
  required token_response map (deny_unknown_fields rejects the example as
  previously written); Gmail/Google Calendar corrected to first_party
  runtimes (their manifests declare kind = "first_party"); the worked
  api_key recipe is github's, not slack's; the tail "Quick implementation
  checklist" and reference list were still v2-era (script lane,
  assets/<extension>/ path, "manifest v2", v2.rs pointer) and now teach
  the v3 shape; composition/CLI package-naming claim narrowed (the binary
  does link slack/telegram adapter crates).
- contracts/extensions.md: legacy-format paragraph no longer claims
  host-bundled packages ship v2 (none do), and origin_gate_matrix is
  attributed to capability.rs + building-a-tool.md instead of
  extension-runtime/overview.md §3, which does not mention it.
- how-to-port banner: `script` manifest authoring is retired; the
  RuntimeKind::Script symbol survives as the process-sandbox lane's kind.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(contracts): repoint delivery_resolution.rs to its family directory

PR #7157 (merged to main 2026-08-07) cited
crates/ironclaw_outbound/src/delivery_resolution.rs in the
communication-delivery-resolution contract; the crate lives at
crates/domains/ironclaw_outbound/. Caught by this branch's docs surface of
check-guidance.py on the first merge of main after the gate landed —
exactly the drift class it exists for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(test-plan): route docs pages to the doc-fact tests that read them

docs/ sat in IGNORED_PREFIXES as a pure-prose class, which this PR's
doc-fact tests falsify: three cargo tests now read published pages, so
a docs-only PR would have selected zero crate tests and merged green,
leaving the failure to land on whichever unrelated change ran the full
plan next.

Published Markdown now selects the registry's schema-version sweep;
docs/using/cli.mdx and docs/api/responses.mdx additionally select
their owning crates. All selections are direct exact test targets —
no reverse-dependency widening, since prose only changes the doc-fact
assertions that read it. Fenced trees (docs/internal/, docs/reborn/,
drafts) and non-page files keep the prose classification.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): harden the docs gate and fix review-surfaced doc drift

Applies the verified findings from the PR #7376 code review:

- The loop-exit and turn-runner contract docs claimed the deleted
  loop_driver_host checkpoint-rejection test had 'moved into the
  module'; it was deleted in #6696 and the fenced verification command
  could not run. Both now cite the real surviving pins
  (planned_driver.rs executor test + the ironclaw_turns projection
  test mapped in scripts/reborn-e2e-rust.sh), with runnable commands.
- An unterminated comment now refuses at EOF like an unterminated
  fence; before, one typo'd closer silently un-scanned the rest of the
  file.
- Markdown links in the re-included corpora are now checked as repo
  paths (they are never published, so the Mintlify-route rationale did
  not apply); this alone added ~165 verified references.
- Each DOCS_REINCLUDED_PREFIXES entry must match at least one tracked
  page or discovery refuses, so the planned docs/reborn consolidation
  cannot silently drop the corpus from the scan.
- The living extension-runtime spec pages (overview.md,
  standard-operations.md) and guidance-conventions.md join the scan;
  guidance-conventions.md now describes the docs surface and the MDX
  marker form, and its one dangling test path is repointed.
- Floors comment corrected (57 rule globs, not 38).

Also fixes four drifted claims from #7375's pages, verified against
live code: the interleaved function_call_output example was rejected
with 400 (resume input must be exclusively function_call_output items
with previous_response_id); model is echoed only on create (GET/cancel
report the 'reborn' placeholder); output_schema_ref is optional; and
the unknown-fields claim now names the two deliberate exemptions.

Self-tests: 43 pass (three new arms — unterminated comment refusal in
both syntaxes, re-included links as repo claims, stale re-included
prefix refusal).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): sync module docstring with re-included link checking

CodeRabbit caught the docstring still claiming the link extractor is
off for all of docs/** — stale since b172f69c7 enabled it for the
re-included corpora. The docstring now states the exception and the
current re-include set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(docs): drop the retired reborn/ entry from the publication-fence mirrors

reborn/ left docs/.mintignore when #7559 consolidated it into internal/;
the fence mirrors in docs_manifest_schema_version.rs and
reborn_pr_test_plan.py still listed it. Fixture paths follow the move.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(docs): tighten doc-fact comments and docstrings

Same behavior; module docs and test docstrings trimmed to the point.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(docs): harden the doc-fact suites per CodeRabbit review

- CLI: validate full documented command paths via `ironclaw <path> --help`
  (immediately caught and removed the nonexistent `extension activate` row)
  and match visible aliases as exact tokens, not substrings.
- Responses: seed a real prior response so `previous_response_id` is
  actually submitted and accepted; document `metadata` in the visible table
  to match the marker.
- Manifest sweep: parse the publication fence from docs/.mintignore instead
  of mirroring it, so a removed fence entry widens the scan with it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(docs): correct the completion syntax and parse the fence in the planner

Review findings (sub-agent /code-review):
- docs/using/cli.mdx taught `ironclaw completion <shell>`; the binary only
  accepts `--shell <shell>`. The contract test stops extracting at flags,
  so it could not catch this.
- The planner's doc-fact arm mirrored the .mintignore fence as constants —
  the same hand-maintained-mirror class the PR removes elsewhere. It now
  parses docs/.mintignore via docs_publication_boundary, and a .mintignore
  edit itself routes to the published sweep.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(test-plan): treat a missing docs/.mintignore as no fence, not a crash

Matches docs_publication_boundary.find_violations(): fence gone means
everything is published, so every page routes to the sweep.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(docs): replace the doc-fact count floors with derived anchors

Same move as #7376's MIN_DOCS_FILES removal: MIN_DOC_COMMAND_ROWS was
redundant with the completeness check (the binary defines the expected
set), and MIN_SCANNED_PAGES is now a docs.json nav-coverage assertion —
every source-backed navigation route must be among the walked pages.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(docs): assert the current schema version instead of scanning for a retired literal

Hardcoding `reborn.extension_manifest.v2` was backward-looking: retiring
v3 would need a hand-edit or the test goes stale. The scan now extracts
every `reborn.extension_manifest.<version>` mention in published pages
and asserts it equals `MANIFEST_SCHEMA_VERSION_V3`, with the family
prefix derived from the same constant — the next schema bump retargets
the test by itself, and typo'd or older versions (v1, v33) are caught
too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 16:40:54 +00:00
Josh Ford
f75947032b release(docs): deploy public docs from a docs-live branch moved by stable releases (doc-truth PR 4/5) (#7379)
* release(docs): deploy public docs from a docs-live branch moved by stable releases

The Mintlify GitHub App deployed docs/ on every push to main, so the public
site described unreleased behavior while binaries shipped from ironclaw-v*
tags — the root cause behind #7317's drift reports. The site now tracks the
latest stable release:

- .github/workflows/ironclaw-release.yml: new publish-docs-live job (after
  host, prerelease-guarded via the dist manifest's
  announcement_is_prerelease) force-updates refs/heads/docs-live to the
  released commit through the GitHub refs API, bootstrapping the branch on
  first run. Forced by design: successive stable tags need not be
  ancestor-related, and docs-live is a pointer, not a history. The Mintlify
  dashboard repoint to docs-live is the one out-of-repo step, documented in
  the release strategy.
- scripts/ci/ws12_workflow_contracts.py: the job, the branch ref, and the
  prerelease guard join REQUIRED_MARKERS so a cargo-dist regeneration that
  drops the hand-added job fails Code Style instead of silently unhooking
  docs publication.
- docs/changelog.mdx (new, in en nav under a Releases group): human-curated
  <Update> entry per stable release, seeded with v1.1.0 and v1.0.0 from the
  GitHub release notes; links tagged docs trees for older releases instead
  of maintaining versioned page sets.
- scripts/ci/cut_ironclaw_release.py: ensure_stable_changelog_entry — a
  stable (non-rc) cut refuses when the candidate commit's changelog lacks
  the release's vX.Y.Z entry, with an actionable message; rc cuts are
  exempt so the freeze/blocker flow is unimpeded. Five new cases in
  test_cut_ironclaw_release.py (missing entry, missing file, present entry,
  rc exemption, malformed-version deferral to the canonical validator).
- docs/internal/weekly-release-strategy.md: Monday checklist writes the
  changelog entry on the release branch; promotion notes the automatic
  docs-live repoint; new "Docs publication" section records the dashboard
  configuration, post-promotion verification, branch-protection
  recommendation, emergency manual repoint, and older-release access.

End-to-end proof of the workflow job rides the next stable release; until
the Mintlify dashboard is repointed the site keeps deploying from main, so
the rollout order is: merge, repoint the dashboard, then the next stable
tag takes over.

Part of #7317 (doc-truth pipeline, PR 4 of 5).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* release(docs): harden the docs-live chain after review

Three gaps found reviewing the pipeline end to end:

- Changelog history loss: the runbook told the release owner to write
  the changelog entry on the frozen release branch, which is never
  merged back — next week's candidate, cut from main, would ship a
  changelog missing the release and docs-live would silently drop it.
  The entry now lands on main before the Monday cut; the gate's
  message and the runbook say so, with a cherry-pick fallback.
- Substring gate false-pass: the stable gate accepted any 'vX.Y.Z'
  substring, so an rc-labeled entry (description="vX.Y.Z-rc.1") or a
  prose mention satisfied it. It now requires the exact
  description="vX.Y.Z" attribute; both cases pinned in tests.
- Backwards repoint: re-running an older release's workflow would
  force-move docs-live to the older commit and silently revert the
  live site. publish-docs-live now moves the pointer only when its own
  tag is the newest stable ironclaw-v* tag, and the guard is pinned in
  ws12 REQUIRED_MARKERS so regeneration cannot drop it.

The runbook also gains the docs-hotfix recipe (publish tag+fix, never
main) and spells out that docs-live branch protection must allow force
pushes or it 422s the automation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* release(docs): match a real changelog Update tag and fail the repoint closed

Post-review hardening: the stable gate now requires the version on an
actual <Update> tag (lookalike attributes and other elements refuse, with
a regression test); the publish-docs-live step runs under set -euo
pipefail and rejects empty tag discovery instead of silently skipping;
ws12 pins the backward-repoint comparison itself; comments trimmed and
the public changelog no longer overclaims while the dashboard still
deploys from main.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* release(docs): close the review-round gaps in the docs-live chain

Seed the changelog with v1.2.0 (already-shipped stable whose branch never
merges back), put ironclaw-release.yml in code_style's has_code scope and
pin that semantically in ws12, bootstrap docs-live via POST only when the
branch truly does not exist so protection 422s surface as themselves,
drive the changelog gate through main() in tests, restore the emergency
paragraph to its own runbook section, and tell contributors docs reach
the live site with the next stable release.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 16:29:08 +00:00
firat.sertgoz
f13d0e3887 perf(events): coalesce runtime milestone writes (#7631)
* perf(events): coalesce runtime milestone writes

* style: format event write reduction

* fix(events): preserve durable lifecycle delivery

* test(composition): scope cancellation event guard

* style(events): format review fixes

* build: account for shared event sink seam

* fix(events): keep runtime emission non-blocking

* test(events): wire nonblocking milestone sink

* fix(events): address review feedback

* fix(events): require nonblocking sink capability

* test(events): wire nonblocking milestone sinks
2026-08-17 11:13:38 +00:00
Benjamin Kurrek
2b88e4309e feat(unbound-turns): complete the switchover to prepared-context turns (#7634)
* docs(internal): agent-execution seam proposal and system architecture

Two design documents for the product-neutral agent execution seam:

- 2026-08-12-agent-execution-seam.md — the AgentExecution port: the
  request contract (ExecutionContext::{Thread, Snapshot}), the
  AgentMessage interface, OutputContract, gates policy, two-plane
  events/observation, the runtime invariants (I1-I5) the design
  preserves, crate placement, open questions.
- 2026-08-12-agent-execution-architecture.md — the system-level picture:
  every surface (channels, WebUI, automations, suggestions, OpenAI-compat)
  submits through the one seam and interprets the output its own way;
  manifest-driven channel reply (stream vs send_reply); boundary rules;
  phased sequencing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): rename Thread context freeze-point field to accepted_message

The field is today's SubmitTurnRequest.accepted_message_ref; naming it
that way makes the 1:1 phase-2 mapping visible and documents why the
freeze-point exists (deterministic replay; late arrivals steer or queue).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): add 'why not materialize everywhere' rationale to seam doc

Pins the answer to the natural simplification (conversation workflow
assembles system_prompt/messages/tools itself and the engine stays
context-kind-agnostic): a conversation run's context is engine-mutated
for the run's whole life — steering, per-iteration skill/memory
re-selection, mid-run compaction write-back, resume rebuild, surface
re-versioning, prompt-bundle anti-forgery, refs-only storage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): pin request-knob x context validity matrix

Resolves an ambiguity Ben caught: 'tools=[]' meant 'no tools' in the seam
doc but 'profile surface' in the architecture example. Pinned semantics:
request knobs are per-invocation choices bounded by the profile's per-class
policy — Thread requires empty tools (surface is profile-derived, per
iteration) and AssistantMessage output; Snapshot names an exact subset of
the profile surface. Violations reject fail-closed at the seam.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): name the actual production default profile for Thread context

Conversations resolve None -> the planned default profile via the
resolver's implicit default; trigger fires are forced onto the
deny-mapped scheduled-trigger profile. The prior parenthetical named
contract-layer ids instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): OutputContract v1 mechanism = schema as forced tool

For JsonSchema contracts the host injects a synthetic host-owned result
tool whose parameters are the registered schema, riding the existing
strict tool-schema path every provider supports; reply admission
intercepts the call as terminal output. Not a capability (no
authorization/dispatch, like capability_info). Provider-native response
modes become later per-provider upgrades behind the same contract.
Prior art: pi's typed-output idiom.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): complete the open-questions inventory

Seam doc gains three implementation questions surfaced in planning
(process-kind encoding + rolling-compat test, ExecutionLimits mapping,
pi-style per-tool crash-replay declaration); architecture doc pins the
three phase-2 questions (origin-metadata home, delivery-observer
transition shim, pins inventory).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): resolve open questions per design review

Schema rides the request inline (registry deleted from the design;
journaled request keeps results interpretable). Strict-only validation.
New ProcessKind variant + legacy-kind audit via rolling-compat path.
ExecutionLimits maps onto existing budget machinery. Observation buffers
and message bounds reuse existing sizing/behavior. Per-tool crash-replay
declaration rejected — detached executions inherit standard recovery.
Open list now: suggestions tool need, detached concurrency cap value,
gate-resolve affordance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): sweep remaining schema-registry mentions

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): converge the seam onto main — threads are the unit of work

Final model after design review with the team: main already has the
agent-execution service (TurnCoordinator + process runtime + canonical
loop), threads are already the universal unit of work, and a conversation
is already a thread plus a binding. The seam is TurnCoordinator's front,
matured: two admission idioms (bound thread = today's submission;
content = mint-seed-submit of an unbound, ownerless thread with an
internal id), plus the genuinely-new pieces (OutputContract, per-run
knobs, subscription, detached profiles, codified taxonomy).

Replaces earlier design elements accordingly: no SnapshotBackedLoopContextPort
(one materialization path), no thread-kind flag (binding-absence +
ownerlessness classify, and owner-scoped listings already exclude), no
kernel scope changes, process journal role unchanged. Adds the what-main-has
section, the TurnCoordinator method mapping table, and the subagent
precedent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): update both diagrams to the converged model

The two mermaid diagrams still showed the pre-convergence picture. Both
now show mint-seed-submit behind the door (Snapshot → unbound, ownerless
thread → the same unchanged turn admission) and the single thread-backed
materialization path; the seam node is labeled as TurnCoordinator's
front, matured. Labels reflowed so renderers that strip <br/> tags keep
readable spacing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): classify the interface inventory

The seam doc defines ~19 type blocks, which reads as a large new surface;
classified honestly it is one trait + request/response DTOs. Adds an
inventory table (new API vs ironclaw_llm extension vs read-side views vs
referenced-unchanged) and labels the observation types as per-execution
views over the existing durable vocabularies and live-hint plane — no new
durable event language.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): drop the AgentExecution port — expand TurnCoordinator directly

Per design review: 'TurnCoordinator's front, matured' meant literally
expanding the coordinator, not naming a new port over it. The docs are
reworked accordingly and renamed (agent-execution-seam → detached-turns;
agent-execution-architecture → one-engine-many-surfaces):

- One new method, submit_detached_turn(SubmitDetachedTurnRequest), beside
  the unchanged submit_turn (precedent: submit_child_run). The request
  enum, AgentExecution trait, ExecutionId, ExecutionCaller, and the
  wrapper/response DTOs are all deleted from the design; the run id and
  SubmitTurnResponse are the handles. Knob misuse becomes unrepresentable
  (knobs exist only on the detached request), so the validity matrix and
  its fail-closed rejection path are gone too.
- Conversations now have NO migration at all — not even an entry-point
  re-plumb; they already call the trait being expanded. The old phase 2
  reduces to an independent SubmitTurnRequest-slimming follow-up (binding
  refs → workflow association state), explicitly hygiene-not-architecture.
- Rich subscribe moves off the kernel trait to a product-tier
  RunObservation façade (a kernel trait must not depend on read models);
  observation types renamed to Run* view vocabulary.
- Both diagrams, all flow pseudocode (submit_turn shown verbatim-as-today
  for conversations), the interface inventory, phases, non-goals, crate
  placement, ownership, and resolved-decisions updated coherently; boundary
  rule 3 corrected to match reality (engine stores binding refs opaquely
  today; slimming is the follow-up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): one submit_turn with optional bindings + one shared detached accept door

Final API shape per design review: the coordinator keeps exactly one
submission method — submit_turn — whose binding refs become Option,
making the request type express the taxonomy directly (thread required =
the unit of work; binding optional = what makes it a conversation).
Per-request data moves to the accept step where conversations already put
theirs: accept_detached_context (threads tier, ONE shared implementation)
mints the unbound, ownerless thread, seeds content, and journals the
declarations (tools/output/limits); every non-channel caller uses it —
suggestions, OpenAI-compat, and subagent spawn, whose hand-rolled
ensure_thread + accept_inbound_message + synthetic placeholder refs
retire (refactor lands as its own follow-up PR). submit_detached_turn is
deleted from the design; ResumeTurnRequest refs optionalize alongside;
model hint stays on submit as requested_model. Diagrams, flows, inventory,
placement, sequencing, and resolved decisions updated coherently; subagent
lane nuance footnoted in the semantics table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): layers-table wording — coordinator is not 'expanded'

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): fix two spec inconsistencies before implementation handoff

Gates: the blanket GateNotSupported policy contradicted the OpenAI-compat
external-tool flow (and our own ResumeTurnRequest note); the ExternalTool
gate is now an explicit exemption — its resolver is the submitting client,
not a human surface. Process-kind: the 'new ProcessKind variant' resolution
predates the final design and is superseded — detached turns are ordinary
AgentTurn processes, and rolling compat reframes onto None-refs +
detached-profile rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(unbound-turns): prepared-context accept door, unbound run lane, kernel binding-ref deletion (#7633)

* test(turns,threads): pin today's behavior before the detached-turns change

Tests only — every pin passes against unmodified code, so the commits that
follow cannot silently change the behavior they capture.

- prepare_turn reservations (new coordinator_prepared_run_contract): the
  prepared-id path had no coverage and no production caller — pin
  consume-once, the cross-scope Unauthorized rejection, the child-run
  exemption, abort_prepared_turn, and the 4096 capacity cap.
- Rolling-compat reader posture (new run_metadata_compat_contract): a
  durable agent_turn metadata row missing the binding-ref keys is rejected
  fail-closed by today's reader (the old-style-reader proof for the
  upcoming Option refs); unknown run-profile ids (e.g. a future
  detached_structured) still rehydrate; TurnRunProfile's lenient legacy
  deserialization gets its first direct pin.
- Taxonomy baseline (both threads backends): a thread whose scope carries
  no owner_user_id is structurally invisible to owner-scoped listings and
  vice versa — the exact-tuple scope filter the detached lane's
  unbound+ownerless threads rely on.

Existing pins verified green as the baseline (not modified): busy
admission DeferredBusy/RejectedBusy + duplicate replay
(inbound_turn_contract, product_surface_contract, steering.rs), submit
idempotency replay (idempotent_replay.rs, process journal contract —
which deliberately replays without payload comparison),
one-active-run-per-thread (generated_gate_sequences, cancel.rs),
cancel-time Queued→RejectedBusy reconciliation (steering_reconcile),
trigger trusted-path submission (triggered_submit.rs), subagent spawn
(subagent_spawn_port tests, subagent_await_edge.rs), lease reclaim
(lease_wedge.rs), model repair/retry (model_recovery.rs, tool_call.rs).

Two doc-vs-code contradictions surfaced by pinning, preserved as-is and
carried to the PR body: subagent child threads are owner-scoped and DO
surface in conversation listings today (spawn is production-disabled, so
this is latent), and submit idempotency replay ignores fresh invocation
identity rather than failing closed on payload mismatch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(unbound-turns): prepared-context accept door, unbound lane, and kernel binding-ref deletion

The kernel stops carrying reply routing. SubmitTurnRequest, ResumeTurnRequest,
RetryTurnRequest, SubmitChildRunRequest, TurnRunState, TurnRunRecord, agent-turn
metadata, and SubmitTurnResponse::Accepted lose source_binding_ref /
reply_target_binding_ref entirely; routing lives in product-side conversation
state. Frozen legacy-shape compat test proves old readers fail closed on new
rows while new readers rehydrate old rows (serde ignores unknown keys).

New primitives:
- ironclaw_host_api::prepared_context — PreparedTurnDeclarations, OutputContract,
  TurnLimits, PreparedContextSource, structured-result capability ids.
- ironclaw_llm::agent_message — provider-neutral AgentMessage vocabulary with
  bounded validation and total ChatMessage conversions.
- ironclaw_threads accept door — accept_prepared_context / read_prepared_context
  on SessionThreadService (in-memory + filesystem): deterministic unbound-thread
  minting, seeded rows via CAS, journaled PreparedContextRecord as commit marker,
  idempotent replay by key.
- Coordinator prepared-context probe — unbound profile derivation +
  declared-limits narrowing at admission; unbound concurrency class.
- Agent-loop unbound families (gate-not-supported, structured-output reply
  admission, structured-result stop) + loop_host structured_result capability
  with strict JSON-schema validation.

Subagent spawn lands on the shared accept door: synthetic per-child binding
refs, mark_message_submitted step, and AwaitEdge ref plumbing are deleted;
child submits reference the accepted seed message.

Product-side rerouting:
- run-delivery observer routes notifications from the conversation binding it
  already resolves (runs carry no reply route).
- model-channel same-origin check asks the durable conversation-binding store
  which thread a sealed reply-target ref is bound to
  (resolve_stored_reply_target) instead of reading kernel run state; the
  late-bound trigger-source turn-state slot this replaced is deleted.
- approval/auth/blocked-auth/webui gate resumes stop minting synthetic refs.
- ProcessGateRecord no longer surfaces resume/reply refs; legacy journal
  migration stops copying them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(unbound-turns): production wiring — profiles, drivers, surfaces, caps

- Register the unbound-default / unbound-structured planned drivers and run
  profiles (steering disabled; the structured profile allows no-reply
  completion — its terminal output is the validated result row); both loop
  families join the production family registry, with re-pinned BLAKE3 digests.
- Wire ThreadServicePreparedContextSource into both production coordinator
  builds so admission derives unbound profiles from the journaled record.
- Capability surface: the `unbound_tools` deny-map strips subagent spawn and
  the trigger mutators; a prepared context that declares its tools narrows the
  surface to exactly that allowlist (read fails closed at host build).
- Unbound structured runs get the synthetic builtin.structured_result tool
  built from the journaled output schema at capability-port assembly.
- Unbound runs skip the skill/identity/memory context lanes and the
  after-turn memory recorder: the prepared context is the complete input, and
  the exchange is caller data, not a user observation.
- New `unbound` concurrency class cap plumbed through RunnerSection →
  IRONCLAW_REBORN_RUNNER_MAX_CONCURRENT_UNBOUND_RUNS → TurnRunnerSettings →
  process concurrency limits (default 4), documented in .env.example.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* test(unbound-turns): whole-path integration coverage + docs to the shipped end state

- New integration bin `reborn_integration_unbound_turns`: accept →
  refless submit → derived profile → planned loop → terminal state, pinning
  the structured-result happy path, the invalid-then-valid repair loop, the
  plain-final default contract, accept/submit idempotent replay, and the
  ownerless-listing exclusion; registered in Cargo.toml + tests/CLAUDE.md.
- Fix surfaced by the whole-path run: reconstructing a `__system__`-slot
  process scope now yields `TurnThreadOwner::Ownerless` (not actor-fallback),
  so an unbound run's thread reads stay on the slot the accept door wrote —
  actor-fallback reconstruction re-pointed them at `owners/<actor>` and the
  loop failed with `host_stage_unavailable_prompt`.
- Group harness: `builtin_tools_with_durable_capability_io()` ctor so the
  capability port reads the SAME thread store the runtime uses (production
  parity for the structured-result declarations read).
- Docs: design drafts renamed to the unbound/prepared-context vocabulary with
  an explicit implementation-delta section (refs deleted rather than
  optionalized; probe-derived profiles; spawn on the shared door), and the
  loop-exit contract's stale MVP gate list amended.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* chore(unbound-turns): gate-pass fixes — clippy matches!, host_api ceiling raise, guidance rows

- `part_allowed` rewritten as `matches!` (clippy, all-features gate).
- ironclaw_host_api size ceiling 19_026 -> 19_274: the `prepared_context`
  contract module (declarations/output-contract/limits vocabulary, the
  admission-probe trait, and the structured-result capability ids) — neutral
  authority vocabulary only; behavior stays in threads/loop_host/turns.
- ironclaw_assistant AGENTS.md module tables drop the deleted webui binding
  helpers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(unbound-turns): clear the production panic baseline

- `chat_messages_from_agent_messages` re-raises the typed
  `ToolMessageResultCount` / `UnpairedToolResult` errors for its
  post-validation lookups instead of `.expect()` — the conversion stays
  total with no panic path.
- The structured-result capability-id constant's `.expect()` carries its
  inline `// safety:` rationale on the invocation line, where the baseline
  scanner reads it (compile-time constant, pinned by a test).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(ci): repoint changed-coverage exemptions displaced by the ref deletion

Two exemption entries named lines beyond their files' new EOF after the
binding-ref deletion shortened them: the await-edge type-repoint exemption
(mod.rs 112/149 -> 105/142) and the steering-allowed serde-default exemption
(metadata.rs 139-141 -> 130-132). Same exempted code, current positions;
`reborn_changed_coverage.py --validate-manifest-only` passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* feat(unbound-turns): seed tool/reasoning history through the accept door

The prepared-context accept door now seeds caller-supplied tool history as
first-class transcript rows instead of rejecting it:

- An assistant tool-call turn follows the live storage shape (no persisted
  assistant tool-call row): each tool result seeds one ToolResultReference
  row whose ProviderToolCallReferenceEnvelope carries the host-owned
  `prepared-context-seed` sentinel identity, the paired call's capability /
  arguments / provider tool name, and the turn's reasoning display text as
  `response_reasoning` — with `signature` forced None so caller-authored
  history can never smuggle provider replay artifacts.
- Full outcome bytes land in the durable tool-result record store under
  deterministic `result:seed.{sha256}` refs (crash-retry convergent), so
  builtin.result_read pages seeded results exactly like live ones; the row
  envelope carries a bounded preview observation.
- The model gateway's replay-identity gate carves out the exact sentinel so
  seeded rounds replay as faithful tool_use/tool_result exchanges on any
  route; every other mismatched identity still degrades to the summary-style
  user message.
- Validation replaces the blanket rejection: provider-token grammar on call
  ids, provider-safe tool-name derivation, and reasoning only on tool-call
  turns (the one storage slot); all checked before any state is minted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(unbound-turns): close the completeness-audit gaps

Findings from the adversarial design-vs-code audit of the unbound lane:

- The stale rename-era assertion in the structured reply-admission test now
  pins the REAL control text (`builtin__structured_result`); it had been
  asserting a string the hint never contained.
- The `ironclaw_threads -> ironclaw_llm` same-layer edge is inventoried
  (with its mirror-DTO-ban rationale) and SAME_LAYER_EDGE_BASELINE moves
  71 -> 72; the manifest edge now pins `default-features = false` like the
  other policed edges.
- `gate_not_supported` failures now carry the aborting gate kind as the
  sanitized failure detail, making the loop-exit contract's "the gate kind
  rides the sanitized failure detail" true in live code; other gate-abort
  classifications keep their pinned bare-category shape.
- Design docs reconciled to the shipped end state: the companion
  one-engine doc's Option-refs phase notes are annotated as
  landed-stronger-than-drafted, and the unbound-turns §4.2 continuation
  claim now states the truth (a fresh key mints a fresh thread; in-place
  appending is a follow-up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(openai-compat): chat completions adopt the prepared-context door

The requests today's conversation path silently mistreats now go through
the engine's shared accept door:

- `response_format` json_schema / json_object maps onto
  `OutputContract::JsonSchema` (strict; json_object gets the permissive
  object schema) and derives an unbound_structured run; the completion
  returns the VALIDATED result payload read in full from the durable
  tool-result record store.
- Assistant/tool history (without declared live tools) seeds faithfully as
  transcript rows through `accept_prepared_context` on an ownerless unbound
  thread (public completion id == thread id) instead of being flattened
  into one JSON string; client tool names map onto the same
  `external_tool.{name}` capability identity the live client-tool lane
  registers.
- The lane is decided in the route crate (`prepared_turn.rs`) and the
  implementation lives behind a composition-owned port
  (`OpenAiCompatPreparedTurnGateway`): accept + refless submit synthesize
  the same `ProductInboundAck::Accepted` shape, so ref-store idempotency,
  replay, and conflict handling are shared byte-for-byte; both lanes
  validate BEFORE the reservation so an invalid body never burns the
  caller's idempotency key.
- Read-back: prepared runs resolve from run state + the unbound thread
  (structured payload or the finalized assistant row); the conversation
  timeline reader is untouched for the conversation lane.
- Declared client tools keep the conversation lane (the external-tool
  park/resume flow lives there); tools + response_format and streaming +
  json output are rejected loudly instead of half-honored — streaming
  previously dropped the schema silently.

Coverage: 7 lane unit tests, 5 workflow contract tests (routing, fail-closed
501 without the port, key-preservation), and 2 composition gateway tests
(full-history seeding + refless ownerless submit shape; structured payload
resolution through the record store).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(loop): structured runs force the result tool on repair retries

`LoopModelToolChoice::ForcedCapability` rides `LoopModelRequest` /
`HostManagedModelRequest` from strategy to provider. The
unbound_structured family's new `StructuredResultModelStrategy` emits it
only while a structured-output rejection is pending — first attempts stay
on provider "auto" so ordinary tools remain usable; the repair retry after
a rejected plain-text final must call `builtin.structured_result`
(family digest re-pinned).

The gateway resolves the forced capability against the visible tool
surface (the same definitions it declares to the provider) and fails
closed — InvalidRequest for a capability off the surface or a forced
choice on a text-only call; no silent dispatch.

Provider encodings fixed in the same pass so a named tool_choice is
honored instead of silently downgraded:
- nearai_chat: named choice now serializes as the OpenAI object form
  (a bare tool-name string is rejected by OpenAI-compatible servers).
- openai_codex + codex_chatgpt (Responses API): named → object form with
  the sanitized provider-facing name; codex previously hardcoded "auto".
- bedrock: named → `SpecificToolChoice` (was: silently "auto").
- gemini_oauth: named → mode ANY + `allowedFunctionNames` (was: AUTO).
- rig adapter: named → `ToolChoice::Specific`; rig providers without
  specific-tool support reject loudly (was: silently dropped).

Pinned by strategy tests (repair-retry-only forcing), two gateway tests
(provider receives the resolved provider tool name; off-surface forcing
rejected before dispatch), and provider serialization tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(loop): enforce per-run budgets — wall clock, model calls, capability invocations

The three `ResourceBudgetPolicy` knobs stop being decoration:

- `max_wall_clock_seconds` (new, serde-default None) joins the policy and
  `TurnLimits`; declared limits narrow the profile ceiling at admission
  (min, never widen — pinned by a coordinator resolver test). The budget
  stage arms `run_started_at` on its first pass (initial state stays
  deterministic; old checkpoints re-arm from resume time) and hard-stops
  past the limit: cancel-check, explained failure, final checkpoint,
  `wall_clock_limit` exit. The dormant `BudgetStrategy::wall_clock_limit`
  seam now participates via min with the policy ceiling.
- `max_model_calls` / `max_capability_invocations` were defined-but-inert
  (the audit's named gap): the state now counts every dispatched provider
  attempt and every batched invocation, and the budget stage hard-stops
  with `model_call_limit` / `capability_invocation_limit`.
- All three failure kinds (plus `gate_not_supported`, which PR #7633 never
  registered) are wired through `ALL_RUN_FAILURE_CATEGORIES`, the
  user-facing failure summaries, and the text-loop driver matcher — the
  driver previously collapsed `gate_not_supported` to `driver_bug`.

Also in this change, three defects the docs-branch merge surfaced (all
red on the base branch, none exercised by CI's lanes):

- Ownerless scopes lost their disposition through the process journal:
  the `__system__` owner slot holds both ownerless and actor-fallback
  runs, so reconstruction guessed. Process metadata now journals an
  `ownerless_thread` marker (absent = legacy actor-fallback) and the
  snapshot reader honors it; both directions pinned.
- `production_registry_binds_default_and_subagent_families` still
  asserted the pre-unbound family count.
- Composition mass gate: the OpenAI-compat prepared-lane behavior moved
  out of composition into `ironclaw_assistant::UnboundPreparedTurnService`
  (charter: composition assembles, behavior lives in an owning crate);
  composition keeps only the wire-DTO mapping, the port adapter, and
  wiring. `ironclaw_threads` re-exports the `agent_message` seed
  vocabulary for accept-door callers (same precedent as `AttachmentRef`).
  `loc_ceiling` re-measured to the merged tree with rationale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(unbound-turns): close the completeness-audit gaps

Findings from the adversarial design-vs-code audit of the unbound lane:

- The stale rename-era assertion in the structured reply-admission test now
  pins the REAL control text (`builtin__structured_result`); it had been
  asserting a string the hint never contained.
- The `ironclaw_threads -> ironclaw_llm` same-layer edge is inventoried
  (with its mirror-DTO-ban rationale) and SAME_LAYER_EDGE_BASELINE moves
  71 -> 72; the manifest edge now pins `default-features = false` like the
  other policed edges.
- `gate_not_supported` failures now carry the aborting gate kind as the
  sanitized failure detail, making the loop-exit contract's "the gate kind
  rides the sanitized failure detail" true in live code; other gate-abort
  classifications keep their pinned bare-category shape.
- Design docs reconciled to the shipped end state: the companion
  one-engine doc's Option-refs phase notes are annotated as
  landed-stronger-than-drafted, and the unbound-turns §4.2 continuation
  claim now states the truth (a fresh key mints a fresh thread; in-place
  appending is a follow-up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(unbound-turns): repair the three red lanes on the merged base

- `production_registry_binds_default_and_subagent_families` still asserted
  the pre-unbound family count; the registry binds four families.
- Ownerless scopes lost their disposition through the process journal: the
  `__system__` owner slot holds both ownerless and actor-fallback runs, so
  reconstruction guessed. Process metadata now journals an
  `ownerless_thread` marker (absent = legacy actor-fallback) and the
  snapshot reader honors it; both directions pinned.
- Cherry-picked completeness-audit fixes: the structured reply-admission
  control-text assertion pins `builtin__structured_result`, and the
  threads→llm same-layer edge is inventoried with the baseline re-summed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* test(arch): allow the prepared-context seed module at the replay boundary

Seeded tool history persists provider-shaped tool-call reference envelopes
under the seed sentinel identity — the same replay boundary the model
gateway and trace capture already occupy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(threads): prepared-context threads stay out of owner-scoped listings

Every thread the accept door mints is stamped with a `prepared_context`
metadata marker (caller metadata must be a JSON object so the stamp can
coexist — the subagent path keeps its crash-reconstruction fields). Both
backends exclude stamped threads from `list_threads_for_scope`, plus the
pre-marker subagent metadata shape (`"kind":"subagent"`) so pre-existing
child threads hide too. The filesystem page cursor now advances over the
FETCHED index rows, so a page consisting entirely of hidden threads makes
progress instead of stalling — pinned by a paging walk in the filesystem
contract suite and a same-scope exclusion pin in the shared suite.

Also folds the gateway's forced tool choice into `ProviderRequestContext`
(clippy: too_many_arguments on `complete_model_request`).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(unbound-turns): mint-helper dedup, seeded-history whole-path scenario, replay-gate fix

- `InMemorySessionThreadService` mints threads through ONE shared
  `mint_thread_locked` for both creating doors (`ensure_thread` and
  `accept_prepared_context`) so the stored shape cannot drift; the
  replay-branch scope-check order is untouched.
- New whole-path integration scenario: seeded assistant tool-call + tool
  result history flows through accept → refless submit → the production
  context build → completion, with the round persisted as a
  `ToolResultReference` row whose full outcome bytes page back out of the
  durable record store.
- The scenario caught a real gap: `validate_provider_replay_identity`
  (the strict gate after the match gate) still required route equality
  for seeded sentinel envelopes, so ANY whole-path seeded replay failed
  with `driver_failed`. The sentinel carve-out now covers both gates.
- Documented the deliberate listing blind spot in the stale-attachment
  sweep (prepared threads carry no attachments today).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(gates): register the budget failure kinds end to end + re-sum ratchets

The Tier-2 failure-summary source-parity pin scans `LoopFailureKind::as_str`
from source: the four kinds this PR introduces or repairs
(`gate_not_supported`, `wall_clock_limit`, `model_call_limit`,
`capability_invocation_limit`) now carry their user-facing summaries in the
assistant table too. Struct-debt ratchet re-summed after main's merge
deleted composition factory/runtime debt (shrink-only baseline + WS0
member baseline 274 -> 270 per its own doc rule).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(gates): register gate_not_supported end to end + re-sum struct ratchets

`gate_not_supported` (PR #7633's typed abort) was never registered with the
run-failure vocabulary: the Tier-2 summary source-parity pin, the canonical
category list, the user-facing summary table, and the text-loop driver's
matcher (which collapsed it to `driver_bug`). All four now carry it.
Struct-debt ratchet re-summed after the main merge deleted composition
factory/runtime debt (shrink-only baseline + WS0 member baseline 274 -> 270
per its own doc rule).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(unbound-turns): close the design-conformance audit — nothing deferred

A 71-clause audit of both design docs against this tree (55 implemented as
drafted); every divergence is now closed in code or recorded as a
landed-vs-drafted delta in the docs' implementation-delta sections:

- **Lane extension (flatten retirement)**: replayed history of ANY shape —
  assistant/tool rows or multiple user turns — takes the prepared lane on
  non-streaming requests without declared tools; the JSON-flatten hack no
  longer serves stateless-replay clients. The message-count bound now runs
  before lane dispatch on both lanes (bounds before dispatch).
- **Run evidence on completions**: the prepared read-back returns the
  effective model (resolved model route) and provider-reported usage
  (incl. cached tokens) from run state instead of echoing the request.
- **Dead terminal-shape validator deleted**: `validate_terminal_output_message`
  had zero production callers — loop reply-finalization owns the guarantee;
  recorded in the doc deltas rather than kept as dead vocabulary.
- **Doc deltas appended** (unbound-turns + one-engine): thread-id-on-the-wire
  decision, seeded-reasoning semantics, AttachmentRef reuse, result-surface
  shape, structured-interception site + repair-retry forcing, TurnLimits
  scope (no USD/output-token seams), gate posture (fixed deny list, typed
  abort, engine-only external-tool exemption until a surface exposes client
  tools on unbound runs), OpenAI-compat lane split incl. streaming staging.
- Composition mass ceiling re-ratcheted to the observed count (42_093) with
  the WS0 record moved in the same commit, per the armed-ratchet pin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(openai-compat): mirror every deterministic door rejection before the key reservation

Review triage (IronLoop): a prepared body the accept door would refuse
deterministically could reserve the idempotency key and then fail, leaving
the key burned for a corrected retry. `prepared_pre_validate` now mirrors
ALL of the door's deterministic rejections, not just the grammar subset:
the 128-seeded-message budget, the 64 KiB text-part and system-prompt
budgets, the 256 KiB request budget, the 64 KiB tool-argument budget
(measured on the same serialized shape the mapping produces), and
tool-call/result pairing in order. The mirrored constants are pinned equal
to `ironclaw_llm::agent_message`'s by test (via the `ironclaw_threads`
re-export), and a workflow contract test proves a 129-message request
fails with 400 before any reservation while a corrected retry on the SAME
idempotency key succeeds.

The review's other finding — the strict replay-identity validator
rejecting seeded sentinel envelopes — was already fixed on this branch
(both gates carve out the sentinel), caught independently by the
whole-path seeded-history scenario.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(unbound-turns): review-triage batch — false-complete stop bug, delete purge, door bounds, seam pins

Triage of the 72 bot-review threads on the base PR, verified against both
trees; the live findings land here:

- **False-complete on failed result attempts (critical)**: the capability
  ERROR path recorded failed calls into `observed_signatures` (write-only
  vestige on main; PR #7633's structured stop strategy became its first
  consumer), so a structured run whose result attempt failed schema
  validation "completed" with NO durable result, and the all-failed abort
  could never trip. Error results are no longer recorded as observed;
  pinned at the strategy tier and by the repair integration scenario,
  which now reads the durable record back and proves the CORRECTED payload
  was stored and the rejected one was not.
- **Delete-then-re-accept replay**: the in-memory backend orphaned the
  prepared-context record on `delete_thread`, replaying a deleted thread;
  purged with the thread, pinned on both backends.
- **Door bounds**: attachment metadata gains byte caps (id/mime/filename/
  storage_key) and `extracted_text` respects the text-part budget.
- **Seam pins**: filesystem concurrent duplicate accepts converge with
  exactly one non-replay winner; a crash before the commit marker retries
  to a fresh accept with no duplicated rows; service-level replay guards
  (key/actor mismatch); admission-class absence default; model-delivery
  same-origin authority-denial branches (same-thread stays denied,
  other-thread proceeds); runner config coverage for
  `max_concurrent_unbound_runs`; listing test gains a positive anchor.
- Diagnostics: "prepared context" naming in door messages, traced (not
  swallowed) admission read errors, redaction-clean stable summaries for
  declarations reads, schema-compile message typo.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(cli): drop a stray unbound-cap assert from the conversation env test

The scripted test edit matched an assertion block that appears in two
tests and inserted the nonzero-round-trip expectation (7) into the
conversation env-override test, where the unbound cap legitimately holds
its default (4). The round-trip test keeps the real coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(ci): re-ratchet the composition mass ceiling after main's #7533 shrink

The main merge deleted composition code (batch-strategy plumbing retired),
leaving 213 LOC of slack — over the 200-LOC nudge window the gate's
self-test enforces. Ceiling, observed record, and the WS0 baseline move to
the measured 42_086 together.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* refactor(unbound-turns): design-audit fixes — one validator, one encoder, subtractive cleanups

- Collapse the mirrored OpenAI-compat pre-validation onto the accept door's
  own validate_prepared_seed_content: the route crate maps wire messages to
  the engine vocabulary once and calls the ONE validator pre-reservation,
  reaching it through ironclaw_assistant's re-exports so the boundary pin
  keeping route facades off domain crates stays intact
- Lane rule final form: every non-streaming request without declared client
  tools takes the prepared door (heuristic deleted); schema+tools and
  schema+stream rejected loudly
- Rename UnboundPreparedTurnService → UnboundTurnService; submission carries
  tools + idempotency key; Internal{reason} with cause-preserving errors
- ProviderToolName::for_capability/encode_capability_str own the `.`→`__`
  mapping; tool_search consumes the owner (allowlisted as protocol boundary)
- thread_id required on the accept door; derived-id helpers deleted
- prepared_context listing marker single-spelling + one-time filesystem
  backfill stamping legacy "kind":"subagent" rows, contract-pinned incl.
  crash-before-marker retry and concurrent-accept convergence
- mint_thread_locked shared by both thread-creating doors; delete purges
  prepared context state
- declaration-read/admission failures trace at debug with stable summaries
- docs: implementation deltas (who-mints ordering, concurrency default,
  lane rule, one-validator, derived-id deletion)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* fix(arch): honest ironclaw_threads edge for the compat prepared lane

The one-validator collapse gave ironclaw_openai_compat its threads-tier
vocabulary through an ironclaw_assistant re-export, which trips the WS5
facade-dissolution pin (product_declares_no_foreign_re_export_facade):
assistant's root re-exports only what assistant declares, and a consumer
imports every other name from the crate that owns it.

Do exactly that: drop the re-export and give the route crate a direct
ironclaw_threads dependency for the accept door's seed vocabulary and
validate_prepared_seed_content — the same edge ironclaw_webui already
holds. The BoundaryRule gains a documented carve-in (thread/turn SERVICES
still arrive as composition-built ports; the edge is vocabulary + the one
door validator), and the crate README/AGENTS boundary prose is trued up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ViVKgEfYb6YCNUUMDoGE4t

* feat(unbound-turns): thread the authenticated caller as unbound thread owner

The product unbound lane (UnboundTurnService) minted Ownerless threads,
landing every deployment's completions in the tenant __system__ storage
slot with no per-user boundary, no owner-scoped deletion path, and no
scope enforcement on run-state reads. The lane now names its
authenticated caller as the thread owner: threads shard per-user like
every other surface, and get_run_state rejects foreign-owner reads.
wait_for_completion gains the actor so read-back resolves the same
owner scope the accept seeded (a mismatch is a hard ScopeNotFound).

Listing behavior is unchanged: prepared threads stay hidden via the
unconditional prepared_context metadata stamp, not via ownerlessness.
The engine-level Ownerless variant remains supported; retiring it (and
ActorFallback) is a follow-up with its own journal-compatibility and
per-owner-concurrency review.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(unbound-turns): record the caller-owned thread delta across guidance

The design doc's implementation-delta list gains the caller-owned entry
(product lane names its authenticated caller as thread owner; Ownerless
stays engine-supported pending the journal-compat and concurrency-cap
follow-up), and the now-false "isolation rests on ownerless scoping" /
"no per-user identity axis" landed claims are corrected. The
PreparedContextRequest scope doc no longer instructs product callers to
pass None; the openai-compat and integration-suite doc comments and the
tests/CLAUDE.md scenario row now describe owner-scoped storage with
stamp-based listing invisibility.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(composition): wire the prepared-turn port into the chat mount test

The lane rule (every non-streaming request without declared client
tools takes the prepared-context door) landed without updating the
composition mount test, which constructs OpenAiChatCompletionsWorkflow
directly and never wired a prepared_turn_port — so the authenticated
request 501'd where the test asserted 200. This was red on CI before
the owner-threading commits (identical failure on the prior branch
head); production is unaffected because build_openai_compat_route_mount
always supplies the real gateway.

The test now mirrors the production shape with a counting port double
and pins the lane rule at the mount tier: the tool-less request drives
one prepared accept and zero conversation-surface submits, with the
auth gate assertions unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(loop): enforce per-run budgets at dispatch via a typed BudgetLedger

Review of the budget-enforcement commit found three binding gaps: a
model recovery-retry burst could dispatch up to max_total_model_attempts
provider calls past max_model_calls (the check ran only at iteration
boundaries), a capability batch was admitted and charged whole even when
it exceeded the remaining allowance (unbounded — no per-turn batch cap
exists), and same-call capability retries re-dispatched without counting
at all. A fourth: retry runs inherited the failed run's exhausted
counters because rebase_for_run never reset the per-run fields.

Rather than patching four sites, spending is now a chokepoint: the
per-run counters move into a BudgetLedger on the loop state (private
fields, typed try_charge verdicts), so no dispatch — first attempt,
recovery retry, or batch member — can bypass accounting. Exhausted
verdicts map onto existing exits only: a model retry converts to the
outer-loop re-entry where the budget stage's ModelCallLimit hard stop
fires; an over-budget batch tail gets paired blocked results through the
denied-calls machinery and the run hard-stops at the next boundary; an
exhausted capability retry is not re-dispatched. rebase_for_run resets
via Ledger::fresh_for_run for a different run and preserves state for
same-run gate resumes.

The checkpoint wire shape is frozen: serde(flatten) keeps the three
fields as top-level keys with their original attrs, pinned by
literal-JSON shape tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(unbound-turns): harden the prepared lane per review findings

Four review findings on the prepared lane, each verified before fixing:

- Output schemas are bounded at the accept door: 64 KiB serialized size
  and depth 32 (iterative walk, no recursion over untrusted input),
  enforced in the threads-tier validator and called by the route BEFORE
  the idempotency reservation — the same one-validator pattern as the
  seed bounds, no mirrored constants. Previously a caller could submit a
  ~14 MiB schema that jsonschema compiled per run.
- The external-tool provider-name mapping is unforked: ProviderToolName::
  for_capability gains the external_tool. carve-out (bare suffix, matching
  what the live lane always emitted) and the live lane derives through it,
  so seeded replay history and declared tool definitions agree by
  construction. Fixed pre-release, so no wrong rows ever persist.
- The capability-id validation error keeps its cause (the map_err ban),
  and prepared chat usage now counts cache-creation tokens in prompt and
  total through the same conversion the Responses surface uses.
- The streaming schema-rejection test drives the real mounted streaming
  route instead of the non-streaming entry that rejects stream=true
  first; the old test is kept renamed since it pins a distinct guard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(unbound-turns): record charge-at-dispatch budget enforcement

The TurnLimits delta now states the enforcement mechanism precisely: a
typed BudgetLedger chokepoint charges every dispatch, with exhausted
verdicts mapped onto the existing exits — replacing the draft's
unqualified "hard stops" wording that the review showed was inaccurate
at iteration-boundary granularity.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(contracts): size interactive budget ceilings above the iteration backstop

CI caught the enforcement landing on real workloads: the interactive
tier's 32-model-call / 64-invocation ceilings predate enforcement (they
were written while the caps were inert) and sit far below the loop's own
1,024-iteration backstop, so a legitimate long tool run — the window-
compaction integration scenario drives 132 model calls and 130 tool
invocations in one turn, and the provider-contract E2E workflows exceed
64 invocations routinely — hard-failed with model_call_limit /
resource-denied tool results. The ceilings are runaway insurance BEHIND
the iteration backstop, so they now sit above it: model calls at 2x the
backstop (per-iteration recovery retries), invocations at 4x (parallel
dispatch width).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(loop): move BudgetLedger test seams into a test-gated module

The struct-debt ratchet (affected-3) correctly rejected the four
cfg(test) members on the production BudgetLedger impl; they live in a
test-gated module now, exactly as the ratchet's guidance prescribes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(loop): box every capability-port decorator delegation

The stacked LoopCapabilityPort decorator chain (~10 wrappers, each
awaiting self.inner inline) compiled into a single oversized poll frame
that overflowed default 2 MiB test-thread stacks and left production's
8 MiB stack (serve.rs) with unproven margin for deeper lanes. Every
port delegation is now Box::pin'd — the same treatment
executor/capabilities.rs already applies to its recursion sites — so
adding a decorator no longer grows one giant frame. Verified by the
previously-overflowing reborn_integration_model_recovery suite passing
at the default stack; a residual, shallower pressure point remains on
the kernel sandbox dispatch path (tracked as follow-up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(unbound-turns): address the automated review wave across the prepared lane

Twenty verified findings, grouped:

Correctness: a standalone tool_choice with no declared tools now 400s
before the idempotency reservation instead of silently dropping the
constraint; the prepared-turn port is a required dependency end to end
(production always wired it — the optional field's dead 501 branch could
have failed plain chat on a mis-wire); the structured-result paging loop
rejects non-advancing offsets and bounds total payload bytes; the tool
disclosure matcher also resolves the generic external_tool__ encoding; a
text content part without a text field is a typed 400, not an empty
string; the ownerless-slot and backfill tests now assert their exact
expected outcomes instead of passing on any error.

Diagnostics and hygiene: the three unbound-turn catch-all error arms
trace their bound cause; the submit idempotency-key error names itself;
silent fallbacks carry silent-ok markers; the runner env guard clears
the fifth (unbound-cap) variable; a redundant test reassignment and a
misleading re-export rationale are gone; the conversation-lane test
helper is named for what it does.

Docs and pins: both design docs carry the caller-owned delta and
supersede pointers at every stale draft claim; the openai-compat
AGENTS.md documents the threads carve-in and the prepared-lane rule;
composition LOC pins re-measured coherently at 41991; the test scenario
map recounts its bins and covers db_write_canonical.

Deferred with issues: typed ToolChoice across providers (#7672), ledger
truncated-launch reconciliation + charge durability (#7673),
symbol-level allowlist for the threads edge (#7674).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Henry Park <henrypark133@gmail.com>
2026-08-15 04:47:11 +00:00
Benjamin Kurrek
57c491af36 feat(telegram): pair linked devices with the bot channel (#7464)
* docs(design): Telegram linked-device — proposal, plan, checklist, ADR

Design-only. Adds the engineering spec for linking a user's personal Telegram
account as a real MTProto linked device, so the agent can read their
conversations and act as them through the standard messaging operations.

Docs only — no production code, no behavior change.

Shape:
- README: overview, architecture, footprint, explicit non-goals
- PROPOSAL: decisions with rationale, per-crate change inventory, review log
- PLAN: 8 PRs with dependency edges and per-PR watch-lists
- CHECKLIST: definition of done, each box naming something to run or read
- ADR: the auth-hook decision and what it costs

Load-bearing decisions:
- Reads are live; no message content is persisted. Telegram is a cloud
  messenger, so history and search are server-side — which removes the mirror,
  the retention policy, the FTS plane, and (because no update stream is
  consumed) the session-sourced ingress work from v1.
- Device-link is an auth method with a narrow adapter hook, taking the
  extension-runtime spec's own "a vendor defeats the descriptor" revisit
  trigger. The hook revokes a stated security invariant; the ADR records the
  real compensation set and the in-process-vs-sidecar trade.
- Custody extends ironclaw_auth (a linked account is a CredentialAccount); the
  only genuinely new persistence surface is a CAS write path for a mutable
  binary secret.
- Sessions live in the existing telegram package behind a contracts-declared
  port; no new crates, no new runtime lane.

Vendor claims are verified against grammers 0.10.0 sources and the reference QR
implementations rather than assumed (PROPOSAL 14.1), and the whole document was
re-verified against origin/main after upstream #7377/#7397 (14.4) — which
removed owner-vs-actor and thereby retired this design's worst finding.

Status: sign-off withheld pending the conditions in the review log. Not
approved for implementation; opened for review.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(telegram): linked-device — device-link auth, session custody, standard-op tools

Implements the design in docs/internal/design/telegram-linked-device/: a user
links their personal Telegram account as a real MTProto linked device, and the
agent reads their conversations and acts as them through the standard messaging
operations. Reads are live — no message content is persisted.

Contracts
- ironclaw_extension_contracts: device_link (DeviceLinkAdapter + its step/input
  vocabulary) and linked_session (SessionBytes, LinkedAccountRef/Grant,
  LinkedSessionPort + factory). VendorAuthRecipe::DeviceLink with every arm,
  including the (DeviceLink, DeviceLink) compatibility case and
  keepalive_idle_threshold -> None, both of which fail at activation rather than
  compile if missed.
- ironclaw_host_api: send_message.output.v2.json as a NEW file carrying the
  sent_unverified branch. .v1 is byte-identical — the standard's schema
  immutability rule forbids an in-place edit. Schema resolution is now
  version-aware so existing bindings keep resolving .v1 forever.

Auth
- Device-link flow: ordered steps with revision CAS so a duplicated poll is
  idempotent and never re-invokes the adapter; two clocks (a step clock that
  re-mints, a flow clock that terminalizes); AwaitingVendor projected explicitly
  as Authenticating rather than falling through to Disconnected.
- link_revision on CredentialAccount with a CAS-bearing opaque-material write.
  Auth owns conflict detection only; it does not parse the session blob.

Host
- Device-link binding slot and its check_binding arms. Retires
  auth_never_binds_is_not_a_binding_field, which encoded the invariant the ADR
  deliberately revokes; the retirement cites the ADR.
- SnapshotDeviceLinkDriver resolves extension -> bound adapter and enforces poll
  rate limits and TTLs host-side.

Package
- MTProto via grammers 0.10.0, exact-pinned: the pin is a security control, not
  hygiene, because 0.10.0 never persists server-pushed DC addresses and that is
  what makes address validation in IronclawSession airtight.
- Sockets confined to transport.rs. NoRetries plus an explicit wrapper, because
  AutoSleep would re-send a write after an I/O error and no retry policy can see
  whether a request is a write.
- QR login by re-export polling (the flow an official client uses), phone and
  2FA paths, per-link mutex, logout on every post-acceptance abort.
- 15 standard ops. send_message returning id == 0 is a confirmed-but-uncorrelated
  send: Completed with sent_unverified, never a failure — a failure is what a
  model retries, and the retry double-sends to a human. Dropped/Io on a write is
  outcome-unknown and maps to vendor_error instead.

Frontend
- One QR/countdown implementation, shared by the existing pairing panel and the
  new device-link card. QR <-> phone switch, 2FA entry, stale-revision guard,
  polling stops on terminal states.

Gates
- Vendor names kept out of generic crates.
- Cross-crate include ratchet 16 -> 17, recorded deliberately in that file: the
  telegram package gained prompt docs when it gained tools, using the same
  include shape Slack already uses. Not a new class of reach-in, and not
  repointable while the layer matrix forbids runtimes -> products.

Local verification: cargo fmt, clippy --all-targets --all-features -D warnings,
and the full ironclaw_architecture_tests suite all pass; 1342 unit/contract tests
green across the touched crates.

NOT complete. The design's checklist is largely unticked — no integration tests,
no live-Telegram verification, and the security conditions in PROPOSAL 14.2-14.4
remain unmet. See the PR body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(telegram): integration harness, supply-chain pin, ownership pins, content bounds

Closes four gaps from the previous commit. All gates green; see the honesty
section below for what this does NOT close.

Integration
- tests/integration/support/harness/profiles/device_link.rs mounts the real
  bundled telegram package (its shipped manifest: channel + [auth.telegram] +
  15 standard_op tools) and writes [admin_configuration] through the real
  capability. New reborn_group_device_link target, registered in the root
  Cargo.toml, driving 3 scenarios against a scripted adapter.

Supply chain (the ADR requires these WITH the dependency, not later)
- All three grammers edges: =0.10.0 exact, default-features = false, explicit
  feature allowlists. grammers-client drops its `fs` default (nothing calls
  upload_file/download_media; attachments ride ironclaw_attachments). The socks5
  `proxy` feature stays off — a proxied dial bypasses Session::dc_option, which
  is the only seam our DC address validation owns.
- New reborn_linked_device_supply_chain_pin gate (13 tests) failing on any
  version or feature-set drift, with the rationale in its module doc.
- dependabot ignores grammers-*; Cargo.lock unmodified.

Ownership
- NewCredentialAccount::for_linked_device pins ExtensionOwned + empty grants;
  bump_link_revision refuses an unpinned account in both the production store
  and the fake. A §4.5 logout-before-unbind family lands in auth::cleanup.

Untrusted read content (§6.4)
- The content bounds become §7.2 constants with zero-checks and relationship
  asserts. @username handles now pass through sanitize_untrusted_text — the
  handle is the identity the model is told to trust.
- New conformance.rs proves every content-returning addendum frames its output
  as untrusted, and that the framing predicate is not inert.

Honesty — this is NOT a working feature yet
The handshake has no production wiring: nothing constructs a DeviceLinkDriver,
session custody resolves to unavailable() in every deployment, the durable
credential store does not implement opaque material (blocked on a CAS-bearing
SecretStorePort::put that was never built), completion cannot mint an account,
LinkedAccountResolver has zero implementations, and the shipped UI calls
/api/reborn/product-auth/device-link/... routes that do not exist. Fourteen
TODO(design) markers record each seam. Nothing here has ever spoken MTProto.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(webui): device link rides the generic product-auth routes, not a new namespace

A build agent invented /api/reborn/product-auth/device-link/{start,flow/{id}/
status,flow/{id}/input,flow/{id}/cancel} and then recorded its own invention as
a missing backend dependency. PROPOSAL §8.12 says the opposite: "additive
flow-status fields (step, revision, display, retry-after); route input
submission to the driver" — extend what exists.

A device link IS an AuthFlowRecord, and flow_status(scope, flow_id) is already
generic over flows; the route is only *named* oauth/... for historical reasons.
So the browser now calls the routes that are actually mounted:

  status  -> /api/reborn/product-auth/oauth/flow/{flow_id}/status
  start   -> /api/reborn/product-auth/extension/oauth/start
  input   -> /api/reborn/product-auth/manual-token/secret/submit
  cancel  -> /api/reborn/product-auth/oauth/flow/{flow_id}/reconcile

That removes "no backend routes exist" as a blocker. What remains is genuinely
additive and much smaller: the status response must carry the device-link frame,
and secret submission must route to the device-link driver — both extensions of
handlers already mounted in product_auth/mod.rs, marked TODO(backend) at the one
place that reconciles them.

Frontend suite green: 143 files, 1264 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(webui): device-link routes — status is shared, start/input/cancel are not

Corrects an over-correction. The previous commit routed EVERYTHING through
existing product-auth routes, which is wrong: extension/oauth/start builds an
authorize URL (a device link has none) and manual-token/secret/submit means
"user pasted an API key" (not "user typed step 3's 2FA code").

The honest shape is a mix:

- STATUS is genuinely shared. flow_status(scope, flow_id) fetches an
  AuthFlowRecord and returns its status with no OAuth-specific logic, and
  PROPOSAL §8.12 asks for additive fields on exactly that response. Polling
  extends the existing route.
  Naming wart recorded: the route is spelled oauth/flow/... though the object
  it serves is generic. Renaming to /product-auth/flow/{flow_id}/status with
  the old spelling kept as an alias is the right follow-up, and is a
  route-descriptor change rather than part of this feature.

- START, INPUT and CANCEL are device-link specific, because the operations
  differ: start takes a link mode (QR vs phone); input carries a typed kind
  plus the step revision it was typed against; cancel must ask the vendor to
  log the device out, or an accepted-but-abandoned link leaves an orphan
  authorization on the user's account (§4.3). Nothing existing does that.

These three are marked TODO(backend) as work THIS feature owes — not, as the
original agent comment claimed, a dependency on another branch.

Frontend suite green: 143 files, 1264 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(telegram): finish the linked device — custody, mint, routes, resolver

The branch shipped a fail-closed skeleton: green gates over fourteen
`TODO(design)` markers and a feature that could not link an account. This
closes the chain end to end and fixes what the extension-unification audit
found on the way.

Audit findings, fixed rather than worked around:

* `LinkedAccountResolver` was declared by the telegram package, so the
  containment PROPOSAL §5.1 requires — rooted in a HOST-minted grant — could
  only ever have been satisfied by the package itself. Moved into
  `ironclaw_extension_contracts` beside `LinkedSessionPortFactory`, supplied
  on `BindContext`, and implemented host-side over the same
  credential-account selection every runtime injection uses.
* `DeviceLinkBinding` carried a bare `user_id`, which is *why* completion
  could not mint: minting needs an `AuthProductScope` and synthesizing one
  from a user id would re-derive security-relevant scope. It now carries the
  durable flow's own scope (`user_id()` is an accessor over it).

The implementation chain:

* `ironclaw_secrets` gains a compare-and-swap write path
  (`put_versioned`/`read_versioned`): the previous last-writer-wins `put`
  would let a concurrent write clobber a rotating vendor auth key, which is a
  silently dead link. Decorator and four test doubles follow the widened
  trait.
* `ironclaw_auth`'s durable store implements opaque-material load/store over
  it — detection only, never a semantic merge: the blob is vendor-private and
  only the package can read it. `complete_linked_device_link` is the one
  place the completion policy lives (reuse-before-create, never resurrect a
  revoked account, the §4.5 ownership pin, load-then-CAS so a crashed prior
  link cannot brick relinking).
* Custody splits by revision: a provisional in-process space for the
  handshake (the blob exists *before* any account does — §4.3's store → mint
  → report) and durable custody behind the credential service, plus the
  ref→account directory that maps a host-issued `LinkedAccountRef` to the
  coordinates the auth domain needs.
* The extension host's driver mints the account at completion and registers
  it with custody, so `DeviceLinkStepOutcome` finally carries `Some(account)`
  and a link can complete.
* Composition wires all of it and attaches the flow driver and the linked-
  device revoker to the product-auth bundle.
* `api_hash` is `secret = true`, so it cannot ride `BindContext` (non-secret
  config only). It resolves at *load* — the one I/O-legal point before bind —
  through a new pre-scoped `LoadTimeAdminSecrets` port
  (`NativeExtensionFactory::load` becomes `async`). Unset, the adapter still
  binds and fails every link attempt closed, so a bot-only deployment keeps
  activating.
* The CLI binds all three Telegram surfaces (channel + device-link + tools),
  the shape `check_binding` has always required.

WebUI: four device-link routes, not three. `poll` is the departure from the
design — a card cannot poll the read-only status route, because a link only
advances when the host re-exports the login token (§4.2) and nothing else
drives it, so a card polling a pure read waits forever on a QR that was
already scanned. Routing the advance through the shared GET would also hide a
vendor call behind a descriptor declared read-shaped. STATUS stays shared,
stays a read, and carries §8.12's additive frame so a re-rendered card
hydrates without disturbing a live link. The ADR's detection control ships in
the completion card: the resolved account plus a count-the-devices ask,
worded to claim only what it catches.

Two failures were real behavior, not test drift:

* A cleanup decorator built at construction captured the account read model
  before it was final and would have broken *every* cleanup. The
  logout-before-unbind ordering moved to the bundle's single cleanup entry
  point, where the read model is settled.
* A lost compare-and-swap surfaced as 503 "retry later" to a card holding a
  stale step revision; retrying a superseded revision can never succeed. It
  maps to 409.

Proof: `scenario_handshake_mints_and_serves` drives composition's real
`DeviceLinkFlowDriver` start → poll → submit → completed, asserts the §4.5
ownership pin on the account the mint produced, asserts custody actually
persisted, and proves a linked tool call resolves to that account. Three
caller-level route tests drive the four routes over the mounted router.

NOT DONE, and not claimed: nothing here has ever spoken MTProto. Every test
drives a scripted adapter, so QR acceptance, DC migration, 2FA and flood-wait
are unexercised, and the `id == 0` / `Dropped`-on-write evidence rules have
never met a real server. PROPOSAL §14.3's withheld security sign-off is
unchanged. CHECKLIST is 51/135 with a note on why the ratio is what it is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(device-link): the browser could start a link but never finish one

Three defects found by driving the real flow in a local stack, each at a
seam between two things that were individually tested and green.

**Blank ids were sent for ids the caller did not have** (frontend). The chat
auth-gate card reads its scope from a gate model that has no `threadId` at
all and leaves `invocationId` null, so it posted `thread_id: ""`. The host
parses each of these into a validated newtype — `ThreadId`, `InvocationId`,
`TurnRunRef`, `AuthGateRef` — every one of which rejects a blank, so `start`
answered `400 invalid_request` before the flow began. Absent optional ids are
now omitted at the one API choke point rather than defaulted to `""`.
Required fields are deliberately NOT filtered: a blank `provider` must reach
the host and be rejected, not vanish into a request meaning something else.

**`start` minted an invocation and never returned it** (wire contract).
`scope_matches` is exact equality over the whole scope, and poll/input/cancel
re-derive that scope from what the browser sends back. A card opened outside
a run — the Extensions configure modal, no run and no gate — has no
invocation to carry in, so the host minted one and stored the flow under it.
With no way to learn that value, every follow-up call built a different scope
and 400'd: the flow could be started and never advanced. `DeviceLinkFlowResponse`
now echoes `invocation_id`, exactly as `ManualTokenSetupResponse` already
does and for the same reason, and the panel prefers it over the prop.

**A still-valid QR was reported as "nothing to show"** (adapter). Telegram
returns the same token bytes on every poll for the whole of a token's
window; `paint_token` treated unchanged bytes as `AwaitingVendor`, which is
defined as "nothing to show", so the card blanked the code about one poll
after painting it and parked on "waiting for the vendor" until expiry. It now
always returns `Display`, with `expires_in` recomputed so the countdown stays
honest. Re-emitting identical bytes repaints identically — there was no churn
to avoid. This left `poll_interval` and its two constants dead (that helper
only ever fed the deleted arm; pacing comes from the host's 3s
`DEVICE_LINK_POLL_INTERVAL_MILLIS`, the cadence the module header documents),
and made `PendingPhase::AwaitingScan`'s `token` field write-only; both are
removed, and `paint_token` — which never touched `self` — is now a free
function so the regression test can drive it directly.

Compatibility: `invocation_id` is a new required field on a response DTO that
no released client consumes; the two frontend changes are additive at the
request boundary and widen what the host accepts nowhere.

Test Strategy
- Crate: `paint_token` re-export test asserts both the first poll and an
  identical re-export paint the scannable code (`ironclaw_telegram_extension`,
  201 passed). Sabotage-checked by reintroducing an `AwaitingVendor` arm.
- Frontend: new `device-link-api.test.ts` (blank ids omitted, required fields
  preserved, `revision: 0` survives the filter) and a `device-link-panel`
  case pinning that the host-minted invocation reaches the follow-up poll.
  Both sabotage-checked; 1271 vitest passed.
- Integration: `reborn_group_device_link` 15/15; `ironclaw_architecture_tests`
  green (wire DTO gained a field); clippy clean on both touched crates.
- Live: `start → poll → cancel` driven against real Telegram MTProto — the
  code is exported, re-exported, and still displayed on the second poll.
  NOT verified: nobody scanned it, so acceptance, DC migration, 2FA, and the
  credential mint on completion remain unexercised at every tier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): repair test targets broken by the #7477 x device-link merge

Three sites neither branch could have seen: our device-link driver tests
and integration harness profile still built the pre-#7477
channel-adapter shape (now ChannelSurfaces), and #7477's new
lifecycle_contract helpers predate the device_link binding slot and the
custody fields on ExtensionHostDeps. Caught by the workspace clippy
gate; the earlier post-merge check piped through tail and masked the
failing exit code.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(linked-device): close the four review findings, with regression tests

- SessionPool: a poisoned lock now reports SessionPoolError::Poisoned
  instead of masquerading as AtCapacity (permanent corruption vs retry
  shortly), matching session_store's taxonomy.
- SessionPool takes Option<i32> api_id: a deployment without an MTProto
  identity now fails acquire closed with NotConfigured before custody or
  any dial (previously dialed with api_id = 0 and failed at the vendor),
  and a cold revoke reports LogoutUnverified immediately instead of
  burning a doomed handshake. The tool-side mapping carries the honest
  not-configured sentence; the CLI drops its unwrap_or(0).
- session_blob errors now carry a bounded serde reason (category +
  line/column only — never blob content), making a corrupt custody blob
  attributable.
- credential.rs relink version probe carries the silent-ok contract its
  fallback relies on (CAS is the authority; a failed load costs one
  conflict round-trip, never a clobber).
- GenericExtensionHost custody params are now required, with the
  fail-closed collapse moved to the composition boundary; test sites
  pass the unavailable shapes explicitly.

Also repairs two merge-tail gaps #7477 exposed: the device-link fixture
manifest now speaks the per-axis channel grammar (was the retired
inbound/outbound booleans, failing 25 extension_host tests), and the
manager field-status expectation includes the two MTProto admin fields.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(device-link): close the audit's blocker and vendor-neutrality findings

A 14-agent design pass found the generic device-link machinery is
vendor-blind in its runtime spine but not in its vocabulary, plus one
functional blocker independent of any second vendor.

THE BLOCKER. The auth tier re-mints a lapsed frame by calling
`DeviceLinkDriver::begin` again on the same flow, and the host driver
refused exactly that as a non-restartable `Internal`. The host flow TTL
(10m) outlives the step clock (60s), so at a lapse the flow is always
still live and the refusal always fired: every link not completed inside
the first frame terminalized with "cannot be completed for this
account". Both halves were tested and both suites passed, because
nothing crossed them.

`begin` now means one thing. A begin naming a live flow is a re-mint:
the stale vendor conversation is cancelled first (the parallel-
conversation hazard the refusal was really guarding), then a fresh one
starts under the same flow, carrying the flow clock and secret-attempt
count forward so a re-mint can neither extend the attempt nor reset an
abuse counter, and skipping the begin budget because it is not a new
attempt. The superseded test is rewritten, not deleted, to pin the
surviving invariant. `ironclaw_auth` now exports a DeviceLinkDriver
conformance suite covering the cross-half obligations, run against the
production driver; sabotage-tested by restoring the refusal.

VENDOR NEUTRALITY. The recipe's mode-label seam was built and then
bypassed, so the generic panel hardcoded one vendor's QR-vs-phone
ceremony in 11 locales and wedged forever against a vendor that
declares no alternate. `DeviceLinkPromptView` now carries
`alternate_available`, both recipe labels, `display_kind`,
`extension_id`, and `vendor_user_ref` (which used to double-book the
`code` slot); the labels ride the durable challenge beside
`display_name`, additive with serde defaults. The card gates and labels
the switch from the wire, resets mode on restart, honours display_kind,
and passes resume_flow_id (previously dead on every caller). Completion
copy no longer names one vendor's settings menu.

The host also gets its own voice: HostThrottled and LimitReached, so a
host budget stops reporting itself as vendor pushback, and NoBinding
maps off AccountUnavailable (an operator condition is not a broken
account).

ALSO: per-user budgets keyed by (user, extension) with counters evicted
on reap; at-most-one device_link recipe enforced at manifest parse (a
second was silently ignored, mis-attributing flows and grants);
PENDING_LINK_REVISION deduped and the provisional cap derived from the
driver's limit rather than agreeing by comment; port obligations
documented where implementors read them and the contradictory
poll-purity sentence reconciled; the specificity gate widened to
build.rs and frontend scripts, which immediately caught two real vendor
names now fixed rather than allowlisted; and a sabotage-tested
sole-consumer assertion on the MTProto stack.

Contracts ceiling re-pinned 10_344 -> 10_512 with the rationale in the
gate: declaration and documentation only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(telegram): design automatic linked channel identity

* feat(telegram): pair linked devices with bot channel

* fix(ci): classify linked-device Dependabot config

* fix(telegram): align linked-device CI evidence

* chore(telegram): document fixture panic invariants

* test(composition): classify channel harness as test-only

* fix telegram device-link setup CI

* fix standard messaging schema parity assertion

* fix telegram model search projection test

* feat(telegram): make the device-link cutover breaking for retired pairings

Reverses the zero-touch upgrade decision in AUTO-CHANNEL-IDENTITY §8 (owner
call, 2026-08-14): a proof-code identity binding written before the cutover
no longer authorizes anything on a device-link channel. Each connection
strategy now owns exactly one identity keyspace — the fallback chain is
deleted, and the device-link-v1 prefix is a fence that keeps retired rows
inert. Previously paired users land back at setup, get the connect-required
notice on their next bot DM, and re-link once: the same first-run ceremony
as a fresh install, and the same missing-credentials UX as every other
extension.

- admission, connection status, command roles, and outbound-target
  validation consult only the strategy's single keyspace; the plural
  channel_identity_lookup_keyspaces API is deleted
- the binder's retired-namespace cross-user veto is removed: an inert row
  cannot block a freshly authenticated link its owner has no way to clear
- ChannelIdentityKeyspace::Legacy renamed to Unversioned — nothing legacy
  about the namespace OAuth/pairing channels still live in
- retired rows stay untouched data (no bulk delete); explicit disconnect
  scrubs both generations, unchanged
- docs: AUTO-CHANNEL-IDENTITY §8 and the telegram package README now
  describe the breaking cutover; rollback stays valid because pre-cutover
  rows are never rewritten

Flipped pins, each watched red then green: resolver ignores a retired
pairing key; connection status reports disconnected; command roles confer
nothing; a stale foreign row does not veto a link; a retired delivery
target is offered only after re-link.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LACC5dKyNy3GNXRNbvgz57

* fix(tests): bind the acme fixture's device-link adapter only when declared

The extension-profile fixture bound ScriptedDeviceLinkAdapter on every acme
bind, but check_binding proves agreement per axis and the stock
acme-messenger manifest declares oauth2_code only — so every bind failed
with UndeclaredDeviceLinkAdapter, activation never completed, install turns
recorded no capability results, and the standard-op tests' approval gates
never existed. Hidden until the merge queue: the PR lane's affected-surface
planner had never selected integration shards 2/3.

The adapter is now bound iff the installed manifest declares a device_link
recipe (the same declared_device_link_recipe test check_binding runs), so a
future acme device-link variant still gets the scripted adapter.

Verified: reborn_integration_extension_ingress 17/17 and
reborn_integration_extension_runtime 25/25 locally, Postgres legs included
(previously 3 + 4 failures reproducing the queue's shard 2/3 ejection).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LACC5dKyNy3GNXRNbvgz57

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 13:35:38 +00:00
firat.sertgoz
75d53fd7c2 feat(automations): add structured execution contracts (#7548)
* feat(automations): add structured execution contracts

* ci: rebaseline composition budget for trigger preflight

* fix(automations): require structured contracts for new triggers

* test(automations): keep trigger fixtures lane-local

* test(automations): migrate structured trigger fixtures

* fix(triggers): address ironloopai/coderabbitai review — fail-closed preflight, one-pass render, policy hardening (#7548)

- Preflight fails closed: missing capability-visibility record denies; synthetic
  bridge ids rejected when tool disclosure is Off (mirrors the runner's
  decorator-attach condition).
- TurnExecutionPolicy denies unknown fields; RequiredSkill adopts the canonical
  validated-newtype shape (serde try_from, shared validate, TryFrom/AsRef/Display).
- Execution-spec prompt rendering is single-pass so field values containing
  placeholder tokens stay verbatim.
- Persisted structured prompt is authoritative: TriggerRecord::validate no longer
  re-renders against the current template, so template edits cannot brick stored
  triggers.
- TriggerCreateHook::validate_execution_policy is mandatory; the preflight-less
  compatibility path rejects restrictive policies and persists nothing.
- Unbound preflight reports TriggerError::Backend (host wiring fault), not a
  caller contract error.
- RequiredSkillUnavailable classifies as PolicyDenied, matching skill_context;
  required-skill resolution extracted to one shared predicate.
- QA tests: every trigger_create payload parsed and validated against
  TriggerExecutionSpec; fired-routine replay asserts the finalized assistant
  reply persisted in the run thread.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tests): restore shared trigger_execution_contract helper for trigger scenarios

The group_triggers scenarios call super::support::trigger_execution_contract,
but the helper only existed as local fns in two root suites, so the
reborn_group_triggers target did not compile on this branch. Hoist it into
tests/support/mod.rs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): map direct tests/support files in the Reborn PR test planner

'Detect Reborn test scope' failed with 'unmapped test or CI path:
tests/support/mod.rs' after the shared trigger_execution_contract helper
moved there. Direct tests/support/*.rs files are compiled into the root
suites and (via #[path]) the integration group targets, so schedule a
representative partition of each tier. Specific INTEGRATION_SUPPORT_OWNERS
mappings keep precedence.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): allow lane-specific trigger test helper

* fix(ci): finalize fired routine replay reply

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 08:07:20 +00:00
Illia Polosukhin
7d77a752f2 feat(documents): edit docx/xlsx/pptx structurally, render PDF from HTML, and fix the #7109 text-log regression (#7163)
* fix(coding): stop the binary-write backstop rejecting ordinary text logs

Follow-up to #7109, which shipped with the binary backstop in
`verify_read_before_edit` using the STRICT probe (any NUL in the first
8KiB) while `read_file` admits text carrying a few stray NULs via
`reject_binary_probe_lenient`.

The two disagree about what "binary" means, so a syslog the model had
just read successfully became unwritable — and was reported as a "binary
document", which it is not. `read_file_tolerates_stray_nul_and_invalid_utf8_in_text_logs`
pins the lenient read; this makes the write path agree with it.
`apply_patch` keeps applying the strict probe itself, where byte-fidelity
for write-back demands it.

Also adds the coverage #7109 landed without:

- `builtin_write_file_still_overwrites_text_log_with_stray_nul` — fails
  before this fix, and is the regression above.
- `builtin_write_file_rejects_extracted_read_representation_at_unlisted_extension`
  — the recorded-read-representation guard had no test at all. The
  existing docx test returns at the extension guard before
  `verify_read_before_edit` is ever reached, so `ReadState`, the
  `read_file` plumbing and the representation check were untested. `.rtf`
  is the one format `read_file` extracts that the extension lists omit,
  making it that guard's only live surface.
- `builtin_apply_patch_rejects_a_binary_document_with_an_actionable_reason`
  — #6898 named apply_patch's opaque failure as part of the bug, and the
  fix addresses it, but nothing pinned it.
- `reborn_integration_document_edit` — the whole journey at the
  integration tier: upload a .docx, ask for an edit, and the bytes served
  by the production `InboundAttachmentReader` (the WebUI download path)
  are byte-identical to the upload.

Refs #6898, #7109

* feat(documents): structure-preserving docx/xlsx/pptx editing and HTML-to-PDF

Issue #6898 item 3: the deferred "real document round-trip capability".
This is the library layer; capability wiring follows.

The governing rule is copy-through: rewrite only the parts an edit
targets, copy every other zip entry byte-for-byte. A generator that
rebuilds a document from the text a model saw drops everything the model
never saw — styles, numbering, headers, images, embedded objects — and
the file still opens, so the loss is invisible. Copy-through makes that
impossible by construction rather than by diligence, which is why this
is not "add docx-rs and write a docx".

- `ooxml`: package read/write plus an event-level XML transform. An
  event the handler does not claim is re-emitted exactly as read, so
  attribute order, namespace prefixes and self-closing forms survive; a
  DOM round trip normalizes all three and churns untouched parts.
- `docx`: paragraphs with `w:ins`/`w:del` surfaced as typed revisions
  (flat extraction shows deleted text as if it were still in the
  contract), and accept/reject. Rejecting a deletion converts `w:delText`
  back to `w:t` — without that Word renders the restored run empty.
  Revisions split across runs coalesce into one span.
- `xlsx`: shared strings resolved so headers read as text, formula edits
  that drop the stale cached `<v>` and set `fullCalcOnLoad`, and new
  cells inserted in ascending column order (Excel repairs, and silently
  drops content from, an out-of-order row).
- `pptx`: slide clone that duplicates the slide's rels part, so the copy
  inherits its layout — style lives in the layout chain, not the slide,
  which is why constructing a slide cannot preserve it. Registers the new
  part in content-types, presentation rels and slide order.
- `html_pdf`: PDF is deliberately not edited. A documented HTML subset
  renders via printpdf's standard-14 fonts. printpdf's `html` feature is
  left off on purpose: it resolves fonts through rust-fontconfig, which
  scans system fonts and would make pagination depend on the host.

51 tests, including the copy-through invariants (unrelated parts stay
bit-identical, a no-op edit is byte-stable) and every trap named above.

Refs #6898

* feat(coding): document_edit and html_to_pdf, and read_file reads OOXML structurally

Wires issue #6898 item 3's library layer to the model. Reads unify into
read_file; writes stay typed. That asymmetry is the design, not a
compromise:

- read_file on a .docx/.xlsx/.pptx now returns the ADDRESSABLE STRUCTURE
  (paragraphs with tracked-change spans, cells with resolved headers and
  formulas, slides) and records `ReadRepresentation::Structured`. Folded
  in rather than offered as a `document_read` tool because a model
  reaches for read_file on whatever path it is handed — a tool it had to
  know to prefer would go unused while read_file kept returning
  tag-stripped text that shows a redline's DELETED words as if they were
  still in the contract.
- write_file cannot be folded: its contract is (path, content: &str), so
  it cannot express "accept the revision in p3". Overloading it means
  regenerating the document from text — the corruption #6898 banned — or
  a sometimes-JSON `content`. apply_patch fails for the same reason plus
  ambiguity: its anchors match extracted text, and mapping a match back
  to runs is undefined when a string spans a revision boundary. The
  binary-write ban stays permanent.
- document_edit takes typed ops and always writes to a NEW path, so a
  bad edit can never cost the user the original. It requires a prior
  Structured read of the source, keeping write_file's mid-air-collision
  guarantee on a fingerprint over the same raw bytes.
- html_to_pdf renders; it refuses to overwrite an existing file, because
  silently replacing a PDF the user uploaded is the same class of loss
  the binary-write guard prevents.

Two findings from writing the tests, both fixed here:

1. The surface test caught that a builtin capability is invisible to the
   model without a published input schema — registration alone is not
   enough. Both new tools now publish one.
2. The xlsx journey caught that set_cell_formula could not create a row.
   A totals row sits just below the data, so "the row is not there yet"
   is the ordinary case, not an edge case; the crate fixture happened to
   have the row already and masked it. Rows are now created in ascending
   order, with regression tests.

Tests: 6 capability tests through real dispatch (including that a
Structured read still does NOT authorize a raw write_file overwrite —
the two guards must not cancel each other), and four integration
journeys on real OOXML fixtures: docx redlines resolved into a clean
copy, an xlsx total under a named column, a pptx slide cloned with the
source's layout, and a PDF produced by authoring HTML and rendering it.

Refs #6898

* test(reborn): register e2e coverage for document_edit and html_to_pdf

`reborn_builtin_first_party_capability_e2e_coverage_is_complete` requires
every always-on first-party capability to name where its Reborn e2e
coverage lives. The two new document capabilities had that coverage —
`reborn_integration_document_edit` drives all four journeys — but were
not registered, so the guardrail failed.

Caught by the pre-push hook, which runs the full workspace suite; the
per-crate runs used while developing never touch this test.

* ci(planner): classify tests/fixtures document fixtures

`Detect Reborn test scope` failed with "unmapped test or CI path:
tests/fixtures/contract.docx". The planner deliberately hard-errors on
any unclassified path under tests/ to force a per-file decision, and
binary document fixtures had no arm — only recorded LLM traces under
tests/fixtures/llm_traces/ were mapped.

These fixtures are consumed by integration tests through `include_bytes!`,
so a changed fixture changes what those tests assert; it schedules a
representative integration lane, matching how shared integration support
is treated.

Adds the matching planner test.

* ci: satisfy the panic check in the test-only fixture builder

`Fast deterministic checks` flagged four unwraps in
`ironclaw_documents::test_fixtures`. The module is `#[cfg(test)]`-gated
(`has_cfg_test_module_declaration` agrees), but the checker's main scan
path does not consult that for crate-root modules, so it reads them as
production code.

Uses the inline `// safety:` suppression the checker documents rather
than relaxing the checker, which guards a real invariant for everything
else. The comments must be on the same line as the call to take effect.

* fix(documents): nested paragraphs and leaked revision flags corrupted docx edits

Two critical defects from CodeRabbit's review of #7163, both confirmed by
tests that fail before the fix.

1. Nested `w:p` lost the outer paragraph and desynchronised ids.
   Word nests a paragraph inside a paragraph when a run holds a text box
   (`w:txbxContent`). The reader kept one `current` slot, so the inner
   Start overwrote the outer and the outer was never emitted; the writer
   meanwhile counted EVERY `w:p` Start. Read ids and write ids therefore
   addressed different paragraphs, so an edit landed on the wrong one.
   The reader now keeps a stack and assigns ids in Start order, matching
   how the writer counts, and both write paths keep a target stack so a
   nested paragraph's End restores the enclosing paragraph's state
   instead of clearing it.

   Note the mechanism: table cells do NOT reproduce this — `w:tbl`/`w:tc`
   paragraphs are siblings in document order. Only a text box nests.

2. Revision flags leaked past a dropped subtree and unbalanced the XML.
   Reject-insert and accept-delete set `dropping = 1` AND the
   `in_insert`/`in_delete` flag. The dropping branch then consumed the
   matching End, so the flag was never cleared and stayed set for the rest
   of the document — deleting a LATER paragraph's `</w:ins>`. Resolving
   revisions in one paragraph emitted `word/document.xml` with an
   unclosed element, which Word rejects.

   The flag is now set only on the unwrap paths, where the End genuinely
   must be dropped by the handler. On the drop-subtree paths the End is
   consumed by the dropping branch and no flag is needed.

Both defects are invisible to single-paragraph, flat fixtures — which is
what the crate's own fixtures were.

* fix(documents): preserve cell styles, resolve sheets by relationship, keep text around comments

Four more defects from the #7163 review, each with a test that fails
before its fix.

xlsx:
- Replacing a cell dropped its `s` style index, silently reverting a
  currency or date column to General. That is the precise "preserve what
  you did not touch" promise the crate exists for, so the style now rides
  across the replacement.
- Sheet names were paired to worksheet parts POSITIONALLY. I shortcut
  this deliberately and said so in a comment; the reviewer was right that
  it is wrong. Sheet declaration order does not have to match worksheet
  file numbering, so an edit could land in the wrong worksheet. Names now
  resolve through `r:id` in `xl/_rels/workbook.xml.rels`, falling back to
  positional pairing only when the rels part is absent.

html_pdf:
- A comment or declaration cleared the buffered text before it, so
  `<p>hello <!-- note --> world</p>` rendered as `world`. This directly
  contradicted the module's claim that wrapping markup never swallows
  content. Whitespace now also collapses across the resulting span join,
  so the repaired text reads `hello world` rather than `hello  world`.
- A stray `&` consumed up to ten following characters looking for `;`,
  swallowing a real entity behind it. The scan now stops at any character
  that cannot appear in an entity name.

* fix(documents): reject duplicate zip entries, stop self-closing tags corrupting reads and rows

Three more from the #7163 review.

- A duplicate zip entry name kept both names but only the last bytes, so
  `write()` emitted the same content under both and silently rewrote a
  package we were asked to preserve. Now rejected at read.
- `Event::Empty` latched `in_value`/`in_formula`. A self-closing `<v/>`
  has no matching `End`, so the flag stayed set and the NEXT cell's text
  was attributed to the empty one, corrupting every later value in the
  row.
- A self-closing `<row r="N"/>` target produced a SECOND row with the
  same `r`, which Excel repairs by dropping content. The existing empty
  row is now replaced in place.

* chore: retrigger Railway preview

* fix(documents): address review findings

* test(reborn): refresh read-file golden payloads

---------

Co-authored-by: serrrfirat <f@nuff.tech>
2026-08-13 22:05:44 +00:00
Josh Ford
d82c9584e5 ci(check-guidance): extend the reference gate to the docs/ surface (doc-truth PR 2/5) (#7376)
* docs: fix live drift in extension, responses API, and channel docs

The public tutorial taught the retired manifest v2 authoring format
([[host_api]] / [capability_provider.tools] / runtime_credentials), which
the v3 parser hard-rejects, and never mentioned origin_gate_matrix; the
Responses API page claimed temperature is rejected (accepted 0.0-2.0 and
forwarded), claimed model must be "default" (any well-formed name <= 256
bytes), claimed max_output_tokens is rejected (accepted and ignored by DTO
policy), and omitted the required model field from every request example;
the channel tutorial pointed at two files that no longer exist.

- docs/extensions/building-a-tool.md: rewrite manifest sections to the v3
  [[tools]] / [[tools.credentials]] / [auth.<vendor>] shape, document
  origin_gate_matrix (origins, policies, ratchet), correct the hosted-MCP
  [mcp] section, packaging via ironclaw_extension_support package modules,
  and v3 test references; drop the nonexistent script runtime kind.
- docs/api/responses.mdx: correct model/temperature/tools/tool_choice
  rejection rules, document unknown-field tolerance, add the required
  model field to all 15 request examples.
- docs/channels/building-a-channel.mdx: replace dead
  crates/ironclaw_first_party_extensions + available_extensions.rs
  registration instructions with the current package-directory mechanism.
- docs/reborn/contracts/extensions.md: state that production manifests
  author v3 (lowering into the v2 resolved model described there); label
  the v2 examples as legacy.
- docs/reborn/how-to-port-tool-to-reborn.md: superseded banner pointing at
  the v3 guides.

Part of #7317 (doc-truth pipeline, PR 1 of 5).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): extend the reference gate to the docs/ surface

The public Mintlify tree had no path-reference validation — a published
tutorial told contributors to edit files that no longer exist and nothing
caught it. check-guidance.py already owned the machinery (tracked-tree
resolution, fence exclusion, suppress markers, shrink-only debt, fail-closed
floors), so the docs surface joins the same gate rather than a fork.

- discover_guidance() now collects every tracked docs/**.md|.mdx: published
  pages, the zh/ locale mirror, and the living contract corpus
  docs/reborn/contracts/. Dated archives (docs/internal/, the non-contract
  parts of docs/reborn/) are excluded as classes — measured 2026-08-07,
  705 of 709 dangling docs references sat in those historical corpora, and
  forcing dated plans/ADRs to track today's tree would either rewrite
  history or drown KNOWN_MISSING.
- docs/ files extract backticked inline paths only; Mintlify markdown link
  targets are site routes (extensionless pages, site-absolute /using/cli),
  a different namespace than the tracked tree, so the link extractor is off
  there by design.
- _reference_lines learns MDX comments ({/* ... */}), including
  {/* check-guidance: path-ok */} as the .mdx suppress-marker form, with the
  same one-reference-per-marker and multi-line semantics as HTML comments.
- Floors re-measured and re-dated (364 files / 2276 references; floors
  180/1100), plus a dedicated MIN_DOCS_FILES=60 floor: the aggregate floors
  sit below the guidance-only remainder, so the docs branch of discovery
  silently breaking needs its own refusal. --json now reports docs_files.
- Fixes the four real dangles the new scan found in docs/reborn/contracts/
  (moved nested_dispatch_stream.rs test home, retired event-store migrations
  directory, loop_driver_host tests->src move). KNOWN_MISSING stays empty.
- Self-tests: 8 new cases (dangling docs path fails; Mintlify links are not
  references; MDX marker suppresses exactly one reference; multi-line MDX
  comment hides content; zh discovered; archives excluded but contracts
  scanned; docs fence fails closed; docs floor refuses).
- ws12_workflow_contracts.py: docs/api/responses.mdx and docs/zh/index.mdx
  join the has_guidance in-scope probes so a narrowed trigger regex cannot
  silently skip the gate for public docs.

Part of #7317 (doc-truth pipeline, PR 2 of 5); stacked on #7375.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: address Copilot and CodeRabbit review on doc-drift PR

- responses.mdx: tool_choice is rejected only without external-tools wiring;
  with external tools enabled it passes validation and is currently ignored
  (validate_responses_supported_fields_with_external_tools never checks it).
- building-a-tool.md: clarify that effect-derived host ports are validation
  vocabulary against the HostPortCatalog allowlist; adapters are built by
  host-runtime services after authorization/obligations, never from manifests.
- how-to-port-tool-to-reborn.md: mark the decision tree's RuntimeKind targets
  historical (v3 accepts only wasm|first_party; MCP is top-level [mcp];
  process/CLI work is the sandbox lane).
- building-a-channel.mdx: document the user install flow — virtual package
  root /system/extensions/<id>/manifest.toml, ironclaw extension search /
  install <extension-id> (ID, not path), WebUI Extensions lifecycle.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(responses): align the limits bullet with the corrected tool_choice claim

The rejection list was corrected in the previous commit (tool_choice is
rejected only without external-tools wiring); the "Limits and quirks"
bullet still said "not supported ... rejected with 400". Same claim, one
wording.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: apply verified code-review findings on the drift PR

A full code review of this PR against live code surfaced claims the
original drift pass got wrong or missed; every fix below was re-verified
against the cited source before editing:

- responses.mdx: standard `ironclaw serve` deployments always wire
  external tools (OpenAiCompatRouteMountPorts requires the store/resume
  pair; mount.rs wires them unconditionally), so `tools` is accepted and
  `tool_choice` is accepted-and-ignored on shipped binaries — the
  conditional 400s apply only to custom compositions without the wiring
  (now a Note). temperature is validated and carried in the submitted turn
  payload but not applied as a provider sampling parameter. Non-streaming
  wait timeout is 30 s (DEFAULT_RESPONSES_WAIT_TIMEOUT), not 120. usage on
  retrieval is read best-effort from persisted run state incl. USD cost
  (read_run_usage), not always zero.
- building-a-tool.md: the [auth.example] oauth2_code recipe gains the
  required token_response map (deny_unknown_fields rejects the example as
  previously written); Gmail/Google Calendar corrected to first_party
  runtimes (their manifests declare kind = "first_party"); the worked
  api_key recipe is github's, not slack's; the tail "Quick implementation
  checklist" and reference list were still v2-era (script lane,
  assets/<extension>/ path, "manifest v2", v2.rs pointer) and now teach
  the v3 shape; composition/CLI package-naming claim narrowed (the binary
  does link slack/telegram adapter crates).
- contracts/extensions.md: legacy-format paragraph no longer claims
  host-bundled packages ship v2 (none do), and origin_gate_matrix is
  attributed to capability.rs + building-a-tool.md instead of
  extension-runtime/overview.md §3, which does not mention it.
- how-to-port banner: `script` manifest authoring is retired; the
  RuntimeKind::Script symbol survives as the process-sandbox lane's kind.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(contracts): repoint delivery_resolution.rs to its family directory

PR #7157 (merged to main 2026-08-07) cited
crates/ironclaw_outbound/src/delivery_resolution.rs in the
communication-delivery-resolution contract; the crate lives at
crates/domains/ironclaw_outbound/. Caught by this branch's docs surface of
check-guidance.py on the first merge of main after the gate landed —
exactly the drift class it exists for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): harden the docs gate and fix review-surfaced doc drift

Applies the verified findings from the PR #7376 code review:

- The loop-exit and turn-runner contract docs claimed the deleted
  loop_driver_host checkpoint-rejection test had 'moved into the
  module'; it was deleted in #6696 and the fenced verification command
  could not run. Both now cite the real surviving pins
  (planned_driver.rs executor test + the ironclaw_turns projection
  test mapped in scripts/reborn-e2e-rust.sh), with runnable commands.
- An unterminated comment now refuses at EOF like an unterminated
  fence; before, one typo'd closer silently un-scanned the rest of the
  file.
- Markdown links in the re-included corpora are now checked as repo
  paths (they are never published, so the Mintlify-route rationale did
  not apply); this alone added ~165 verified references.
- Each DOCS_REINCLUDED_PREFIXES entry must match at least one tracked
  page or discovery refuses, so the planned docs/reborn consolidation
  cannot silently drop the corpus from the scan.
- The living extension-runtime spec pages (overview.md,
  standard-operations.md) and guidance-conventions.md join the scan;
  guidance-conventions.md now describes the docs surface and the MDX
  marker form, and its one dangling test path is repointed.
- Floors comment corrected (57 rule globs, not 38).

Also fixes four drifted claims from #7375's pages, verified against
live code: the interleaved function_call_output example was rejected
with 400 (resume input must be exclusively function_call_output items
with previous_response_id); model is echoed only on create (GET/cancel
report the 'reborn' placeholder); output_schema_ref is optional; and
the unknown-fields claim now names the two deliberate exemptions.

Self-tests: 43 pass (three new arms — unterminated comment refusal in
both syntaxes, re-included links as repo claims, stale re-included
prefix refusal).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): sync module docstring with re-included link checking

CodeRabbit caught the docstring still claiming the link extractor is
off for all of docs/** — stale since b172f69c7 enabled it for the
re-included corpora. The docstring now states the exception and the
current re-include set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): drop the docs/reborn re-include machinery after the docs/internal migration

The docs-surface scan carried a double negative — exclude docs/reborn/ as
an archive class, then re-include its living pages via
DOCS_REINCLUDED_PREFIXES — because the old tree mixed dead archives with
living specs. #7559 moved everything under docs/internal/, so the structure
is now: one excluded archive class (docs/internal/), and the living spec
pages (the contract corpus, the two extension-runtime spec pages,
guidance-conventions.md) named in INTERNAL_GUIDANCE_PREFIXES and scanned as
first-class guidance files — full link checking, guarded by the same
per-prefix zero-match refusal. The published-docs floor now counts only the
Mintlify surface (measured 2026-08-13: 82 pages; floor re-halved to 40).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): validate docs discovery against docs.json navigation instead of a count floor

MIN_DOCS_FILES was an arbitrary magnitude tripwire (half of last measured,
hand-re-dated) that only caught the docs branch of discovery losing ~half
its pages. The published surface already has an independent definition —
docs.json navigation, owned by docs_publication_boundary.py — so the gate
now asserts every navigation page's source file is in the reference scan
(reusing the boundary script's nav walker and OpenAPI pseudo-page filter).
Discovery breaking refuses on the first missing published page, unreadable
or page-less navigation refuses rather than passing vacuously, and there
is no docs count floor left to tune. --json reports nav_pages_covered.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): count the living internal spec pages in the docs_files metric

CodeRabbit: docs_files under-reported the scan — the living internal spec
pages are scanned docs files but were excluded from the count, a leftover
of the deleted MIN_DOCS_FILES floor's published-only semantics. The metric
now reports every scanned file under docs/ (131 at measurement); published
surface health has its own signal in nav_pages_covered.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(check-guidance): tighten comments and docstrings

Same behavior; the docs-surface comments and test docstrings were carrying
paragraph-length rationale better kept in the PR description.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 20:27:31 +00:00
Benjamin Kurrek
b215bf0651 feat(slack): bind the remaining eight core standard messaging ops (#7515)
* feat(slack): bind the remaining eight core standard messaging ops

Slack bound 8 of the 16 core standard messaging operations. The other
eight were a named fast-follow the framework design spec deferred
(§13 "Deliberately not built"); this lands them, so Slack now covers
the full core vocabulary.

Manifest: eight new standard_op-bound [[tools]] — edit_message,
delete_message, add_reaction, remove_reaction, open_dm, get_message,
resolve_user, list_members — appended after the v2-era eight so the
projection-parity gate's positional comparison still holds.

Guest: vendor mechanics over chat.update, chat.delete,
reactions.add/remove/get, conversations.open/members/history/replies,
and users.list. Three behaviours worth a reviewer's attention:

- get_message has no Slack endpoint. It reads the conversation at the
  exact ts and falls back to the thread when the message is a threaded
  reply, accepting only an exact match — a near miss is unknown_message,
  never the neighbouring message.
- remove_reaction's optional emoji is implemented, not rejected. Slack's
  endpoint requires a name, so the omit-emoji variant reads the message's
  reactions and removes each one the connected account added. That is why
  reactions:read joins the grant, and why that variant returns no emoji.
- already_reacted / no_reaction are treated as success. Slack returns
  them only for a message it resolved, so the requested end state holds;
  an error would push the model into retrying a no-op. It also makes the
  removal loop converge on retry after a partial failure.

Scopes: the [auth.slack] union gains reactions:read, reactions:write and
im:write.

COMPATIBILITY: there is no scope-upgrade re-consent flow, so an account
connected before this change holds a token without the new scopes, and
Slack answers those three tools with missing_scope -> mapped to
messaging.permission_denied (a model-visible denial, deliberately NOT an
AuthRequired re-auth gate, since re-running OAuth would request the same
manifest scopes). Widening an existing grant means disconnecting and
reconnecting Slack; the other 13 tools are unaffected. ROLLBACK: reverting
this commit restores the 8-tool surface and the narrower grant; already
-widened user tokens keep unused scopes, which is inert.

Gates: two pins were evolved rather than relaxed. The v2->v3 parity suite
gained a PackageAdditions declaration, so the v2-era eight still project
identically and positionally while the additions and the exact scope delta
are declared explicitly; the catalog scope pin moved from
assert!(matches!(..)) to assert_eq! so a future drift prints which scope
moved.

Tests: 24 guest unit tests — canonical input serde, output shapes, the
reaction authorship filter, exact-ts matching, the idempotent-success
arms, directory matching, the error taxonomy, and capability-id dispatch
— plus manifest-projection assertions that all 16 bind exactly the core
set with host-synthesized schema refs and external_write on writes. The
authorship filter and exact-ts matching were sabotage-verified, as were
both evolved gate assertions.

Not covered: the I/O orchestration inside each operation (which calls
run, in what order, and that a failed auth.test aborts the omit-emoji
removal) is unreachable without a host seam in the WASM guest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(slack): correct four guest edge cases and pin all eight new ops at the dispatch seam

Guest fixes (artifact rebuilt, digest re-recorded):
- reactions.get now sends full=true — without it Slack truncates per-reaction
  users arrays on popular messages and the omit-emoji remove_reaction could
  skip the connected account's own reaction while reporting success.
- open_dm validates user_ref with is_slack_user_id before calling Slack:
  conversations.open takes a comma-separated list, so an unvalidated
  "U1,U2" silently opened a group DM against the 1:1 contract.
- get_message's thread fallback pinches the range to the exact target ts
  (oldest=latest=ts&inclusive=true&limit=5) instead of scanning a fixed
  999-message page — finds a reply at any depth and stops shipping ~1MB
  pages to keep one message.
- resolve_user's users.list page size now equals the match limit (default
  raised to the full 200): Slack cursors are page-granular, so the old
  mid-page break dropped matches between the cap and next_cursor that no
  amount of paging could recover.
- next_cursor extraction deduplicated into one helper across all four
  paging ops; add_reaction's per-tool scopes drop reactions:read (it only
  calls reactions.add — the read scope belongs to remove_reaction alone,
  exactly as the scope table documents).

The manifest's scope-widening compatibility comment now describes the real
mechanism: a pre-widening account is refused by the HOST's provider-scope
gate at credential staging and lands on the AuthRequired re-auth gate, which
reconnects the same account with the widened union (binding skips the scope
gate). Slack's own missing_scope only fires for server-side drift after
staging passed, and that maps to messaging.permission_denied.

Coverage (contradicting the "no host seam exists" claim this PR shipped
with): the existing host-runtime WASM harness drives the committed
slack_user_tool.wasm through invoke_capability with scripted egress, so the
conformance sweep now covers all sixteen ops, plus ten behavioral pins —
missing_scope-vs-AuthRequired layering, both reaction end-state codes, the
omit-emoji orchestration (ordering, full=true, ownership filter, fail-closed
identity), near-miss/threaded get_message flows, zero-egress input
rejections (blank query, empty emoji, malformed open_dm user_ref), missing
provider evidence, the list_members clamp, and loss-free resolve_user
paging.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(e2e): classify the eight slack ops in the product-surface coverage map

The eight new slack.* capability ids were shipped unclassified, which is
what turned the "Validate product-surface evidence contracts" step red
(summary.missing == 8) and cascaded into the Reborn E2E roll-up.

- classifications.tested gains the eight ids.
- Eleven typed ProviderOperationCases: five writes with provider readback
  and cleanup against the never-reset Slack world (edit and delete seed
  their own subject message; remove_reaction seeds its own reaction as the
  connected account; open_dm proves idempotence by re-opening directly),
  get_message driving the history-miss -> thread-fallback path at the
  seeded reply, resolve_user with a natural empty, and proxy-served
  list_members pages (Emulate answers conversations.members POST-only at
  the pinned ref while the guest reads via GET, as real Slack allows).
- get_message's canonical output requires the message, so its empty class
  is the typed model-visible miss: ProviderOperationCase gains
  expected_status ("completed" default, "failed" for exactly this shape)
  and the runner asserts the declared status instead of hardcoding
  completed.

All three gate files pass under the hermetic wrapper (98 passed), and the
product-surface generator reports 131 capabilities, 0 missing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: update guidance for the sixteen-tool slack surface

- reborn-extension-surfaces no longer hardcodes a tool count for the slack
  manifest (the grep recipe is the source of truth) and names github as the
  plain schema-declaring exemplar now that every slack tool is
  standard_op-bound.
- standard-operations.md repoints slack_error_to_standard_code at the live
  crates/extensions path instead of the dead pre-restructure assets/ path.
- tests/e2e/CLAUDE.md stops pinning a literal capability count and defers
  to the generated product-surface report (131 as of this change).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(e2e): declare the expected failed tool result on the get_message empty case

The mock-LLM trace replayer treats any failed capability result as a replay
error unless the step's request_hint names the expected failure — the same
mechanism the provider fault cases already use. slack_get_message_empty
declares expected_failed_tool_result_contains="messaging.unknown_message",
and the operation-case runner forwards it onto the synthesized trace; the
adjacent slack_list_members_empty failure in the provider-2-3 lane was
collateral from the poisoned replay state and needs no change of its own.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tests): widen the seeded slack accounts to the manifest scope union

The merge queue runs the root integration lanes that skip on PRs, and every
one of them seeded the slack account with the pre-widening eleven-scope
union. The manifest's [auth.slack] ceiling now includes reactions:read,
reactions:write, and im:write, so slack installs parked on the auth gate
(BlockedAuth) — the exact provider-scope-gate mechanism this PR documents,
biting its own lockstep fixtures. All six seeds (extension_delivery,
extension_runtime, delivery_user_journeys, tool_call, the slack lifecycle
group scenario, and the QA harness profile that reborn_qa_smoke_scenarios
composes) now carry the full union, and the group scenario's lockstep
comment names the write additions.

Locally green with Docker-backed postgres legs: extension_delivery 23/23,
extension_runtime 25/25, delivery_user_journeys 25/25, group_extensions
16/16, tool_call 38/38, and the bundled-extension-surface QA smoke.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:34:09 +00:00
Josh Ford
318a6e6748 docs: consolidate docs/reborn/ into docs/internal/reborn/ (#7559)
* docs: consolidate docs/reborn/ into docs/internal/reborn/

Move-only migration; no content changes beyond path references. Executes
the follow-up that PR #7259 left open: docs/.mintignore's reborn/ entry
was kept only because the path was load-bearing, and its comment
documented that it moves under internal/ once its consumers move with it.

- git mv docs/reborn docs/internal/reborn (115 files, history preserved)
- rewrite docs/reborn -> docs/internal/reborn across every consumer
  (crate AGENTS/READMEs and doc-comments, .claude/ skills and rules,
  AGENTS.md, CI scripts, reborn-e2e.yml path filters, Dockerfile, tests,
  docs/internal plans)
- fix six relative internal/adr/ links inside the moved tree for the
  added directory level
- drop reborn/ from docs/.mintignore and FROZEN_MINTIGNORE_PATTERNS in
  scripts/ci/docs_publication_boundary.py (the frozen list only ever
  shrinks); internal/ already fences the new location

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: classify tests/dockerfile_runtime_home.rs and shrink boundary self-test fixture

Two CI gates failed on the docs/reborn consolidation and forced decisions
this commit records:

- The Reborn PR test planner failed closed on tests/dockerfile_runtime_home.rs
  (its path-rewrite edit is functional: the test reads the moved deploy doc).
  The file was deliberately unmapped because no lane inventoried it. Decide it
  now: _root_test_partitions() and run-reborn-root-partition.sh both inventory
  it alongside support_unit_tests.rs, so the hermetic root-partition lanes run
  it (they previously ran it nowhere) and a change to it selects its partition.
  With the reader laned, map the two config.hosted-single-tenant*.toml readers
  it owns in DOCKER_RUNTIME_CONFIG_OWNERS — root-test owners select their root
  partition, completing the per-file decision set the planner comments left
  open. docker/process-sandbox-entrypoint.sh stays fail-closed.
- test_docs_publication_boundary.py's subset fixture still listed reborn/ in
  the frozen mintignore list; use the surviving entries.

Verified: both self-test suites pass (77 planner + boundary), the planner
emits a valid selected plan for this PR's full 342-path diff, shell and
Python inventories agree on partition assignment (index 0), and
dockerfile_runtime_home passes (19 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 09:39:06 +00:00
Benjamin Kurrek
88fe2d01e9 refactor(channels): normalize ingress and split reply from delivery (#7477)
* fix(webui): only offer the web-app notification channel when a browser is enrolled

The "Web app" row in the notification-channels picker rendered READY with a
selectable checkbox even with zero enrolled browsers (no push subscription),
so a user could pick a channel that has nowhere to deliver. Selectability and
the pill now follow the account's enrollment count: with no enrolled browser
the web-app checkbox cannot be SELECTED and its pill drops from Ready to
Unavailable, while the nested "Enable notifications in this browser" affordance
shows how to fix it. An already-stored selection stays deselectable (disabled
only when unchecked), so a browser that unsubscribes never leaves a locked-on
checkbox.

The web-push row now owns its device hook (WebPushChannelRow) so the account
status query still mounts only when the row is present; the shared row label
was extracted (renderChannelRowLabel) so every other channel renders its
checkbox inline, unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* wip(ingress): declare authenticated-session ingress recipe

Add IngressVerificationRecipe::AuthenticatedSession — the trust class for a
channel whose caller the host's authenticated transport (T1) already verified,
so it needs no webhook signature. Handle it fail-closed in the two webhook host
sites: no evidence mint in channel_host, and a NotWebhookVerifiable rejection in
the ingress verifier (a session channel mounts no webhook route and must never
be attested verified-inbound through the webhook path).

Incremental checkpoint toward the generic-inbound pipeline (PR2).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(ingress): make channel route_suffix optional, paired with trust class

A channel's ingress mount now depends on its trust class. Webhook recipes (T2:
hmac/shared-secret/none) mount /webhooks/extensions/{id}/{suffix} and MUST
declare a route_suffix; an authenticated_session recipe (T1) is verified upstream
by the host transport, mounts no webhook route, and MUST NOT declare one.

- route_suffix becomes Option<RouteSuffix> (serde default+skip; existing webhook
  manifests parse unchanged into Some).
- ChannelDescriptor::validate pairs the recipe kind with route_suffix presence,
  fail-closed both ways (SessionIngressWithRouteSuffix / WebhookIngressWithoutRouteSuffix).
- Every mount/route-table consumer (active snapshot build + resolve + conflict,
  deployment channels resolve, lifecycle reserved-route check) fails closed when
  a session channel carries no route_suffix — it can never match a webhook route.

Groundwork for routing the web app's authenticated session through the one
generic inbound pipeline (PR2).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(design): materialize the unified channel model (target architecture)

Every channel (web-app, Slack, Telegram) becomes one ChannelAdapter that
implements inbound + outbound/reply + notifications the same way. The only
per-channel variation is the declared ENTRYPOINT (webhook / api_key /
authenticated-session POST) and declared DELIVERY capabilities (reply mode
streaming|batched, optional max_message_chars, markdown, threads). Everything
between entrypoint and delivery is one abstract, channel-agnostic core:
idempotency -> bind(OwnedThread|ExternalRef) -> submit_turn -> durable reply
events -> per-mode reply sink.

Removes the current smell (two post-ingress cores; the web-app special-cased on
both inbound and reply). Records the migration deltas, the trust/security
invariants, and what stays on ProductSurface (the web-app's rich non-messaging
client API). Authoritative target for future agents touching channel code.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(design): no channel-specific code — generic notification setup, kill web-push routes

Strengthen the unified channel model per direction: the hard invariant is that
NOTHING in the codebase is specific to a given channel for inbound, outbound, or
notifications. Every route is generic and extension_id-parameterized; channel
behavior lives only in the adapter (packages/*).

- Web-push enrollment becomes GENERIC channel notification setup: a channel
  declares notifications_require_setup; a generic status/enable/disable surface
  (by extension_id) dispatches to the adapter. VAPID/endpoints/subscription store
  move behind the web-app adapter.
- Delete /web-push/{subscribe,unsubscribe,status} and the web-app-specific
  message route; replace with generic session-inbound + notification-setup routes.
- Extend the specificity gate: zero channel names / channel-specific routes in
  generic crates. Rename web-push -> web-app (id/routes/constants) is in-scope.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(design): notification send is a generic facade over ChannelAdapter::deliver_notification

Channels implement their notification logic in the adapter
(ChannelAdapter::deliver_notification); the DeliveryCoordinator is the generic,
any-caller facade that dispatches to it by extension_id. Routines are one caller
among several (the model's outbound_deliver already is another) — callers own
WHEN/WHAT, never HOW or which channel. Delivery is already adapter-based, so this
is exposing the facade + adapter method, not a rebuild. Setup stays a separate
generic surface (7b). Migration renumbered 8-11 accordingly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(inbound): trust-class + binding enums on the channel inbound contract

The unified-channel-model inbound vocabulary (§12.1 of
docs/internal/design/2026-08-10-unified-channel-model.md):

- ChannelInboundSurfaceRequest carries a trust-class enum
  (VerifiedInbound { evidence } | SessionCaller { caller }) and a binding
  enum (ExternalRef | OwnedThread { thread_id }) instead of bare webhook
  evidence, plus the session transports' requested_model hint.
- ProductInboundEnvelope carries the same pair (ProductInboundTrust /
  ProductInboundBindingDirective); auth_claim() is now Option, with
  require_verified_auth_claim() failing closed for session envelopes on
  every external-ref path (binding requests, command context, projection
  subjects).
- TrustedInboundContext::from_session_caller mints the session-arm context;
  the webhook constructors are unchanged in behavior.
- ProductInboundAck::Accepted gains optional submit-time metadata
  (AcceptedTurnSubmission) and the busy variants gain an optional
  BusyRunSnapshot, both serde-defaulted so ledger rows settled before this
  change still deserialize (pinned by ack_rows_without_submit_metadata_
  still_deserialize).
- ChannelInboundProductSurface gains a default-fail-closed inline-attachment
  admission door for session transports.
- ProductSurfaceRejectionKind gains DuplicateAction and ReplayUnavailable
  for the session-lane replay taxonomy; every exhaustive matcher classifies
  them explicitly.

Mechanical fallout: constructors updated across extension_host, openai_compat,
composition and the integration/parity harnesses; no behavior change on the
webhook lane.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017DectwBVo9eDqV5dhRDGkb

* feat(inbound): owned-thread session lane inside the one inbound core

The webhook core's InboundTurnService gains the session lane (§12.2): the
envelope's binding directive selects the arm, and everything below
TurnCoordinator::submit_turn stays shared.

- OwnedThread prepare: the authenticated caller is the binding authority —
  ownership-probed through SessionThreadService (missing and foreign threads
  are indistinguishable, no existence oracle), never created implicitly, and
  the external binding resolver never runs.
- Session replay probes the exact persisted browser binding-id schemes
  (caller-scoped primary + thread-scoped legacy) so messages accepted by
  earlier builds replay instead of double-accepting; a client action id
  replayed against a different thread fails as ClientActionReplayMismatch.
- The submit tail is lane-parameterized: webui-src/webui-reply ref prefixes,
  the raw client action id as the coordinator idempotency key, and the WebUi
  product context are preserved byte-for-byte for session turns; webhook
  submissions are unchanged.
- Fresh submissions carry AcceptedTurnSubmission metadata; busy outcomes
  carry the blocking-run snapshot; session busy replays report no run
  metadata (the dedicated browser path's exact shape).
- Session skill-activation hooks record between acceptance and submission
  and clear on busy/error, matching the browser path's ordering.
- New session-lane failures (OwnedThreadUnavailable 404,
  ClientActionReplayMismatch 409/duplicate, ReplayUnavailable 409,
  SkillActivationFailed internal, AttachmentLanderUnavailable 503) never
  settle the idempotency ledger.
- submit_inbound_inner admits only user-message payloads from session
  callers, and build_channel_envelope rejects mixed trust/binding arms fail
  closed: webhook trust/pairing machinery can never run for a browser
  message and vice versa.
- CapacityExceeded submissions now surface non-retryable, matching the
  workflow's own settle decision (turn_error_is_retryable).

Covered by the new session_lane suite in inbound_turn_contract (ownership
probe and cross-thread guards sabotage-verified) plus the serde-compat pins
from the previous commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017DectwBVo9eDqV5dhRDGkb

* refactor(inbound): route browser + OpenAI-compat submit_turn through the one core

RebornServices::submit_turn (the SUBMIT_TURN_COMMAND implementation both the
browser route and the OpenAI-compatible transport invoke) now builds the
neutral session inbound request and admits it through the same
DefaultProductSurface core webhook channels ride — durable idempotency
ledger → owned-thread binding → TurnCoordinator::submit_turn (§12.2–3) —
then renders the acks back into the unchanged RebornSubmitTurnResponse wire
shape (fresh Submitted from submit-time metadata; replays as
AlreadySubmitted with the run's current state; busy shapes with their
decision-time snapshots; ledger-replayed busy without run metadata).

The duplicate browser tail is deleted: replay_webui_send_message,
replay_accepted_message, AcceptedWebUiMessage, mark_message_submitted_or_
replay, reconcile_terminal_duplicate, resolve_webui_thread_metadata,
parse_replay_run_id, and the webui binding-id scheme fns now live only as
the session lane of the shared core (the schemes byte-identical, with
legacy replay fallback). The reborn_services module-charter map is updated
in the same change.

Composition wires the durable session ledger
(build_session_inbound_ledger over the extension filesystem, mirroring the
per-extension channel ledgers' mount/bounds/CAS discipline) into every
product-surface instance; standalone/test builds keep the in-memory
default. SessionLaneRejectingBindingResolver guards the session core's
external-ref door fail closed.

The full reborn_services_contract suite (278 tests) passes unchanged
through the re-plumbed path — caller-owns-thread, no implicit thread
creation, client_action_id replay (including legacy binding-id rows),
cross-thread reuse rejection, busy/deferred/steering shapes, attachment
landing, and skill-activation ordering all preserved.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017DectwBVo9eDqV5dhRDGkb

* feat(ingress): generic session-inbound route keyed by extension_id

The web-app-specific browser message route is deleted and replaced by the
generic session-inbound door (§12.4, §8):

- POST /api/webchat/v2/channels/{extension_id}/messages replaces
  POST /api/webchat/v2/threads/{thread_id}/messages. No route names a
  channel; the path extension_id overrides the body and the thread rides the
  body (the caller owns it). Route descriptor policy is unchanged
  (14 MiB body, 60/60s per-caller, TurnCoordinator effect path).
- ProductSubmitTurnRequest/SendMessage carry the optional extension_id; the
  product surface validates it against the new
  SessionChannelDirectory port (declared in
  ironclaw_product_contracts::session_ingress, implemented by the extension
  host over the deployment channel registry — manifest-derived, install-state
  free). Unknown or non-session extensions are 404, indistinguishable from an
  absent route; a missing directory fails closed as 503. Transports that
  predate the parameter (OpenAI-compat) submit under the legacy session
  surface identity, unchanged.
- The web-app manifest declares its entrypoint: inbound = true with the
  authenticated_session verification recipe, no route_suffix (a browser
  request can never reach the webhook mount), conversation_model isolated.
  The manifest-lockstep pin now asserts exactly that.
- The deployment's session channel is advertised to the SPA on
  GET /session (session_channel_extension_id, derived from the registry —
  exactly-one resolves, otherwise none and sends fail closed client-side);
  the frontend plugs it into the generic route and carries no channel name.
- e2e harness + raw-route scenarios read the session channel from
  GET /session; Playwright mocks match the generic pattern.

Caller-level coverage: directory-missing 503 / unknown-extension 404 /
declared-channel admit in reborn_services_contract; the session-channel
directory contract in extension_host; route-table, handler, and charter
gates updated in the same change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017DectwBVo9eDqV5dhRDGkb

* feat(reply): channel-declared reply mode — streaming sinks never batch

The reply model's declaration half (§12.5–7): every channel declares how its
reply sink consumes the durable reply-event stream.

- ChannelDescriptor gains reply_mode = streaming | batched (default batched;
  validation pairs streaming with the authenticated-session entrypoint —
  a webhook vendor has no projection stream to consume).
- The web-app manifest declares streaming: the existing SSE/WebSocket
  projection forward IS this channel's reply sink, exactly as it runs today —
  a consumer of durable reply events, never a replacement (the gateway-events
  layering rule). Its max_message_chars is now undeclared: a streaming sink
  never batches or splits, so the channel is unlimited (§6). Slack and
  Telegram declare batched explicitly; their declared bounds are unchanged.
- ResolvedChannelDelivery carries the declared mode from the same
  generation-pinned snapshot read, and the DeliveryCoordinator gates both
  delivery doors: conversation-reply intents for a streaming channel return
  NoDelivery before any attempt is persisted (the projection stream is the
  delivery), while notification-class sends (BackgroundRunNotice,
  ModelDelivery) flow regardless of mode so the notifications capability
  keeps working. Pinned by
  streaming_channel_conversation_reply_skips_batched_delivery and
  streaming_channel_still_receives_notification_class_deliveries.
- max_message_chars stays adapter-enforced at render time (channel-specific
  splitting is adapter behavior by charter); the declaration remains the
  model-facing hint. The batched sink itself never splits for a streaming
  channel by construction.

No behavior change for any existing delivery: no streaming channel receives
conversation-reply deliveries today, so the gate is the fail-closed
materialization of the current structure.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017DectwBVo9eDqV5dhRDGkb

* feat(notify): ChannelAdapter::deliver_notification + the generic notify facade

The §7a send half of notification generalization (§12.8):

- ChannelAdapter gains deliver_notification(envelope, egress) — the
  channel-specific notification send, defaulting to the channel's ordinary
  delivery (a conversational channel's notification is a message; a
  notification-only channel's whole delivery IS this send). Only the generic
  DeliveryCoordinator calls it, never feature code.
- The coordinator classifies each policy-lane delivery before the request is
  consumed: a run-notification that is not source-routed (it targets a
  notification channel, not the originating conversation) and is not an
  explicitly routed final answer rides the adapter's notification send;
  everything else rides ordinary delivery. Pinned by
  notification_class_delivery_rides_the_adapters_notification_send /
  conversation_reply_rides_the_adapters_ordinary_delivery. Zero behavior
  change for shipped adapters — all three inherit the delegating default.
- run_delivery::notifications is the named any-caller facade over the
  coordinator: notify(target, content) for one explicit catalog-resolved
  channel target, notify_user(user, content) fanning out over
  resolve_user_notification_targets (the picker set). The routine driver's
  own notification internals now delegate to it — one send path, with the
  routine lane as one caller among any number. Callers own WHEN/WHAT, never
  HOW, and never name a channel.
- The coordinator's streaming-reply gate now reads the new lightweight
  ChannelDeliveryResolver::channel_reply_mode lookup instead of performing a
  second full resolution, preserving the single generation-pinned
  resolve_channel_delivery read the OUT contract pins.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017DectwBVo9eDqV5dhRDGkb

* feat(notifications): generalize notification setup behind ChannelAdapter (§7b)

Replace the bespoke /web-push/{status,subscriptions,subscriptions/remove}
routes with one generic per-channel surface keyed by extension_id:
GET/POST /api/webchat/v2/channels/{extension_id}/notifications{,/enable,/disable}.

- contracts: web_push descriptor module deleted; notification_setup module
  (status view + enable/disable command descriptors) and the
  RebornNotificationSetup* wire family replace the RebornWebPush* DTOs;
  body extension_id serde-defaulted (route path is canonical).
- assistant: reborn_services/web_push.rs deleted; notification_setup.rs adds
  ChannelNotificationSetupService + fail-closed Unsupported default +
  AdapterChannelNotificationSetupService dispatching to the channel adapter
  via ChannelDeliveryResolver (unknown extension -> 404, no-setup channel ->
  enabled:true + mutation 400, payload/detail byte bounds enforced).
- delivery coordinator: the streaming-reply gate now keys on the ROUTE, not
  the intent — a notification-routed send (RunNotification + non-live-source
  origin) flows to a streaming channel even with a conversation-shaped
  intent; pinned at the contract tier and by the blocked-fire push journey.
- web-push package: adapter implements the three setup operations over the
  slot runtime (scope byte-identical to the retired product service; detail
  carries vapid_public_key/subscription_count/subscriptions with
  endpoint_digest correlation).
- composition: wires AdapterChannelNotificationSetupService over the channel
  delivery resolver; WebPushComposition handle family deleted (the slot
  install inside assemble_web_push is now the single consumer).
- webui: route descriptors/router/handlers swapped to the generic surface;
  CONTRACT.md route table + outbound charter row updated.
- frontend: api.ts gains getNotificationSetupStatus/enable/disable keyed by
  extensionId; web-push.ts -> device-push.ts and useWebPushDevice ->
  useDevicePush re-read the channel-opaque detail; the notification panel's
  device row is matched by the GET /session-advertised session channel id —
  no channel name remains in the frontend; webPush.* i18n keys renamed
  devicePush.* across all 11 locales.
- tests: 5 new setup-dispatch contract tests + streaming-notification
  regression pin; product-api round-trip and delivery journey rewritten onto
  the generic surface; frontend suites updated (1241 pass).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(channels): rename web-push -> web-app and retire the old spelling (§12.11, §13)

Product identity rename: extension id / channel name / catalog target id are
now 'web-app'; package dir crates/extensions/packages/web-app (crate
ironclaw_web_app_extension); domain crate crates/domains/ironclaw_web_app;
WEB_PUSH_* constants -> WEB_APP_*, WebPush* types -> WebApp*. PROPOSAL §5
tree updated (check-target-tree: 66/66 OK). Space-separated 'Web Push'
protocol prose stays — the protocol keeps its RFC name; the CHANNEL does not.

Persisted coordinates deliberately keep pre-rename bytes, each commented in
place and pinned by the new gate's allowlist:
- secret-store credential handle value 'web_push_vapid' (renaming = VAPID
  rotation = every existing browser subscription breaks cryptographically);
- enrollment document path /web-push/subscriptions.json plus composition's
  /web-push per-user mount alias (the alias resolves to a physical subpath;
  renaming would orphan enrollments);
- binding-ref grammar mints web-app/v1/ and decodes legacy web-push/v1/
  forever (regression test added).
Documented residue, no migration: stored notification-channel selections
carrying the old 'web-push' target id render Unavailable until re-selected
(population ~QA-only; the channel shipped 2026-08-09).

The install-catalog hide for the host's own surface is no longer an id
match: is_builtin_host_surface consults the SessionChannelDirectory (the
manifest-derived authenticated_session fact), failing OPEN on an absent
directory; the production round-trip test covers the hidden-listing behavior
end-to-end.

Enforcement (§13): new architecture gate
reborn_web_push_vocabulary_retired.rs pins web-push/web_push/WebPush/
webPush/WEB_PUSH at zero occurrences across crates/ (frontend sources
included), tests/integration/, and skills/, with an exact-term shrink-only
allowlist over the five persisted-compat files, a stale-sanction check, and
an assertion that the session + notification-setup routes stay
{extension_id}-parameterized. The specificity gate's web-app carve-out doc
records the rename.

E2E journey vocabulary renamed on both the Rust and Python sides
(case ids, test names, delivery-target enum member).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(deps): bump lru 0.18.1 -> 0.18.2 (RUSTSEC double-free advisory)

cargo-deny advisories began failing on every head when the lru advisory
published; 0.18.2 is the fixed release (lru-rs#238).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(ci): raise composition arc_dyn ceiling 816 -> 818 (unified channel model)

Two net-new dyn seams wired at assembly, both genuine inversion ports:
the session-inbound lane's SessionChannelDirectory + durable
IdempotencyLedger, and the generic ChannelNotificationSetupService —
offset by the deleted WebPushComposition handle family. Observed on the
merged branch: 833 = 818 + 15 tolerance exactly, no slack.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): migrate the last pre-unification callers and re-pin ratcheted ceilings

Everything here is fallout of surfaces this PR deliberately changed:

- smoke + composition webui_v2_e2e: the raw-HTTP browser-send helpers now
  discover the session channel from GET /session and post the generic
  /channels/{extension_id}/messages route (thread_id in the body) — the
  same flow the SPA ships.
- InboundUserMessageDispatch::Accepted is boxed (clippy large_enum_variant:
  the merged InboundTurnOutcome grew past the threshold; a rejection stays
  slim).
- journey coverage: a channel whose ingress verification is
  authenticated_session has no webhook mount — its inbound IS the WebUI
  session route, so it maps onto the webui journey evidence instead of
  demanding a per-channel label.
- attachments no-lander test pins the sharpened
  AttachmentLanderUnavailable variant (503, never settles the idempotency
  reservation) instead of the old generic rejection.
- body-limit contract test pins webui.v2.session_channel_message (14 MiB)
  after the route rename.
- contracts size ceilings re-pinned to measured merged values with
  rationale: extension_contracts 8_157 (AuthenticatedSession trust class,
  reply modes, §7b setup adapter surface), product_contracts 16_119
  (trust/binding enums, SessionChannelDirectory, setup descriptors + wire
  family), host_api 19_003 (doc churn referencing the renamed crate).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): keep the catalog target id persisted, migrate e2e callers, pin notice routing

Review triage for #7477 (IronLoop + multi-agent review).

Persisted identity (IronLoop Medium, also flagged by the review): the
catalog target id now keeps its pre-rename `web-push` bytes. The
notification-channel picker stores its selection as target ids in each
user's communication preferences, so this is a persisted per-user identity
exactly like the VAPID handle, the mount alias, and the binding-ref prefix —
the three the PR already kept. Renaming it resolved every stored selection to
Missing and dropped those users from notification fan-out. Applying the PR's
own rule uniformly removes the documented residue rather than shipping it;
the integration test now pins the split (target id `web-push`, channel
`web-app`) so the two can't be conflated again.

E2E callers of the retired send route (the browser-lane CI failure): four
Playwright interceptions and one API helper still targeted
/threads/{id}/messages, so failure injection never fired. All now use the
generic /channels/{extension_id}/messages route, and every mocked GET
/session advertises session_channel_extension_id the way a real deployment
does — the SPA fail-closes without it.

deliver_notice asymmetry (review Medium, correctness): confirmed correct and
now pinned. Notice-class intents are source-routed, so none is ever
notification-routed and `deliver`'s carve-out cannot apply; for a streaming
channel the originating conversation IS the projection stream, and Retract /
React have no counterpart there (the adapter reports both unsupported). The
new test drives all seven notice intents plus the notification path in one
breath so they cannot drift.

Session surface is built once (review Low/Medium, hot path): submit_turn
rebuilt DefaultProductSurface plus ~5 Arc'd services per browser message;
every input is an immutable builder-wired Arc, so it is memoized behind a
OnceLock.

Docs the rename sweep left stale: tests/CLAUDE.md cited a test name that
never existed, the web-app README cited a VAPID handle value that doesn't
exist (the constant deliberately keeps the old value), the extensions
package-inventory row still called the channel outbound-only with no ingress,
and a merge left a duplicated comment block in inbound_turn.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(review): cover the notify_user fan-out and the web-app adapter's setup errors

Closes the two review findings that were test gaps rather than follow-ups —
the repo's own rule is that production-wired behavior ships with its
caller-level test, and both of these were new production surface with none.

notify_user (crates/product/ironclaw_assistant/tests/run_delivery_contract.rs):
the driver only ever calls the single-target notify, so the fan-out loop was
untested. Two contracts now pinned through the real facade: an unconfigured
user yields an empty result rather than an error, and a target whose channel
no longer resolves surfaces its own Err while its healthy sibling still
delivers (asserted at the adapter, not just the return value).

web-app adapter (crates/extensions/packages/web-app/tests/notification_setup_contract.rs):
the generic service tests drive a scripted adapter, so the real parse →
validate → store path had no error coverage. Six cases: non-JSON document,
missing key material, undecodable base64url keys, an endpoint on an
undeclared push host, a malformed unenrollment document, and every setup
operation with no runtime installed. Each asserts the store was never
touched, so a rejected payload can't leave the browser believing it is
enrolled with no server record behind it. All six passed on first run — the
arms were correct, just unproven.

observer.rs: the repeated fallible-from_envelope fallback is now one
`degradable_binding` helper — but only for the two sites that genuinely
merge 'no request' and 'no binding' into the same degrade. The delivery path
still propagates (a send with no binding is a fault), and the
rejection-hint path still distinguishes them (posted nothing vs handled by
staying silent); both reasons are documented on the helper rather than
flattened away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channels): close the audit findings — dead per-channel routes, gate gaps, session-surface fallback

Fallout from the four channel-specificity audits, plus two defects those
audits found in MY OWN branch that would have kept CI red.

Introduced by this branch, now fixed:
- a contracts-crate doc comment named Slack/Telegram, an untracked
  specificity-gate violation (the gate strips #[cfg(test)], and this was
  production code)
- deleting the dead telegram frontend modules made five ALLOWLIST rows
  stale; that list is an EQUALITY ratchet (both <= and >=), so the rows are
  removed and the baseline drops 117 -> 112

Pre-existing, found by the audit:
- 871 lines of orphaned telegram-setup frontend code holding the only two
  channel-named URL literals in the SPA, pointing at routes the backend no
  longer serves. Deleted.
- a dead ironclaw_assistant -> ironclaw_web_app dependency edge: a generic
  product crate holding a compile-time edge to one channel's domain crate
  for nothing
- telegram_extension_gates.rs still documented the retired per-channel
  pairing route as live

New gate (§13's structural half): no source file may name a channel in a
/api/webchat/v2/channels/ route. It scans SOURCE rather than the descriptor
table, because the defect it exists to catch lived entirely in callers the
route table never knew about — the table was clean while 871 lines of
channel-named client code sat beside it. Sabotage-tested against the deleted
file, which it flags. Placeholders and test fixtures are exempt (tests may
name channels, per the extension-runtime overview).

Session-surface regression, found by CI on the composition e2e suite:
a deployment that installs no channel extension had NO route to submit a
browser turn, because the old /threads/{id}/messages route is gone and
/session advertised no channel id. That is a supported deployment shape
(assemble_web_app treats its slot as optional), so browser chat must not
depend on an installed extension. BUILTIN_SESSION_SURFACE_ID now lives in
product_contracts::session_ingress; composition advertises it when no channel
claims the surface, the product gate accepts it, and WebuiServeConfig
defaults to it rather than None — the transport defaulting the surface to
'absent' was the actual defect. 15/15 composition e2e tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): channel output has two axes — reply and delivery

Design record for the follow-up train. The unified channel model unified the
pipeline; this reshapes the contract it drives. Recorded here rather than in
the follow-up PR so the decisions survive the conversation that produced them.

The root finding is not the eleven-method ChannelAdapter — that is the
symptom. It is that two independent concepts share one vocabulary:

- reply    = answering the run's input, SOURCE-routed, never without a run
- delivery = reaching someone out-of-band, TARGET-resolved, runs optional

They are orthogonal, not alternatives: one run can stream an answer into an
open tab AND push a notification because the user is not looking. Dispatching
on intent rather than on this axis already produced a real defect on this
branch — a gate prompt is a reply when a human is in the thread and a delivery
when a 3am routine is blocked, and keying the streaming skip on the intent
silently dropped the second case.

Decisions: three manifest sections (ingress/reply/delivery) replacing the
inbound/outbound/notifications booleans; OutboundRoute plus two transport
enums so nonsense combinations are unrepresentable; DeliveryOrigin keeping
model-chosen targets from inheriting user-configured trust; a streaming
delivery returns a projection cursor as evidence instead of NoDelivery,
closing an audit hole where browser replies produce no record at all;
activate/cleanup become an ingress-registration recipe; the attachment fetch
moves AFTER the ack (the durable write currently depends on it, which is what
puts a vendor round-trip on the webhook deadline path); enrollment moves
host-side with no adapter method, keeping one generic pre-storage check that
exists to prevent an SSRF primitive.

Five open questions and a six-step sequencing table are recorded; step one is
the smallest and closes both the no-op and the audit hole.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): repair fmt, the last retired-route callers, and the contracts ceiling

Three red checks on 3e8d40bc74, all mechanical:

- cargo fmt: the export-list edit in ironclaw_assistant/src/lib.rs left an
  unformatted line (fmt reflows multi-item use blocks; I removed a symbol
  after the last fmt pass).
- composition webui_v2_serve: two tests still posted to the retired
  /threads/{id}/messages route and got 404. Migrated to the generic
  /channels/{extension_id}/messages with thread_id in the body. One of them
  exists specifically to pin 'the shape api.ts builds', so it has to track
  the SPA; the other pins the 14 MiB descriptor cap against Axum's 2 MiB
  Json default, which is unchanged by the route move.
- product_contracts size ceiling 16_119 -> 16_132: BUILTIN_SESSION_SURFACE_ID
  plus its doc, the built-in session surface that keeps the generic session
  route from depending on an installed channel extension.

Verified: cargo fmt --check clean. The suites are left to CI — another agent
is working in this worktree and a local battery would block it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* wip(channels): reply and delivery become two declared axes

Contract half of the channel-output redesign
(docs/internal/design/2026-08-11-channel-adapter-contract.md §1-§3).
INCOMPLETE — see the PR body / handoff for the stale-reference list.

- ChannelReplyMode -> ReplyTransport{Stream,Message} +
  DeliveryTransport{Push,Message}. Two enums so Stream-for-delivery and
  Push-for-reply are unrepresentable; a third transport joins as a
  variant rather than a reshape (§10.5).
- [channel.reply] and [channel.delivery] manifest sections replace the
  inbound/outbound/notifications/notifications_require_setup booleans
  and reply_mode. Absence of a section means the axis is unsupported,
  so a declaration can no longer say *that* a channel does something
  without saying *how* (§2, §9).
- max_message_chars moves from [channel.presentation] to
  [channel.reply]: a split bound is a property of the reply transport
  and is meaningless for transport = stream.
- ChannelDeliveryResolver::channel_reply_mode ->
  channel_reply_transport; notifications_require_setup ->
  requires_enrollment.

Fixes a live defect found while reshaping, not a rename: the stream
reply/session-ingress pairing check sat inside 'if let Some(ingress)',
so a channel declaring a stream reply with NO ingress validated
silently. The check now sits outside that block and the no-ingress arm
is pinned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(channels): one derived projection for channel output facts

Replaces the three loose `max_message_chars` scalars an earlier pass in
this branch added beside `channel_presentation` at each carrier.

The bound legitimately moved out of [channel.presentation] into
[channel.reply] (it is a property of the reply transport, and is
meaningless for transport = stream). But pulling it out of a struct that
was ALREADY threaded through three carriers turned 'one type threaded
three times' into 'one type plus a loose scalar threaded three times' —
re-declaring the value at three layers with nothing keeping them in
agreement (.claude/rules/architecture.md §3).

ChannelOutputFacts is the fix: presentation + the reply bound, assembled
once by ChannelDescriptor::output_facts(), threaded exactly where
ChannelPresentation was. One manifest home per field, one projection,
carrier field count unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* wip(channels): three traits + declarative vendor-call recipes

Steps 2, 4, and 6 STARTED, NOT FINISHED — the ~20 consumers of the
removed ChannelAdapter are not yet updated. Does not compile.

- ChannelAdapter's 11 methods -> ChannelIngress::receive (async),
  ChannelReply::send_reply, ChannelDelivery::deliver/list_targets, held
  as ChannelSurfaces { ingress, reply, delivery }. A None is the same
  fact as a missing manifest section. A stream-reply channel implements
  no reply half at all.
- ChannelVendorCallRecipe: per-channel data, generic execution. Replaces
  activate/cleanup as [channel.ingress.registration]/[deregistration]
  and the attachment fetch as [channel.attachments], run post-ack.
- Telegram's setWebhook/deleteWebhook become manifest data; both method
  bodies go to zero.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* wip(channels): wire the three-trait split, the two-axis router, and host-owned enrollment

The workspace compiles again. `cd167ab8ed` defined ChannelIngress /
ChannelReply / ChannelDelivery and deleted ChannelAdapter without updating
~60 consumers; this wires them and lands the contract changes those consumers
were waiting on.

Step 6 — the split, with the check that earns it. ChannelSurfaces replaces
`Arc<dyn ChannelAdapter>` on ExtensionBindings, ActiveExtension,
DeploymentChannelBinding, ResolvedChannelDelivery and the composition binding.
`check_binding` now proves each `[channel.*]` section against its implementing
half at activation: webhook ingress <-> ingress half, `transport = "message"`
<-> reply half, `[channel.delivery]` <-> delivery half. Two axes are required
ABSENT and that is the point — a `stream` reply is published by the host and
`authenticated_session` ingress is normalized at the session door, so binding
a half there is dead code that reads as live. Without this check the three
Options would be a second copy of manifest facts with nothing keeping them in
agreement (architecture.md §3); with it, declaration and code cannot disagree
past activation. web-app now binds delivery ONLY.

Step 1 — OutboundRoute. The axis is computed once in DeliveryCoordinator from
the resolved routing decision and threaded through the drive chain in place of
the `as_notification` bool. The streaming gate keys on the route, not on
DeliveryIntent::is_conversation_reply — which is the exact conflation that
silently dropped blocked-routine pushes. A stream reply is no longer a silent
NoDelivery: `record_stream_reply` persists a full attempt row and returns
StreamDelivered { cursor }, so "was the user's answer delivered?" has one
answer and web-app stops being invisible in delivery audits (§4.1, §10.4).
Evidence is the projection ref the turn already wrote — §4.4's verify, not own.

Step 2 — activate/cleanup are gone. `[channel.ingress.registration]` /
`[channel.ingress.deregistration]` are executed generically by
`channel_vendor_calls`: `{handle}` substitution from non-secret config,
unresolved placeholders left for egress credential injection, body_credentials
forwarded by handle, single-pass substitution, JSON keys never templated.
Telegram's two method bodies became zero lines. Their assertions move with the
behavior to the host executor.

Step 5 — enrollment is host-owned. `ironclaw_auth::delivery_registrations`
stores an opaque, size-bounded document keyed (tenant, user, extension) with
the one security-critical check generic and pre-storage: the endpoint must
target a host declared in `[[channel.egress]]`, read from the same resolved
manifest egress policy enforces with. Without it enrollment is an SSRF
primitive. Placement is ironclaw_auth over ironclaw_outbound because the
adapter-facing view must live in extension_contracts and auth already names
it. Registrations ride the envelope and the adapter reports prunes — it holds
no store. A channel with zero registrations is a resolvable "no target" before
any adapter call. Pre-§8 documents migrate forward on read; `/web-push/
subscriptions.json` and its mount alias keep their exact bytes.

Still to do: --all-targets (test doubles, integration suites), step 4's
post-ack attachment fetch, docs, ratchets, PR body.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* wip(channels): carry the three-trait split through every test double and fixture

`cargo check --workspace --all-targets` is clean. The prior commit got the
lib and binary compiling; this carries the same change through the test
surface, which is where the behavioural pins live.

- Lifecycle: the three deleted Telegram `activate`/`cleanup` tests are
  re-pinned against the generic recipe executor, verbatim in what they assert
  — the bot token travels as a HANDLE and never as bytes, the shared secret
  rides `body_credentials` so the host inserts its VALUE at the manifest's
  declared pointer, the rendered body carries `url` but never `secret_token`
  nor the handle name, a missing config value and a vendor 5xx both fail
  activation, and deactivation calls deleteWebhook. Adds the arms the adapter
  tests could not reach: deregistration is best-effort and cannot strand a
  deactivation, and a channel declaring no recipes makes no vendor call.
- Binding: per-axis tests drop or add exactly one half against a manifest
  declaring the other two, so a failure names the axis. Plus the two absences
  that are the point — a `stream` reply and `authenticated_session` ingress
  must bind NO half, because the host publishes and the session door
  normalizes.
- web-app: `notification_setup_contract` becomes
  `registration_parsing_contract`, re-aimed at where the behaviour went.
  Endpoint admission and storage bounds are generic now and pinned in
  `ironclaw_auth`; what stays this package's is interpreting the opaque
  document at delivery. New coverage the old shape could not express: one
  unusable registration is pruned WITHOUT costing its siblings their
  notification, because the host owns the list and the adapter no longer
  reads its own store.
- outbound_delivery_contract: the §7b adapter-dispatch block becomes §8
  enrollment coverage. The security-critical arm is explicit — four hostile
  endpoint shapes (undeclared host, http, userinfo smuggling, suffix
  lookalike) are refused BEFORE storage, and the recording store proves
  nothing was written.
- Test doubles across assistant/host/composition/integration bind the halves
  their fixture manifests declare; the Acme fixture gains reply+delivery over
  one shared `send`, as a conversational vendor really behaves.

Reverted in this commit: an in-flight change making `receive` return a
COMPLETE message (attachment bytes + conversation context) so the two fetch
handles could leave the trait entirely. The design is right and is written up
for a fresh pass; landing it 70% done would repeat the breakage this branch
started from.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(channels): return complete inbound messages

Make ChannelIngress::receive the only vendor ingress call and return complete attachments and conversation context through manifest-restricted egress. Delete the host/product late-fetch callbacks while preserving exact byte validation, attachment budgets, policy reconciliation, batch recovery, and ack-after-commit semantics.

Keep Slack URL/history and Telegram two-hop file validation inside their packages. Also carry the selected manifest egress credential into host-owned lifecycle calls so Telegram setWebhook receives its declared token injection; the production libSQL journey covers activation, inbound bytes, dedupe, reply, and refresh.

* docs(channels): align capability contracts and ratchets

Document the ingress/reply/delivery capability split, host-owned session/stream modes, complete receive boundary, and delivery-registration ownership across the contract, package, product, and runtime guides. Amend the design record with the measured pre-ack order and the reasons a declarative attachment recipe cannot model Slack or Telegram safely.

Recapture the measured contract ceilings at extension_contracts 8,594 and loop_contracts 13,307, and the composition budget at 41,751 LOC / 837 Arc<dyn> sites. Widen the retired web-push scanner to Cargo.toml, E2E, and Python while preserving exact persisted-coordinate exceptions.

* fix(cli): warn when session channel is unavailable

Emit an operator-visible serve warning on the documented tracing target when composition resolves no session channel. Capture both the target and message in a regression test so the authenticated session route cannot disappear silently.

* fix(channels): align post-merge contracts

* fix(channels): finish normalized channel boundaries

* fix(channels): harden egress and notification setup

* fix(channels): make stream evidence and wiring explicit

* fix(review): close the verified audit and review findings

Critical — OpenAI-compat lane restored. submit_turn hard-404'd
extension_id: None while both compat workflows send None (the documented
lane: headless SDK clients cannot learn a channel id from GET /session).
Restores the None => BUILTIN_SESSION_SURFACE_ID arm exactly as the wire
doc specifies, keeps every Some strictly directory-gated, and pins the
builtin id as not route-addressable. Two-sided seam pins: the surface
half in reborn_services_contract, the caller half in the compat handlers
contract, plus a drift pin equating the contracts constant with the turn
kernel's WEBUI_SOURCE_CHANNEL.

Delivery reliability:
- Adapter report gate is coverage, not equality: vendor chunking reports
  one outcome per chunk (conformance legalizes >=); requiring == settled
  fully delivered chunked replies Unknown/Failed and invited duplicate
  resends. Under-reporting still settles Unknown, never retried.
- A crash-orphaned Prepared row re-validates and re-authorizes on replay
  (no vendor egress happened; the claim CAS stays the one transition
  authority) instead of wedging AlreadyInFlight forever. A revoked
  replay rejects via a distinct audit row, leaving the stable row for
  the claim fence. Sending-row recovery stays explicitly fail-closed
  per OUT-6; startup wiring needs a status index and is deferred with
  rationale in the PR discussion.
- Working-indicator notice refs are monotonic per run: the stable ref
  made every post-gate re-post settle AlreadyDelivered, so the
  indicator vanished after a gate cycle (and nudge refs reset the same
  way). First-post bytes are preserved.
- Reply-context store failures are logged before mapping to the unit
  port error (both the host source and the coordinator site).
- Partial web-app fan-out reasons carry the failing cause; the push
  status classification matrix (401/403/413/429/5xx/transport/mixed)
  is pinned.

OAuth binding compensation follows the credential: a terminally-failed
lifecycle activation revokes the extension credential, so the identity
binding now rolls back on exactly that arm instead of committing a
"connected with no usable credential" state; retryable dispatch
failures keep the binding (the credential remains valid and the replay
path never re-runs the hook). ContinuationDispatchFailure carries the
terminalization fact to the callback site.

Browser push enrollment un-broken (two-sided wire drift): the client
read the retired flat web-push detail shape while the backend emits
registrations/bootstrap — enroll was permanently dead and enrolled
browsers derived "another account". The client now reads the canonical
shape, project() emits per-registration endpoint_digest (lowercase hex
SHA-256 via ironclaw_common::hashing, matching endpointDigestHex), and
incomplete digest coverage reads correlation-unavailable, never
"not mine". Pinned by vitest parsing tests, the api mock now mirroring
the real shape, and a digest assertion in the integration round trip.

Session-ledger and feedback correctness:
- LlmConfigServiceError::Internal no longer settles a durable permanent
  PolicyDenied: a backend fault is transient, and the same
  client_action_id succeeds after recovery (pinned).
- Duplicate/replay rejections settle silently again instead of
  rendering the false DM-only command copy.
- ProductInboundTrust / ProductInboundBindingDirective persist
  snake_case tags (pinned before the first ledger row ships).
- session_inbound_request and sibling sites use the cause-logging
  internal_from constructor instead of dropping constructor errors.
- Attachment kind classification case-folds MIME at the boundary.
- The persisted webui-src/webui-reply prefixes are defined once.

Test-support honesty: the harness StaticSecretStore stores what
put_if_absent claims to create (and leases remember their handle), so
first-time VAPID bootstrap flows are testable; the SSRF reserved-key
strip in delivery_registrations is pinned; the session-channel catalog
hiding now has a directory-present test; web-app manifest label reads
"Web app".

Refuted with evidence (no change): the outbound record layer is already
CAS insert-if-absent + first-write-wins with the lost-race shape pinned
in outbound_state_store_contract; the WebUI session route needs no
route-level channel check because the product surface enforces the
directory fail-closed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gates): reword a test comment out of the retired vocabulary scan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): stack headroom for the crate-bucket lane; evidence-driven final-reply pin

The composition-core bucket SIGABRTs on Linux: the composed-runtime skills
turn overflows the 2 MiB default test-thread stack (first seen in
the_model_runs_a_skills_script_from_the_workdir_the_body_advertises).
Give the bucket lane the same 8 MiB headroom the integration lanes document;
deep subtrees stay Box::pin'd — this is headroom, not a substitute.

The webui grouping pin asserted a hardcoded 'isFinalReply: false' literal;
the stream-evidence rework made the marker evidence-driven
(isFinalReply: finalizedText from the durable projection's finalized bit).
Pin the derivation — the same in-flight guarantee, stated against the
stronger shape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 18:46:30 +00:00
Benjamin Kurrek
203cbecf15 ci: tolerate sccache install outages (#7552)
* docs: design sccache install fallback

* ci: tolerate sccache install outages

* ci: classify shared sccache action tests
2026-08-12 18:20:43 +00:00
Benjamin Kurrek
173f078bba Gate & ratchet audit: full-inventory report, five fail-opens armed, dead gates deleted (#7373)
* test(architecture): drop the dead ironclaw_storage row and arm the substrate list

Gate-audit finding (open-and-shut): SUBSTRATE_CRATES in
reborn_composition_boundaries.rs carried three rows of rot, all invisible
because the loop's `let Some(..) else { continue }` silently skipped any
entry that resolves to no workspace package:

- "ironclaw_storage": no such package exists (verified against
  `cargo metadata --no-deps`; the only MISSING name of the 29 listed).
- "ironclaw_approvals" and "ironclaw_assistant" were each listed twice.

The silent skip is replaced with a panic naming the stale entry, so the
list can no longer rot invisibly. Verified by sabotage: adding a bogus
"ironclaw_zzz_probe" row now fails the test with
"is listed in SUBSTRATE_CRATES but is not a workspace package"; the
clean list passes (23/23).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): prune the dead sanctioned path from the specificity gate

Gate-audit finding (open-and-shut): SANCTIONED_PATHS in
reborn_extension_specificity.rs still exempted
`extension_host/extension_installation_store.rs` — a file deleted by
#6430. No scanned path matches the fragment (verified with rg across
crates/), so the entry exempted nothing; it is also the one exclusion
surface in this gate with no staleness check, which is how it outlived
its file. Full specificity suite green after removal (8/8).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): drop the v1 ironclaw_gateway/static exclusions from the telegram gates

Gate-audit finding (open-and-shut): both cross-tree scans in
telegram_extension_gates.rs still carved out `ironclaw_gateway/static`
— the v1 monolith's embedded UI, whose crate was deleted with the src/
monolith (no crates/*/ironclaw_gateway directory exists). The exclusions
matched nothing; scans now cover the whole tree with no dead carve-outs.
Suite green after removal (12/12).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): make the dto-collapse gate's header describe the gate that exists

Gate-audit finding (open-and-shut doc rot): the module doc still
described the pre-#6447 freeze design — a dangling doc-link to
FROZEN_COLLAPSE_DTOS (renamed RETIRED_COLLAPSE_DTOS in #6447), a
promised delete-without-trimming failure and an empty-allowlist
assertion that do not exist in the file, and a named owner for a
collapse that completed. The mechanism itself is armed and untouched;
the header now describes the permanent zero-gate it became, and records
the two originally-frozen names that deliberately left governance
(CapabilityOutcome via #6299 deletion, CapabilityDispatchRequest blessed
as the canonical port type). Suite green (2/2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): repoint the manifest-reparse allowlist note at the colocated asset

Gate-audit finding (open-and-shut doc rot): the BundledAsset allowlist
entry's justification still cited include_str! of
assets/memory_native/manifest.toml — a path retired when WS2 (#7037)
colocated packages; the live include in memory_native_extension.rs
reaches crates/extensions/packages/memory-native/manifest.toml. Comment
only; the gate's mechanism and counts are untouched. Suite green (2/2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): give the memory-vocabulary gate the partial-tree floor its twin has

Gate-audit finding: reborn_memory_retired_vocabulary.rs had no
MIN_SCANNED_FILES floor, unlike its explicit twin
reborn_retired_taxonomy.rs — so a partially-moved tree (the CHECKLIST
WS0 / #6963 'green while measuring nothing' shape) would scan a
fraction of the files and still report the vocabulary clean. The gate
was in fact born with an already-dead sanctioned path (its own header
records this), so the rot class is not hypothetical for this file.

Adds the same 500-file floor (real count ~4000), asserts it in the main
gate, and pins the premise on a fixture: a 10-file partial tree scans
clean and is rejected by the floor. Suite green (4/4); clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): close the transport gate's nested-use-group fail-open

Gate-audit finding (sabotage-verified): product_symbols_in's braced-group
branch closed at the FIRST '}' (group.find('}')), so a nested group —
use ironclaw_assistant::{m::{X}}; — truncated mid-element and recorded
zero symbols. Probed live before the fix: appending
use ironclaw_assistant::{zzz_audit::{ZzzProbe}}; to webui's lib.rs left
transports_name_only_the_frozen_residue_of_product_symbols GREEN, while
the plain-path spelling of the same import correctly failed. The same
truncation dropped qualified elements inside flat groups
({qualified_module::X} recorded nothing).

The group branch now does a balanced-brace walk, splits elements at
depth-0 commas only, and records a qualified/nested element's leading
path segment — the same key the single-path branch records for
ironclaw_assistant::module::X. Flat-element semantics are byte-for-byte
unchanged, so the frozen 100-row webui inventory is untouched (suite
green 6/6 on the live tree). Regression fixtures added to
import_scanner_reads_symbols_out_of_real_use_shapes; the original
sabotage now fails with the gate's own message (re-verified).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: delete check-e2e-matrix-files.sh — a gate for a workflow that no longer exists

Gate-audit finding (provably inert): the script's default target is
.github/workflows/e2e.yml, deleted when the v1 e2e suites were retired
(git log --diff-filter=D shows the removing commit); no workflow, script,
hook, doc, or guidance file references check-e2e-matrix-files.sh
(verified with rg across the repo including .github and .githooks).
A checker nothing runs, pointed at a file nothing provides, is dead
weight that reads as coverage.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: delete the measured-broken check-boundaries.sh and its guidance references

Gate-audit finding (provably inert, previously measured): crates/AGENTS.md
recorded on 2026-08-05 that the script fails on a clean tree (check 5
false-positives on live test files) and that checks 1/2/3/6 target the
deleted v1 src/ tree, passing vacuously. No workflow or hook runs it; its
only callers were guidance files, two of which claimed it 'enforces'
root-tests feature gating — an enforcement claim the skill-maintainer
rules forbid for a check nothing executes.

Removed the script and every live reference: the crates/AGENTS.md warning
row becomes a tombstone note; the testing skill + exemplar reference drop
the false enforcement parenthetical; the architecture-review skill's
Verify line drops the dead command; deslop-reborn's allowed-tools drops
the permission; .coderabbit.yaml's driver-leak instruction now points at
the live enforcement (reborn_persistence_driver_boundary). Two dated
docs/internal/ plan snapshots keep their historical mentions.

Verified: python3 scripts/ci/check-guidance.py OK (2084 path references)
and its self-test OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(product): stop hardcoding charter sub-owner counts in the family map

Gate-audit finding (stale prose): crates/product/AGENTS.md said
'19-sub-owner reborn_services charter map' — the enforced map has had 20
sub-owners since #7235 added the inspector row (counted from the live
table). Rather than chase the number, drop both inline counts: the
owning maps and their gates are authoritative, and the re-verify
commands are already inline (skill-maintainer rule: no counts without a
regeneration recipe). check-guidance.py OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): correct the scanner-fixture file's name-filter claim

Gate-audit finding (doc rot with a false coverage claim): the header
said naming the FILE reborn_* makes code_style.yml's
'cargo test -p ironclaw_architecture_tests reborn' see it — but that
argument is a test-NAME filter (the measurement is documented in
reborn_contracts_vendor_census.rs), and none of this file's test fns
contains the substring, so that smoke lane runs 0 of them (11 collected
by the full plan). Comment-only; the note now records the real semantics
so file names are not trusted for lane coverage. Suite green (11/11).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): gate & ratchet audit report + proposed preflight gauntlet

The audit the owner asked for after PR #7157 went red six times across
four gates: every architecture-test gate, module charter, CI script, and
committed baseline inventoried with a verdict and evidence; the handful
worth acting on ranked by friction x weakness; the CI-ergonomics analysis
(why failures surface one per ~1h round-trip: no --no-fail-fast anywhere
in CI, cancel-in-progress on push, sequential fast-checks steps —
measured: two broken gates report 1 failure in 18s under the CI shape vs
both in 211s with --no-fail-fast); and the sabotage log for every probe.

scripts/preflight-gates.sh is the concrete pre-push proposal: the
deterministic-gate classes only (script gates ~10s + architecture suite
--no-fail-fast + changed-crate charter tests), covering all four #7157
gate classes locally in one command. Unwired — nothing invokes it.
Validated end-to-end on this branch: exit 0, 'every deterministic gate
green', 402.8s including gate-binary recompiles.

Placement verified: python3 scripts/ci/docs_publication_boundary.py OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(planner): classify preflight-gates.sh and the deleted check-boundaries.sh

The gate audit's own PR hit the planner's fail-closed arm — 'unmapped
test or CI path: scripts/check-boundaries.sh' — exactly the class the
arm exists to force a decision on (and the audit's report documents).
Per the PR_STATIC_CONTROL_PATHS membership rule (no Reborn test lane
exercises either file):

- scripts/preflight-gates.sh — the audit's proposed local pre-push
  gauntlet; referenced by no workflow.
- scripts/check-boundaries.sh — deleted by the audit; the entry lets the
  deletion diff (and any revert) classify instead of failing every
  downstream Reborn lane.

Verified: the planner now produces mode=selected with the
architecture-misc bucket for this branch's diff, and
python3 scripts/ci/test_reborn_pr_test_plan.py is OK.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): add the fold-tripped asymmetric-tolerance exhibit to the audit

The strongest single exhibit for shortlist item 2, contributed by the
#7157 branch steward after this audit's cutoff and verified against the
gate's code: TOLERANCE = 400 is consulted in exactly one direction (the
banked-slack check, ceiling.saturating_sub(lines) > TOLERANCE); the
growth check is a bare lines > ceiling. With the in-file 'set to
current, not padded' instruction, every ceiling is a hard cap at the
observed count — so one line landing on main in any contracts crate
reds every open branch at its next fold until someone re-captures.

Measured recurrence on #7157: loop_contracts re-captured four times,
~once per fold (14,479 -> 13,850 -> 13,949 -> 13,115 -> 13,181), the
last tripped by main's #7361/#7363 adding 66 lines to
instruction_bundle.rs — nothing the branch wrote. All four deltas were
<= 105 lines: either repair shape in §3.2 (one-line upward tolerance
using the existing constant, or mid-window pinning) would have absorbed
every one with zero red builds. This audit's own sabotage already
proved the jaws (+1 line host_api red / -1 line common red); the fold
history shows the operational cost. The repair stays a recommendation —
adding growth headroom to a ratchet is the owner's call, not this PR's.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(architecture): give the contracts size ceiling upward working slack

Owner-directed repair of the audit's sharpest finding (report §3.2): the
gate's TOLERANCE = 400 was consulted in exactly one direction — the
banked-slack check — while the growth check was a bare lines > ceiling.
Combined with 'set to current, not padded' pins, every ceiling was a hard
cap at the exact observed count, so one line landing on main in any
contracts crate redded every open branch at its next fold until someone
re-captured. Measured on #7157: four loop_contracts re-captures, roughly
once per fold, every delta <= 105 lines — the gate generating its own
busywork.

The growth check now allows GROWTH_TOLERANCE = 150 of working slack
above each pin (sized to composition-budget precedent; the reviewed
raises this gate has caught were +1,069 and +1,214 lines, far above it),
and all six ceilings are re-pinned to the counts the test itself
reported with every ceiling at 0 — which also removes the +400 seed
padding on common/loop_contracts/prompt_envelope that contradicted the
capture rule and put those crates one deleted line from the banked jaw.

Sabotage-verified both ways: +1 line in host_api and -1 line in common —
both red before this change — now pass; a +151-line probe still fails
with the effective-ceiling arithmetic in the message. Full
reborn_dependency_boundaries binary green (41/41); clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(budget): re-equalize composition pins to observed — restore the working window

Owner-directed companion to the contracts-ceiling repair (same annoying
class, other mass gate): merged main-side growth since the 2026-08-05
equalization had drifted +101 LOC and +5 Arc<dyn> sites through the
tolerance windows, leaving 49 LOC / 10 sites of live headroom — the next
routine composition PR would have gone red on wiring alone (the gate
audit measured this the same day it was pinned).

Per the TOML's own maintenance instructions: loc_ceiling/loc_observed
40423 -> 40524 and arc_dyn 814 -> 819, measured with the gate's --print,
set to current not padded, dated notes appended (not overwritten), and
the arch-test record (COMPOSITION_ABSOLUTE_SRC_LOC) moved in the same
commit as its file requires. ceiling_bp stays 658 — the WS0 floor is
deliberately not re-set.

Verified: check-composition-budget.sh OK; its 76-case self-test green;
reborn_restructure_baselines green; probe +100 LOC now passes (was red
at 49 headroom), probe +160 LOC still fails.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(internal): record the landed zero-slack repairs in the audit report

The §3.2 repair moved from recommendation to landed at owner direction;
the report's answer, inventory rows, and §7 ledger now say so, with the
counting-rule fix promoted to the top remaining recommendation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* gates: pin the ceiling-window arithmetic; fail preflight discovery closed

Two review-round hardenings (the open CodeRabbit Majors):

- reborn_dependency_boundaries.rs: extract the size-ceiling comparison into
  contracts_ceiling_verdict() and pin its four window edges with a committed
  regression test (contracts_size_ceiling_window_edges_hold) — accept at
  ceiling+GROWTH_TOLERANCE, reject one line past, accept at
  ceiling-TOLERANCE, reject one banked line further, and a zero-measure scan
  reads Banked, never a silent pass. The pre-repair asymmetry (tolerance
  consulted only downward) can no longer return silently. Live-gate behavior
  re-probed unchanged after the rewiring: +1 line to host_api passes, +151
  fails with the same effective-ceiling message.
- preflight-gates.sh: setup and changed-file discovery now fail closed — a
  missing repo root exits 2, and a failed merge-base/diff widens the charter
  run to all five crates instead of silently skipping them (the same
  fallback the missing-base branch already used). A broken setup may cost
  compile time, never a silent skip.

Full boundary binary 42/42 green; clippy clean; preflight-gates.sh
end-to-end green on this tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 15:04:35 +00:00
firat.sertgoz
09ac3c8af0 feat(memory): memory-save guidance + always-on MEMORY.md prompt lane (#7185) (#7365)
* feat(memory): tell the model when to save durable user facts

Nothing in the system prompt explains that persistent memory exists, that
the memory block it sees came from earlier conversations, or when a stated
user preference is worth saving. Add a `memory_protocol.md` prompt asset,
appended in memory on every resolve like the self-knowledge section, so
existing installs get it rather than only freshly seeded ones (#7185).

Also point `ironclaw.memory.write`'s prompt doc and both provider manifest
descriptions at the durable-fact use case, so the tool description agrees
with the protocol.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(memory): inject MEMORY.md into every prompt without an FTS match

Both proactive-recall lanes are full-text search over the current turn's
query, so a fact saved in conversation A only reached conversation B when
B's opening message happened to share vocabulary with it (#7185).

Add a `read_curated` memory lifecycle hook and `MemoryService::read_curated`,
implemented by the native provider as a plain read of the scope's `MEMORY.md`
(absent or empty = an empty lane, not an error). The host queries it on every
run with no query at all and admits it FIRST, ahead of the search lanes.

The lane is split on line boundaries into per-snippet-sized chunks rather
than admitted as one oversized snippet: a snippet's model-visible text is
validated as a 512-byte `LoopSafeSummary`, which is also where the prompt
denylist runs, so a single wide snippet would mean denylist-checking only the
head of the document. Chunks are capped at 4 snippets / 2 KiB — half the
aggregate — so the standing document can neither starve nor be starved by the
search lanes, and a clipped document is marked truncated.

The hook is opt-in per manifest, so mem0 (no standing-document concept) is
never called on it and is unaffected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(memory): pin cross-conversation recall and the memory-write gate

Integration (real composition + libSQL): conversation A's model saves a
durable fact with append mode, and conversation B — opening on an unrelated
topic, sharing no content word with the fact, making no memory tool call —
still finds it in its system prompt. Another user's MEMORY.md on the same
composite does not. This is the case the existing proactive-recall scenario
cannot cover, because both search lanes need the query to match.

The lifecycle scenario now counts read_curated too, so "a full declaration
drives every hook" stays true rather than silently skipping the new one.

Composition: pin the approval posture the issue is really about — with the
shipping default settings (global auto-approve on) ironclaw.memory.write is
allowed, so saving a fact does not stop the turn on a prompt; with
auto-approve deliberately off it still gates. The approval seam is
per-capability (the gate policy sees the descriptor, never the invocation
input), so "ungated for curated targets only" is not expressible there, and
an unconditionally ungated save tool would override the choice of exactly
the users the gate currently protects.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(architecture): raise the extension-contracts size ceiling for read_curated

The `read_curated` lifecycle hook adds one enum variant, its wire token, its
`ALL` entry, and a doc comment to ironclaw_extension_contracts — 13 lines of
declaration vocabulary. The lane's behavior (reading the standing document,
line-aligned chunking, budgets, sanitization) lives in the native provider
and ironclaw_host_runtime, so no logic reached the contracts tier.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(memory): teach the model how to phrase a memory worth keeping

The memory protocol told the model *when* to save a durable user fact but
nothing about what a good one looks like, and three failure modes showed up
in practice:

- A memory saved as an imperative ("Always respond concisely") is re-read as
  part of the prompt on every later turn — and the new always-on lane
  re-injects it every turn — so it becomes a standing directive that can
  override what the user is asking for right now. Memories must be
  declarative facts ("User prefers concise responses").
- Task progress, session outcomes, PR numbers, and commit SHAs are stale
  within days and crowd out the durable facts the lane exists to surface.
- With no priority framing the model saves whatever is nearest to hand
  rather than the fact that stops the user repeating themselves.

Adds those three rules to `memory_protocol.md` and the declarative-form and
staleness rules to the shared `ironclaw.memory.write` prompt doc, and
clarifies that honoring a forget request means actually rewriting the
document rather than only acknowledging it (raised in review — `write` with
`append` unset replaces the document, so the guidance is executable).

Two asset tests pin it: one for each doctrine phrase models actually copy
(including both halves of the good/bad example pair), one keeping the
protocol inside an 18-line budget, since it is appended to every prompt on
every turn.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory): address review findings on the always-on recall lane

Four bot-review findings on #7365, all of them real:

**The memory protocol claimed a surface a disabled deployment does not
have.** `MEMORY_PROTOCOL_PROMPT` was appended unconditionally, but a
`Disabled` memory binding resolves to no provider and registers no memory
package, so the model sees no `ironclaw.memory.*` tools at all — the prompt
was telling it persistent memory exists and to call tools absent from its
surface. Composition now carries resolved memory availability into prompt
assembly, gated the same way the tool-disclosure protocol is gated on the
bridge tools existing. The three positional `bool`s on
`DefaultSystemPromptIdentitySource::try_new` become a named
`SystemPromptProtocols` struct: each one licenses the prompt to claim a
capability, and a swapped pair would fail silently. Pinned by a
production-caller test asserting the unbound prompt carries neither the
section nor a memory tool id, with a bound arm so it cannot pass vacuously.

**Two consecutive saved facts ran together.** The backend append is
byte-exact and the protocol asks for one self-contained line per fact, so
"likes tea" then "lives in Berlin" persisted as `likes tealives in Berlin`
— and the curated lane splits `MEMORY.md` on line boundaries, so both facts
reached later turns as one corrupted fact. The native service now
terminates every appended entry with exactly one newline. Regression test
drives two guided appends through the real service and asserts the curated
lane reads back two lines.

**The curated split chunked the whole document before truncating.**
`MEMORY.md` is user-controlled and re-read on every run, so a large standing
document allocated a `String` plus a snippet clone per ~400 bytes on the
retrieve-before-run path and then discarded all but four. `split_curated_text`
now takes a chunk cap and stops there; the caller passes `budget + 1`, the
one extra chunk being what proves the document was longer than admitted.

**`tests/CLAUDE.md` omitted the new scenario.** Added to the Memory section;
group totals corrected 51 → 53 against the tree.

Not applied: CodeRabbit also asked for a `coverage-floor.toml` row for the
new scenario, but that file gates per-crate coverage and has no per-scenario
rows. IronLoop asked to bound the curated *read* itself; that needs a
byte-limited primitive on the repository/filesystem contract and is a
separate change — the host-side cap above bounds the work that follows it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory): compose the always-on lane over the document-store read op

The bespoke MemoryLifecycleHook::ReadCurated hook was unearned API: it
duplicated a read every document-backed provider already serves as the
ironclaw.memory.read tool, just under a new lane-specific contract. Replace
it with a general MemoryService::read_document trait method (fail-closed
`unavailable` default) that native and mem0 implement by delegating to
their existing `read`. The host composes the always-on curated lane itself
out of an ordinary document read of MEMORY.md, run unconditionally
(gated only by the existing memory-disabled context-profile check, never by
a manifest declaration) — the whole point of #7185 is that a user's
standing facts do not depend on a provider opting in to a curated-specific
hook. ReadCurated is removed everywhere: the enum variant, both manifests'
lifecycle tokens, and the extension-contracts size-ceiling bump it required.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(memory): rename curated-lane test off retired vocabulary

The reborn_memory_retired_vocabulary ratchet (landed on main) pins the
literal `document_store` at zero occurrences; the curated-lane test name
tripped it after the merge. Rename only — behavior unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory): name the rewrite mode on the forget path, cover the gate at the caller

Three review findings on #7365, all narrow:

- `memory_protocol.md` established `append: true` as the save mode, then
  told the model to "rewrite the memory document" for a forget request
  without naming the mode. A model carrying `append: true` forward appends
  the correction and leaves the entry the user asked to drop, which the
  always-on lane then re-injects alongside it. Say `append: false`, and
  pin both modes in the asset test.
- The persistent-memory gate is composed in `runtime.rs` from whether a
  provider actually resolved; the unit tests only prove the flag is
  honored once someone sets it. Assert at the production construction
  site that a runtime with a bound provider really does inject the
  protocol.
- `tests/CLAUDE.md`: the §3 summary still counted five memory scenarios
  against the seven §3.4 lists, and the lifecycle row did not mention
  that the curated standing-document read runs ungated by design.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory): serve the standing document inside native's long-term lane (#7185)

A fact the user states in one conversation was invisible in the next one
that opened on an unrelated subject. Both proactive lanes were full-text
search over the current turn's message, so recall depended on the reader
happening to reuse the writer's vocabulary.

The native provider now serves its standing `MEMORY.md` at the head of its
own `read_long_term`, ahead of the full-text hits and independent of the
query. That is the lane's stated job — "the user's general, durable
memory" — and search alone cannot deliver it.

Provider-internal, deliberately. An earlier revision made this a third
host lane composed over a new `MemoryService::read_document` trait method;
both are gone. `MemoryService` is byte-identical to main again and the
host's prompt-context service is back to two lanes with no document paths
in it. What the host sees is ordinary lane snippets it cannot tell from
search hits, which is what keeps the whole safety envelope — scope filter,
untrusted envelope, prompt denylist, per-snippet cap — applying unchanged.

Moving with the retrieval, because they are the same decision:

- The line-boundary splitter, the 4-snippet budget, and the truncation
  marker are now this provider's policy. A whole document as one wide
  snippet would be denylist-checked only at its head, so it is cut into
  chunks that each pass the host's per-snippet contract; a line carrying a
  denylisted secret is dropped on its own and the surrounding facts still
  reach the model.
- The marker reserves its own room instead of pushing a chunk past its
  cap, so a clipped document cannot read as a complete one.
- Append-mode writes terminate each entry with exactly one newline. The
  backend append is byte-exact, so without it two correct guided saves
  ("drinks tea", then "lives in Berlin") persist as one run-on line and
  reach later turns as one corrupted fact.
- A memory-disabled context profile short-circuits before the document
  read, not only before the search, so a disabled profile still issues no
  provider read at all.

mem0 is untouched: it serves no standing document, so its lane keeps its
existing behavior and issues no extra read. That also takes the #7505
target-alias divergence off this path — it remains a real tool-path
contract issue, but no lane code is involved in it any more.

The libSQL integration scenario is unchanged and is the behavior-neutrality
proof: the same user-facing property, asserted through real composition,
survived the mechanism moving from the host into the provider. The
lifecycle scenario returns to pure declared-hook counting, because there is
no host-composed document read left to except from it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(memory): let the memory extension ship its own model guidance (#7185)

Nothing ever told the model that persistent memory exists or when a stated
preference is worth saving, so it usually never called
`ironclaw.memory.write` at all. That guidance now ships with the memory
extension instead of the loop tier.

`[memory]` gains one optional field, `guidance_doc`, naming a bundled asset
in the provider's own package. memory-native declares
`prompts/memory-guidance.md`; composition appends whatever the BOUND
provider declares and writes none of it.

The text was previously `ironclaw_loop_host`'s `memory_protocol.md`, which
was the wrong owner: it names concrete `ironclaw.memory.*` tools and
describes one provider's recall behavior. mem0's recall is search-first and
it serves no standing document, so native's "save it to `memory`, it comes
back every turn" advice would be actively misleading under a mem0 binding —
and there was no way to say so short of host code branching on provider
identity. Declaring the doc in the manifest that already declares the tools
puts the wording next to the thing it describes, and makes "ships no
guidance" a first-class answer. mem0 declares none, with the reason in its
manifest; absent means nothing is appended, never a fallback to another
provider's wording.

Consequences worth naming:

- `SystemPromptProtocols` carries provider CONTENT (`Option<String>`), not
  a host flag, and `ironclaw_loop_host` is byte-identical to main again —
  the loop tier owns no memory prompt text.
- Two conditions gate the append, both necessary: a provider must actually
  be resolved (a `Disabled` binding registers no package, so the model sees
  no memory tools and must not be told they exist), and that provider must
  declare guidance.
- The ref is the existing validated `CapabilityProfileSchemaRef`, the same
  newtype `prompt_doc_ref` uses, so a path escaping the package fails the
  manifest parse. A valid ref no bundled provider claims is fail-quiet:
  guidance carries no authority and gates nothing, so an unknown or future
  declaration appends nothing rather than failing a boot.
- The host resolves the ref through `ironclaw_memory_native`'s public API
  rather than `include_str!`-ing its asset tree. The first cut copied the
  inline-schema precedent and `reborn_cross_crate_include_scan` correctly
  rejected it — §11.2.7 is shrink-only. Exporting the ref and the text from
  one file also makes it impossible for the manifest and the asset to drift
  apart.
- `ironclaw_extension_contracts`' §11.2.3 size ceiling is raised 7_892 ->
  7_947 for the field, its doc, and two parse tests. Declaration vocabulary
  only: the text ships with the package, resolution lives in
  ironclaw_host_runtime, assembly stays in composition.

Guidance content is unchanged from the reviewed version, including the
`append: false` rewrite mode on the forget path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(memory): pin the guidance content where the guidance lives

The guidance-content tests moved with the asset. Composition now asserts
only what composition owns — that the bound provider's declared text is
appended verbatim, survives a user-edited SYSTEM.md, is never seeded into
the user's file, and is absent with no provider bound — against its own
fixture rather than any real provider's wording. Asserting native's exact
sentences there was testing one provider through the layer that is supposed
not to know about it, and it silently became the only thing keeping the
write-quality doctrine alive.

The content pins themselves are re-homed in the memory-native package,
where the text ships: the tool ids it names, the curated target and both
write modes, the never-save carve-out, the declarative-form rule with its
worked example pair, the staleness skip-list, the priority framing, and the
heading + line budget it has to keep to be worth appending on every turn.
Each of those is a failure the compiler cannot see.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory): keep the standing document out of its own lane's search half

Two review findings on the curated prefix, both real:

- A query whose words happen to match `MEMORY.md` re-admitted the standing
  document as a full-text hit behind the curated copy the model can already
  see. That spent a second snippet slot on a duplicate and displaced a
  different document that matched. The lane now excludes `MEMORY_PATH` from
  the search remainder for the same reason it excludes thread scratch: the
  prefix already owns it. Regression drives `max_snippets = 2` with both
  documents matching, so the displacement would be visible.
- A single line longer than a chunk became an oversized chunk. Previously
  the host clipped it and stamped the marker; now that the prefix rides the
  ordinary lane, the host sanitizes it exactly like a search hit and
  truncates SILENTLY — so an over-long line reached the model shortened
  with nothing saying so, and the documented per-chunk limit was not
  actually enforced. The splitter now clips such a line at a char boundary
  and marks it before admission, and `clip_and_mark` is the one place that
  reserves room for the marker.

One existing fixture moved off `MEMORY.md` onto an ordinary path:
`native_context_retrieve_excludes_thread_scratch_from_long_term` used the
standing document as its generic "durable doc", which after the first fix
would have tested the new exclusion instead of the thread-scratch one it
exists to pin. Intent and assertions unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): keep composition's measured mass honest by splitting its prompt tests

`Fast deterministic checks` was red on the composition budget gate: 41868
production LOC against an effective ceiling of 41732, 136 over. The gate
counts LOC and cannot parse an inline `#[cfg(test)]` module out of an
otherwise-production file, so this branch's prompt-assembly tests were
being counted as assembly code.

Split `root/default_system_prompt.rs`'s inline test module verbatim into a
`root/default_system_prompt/tests.rs` sibling, which the gate excludes —
the same move #7151-era fixes used for `runtime_context.rs` and
`host/run_context.rs`. No ceiling raise: composition now measures 41393,
which is 302 LOC BELOW main, so the branch ratchets the crate down rather
than spending budget on test code.

Not a rename of the problem: the tests are unchanged and all 10 still pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory): resolve guidance from the bound provider's own bundled assets

Review finding on #7365: guidance_doc was declared per provider but the
host resolved it through a hard-coded match against memory-native's
exported constants — a non-native provider declaring guidance would
silently resolve to nothing.

Resolution is now generic host code over provider-supplied data: each
provider crate exports its own MEMORY_GUIDANCE_ASSETS table, the shared
bundle constructor resolves the declared ref against the bundled
provider's own table at construction time, and BundledMemoryProvider
carries the resolved content. host_runtime cannot name the mem0 crate
(sanctioned-residue dependency ratchet), so mem0's bundle constructor
takes the table as a parameter and composition — which legitimately
depends on both providers — supplies it at the call site.

Failure semantics strengthen from fail-quiet to fail-loud: a declared
ref that does not resolve within the provider's own assets is a
manifest/asset desync and fails bundle construction, same posture as
the existing missing-[memory] check. Absent guidance_doc remains a
normal no-guidance state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(ci): ratchet the composition mass ceiling down to this tree

The memory-save guidance and its content pins moved out of composition
into the memory-native package that owns them, and the prompt tests were
split out of the production file. The gate's NUDGE fired at 283 LOC of
slack, and its C11 self-test fails a ceiling that no longer binds — so
lock the improvement in rather than bank it as headroom.

Measured on the merged tree with scripts/ci/check-composition-budget.sh.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): give the QA replay job the stack its own suite documents

reborn_qa_recorded_behavior's module docs require RUST_MIN_STACK=67108864:
its replay tests build a full Reborn runtime and drive whole turns, so the
composed debug async frames sit within a few percent of whatever stack they
get. The root-tests and group-suite jobs already set exactly this value and
cite this binary's docs; the job that actually runs it never did.

Measured on this suite (macOS, debug): it overflows at 1984 KiB and clears
at 2048 KiB — main and this branch to the byte, so the libtest default IS
the boundary and any layout shift decides it. That is the same 'unrelated
layout shift' the group-suite comment already records.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): move the composition mass record with its ceiling

The absolute-mass ratchet is two committed numbers that must agree: the
gate manifest's loc_ceiling and the arch-test record that asserts it still
binds. The previous commit lowered the ceiling for this branch's eviction
but left the record at its pre-eviction value, so the record then exceeded
the effective ceiling and reborn_restructure_baselines went red.

Both now sit at 41533, re-measured on the merged tree with
scripts/ci/check-composition-budget.sh after merging current main.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(memory): refresh golden payloads for the memory.write description

The memory tool's description now names when to save a durable user fact,
so the two golden payloads carrying the tool surface move with it — the
sentence itself, and the surface sha256 derived from those descriptions.

No other bytes changed in either snapshot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 14:38:59 +00:00
firat.sertgoz
307521f155 fix(processes): lease expiry recovers safe runs instead of failing them; isolate the journal heartbeat pool (#7471)
* fix(processes): resume runs whose lease expired at a safe checkpoint

A hosted run that lost its lease died as a user-visible failure, even when
it had committed nothing and was sitting idle waiting on the model. Lease
recovery could only tell "has a checkpoint" from "has none", so every
checkpointed run was treated as possibly-mid-side-effect and failed
terminally with `lease_expired`.

Record what the checkpoint actually was. `ProcessCheckpointKind`
(`BeforeModel` / `BeforeSideEffect` / `BeforeBlock`) now rides on the
process snapshot beside `checkpoint_ref`, because a recovery sweep reads
process rows without loading checkpoint rows. Recovery requeues a run whose
latest checkpoint replays no external effect, under the same bounded
crash-reclaim budget; `BeforeSideEffect` and unknown kinds stay terminal,
since no durable idempotency exists for a replayed capability call.

The requeue waits one full lease TTL past expiry before acting. A worker
starved of heartbeats and a dead worker look identical from the journal;
a worker still running would have renewed its lease inside that window, so
anything still expired afterwards is genuinely gone. Nothing else changes
timing: cancellation and the checkpointless requeue stay immediate.

Old journals deserialize with no kind, and an unrecognized kind degrades to
"unknown" rather than failing the whole snapshot — both read as
side-effecting, so the fail-closed path is the default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(composition): give the process journal its own Postgres pool

The journal heartbeat is the liveness signal a run's lease depends on, and
it was sharing one max-size-2 connection pool with every other Postgres
consumer — the event store, triggers, result reads. One turn's read burst
could saturate that pool and starve the run's own heartbeat until the lease
expired underneath it.

Two changes, both about not putting the heartbeat behind other traffic:

- A Postgres deployment opens a small second pool (2 connections) and
  mounts the journal's filesystem over it. The mount set is byte-identical
  to the data plane's, so the journal addresses the same rows over a
  different connection — only the pool differs. libSQL and in-memory arms
  are untouched; libSQL is single-writer by design.
- The default data-plane pool goes 2 -> 8. Two was small enough that one
  turn queued behind itself.

Operators sizing connections should budget `pool_max_size + 2` per
instance; docs, the shipped Docker config, and its smoke assertion follow.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(turn_runner): let a turn survive one slow store checkout

Hosted turn runs heartbeat every 5 seconds, and the supervisor uses that
interval as each heartbeat's timeout, so the generic budget of 3
consecutive failures tolerated about 30 seconds of stall — no more than a
single Postgres connection-checkout timeout. One slow checkout was enough
to abandon a healthy run.

Turn runs now get 8, which clears a checkout stall while still giving up
inside the 90-second lease TTL. That bound is the point: a worker that
stops on its own leaves a live lease behind, whereas one still retrying
past the TTL gets its run reclaimed out from under it. The budget is
therefore derived from the configured heartbeat interval rather than fixed,
so widening the interval shrinks the budget instead of producing an abandon
window that outlives the lease.

The generic `ProcessSupervisorConfig` default stays at 3 — the capability
path heartbeats every 30 seconds, where 8 would run far past the TTL.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): route the shipped Docker configs to the test that pins them

`Detect Reborn test scope` failed closed on
`docker/reborn/config.production.toml`, skipping every downstream Reborn
lane. The planner's comment claimed the runtime configs have no owning
lane, but `crates/app/ironclaw_cli/tests/smoke.rs` parses both
`config.toml` and `config.production.toml` and pins their boot profile,
storage backend, pool sizing and runtime policy — so a lane does read
them, and static control would have skipped exactly the assertions such
an edit can break.

Add `ROOT_FIXTURE_TEST_OWNERS`, the read-at-test-time counterpart of
`EMBEDDED_ASSET_OWNERS`, mapping each config to the `smoke` test target
that asserts it. The two hosted-single-tenant configs stay unclassified:
no test parses them, so they must keep failing closed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(processes): pin the serde and grace-window edges review flagged

Three review findings on the recovery path, none behavioural:

- `lease_duration_millis` now converts the TTL once instead of three
  verbatim copies in claim, heartbeat and expiry recovery, so the bound
  and its rejection message cannot drift between them.
- The in-grace sweep now runs strictly after expiry. Expiry is inclusive
  (`lease_expires_at <= now`), so the old instant did select the process
  and the assertion was not vacuous — but pinning the grace hold on a
  boundary instant made that a property of the comparison rather than of
  the grace window.
- A persisted snapshot carrying an unrecognized `checkpoint_kind`, or
  none at all, now has coverage at the serde boundary: it degrades to
  `None` rather than failing the whole read, and `None` already recovers
  as side-effecting, so an older host fails closed.

The retry-projection fixture also derives `kind` the way
`put_loop_checkpoint` does instead of claiming `BeforeModel` in metadata
while storing `None`, and now asserts the retried snapshot inherits the
kind — the propagation was previously unasserted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(turn_runner): cap the heartbeat interval the lease cannot afford

Both reviewers found the same hole. `heartbeat_failure_budget_within_lease`
computed how many failures the lease TTL could pay for and then clamped
the result up to 1 — but at an interval past half the TTL, even one
failure costs more than the lease. An operator setting
`heartbeat_interval_secs = 120` against the 90s default got a budget of
1 whose abandon window was 240s: the journal would declare the lease
expired while the worker was still waiting on its first heartbeat, which
is exactly the "worker is provably gone" premise recovery relies on.

Cap the interval at half the TTL instead of clamping the budget, so
`budget >= 1` is honest for every configurable value and the scheduler
errs toward heartbeating more often than asked. The explicit budget
setter is capped by the same lease-derived ceiling — it was the other
way to construct a config past the TTL.

The test's `|| budget == 1` escape hatch was the hole itself; it is gone,
and the 120s case plus a 10x-TTL case are now asserted for real.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(composition): prove the journal pool reaches the data plane's rows

The only coverage of the journal's separate Postgres pool built two
independent `InMemoryBackend`s and wrote through each, which cannot
prove the two pool-backed handles reach the same rows — a wrong
connection config or a shared-pool fallback passed it unchanged.

Drive `postgres_from_config_and_env` (the only public constructor that
resolves a connection config, and so the only one that opens the second
pool at all), submit a turn through the turn coordinator so the journal
writes its process row over its own pool, and read that row back over a
connection neither build pool owns. Docker-gated through the existing
`postgres_pool_or_skip` harness; the in-memory test stays as the cheap
mount-set parity check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(contracts): describe the expired-lease machine that ships

Both contracts said an expired lease transitions to `RecoveryRequired`.
Recovery converges directly to a settled state instead: cancellation to
`Cancelled`; no checkpoint or a replay-safe one (`BeforeModel`,
`BeforeBlock`) back to `Queued`, the checkpointed case only after a full
lease TTL of grace; a side-effecting or unrecognized checkpoint, or an
exhausted reclaim budget, to `Failed`. Code is the deliberately-reviewed
behavior here, so the docs follow it.

`turn-runner.md` §3 carries the full transition table and notes that the
legacy `RecoveryRequired` status still exists in the vocabulary but is no
longer produced by expiry; `turn-persistence.md` §6 gets the summary and
points at it, so the two cannot drift into two half-descriptions again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci(planner): classify the shipped docker/reborn configs by their parsing test

`Detect Reborn test scope` failed closed on
`docker/reborn/config.production.toml` (`unclassified pull-request path`),
cascading into the whole `Tests (Reborn)` roll-up on #7471.

The planner's comment asserted these configs have no owning lane. They do:
`crates/app/ironclaw_cli/tests/smoke.rs` parses both `config.toml` and
`config.production.toml` through `RebornConfigFile::parse_text` and asserts on
the profile, storage backend and policy. So they are not static control (whose
membership rule is "no Reborn test lane reads the file") and not prose —
either would silently under-select the one lane that catches a broken
production config. `DOCKER_RUNTIME_CONFIG_OWNERS` routes each to that test
target instead.

The two `config.hosted-single-tenant*.toml` siblings stay fail-closed: their
reader is `tests/dockerfile_runtime_home.rs`, which `_root_test_partitions()`
does not inventory, so no lane can be selected for them.

Verified against #7471's real diff: the planner exited 1 before and exits 0
after, naming the owner in its reasons; all 75 planner self-tests pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(processes): fence stale executors on lease reclaim; address review comments

- supervisor: never start a replacement executor while the reclaimed
  process's prior executor is still running; a definitive lease-lost
  heartbeat (InvalidLease/InvalidTransition) now cancels the stale task
  at the next tick instead of waiting out the failure budget
- turn scheduler: clamp heartbeat intervals past half the lease TTL
  (budget-of-one no longer masks an unaffordable interval); CLI config
  rejects them with a clear error
- extract lease_duration_millis helper shared by claim/heartbeat/recovery
- tests: fence regression (no executor overlap), unknown checkpoint-kind
  wire degradation, in-grace sweep strictly past expiry, retry kind
  propagation, stale-worker reply assertions, Postgres pool isolation,
  planner owner pin, smoke test for the interval bound
- docs: turn-runner expired-lease contract aligned with the state machine

* chore(deps): bump lru 0.18.1 -> 0.18.2 (RUSTSEC-2026-0253)

cargo-deny fails fast-checks on the lru 0.18.1 panic-safety advisory
(use-after-free in LruCache::pop, patched in 0.18.2, issued 2026-08-11).
All four dependents (composition, webui, extension_host, hooks) already
require `lru = "0.18"`, so this is a lock-only patch bump.

* fix(loop): lease-fence transcript writes so a reclaimed worker cannot ghost-reply

Lease recovery requeuing a safe checkpoint opened a window the journal alone
cannot close: run transitions are lease-fenced (ensure_lease), but transcript
writes were not, so a worker whose lease recovery already reclaimed — starved
of heartbeats while blocked in a model call, or suspended past the grace
window — could wake and append a second assistant answer beside the
replacement worker's. Time-based fencing (abandon window + one TTL of grace)
bounds only a worker whose runtime is live to observe its heartbeat failures;
it can never be total.

The other half of the guarantee: ThreadBackedLoopTranscriptPort now carries
the lease the run was claimed under and asks the journal — the only authority
on ownership — before every transcript write (draft begin/update, finalize,
capability-result append). A stale or unverifiable lease refuses the write as
an explicit transcript-write failure; the zombie's loop exit then fails
through its own lease-fenced claim, so nothing it produces can land on the
run the replacement completed. The turn runner's host factory installs the
fence for every claimed run.

The recovery branch comment in the process journal now states the two-part
guarantee honestly instead of over-claiming that the old executor has
"provably given up".

Regression coverage: the lease_wedge integration test releases the stale
worker after the recovered run completes and asserts its output never reaches
the transcript; a seam test pins that the production host build installs the
fence at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(processes): drop duplicated lease_duration_millis from the merge

Both reconciled lines extracted the same helper; the merge kept both
copies and E0592'd. One definition remains, with the fuller doc.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(runtime): address lease recovery review feedback (#7471)

* ci: reseed composition budget for merged tree (#7471)

* docs(turns): clarify legacy recovery lock release (#7471)

* fix(architecture-tests): move the absolute-mass record with the reseeded ceiling

The 2026-08-11 budget reseed (41582 -> 41810 for #7471's merged tree)
updated the manifest ceiling but not the paired record constant this
ratchet compares it against, leaving 228 LOC of apparent headroom — past
the 200-LOC nudge window, so the merge-queue run failed
reborn_restructure_baseline_ratchets_stay_armed.

Re-measured on this tree: 41731 (`check-composition-budget.sh`). The
41810 ceiling was seeded from the merge-queue commit, where concurrent
mainline growth adds ~79 LOC on top of this branch — inside the nudge
window of the corrected record.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 21:54:34 +00:00
firat.sertgoz
6f1ae709d5 feat(tool-search): complete fair discovery and benchmark arms (#7410)
* test(tool-search): add large-catalog baseline

* feat(tool-search): return and use bounded signatures

* feat(tool-search): add fair discovery benchmark arms

* test(tool-discovery): add live benchmark harness

* feat(tool-discovery): default to namespace summaries

* ci(tool-discovery): classify benchmark harness

* feat(tool-discovery): use semantic namespaces

* fix(tool-search): harden discovery benchmark and mode wiring

* docs(tool-search): record corrected benchmark verdict

* fix(tool-search): address review feedback
2026-08-11 08:23:03 +00:00
firat.sertgoz
419807c919 feat(stress): add durable memory parity matrix (#7426)
* feat(stress): add durable memory parity matrix

* refactor(stress): split native search behavior

* fix(stress): fail closed on delayed retries and parity

* fix(stress): keep delayed retries in-flight

* fix(stress): address final review edge cases
2026-08-11 06:46:30 +00:00
Benjamin Kurrek
2d64363101 fix(qa): stop the agent asserting unverified state — automation status, per-caller extension auth, recalled memory (#7246, #7247, #7294) (#7474)
* fix(assistant): ground automation-status claims in actual state checks (#7246)

The agent confidently reported a BTC-news-digest automation as running and
delivering to Telegram while the Automations page showed "No automations
yet" — status fabricated from conversation history instead of checked.

The read path already exists: builtin.trigger_list is model-callable
(PermissionMode::Allow), core-tier always-advertised, granted in the
interactive policy, and deliberately retained for scheduled fires. The
failure is grounding: its description said only "List scheduled triggers
owned by the current caller scope" — no bridge from the user vocabulary
("automation", the Automations page; "routine") to the trigger capability,
and no instruction to consult it before asserting status.

Mechanism (mirrors the proven builtin__outbound_delivery_targets_list
grounding pattern — description-level positive rule tied to the exact
assertion the model must not fabricate):

- trigger_management.rs: TRIGGER_LIST_DESCRIPTION now names the surface
  ("the automations shown on the Automations page"), declares the listing
  the authoritative current state, instructs calling it before answering
  which routines/automations exist or saying one is running, paused,
  already set up, delivering, or missing, forbids reporting status from
  conversation history or memory, and grounds the empty result as "the
  caller has no routines".
- schemas.rs: trigger_list.input.v1.json gains a root description carrying
  the same authoritative-state framing into the model-visible schema.

Regression test (red on the old description, green now):
builtin_trigger_list_surface_grounds_automation_status_claims in
first_party_builtin_tools.rs, driven through the production
visible_capabilities surface assembly — not the constant — asserting the
vocabulary bridge, the check-before-assert rule, the memory ban, the
empty-state grounding, and the schema root description.

Validated: ironclaw_host_runtime suite + clippy -D warnings green;
ironclaw_loop_host, ironclaw_turn_runner, ironclaw_composition green;
reborn_group_triggers integration group green, proving the new description
survives the prompt-build descriptor validation chain (VerifiedCatalog
surface, 4096-byte cap) under production wiring.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(memory): frame recalled memory as recollection, not live state; pin cross-thread transcript containment (#7294)

Investigation verdict: the reported "agent remembers a Telegram routine
from another scope or thread" is NOT a retrieval leak. Every isolation
seam holds: providers scope-filter (native retains only scope-equal
results and excludes threads/ scratch from the long-term lane; mem0
partitions by the composed user namespace), and the host re-applies the
ExpectedScope drop filter in ironclaw_host_runtime::memory_context.
Durable memory crossing conversations for the same user is the
contract's design. The defect is presentation: recalled snippets entered
the prompt as bare "Untrusted memory content: ..." system messages, so
the model read a recollection ("user asked for a BTC news routine") as
verified current state ("you already have this set up").

Fix, at the prompt-assembly seam (InstructionBundleBuilder): whenever at
least one memory snippet is admitted, the memory section now opens with
a recall-framing system message (prompts/memory_recall_framing.md,
include_str!) telling the model these are recollections from earlier
turns/conversations that describe past state and must be verified with a
tool before being asserted as currently configured. No rows are
filtered, deleted, or re-scoped.

Regression coverage making the scoping contract explicit (all green
today; sabotage-verified to arm):
- shared conformance suite (ironclaw_memory::test_support, runs for
  native + mem0): a recorded conversation transcript is invisible to
  another thread's short-term AND long-term lanes and to thread-less
  (trigger-shaped) long-term retrieval; durable memory written during
  one conversation stays retrievable from another (cross-thread recall
  is by design, not a leak).
- instruction-bundle unit tests: framing precedes the snippets; absent
  when no snippets are admitted.
- scenario_proactive_prompt_recall_libsql (production wiring, libSQL):
  the writer conversation's after-turn transcript (proven recorded via
  its own short-term lane) never surfaces in the reader conversation's
  prompt, and recalled durable memory arrives behind the framing.

Consumer contract tests in ironclaw_turns / ironclaw_loop_host updated
for the new memory-section header; tests/CLAUDE.md coverage row updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(assistant): communication context reports per-caller credential truth (#7247)

The model-facing communication runtime context asserted connection state it
never verified: RuntimeCommunicationContextProvider hard-coded
`authenticated: true` for every host-Active channel-surface extension, and
carried no per-caller credential truth at all for credentialed tools-only
extensions (the GitHub repro — the model saw 49 github.* tools plus
"installed and active" catalog state and told the user no further connection
was required, right before the next call raised Authentication required).

Mechanism:

- ironclaw_assistant: the provider now takes the existing readiness ports —
  ExtensionCredentialSetupService (the same scope-gated credential_status the
  extensions card and the runtime auth gate resolve through) and
  ironclaw_auth::ChannelConnectionService — and classifies each Active
  installed extension with a new shared caller_extension_auth verdict in
  reborn_services/extensions.rs. The extensions card's channel-unconnected
  computation is extracted into caller_channel_connection and reused by both
  paths, so the card and the prompt can never diverge on "connected for this
  caller".
- Channels: `authenticated` is now the per-caller truth. A channel needing a
  personal OAuth/pairing binding with no proof for this caller renders
  "unauthenticated"; a genuinely paired/connected channel (or one requiring
  no personal binding, e.g. admin-managed) still renders "authenticated" —
  the #6478 truthful positive is preserved and pinned.
- ironclaw_loop_contracts: CommunicationRuntimeContext gains
  PendingExtensionAuthState; a bounded, sanitized render line names installed
  extensions the calling user has NOT authenticated and forbids the
  "already connected" claim. Unknown/empty render nothing.
- Fail-closed: when a needed verdict is unknowable (ports unwired, lookup
  failed, budget expired) both states degrade to Unknown — the slice claims
  nothing in either direction. Tools stay visible; the auth gate still owns
  enforcement at dispatch.
- ironclaw_composition: wires ProductAuthExtensionCredentialSetup and the
  generic channel-connection facade (assembly extracted into
  build_generic_channel_connection_facade, shared with the product surface)
  into the provider.
- Architecture: ironclaw_loop_contracts size ceiling raised 13112 -> 13172
  for the declaration/render vocabulary (reason recorded at the ceiling).

Regression tests: communication_context::tests
{active_channel_requiring_connection_is_not_claimed_authenticated_without_proof,
credentialed_tool_extension_without_caller_credential_is_pending_auth,
credentialed_tool_extension_with_expired_credential_is_pending_auth,
oauth_channel_not_connected_by_caller_reads_unauthenticated,
oauth_channel_connected_by_caller_reads_authenticated,
credentialed_extension_without_credential_port_degrades_to_unknown} plus
render pins in ironclaw_loop_contracts runtime_context/tests.rs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(arch): re-measure loop_contracts size ceiling for the #7247+#7294 union

Each fix measured the ceiling alone against main (#7247 raised it to
13,172 for its additions); batching both crate-growing commits onto one
branch requires the union measurement, 13,306 — read from the gate's own
failure message, pinned exactly per the #7147 no-untracked-slack lesson.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(ci): raise composition arc_dyn ceiling 814 -> 816 for the #7247 context-provider ports

Two genuine dyn seams (ExtensionCredentialSetupService + the channel-
connection facade) wired into communication-context assembly; observed
831 sites, effective ceiling re-pinned to exactly that — no slack.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: apply the 8 CodeRabbit findings (#7474)

- memory_recall_framing: tool results are authoritative for tool-queryable
  current state; conversation text no longer outranks a fresh tool result.
- trigger_list: limit 0 rejected (schema minimum: 1) so an empty result is
  always proof of absence; regression tests at the handler and the
  boundary suite (the old boundary test pinned the buggy empty-success).
- runtime_context: worst-case SafeSummary fixture now saturates the
  pending-auth arm (fits within 4 KiB); byte-budget truncation test added.
- agent_loop_host_contract: recall-framing filter keyed by content_ref.
- communication_context: per-extension credential lookups run under
  bounded concurrency (8, matching the extensions card) instead of
  serially inside the 500 ms budget; account-backed unconnected test
  (expired / refresh-failed rows read unauthenticated).
- composition-budget: arc_dyn_observed re-measured to 831 with the
  rationale corrected (observed == effective ceiling exactly).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 02:02:55 +00:00
firat.sertgoz
38f8de4486 fix(run_delivery): deliver triggered run failures to the creator (#6896) (#7131)
* fix(run_delivery): deliver triggered run failures to the creator (#6896)

Scheduled/triggered runs that ended in Failed, Cancelled, or
RecoveryRequired produced no user-visible notification: the triggered
delivery driver minted notifications only for Completed /
BlockedApproval / BlockedAuth and recorded every other terminal status
as Skipped. A run that timed out before reaching an actionable state
only logged a warn and recorded Failed, leaving the creator in silence.

Delivery:
- triggered_notification_for_state now mints a FinalReplyReady
  notification for Failed and RecoveryRequired using the existing
  per-category failure summaries (reborn_failure_summary_for_category)
  over state.failure.category(), with a generic fallback when no
  category is present.
- Cancelled mints the same notification, preferring a failure-category
  summary when one is present and falling back to a fixed cancellation
  notice otherwise.
- The RunWaitTimedOut branch with no prior blocked marker now delivers
  the timeout notice as a terminal reply instead of recording Failed.
- The wildcard arm is replaced with explicit non-actionable statuses
  (Queued, Running, CancelRequested, BlockedResource,
  BlockedDependentRun, BlockedExternalTool) so a future status fails to
  compile rather than silently skipping.

Observer:
- TriggerFireSettlementObserver gains on_failed_fire_settled as a
  default no-op method, plus a TriggerFailedFireSettlement event
  carrying tenant/trigger/fire-slot/run-id/history-status. Noop and
  existing implementors keep compiling.
- The active-cleanup sweep fires on_failed_fire_settled when
  clear_active_fire succeeds with TriggerRunHistoryStatus::Error, so
  post-accept failures are observable for automation health. Ok,
  Running, and already-cleared fires do not fire the hook.

Tests:
- run_delivery_contract: Failed+model_error, Failed without category,
  Cancelled, and timeout-before-actionable all assert a Delivered
  outcome with the expected notice text and footer.
- worker tests: a terminal-Error active fire fires exactly one
  on_failed_fire_settled; a terminal-Ok active fire fires none.

The larger retry/redrive budget for failed post-accept fires
(retry_disposition has zero production callers) is intentionally left
for a follow-up; it is out of scope for this surgical delivery fix.

* style: cargo fmt the #6896 delivery fix

* fix(triggers): address terminal delivery review feedback

* fix(assistant): drop unused UserId import after merge

* fix(run_delivery): address multi-agent review findings

- Extract shared terminal-notice helpers (final_reply_notice,
  outcome_for_delivery_failure, deliver_terminal_notice) so the
  timeout, OAuth-backstop, and generic failure arms share one notice
  shape and outcome taxonomy instead of a third hand-rolled copy.
- Add a bounded race-grace window after the wait backstop: a run that
  crosses into a terminal state during the final wait (cancellation in
  flight, failure landing after the last poll) now delivers the correct
  terminal notice instead of the timeout copy.
- Cancelled runs always deliver the fixed cancellation notice; the
  failure-category branch was unreachable in production and would have
  mislabeled a host/operator cancel as a failure.
- Update the stale invariant doc, the five-output surface contract
  count, and the exhaustiveness-only comment on the non-actionable arm.
- Document the cheap/non-blocking contract on
  TriggerFireSettlementObserver (the worker awaits it inline in the
  poller sweep) and note it at the active-cleanup call site.
- Add contract coverage for the timeout arm's delivery-failure outcome
  (Failed) and a regression test proving the race-grace path delivers
  the cancellation notice; the cancelled-with-category test now asserts
  the cancellation notice wins.

* fix(run_delivery): address review comments and restore CI gates

Review fixes (CodeRabbit on 01e887f/f8af109):
- Grace loop fails loud: log the bound TurnError on state-poll failure and
  the RunDeliveryError on terminal-notice build failure before falling back
  to the timeout copy, with silent-ok markers on both intentional fallbacks.
- Hoist TriggeredReplyTargetAuthority, CodecChannelTargetResolver, and
  TriggeredNotificationContext to one construction before the watcher loop;
  the race-grace arm, timeout arm, and loop body now share it.
- Collapse the duplicated failure-summary expression into one closure and
  name TurnStatus::Failed explicitly so future statuses are compiler-visible.
- Drop the stale "Only three states" count from the surface-contract doc.
- Test fixture: encode the late-terminal flip as one Option<(usize,
  ScriptedRunState)> field instead of two correlated Options with an expect.
- Terminal-crossing test: document why flip_after=30 deterministically
  outruns the wait poll budget and assert the grace loop issues no
  cancellation (cancel_calls == 0).

CI:
- composition-budget: re-seed loc_ceiling 40432 -> 40593 (measured on the
  merged tree; the #7131 settlement observer adds +161 governed LOC of
  wiring) and move the arch-test record with it.
- trigger_poller: use the colon-form tracing target required by #7146.

* ci: re-trigger pull_request workflows for c2460ed9e

* fix(composition): capture the settlement health warn in the observer test

The traced_test default filter is {crate}=trace, which drops events whose
metadata target is `ironclaw::reborn::…`. The observer warning is emitted
with the colon-form target (required by #7146 — the equals form recorded a
field and never matched RUST_LOG target filters), so the test saw an empty
buffer. Enable tracing-test's no-env-filter feature, the same pattern the
capabilities/host-runtime/mcp/loop crates use for cross-target assertions.

Re-seed the composition budget to the merged-tree measurement (40747 ->
40867): #7131's observer wiring lands on top of post-measurement mainline
inflow; measured with the gate, set to current. The arch-test record moves
with the manifest.

* fix(run_delivery): merge main and adapt to notice_discriminator String

- Merge origin/main (#7377 run-acts-as-invoker, #7323, #7382, #6938,
  #7280, #7393, #7389, #7364, #7228, #7371, #7399).
- main's #7377 landed a narrower terminal arm (generic failure notice for
  TurnStatus::Failed only); keep the #6896 arm, which covers Failed and
  RecoveryRequired with sanitized per-category summaries plus Cancelled
  and the timeout grace path, and adapt to the Option<String>
  notice_discriminator main introduced.
- Re-seed the composition budget to the merged-tree measurement
  (40811 -> 40861, the run-failure settlement observer lands +50 governed
  LOC); the arch-test record moves with the manifest.
2026-08-10 13:57:31 +00:00
firat.sertgoz
b230c326e9 fix(release): resolve candidates by package identity (#7433)
* fix(release): resolve candidates by package identity

* fix(release): address bot reviews — validate workspace manifests (#7433)
2026-08-10 12:20:02 +00:00