* docs(target-arch): resolve the await-edge design question by measurement (D-S) and re-walk the WS9 verify row Appends §12.13 D-S under delegated authority at owner direction, flagged for post-hoc review by Illia Polosukhin (#6696's author): the await-edge store is measured to be a pure projection over ProcessDependencyPort (that half of the shed happened inside #6696 itself), and the resolver is a genuine loop-tier responsibility journal edges cannot express (owner recovery, sanitized transcript result materialization, batch-gate resume-once drain, BlockedDependentRunGate resume policy). §6.7.3 is amended (scheduler DONE / store DONE / resolver KEEP) instead of the shed being executed; the 2.9k figure is corrected to 1,459 production + 1,448 cfg(test) lines. The §12.10 bullet, §2 divergence flag, §9 row 49, §13 validation row, CHECKLIST header/WS4 pointer, README and PLAN all carry the dated resolution. WS9 verify row ticked with evidence: one lifecycle authority (the process journal; TurnRunState/TurnRunRecord are projections via AgentTurnProcessRuntime, ProcessRecord is a capability-invocation view, no bare RunRecord exists) and §7 T4 re-walked clause-by-clause against merged code — matches, including the checkpoint-gated no-auto-retry mechanism (BeforeModel precedes ModelStage; requeue only when checkpoint-free under the 3-claim cap). Docs-only; no code, no tests, no gates touched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): rows 1-2 — package-set tick (64==64/1/0, gate+selftest+independent rederivation) and the 74-row §9 mapping audit (45 L / 15 L-A / 14 OBD / 0 NOT-LANDED; 3 findings recorded) Row 1: check-target-tree.py reports 64 workspace members == 64 documented packages, 1 documented exclusion (tools/ironclaw_silk_decoder), 0 owned exceptions (EXCEPTIONS table empty — §5 steady state); self-test 17/17; cargo-metadata name set diffed empty against an independent §5 parse. Row 2: docs/reborn/target-architecture/ws12-mapping-audit.md is the audit record — per-row executed-evidence, delete-clauses read against WS8's execution notes, all 14 open rows cite their owning CHECKLIST/PROPOSAL row or issue. Findings (recorded, not fixed): F1 prompt_envelope manifest-description fix has no owner row; F2 WS6:429's '#5618 residue deleted' overstates vs the live adopt_migrated_identity + open WS8:523; F3 stale-docs cluster where the tree is ahead of the prose (trace re-export drop, TurnRunTransitionPort, processes->resources). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fold(7154): squash-port fix/red-main-7119 onto family-world main — defect train #7146/#7115/#7104/#7103/#7144 (+#7119 CI lane), 34-hunk contribution.rs port into the split modules, planner entrypoint classification, D-R loopback exception on the widened HTTPS credential guard Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(extractors): issue-number + assertion-rationale doc refinement (rescued844964fb8from rescue/7154-parked-guard) Ports only the doc/assertion refinement commit; the guard-parking commite8f5a31a2on that branch is deliberately NOT taken — superseded by the D-R loopback ruling. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): record D-R — the loopback credential-guard ruling, wiring choice, and regression pins (PROPOSAL §12.13, 2026-08-05) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7154): CodeRabbit round-1 triage — fail-closed tracing-target scan traversal (+node_modules), bounded sidecar output draining (capped capture + discard drain), deadlock regression asserts successful redaction (no seq), XLSX/DOCX empty-classification via extract_document, raise_for_status annotations Threads already addressed by the fold: latency.rs caller-contract wording (merged doc scopes the requirement to latency-trace callers), BodyJsonPointer coverage (the plaintext-refusal test drives all four injection shapes). Deliberately not taken: un-xfailing the four Slack-catalog projections — the xfail is a documented tripwire (unexpected-pass goes red) and clearing them is the #6520 projection-modeling follow-on its comment specs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(assistant): re-point the one field-form tracing target the #7146 gate caught — main's relocated triggered_run_delivery_services carried the drift the PR fixed at its old channel_host address Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * closure fixups: execute the mapping-audit findings — prompt_envelope manifest description (F1), dated ✎ corrections for the #5618 overstatement (F2) and the stale-prose cluster (F3) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): second-reviewer security spot-audit + extension-journey re-verification (rows 5-6) Adversarial second-reviewer pass over PROPOSAL §12.1a/b/c and the batch's own §12.13 D-R loopback carve-out, plus a re-run of the five extension journeys. Attacks were executed rather than argued: two sabotage files and a 38-shape hostile-URL probe were planted, run, and reverted. Verdicts — mint consolidation HOLDS-WITH-RESIDUAL, secrets tightening HOLDS-WITH-RESIDUAL, host/verifier colocation HOLDS, D-R HOLDS. No HOLE. Four findings recorded rather than fixed (report-not-repair): - F1 test_verified/_for_tenant are ungranted mint constructors gated only by the `test-support` feature, in no mint-name table, with nothing pinning the feature to [dev-dependencies]; the shipped binary is measured feature-free. - F2 §12.1b's products-layer residue undercounts by one (ironclaw_assistant). - F3 journey coverage hole: gsuite-with-credential-injection is proven in two halves that no committed test joins. - F4 both recorded census evasions and both fail-open reads are CLOSED on this tree, so §11.2.5/§12.1a/CHECKLIST:552/:597 now understate the seal. Rows 5-6 ticked; only lines 631-632 of CHECKLIST touched so the concurrent rows 3-4 edit folds cleanly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ratchet(closure): lock the budget gate at the program's end state Dispatch ceiling 1122 -> 814 (today's observed, nudge taken; WS0 record 827 stays within effective 829). Mass-share ceiling 2398 -> 658 bp (the WS0 baseline floor — the arch-test assert refuses lower, and observed 578 bp sits inside the nudge window). Absolute LOC re-equalized at 40423: #6831 added 4 governed LOC through the queue's tolerance window; ceiling, observed, and COMPOSITION_ABSOLUTE_SRC_LOC move together here. Both tightenings sabotage-verified red (dispatch 9-over at 790; abs 73-over at 40200). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): gauntlet report — row 3 ticked (full gauntlet green, 0 REAL in scope), row 4 verified-but-open on two pre-existing Postgres-leg test-isolation defects WS12 rows 3-4 verification on the assembled batch tip0c6c0cfb9d: Row 3 (ticked): fmt, clippy default/all-features/--lib --bins, workspace tests (495 targets, 15,203 passed, 0 failed; the smoke.rs:3132 CPU-saturation flake passed first try), arch suite 285/0, the integration-feature lane 1,665/0, recorded-fixture QA (61 fixtures clean, 41/0), frontend (typecheck 1,588 files; vitest 1,088/0; build + bundle budgets), e2e smoke = the CI browser lane under the hermetic wrapper (50 + 21 + 5 passed), and all 41 scripts/ci self-tests (two mapfile/bash-3.2 casualties green under bash 5, the CI shape). Row 4 (stays open, dated note added): both-backend parity proven with legs demonstrably executed for the fabric (57 pg + 81 libsql), triggers (ADR 0003, REQUIRE_POSTGRES), hooks (ADR 0004, all three backends), composition, processes journal, extension-registry, host-runtime libSQL restart, and the backend matrix; fabric-delegated domains enumerated. Two REAL blockers (one class): the Postgres legs of the event-store and assistant-ledger contract suites assert against shared-database state and cannot pass as-written (each failing test passes alone on a virgin database; files byte-identical to origin/main; no CI lane sets their env vars). Full evidence: docs/reborn/target-architecture/ws12-gauntlet-report.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): set the crate/family guidance convention The base commit for the family-guidance program: one canonical home per fact, measured-not-aspirational claims, boundaries stated as exclusions, and the note that guidance files can be gate-pinned. Every family/crate document written on top of this branch follows this shape. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): per-test isolated Postgres databases for the two WS12 parity-blocking contract suites The WS12 gauntlet (ws12-gauntlet-report.md §P6/§P8) measured the Postgres legs of ironclaw_event_store's durable_event_store_contract and ironclaw_assistant's durable_ledger_contract as test-isolation-defective: absolute database-global asserts (event cursors; settled-entry prune bookkeeping) run against the single external database named by their IRONCLAW_*_POSTGRES_URL env vars. Every failing test passes alone on a virgin database - store semantics correct, suites not self-isolating (PROPOSAL §12.13 D-T). Fix: each affected test provisions a private database on the configured server - the fabric contract's IsolatedDatabase pattern (db_root_filesystem_contract.rs) ported locally into each suite: CREATE DATABASE per test, store/pool + migrations against it, courtesy DROP ... WITH (FORCE), and a once-per-binary stale-name sweep. Every assertion preserved byte-identical; libsql/jsonl twins untouched. In the ledger suite only the two retention tests move - the other six Postgres tests keep their proven fingerprint-suffix isolation. Regression pins are the fixed tests themselves: - postgres_replay_advances_next_cursor_past_trailing_filtered_records - postgres_runtime_and_audit_logs_survive_rebuild_with_filtered_cursor_semantics - postgres_settled_entry_limit_prunes_oldest_when_configured - postgres_settled_prune_interval_defers_until_interval_when_configured Green proven on a shared dirty database twice in a row (parallel default threading) and serially on a virgin database; red-first reproduction captured before the fix. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(reborn): record §12.13 D-T (parity-suite isolation ruling) and close CHECKLIST WS12 row 4 D-T (after D-S): the WS12 gauntlet's two REAL findings were one defect class — absolute database-global asserts against the single shared env-var Postgres database — in two suites (event store cursor contract, assistant settled-ledger retention). Ruling executed in commit864d93ee9: per-test isolated databases via the fabric contract's IsolatedDatabase pattern, assertions preserved; alternatives (baseline-relative asserts, serial-only, leave-open) recorded with why they lost; regression pin = the four fixed tests themselves. CHECKLIST WS12 backend-parity row ticks [x] with a dated addendum: red-first reproduction, the three green isolation runs (dirty shared DB twice in parallel; failing pairs serial on virgin), parity now green 10/10. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(extensions): family guidance layer — AGENTS.md rewrite to the guidance-conventions shape, READMEs for all 4 family crates and 14 packages, duplicate-guidance consolidation The family AGENTS.md now teaches the unified extension model (extension = the only product object; channel/tool/auth are manifest surfaces; runtime is loading, never taxonomy; ExtensionId vs VendorId; retired vocabulary pinned by reborn_retired_taxonomy.rs), carries the self-containment and package-to-crate rules from families/extensions.md, the four-responsibility lookup, the measured package catalog, the exclusion list, and the armed gates by test name. Every crate and package gains a README.md (ironclaw_extension_host had no guidance of any kind). ironclaw_extension_registry and memory-native each had both an AGENTS.md and a CLAUDE.md saying overlapping things: AGENTS.md is now canonical, CLAUDE.md a pointer, and memory-native's stale v1 references (src/workspace, src/db/libsql) are dropped in the merge. The slack/telegram agent maps get package framing and a contracts-tier pointer in place of the stale ironclaw_assistant one. Every path literal verified to resolve on disk; all figures (tool counts, dep sets, consumers, layer declarations) measured from the tree at8d13454a1d. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): substrates + lanes family guidance per guidance-conventions.md Family AGENTS.md rewritten to the spec shape for crates/substrates/ and crates/lanes/: boundary, crate table, exclusion lists (mechanism-not-authority for substrates; kernel-decides-lane-executes for lanes), armed gates by test name, and measured deviations stated as deviations (sandbox's three substrate deps, script.rs direct spawn). The lanes wit/-is-load-bearing note is kept. A README.md for every crate in both families (10 new), measured against cargo metadata 2026-08-05: public surface, workspace edges, consumer counts, and enforced invariants each citing their gate. ironclaw_libsql_runtime and ironclaw_wasm_limiter previously had no guidance of any kind; their READMEs carry the sole-pool-home rule (ADDITIONAL_DRIVER_ALLOWLISTS: deadpool = {filesystem, libsql_runtime}) and the outbound-only limiter gate (wasm_sandbox_core_module_stays_domain_free_v1_parity_kernel; no BoundaryRule names the limiter). Duplicate guidance consolidated per rule 1: for the six crates holding both AGENTS.md and CLAUDE.md (filesystem, network, secrets, mcp, sandbox, wasm), CLAUDE.md stays canonical (module spec for filesystem; gate-pinned wording for mcp and wasm) and AGENTS.md becomes a short pointer. No gate-pinned file was edited. Stale reference removed: safety AGENTS.md pointed at src/NETWORK_SECURITY.md, which exists nowhere in the tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(crates-map): rewrite the three top-level maps family-first after the restructure crates/AGENTS.md (264 -> 175 lines): routing map only — the ten families, the read order (family AGENTS.md -> crate README.md -> working rules/module spec -> docs/reborn/contracts/), the enforced seven-layer matrix with the family/layer divergences, measured workspace facts (64 packages, 1 documented exclusion, 0 owned exceptions per scripts/ci/check-target-tree.py), and a verified command block. The 40-row per-crate map is gone: family AGENTS.md files own crate routing per docs/reborn/guidance-conventions.md. crates/README.md (141 -> 119 lines): human map — mental model in family vocabulary, the ten families with measured crate counts, the 14 extension packages (4 crates + 10 data-only), and the two workspace members outside crates/. crates/Architecture.md (1019 -> 1059 lines): audited against the live tree; every named symbol/path re-verified 2026-08-05. Corrected: retired ProductAdapter vocabulary (zero residue in code), the stale pre-rename dependency ladder that still cited the deleted gateway/TUI crates, run-state store mentions, lane-table crate anchors (sandbox/extension_support), declared-in vs minted-by owners in the core data model, and the subagent deny-filter status note (re-verified). Marked the pre-restructure 'partial or evolving' list as unmeasured rather than asserting it. Also documents that scripts/check-boundaries.sh fails on a clean tree (check-5 grep false positives) and greps the deleted v1 src/ in 4 of 6 checks — boundary enforcement for crates/ is the architecture suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(crates-map): package directories carry their own README.md (coordinator sync with extensions-family agent) Every extensions package dir — the 10 data-only ones included — now ships a README.md, so both maps extend the read order to package level. The sibling branch also confirmed what this map already derived per-crate: packages/ is not uniformly products-layer (memory-native and mem0 declare substrates). The other two coordinator corrections targeted rows of the old per-crate map, which this rewrite deleted wholesale; nothing here cites memory-native's CLAUDE.md or claims ironclaw_extension_host lacks guidance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): contracts + events family guidance layer per docs/reborn/guidance-conventions.md - crates/contracts/AGENTS.md and crates/events/AGENTS.md rewritten to the family shape: exclusion lists with destinations, armed gates by test name, layer-matrix rows, crossing guide, measured header counts. - README.md added for all 10 crates (ironclaw_prompt_envelope previously had no guidance of any kind — the CHECKLIST WS11 gap). - One canonical guidance file per crate, other file a pointer: A+C merges for ironclaw_host_api, ironclaw_event_log, ironclaw_event_projections, ironclaw_event_streams; CLAUDE-only content moved to AGENTS.md for ironclaw_loop_contracts, ironclaw_extension_contracts, ironclaw_product_contracts (none of these are root module-spec crates, so AGENTS.md is the working-rules home). - Stale guidance fixed against the live tree: * loop_contracts dep list contradicted the enforced allowlist (manifest is host_api + extension_contracts; common/prompt_envelope are permitted, unused). * event_log still documented the deleted jsonl parse/replay helpers. * event_projections still claimed EventStreamManager, DurableMemoryAuditSink, MemoryAuditProjectionMetadata, and PendingGateProjection — all deleted per PROPOSAL 6.3.3. * product_contracts still carried the pre-D-E open vendor decision under the nonexistent module name llm_config, and a Deferred section contradicting its own operator_llm/operator_service rows. * extension_contracts module table was missing the WS3 runtime module while counting 18. * common's llm_costs note carried the ModelCostTable seam claim refuted by PROPOSAL 12.11 D-F; now cites the pricer-port ruling and the vendor census residue. - Deleted crates/events/ironclaw_event_projections/PENDING_GATE_PROJECTION.md: every claim in it referenced deleted symbols or the removed v1 src/ tree, and its only inbound reference was the crate's own CLAUDE.md. Verified: all consumer counts reproduce via the printed grep commands; 147 path literals across the 28 touched files resolve on disk; no architecture test reads any of these files by name; conflict-marker scan clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): three measured corrections surfaced by the guidance program memory packages are substrates-layer, not products (families/extensions.md); memory_native declares no extension_contracts dep (PROPOSAL §6.8.4); wasm's extension_contracts edge is dev-only and the wasm 'never depends on' bullet is lane-scoped, not family-wide (families/lanes.md). Three further reported defects were checked and NOT corrected — they were misreads: the sandbox 'never above the runtime tier' rule holds (substrates sit below it), and PROPOSAL's safety consumer count already reads 17, matching the tree. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(kernel): family guidance layer — perimeter AGENTS.md, nine crate READMEs, AGENTS/CLAUDE consolidation Family-guidance program, kernel family (guidance-conventions.md shape): - crates/kernel/AGENTS.md rewritten to the family shape: the nine-stage effect pipeline with stage ownership, the sealed-mint table (witness / trust ceiling / approval lease / verified-inbound evidence, each with its mint site and its seal mechanism), the per-stage fail-closed table with file:line or test citations, the sharp exclusion list, and the armed gates by test name (authorized-seal ratchet, sealed-evidence mint ratchet, BoundaryRules, same-layer edge inventory at 21 kernel edges, empty LAYER_MATRIX_EXCEPTIONS register, driver boundary, process storage scan, origin-gate matrix ratchet). - A README.md for each of the nine crates, per the crate shape: measured workspace deps and consumer counts (cargo metadata), public surface with verified citations, enforced invariants naming their gates. ironclaw_processes states the single-lifecycle-authority direction of truth (journal = store; TurnRunState/ProcessRecord/await-edge = projections; PROPOSAL §12.13 D-S); ironclaw_host_runtime documents the D-R literal-loopback carve-out and names its two regression tests. - Duplicate guidance reconciled in all nine crates: AGENTS.md is canonical (guardrails absorbed), CLAUDE.md reduced to a pointer; ironclaw_trust's CONTRACT.md untouched as the co-located cross-crate contract. - Stale references fixed inside owned paths: the deleted capability-profile conformance module (evaluate_profile_conformance — zero hits workspace-wide) removed from ironclaw_capabilities guidance; trust's 'staging branch' / 'PR3' phrasing updated; capabilities' 'later obligation slices' updated to the landed host_runtime obligations split; cross-crate path mentions fully qualified. Every path literal in all 29 kernel .md files verified to resolve on disk; every named symbol swept against crate sources. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(domains): family guidance layer — AGENTS.md boundary doc, 12 crate READMEs, duplicate-guidance consolidation, stale-path fixes Family guidance for crates/domains/ per docs/reborn/guidance-conventions.md: - crates/domains/AGENTS.md rewritten to the family shape: charter table with go-here-when routing, the exclusion list, every armed gate named by test (BoundaryRules + identity/memory allowlists, the 5-entry in-family edge inventory, the naming gates, trusted-trigger ownership, the memory-provider residue ledger, persistence-driver boundary, the two module-charter gates). - A measured README.md for each of the 12 crates: charter, use-when / don't-use-when routing, public surface, measured normal deps + named consumers, enforced invariants with their gates, exact test commands. ironclaw_attachments and ironclaw_identity had no guidance of any kind; identity's README points at CONTRACT.md (the module spec), llm's at its CLAUDE.md module spec. - Duplicate guidance consolidated to one canonical file + pointer per crate: threads/conversations/memory/outbound rules now live in AGENTS.md (CLAUDE.md is a pointer); auth/llm keep CLAUDE.md canonical because their tests/module_charter.rs gates read it (AGENTS.md is the pointer). One misstatement fixed in the conversations merge: transcript content belongs to ironclaw_threads' SessionThreadService, not InboundConversationService. - Staleness fixed inside the family: identity CONTRACT.md two-edge allowlist claim reconciled with D-Q's three entries; trace_commons CLAUDE.md gains the capture module row and strikes its two discharged Known Gaps (recording/paths shims deleted, rename done); llm CLAUDE.md reasoning.rs caller corrected to crates/loop/ironclaw_loop_host; triggers lib.rs 'feature-gated' repo doc comments corrected; pre-family path literals in comments repointed (kernel/approvals+processes, loop/hooks, app/architecture_tests, domains/auth) and the deleted-v1-engine references in skills marked historical. Verified: cargo test -p ironclaw_llm --no-fail-fast (922 passed, exit 0 — CLAUDE.md is gate-pinned); cargo check --all-targets on all six crates with source edits; every cited path literal resolves on disk; conflict-marker scan clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): repair the corrupted kernel bullet and correct two family laws kernel.md: ironclaw_authorization's 'Security & authority role' bullet has been textually corrupted since #6918 — an approvals sentence was spliced into it mid-clause, orphaning its continuation line. Reconstructed, with the spliced sentence restored to the approvals entry where it is true. lanes.md: 'a lane never depends on a substrate' is false as a family-wide law (ironclaw_sandbox holds network/safety/secrets normal deps, which its own entry licenses); the accurate law is the layer ladder, and the narrow claim holds for ironclaw_wasm alone. lanes.md + events.md: the 'every crate ships both an AGENTS.md and a CLAUDE.md' requirement is superseded by docs/reborn/guidance-conventions.md — two files restating one rule is the drift the guidance program removes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(arch): govern the ProtocolAuthEvidence test seam — WS12 audit F1 Two new gates in reborn_sealed_evidence_mint_ratchet (closed paths #12/#13), per the audit's remedy spec: (a) TEST_SEAM_MINT_FNS governs test_verified/test_verified_for_tenant — any production-text call site outside ironclaw_host_api is an offender (comments/strings stripped, #[cfg(test)] blocks stripped, tests.rs / *_tests.rs and cfg-test-only files excluded via the shared census); (b) test-support may appear in no normal dependency table workspace-wide (dependencies / build-dependencies / target.* variants / workspace.dependencies), and no [features] key other than test-support may forward to it — the laundering shape that would evade (b) by one rename. [dev-dependencies] enablement stays legal (cargo-features.md bar 4, the sanctioned dev seam). Measured zero offenders on this tree in both directions before pinning; sabotage-proven red->green both ways (planted production call named with file:line-text; [dependencies] enablement named with its table path). Self-tests drive the same pipelines the gates run (zero-match principle); the definition-location and partition tests now cover the new table. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(integration): join the gsuite credential-injection journey — WS12 audit F3 WS12 row 5 leg 3 was verified in two halves no committed test joined: gsuite handler -> staged credential (crate tier) and staged obligation -> wire (GitHub/Slack only). Scenario 5 already drives gmail.list_messages through production dispatch on a Google-OAuth-configured group; it now also asserts the JOIN: the seeded google account's token (itest-google-token) lands on the recorded outbound gmail.googleapis.com request as 'authorization: Bearer ...', injected at the host egress chokepoint (apply_credential_injection) per the gmail manifest's declared recipe — store -> dispatch-time staging -> chokepoint -> wire, through the caller. Sabotage-proven: disabling the Header injection arm reds exactly this scenario with 'no network egress request matching url gmail.googleapis.com has header authorization' while the request itself still reaches the wire (headers seen: content-type only) — the injection reason, not a setup error; restore -> green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): correct five measured dependency claims in families/domains.md conversations does not depend on safety (its BoundaryRule now forbids it); triggers depends on libsql_runtime + safety and NOT filesystem, so its 'filesystem-routed persistence path alongside SQL' is one path, not two; memory's live set is host_api alone (prompt_envelope is allowlisted, unused); auth was short by extension_contracts + product_contracts. Each verified against the manifest before editing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(guidance): family AGENTS.md + crate READMEs + guidance consolidation for loop/product/app Family-guidance program, families 8-10 (the top of the stack), per docs/reborn/guidance-conventions.md: - Rewrite crates/{loop,product,app}/AGENTS.md from routing stubs to the spec's family shape: exclusion lists, armed gates by test name, layer rows, crossing guides. Loop carries the trust story + the declared Loop*Port decorator chain; product carries the frozen-surface rule, the transports-consume-contracts rule (with the D-B frozen-constant qualification), the evidence-mint prohibition, and the two vendor exceptions; app carries the wires-owners-never-becomes-one charter, the binary-names-packages rule, config's zero-dep guarantee, and the composition mass ratchet (loc 40423 / Arc<dyn> 814). - Add a README.md to all 13 crates (12 new; webui's rewritten to the spec shape) with measured public surface, deps, and consumer counts. - Consolidate duplicate AGENTS.md/CLAUDE.md per spec rule 1: AGENTS.md is canonical and CLAUDE.md a pointer for agent_loop, loop_host, turn_runner, hooks, host_ingress, openai_compat, operator, and architecture_tests; CLAUDE.md stays canonical (module spec / gate-pinned) for webui, composition, and assistant, with composition's AGENTS.md reduced to the pointer. - Fix stale references in owned paths: hooks' dependency diagram and AgentLoopDriver home (ironclaw_loop_contracts, not ironclaw_turns), loop_host/agent_loop port-home claims, turn_runner's pre-#6696 scheduler description, webui's ProductSurface path (product_contracts, not host_api), route count (93, measured), and webui's allowed-dependency list (7 of 10 were listed), the D-S await-edge ruling reflected in turn_runner guidance, composition's llm_admin residue (nearai_login_serve left for operator). Verified: cargo test -p ironclaw_architecture_tests --no-fail-fast (39 binaries, 0 failures — covers the CLI AGENTS.md phrase pin and the composition guidance-markdown scan), scripts/ci/check-target-tree.py, path-literal resolution over all 37 changed files, conflict-marker scan. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): record the closed scan evasions (F4) and the secrets-consumer correction (F2) The sealed-mint census weaknesses PROPOSAL §11.2.5/§12.1a and CHECKLIST recorded as live and owed to WS10 are all closed on this tree, verified by re-attacking the seam with both evasions at once; the docs understated the seal. Ratchet is 23 tests. One residual replaces them: the test_verified test-seam constructors, now pinned by two gates. §12.1b's 'only products-layer crate with the edge' is false by one — ironclaw_assistant carries ironclaw_secrets as port-declaration vocabulary with no expose_secret call. Not a value-reach bypass; joins #7095's inventory. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): correct the app-family layer, config's consumer set, and the webui route count ironclaw_config declares layer=substrates while living in crates/app/; its consumers include operator, extension_manager and extension_host, not just the assembly crate and the binary; webui is 93 contract-locked routes, not 92 (#6780 landed after the last recount). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(stale-sweep): fix agent guidance outside crates/ for the family restructure Audit-and-fix pass over every stale document outside crates/ (PR 2 of the family-guidance program). Live guidance verified against the tree; records kept with dated notes instead of rewrites. Guidance fixes (verified against HEAD before writing): - .claude/commands/trace.md: MCP tool prefix codebase-memory -> codebase-memory-mcp (allowed-tools never matched the real server), ProductSurface home -> ironclaw_product_contracts, capabilities host.rs -> host/ module split, scripts lane -> script-sandbox; deleted the redundant v1-anchors section. - .claude/commands/add-sse-event.md: deleted the banner-quarantined v1 scaffold steps (every path deleted with the monolith); now an honest redirect to the Reborn projection/SSE path. Frontmatter no longer advertises a working scaffold. - .claude/commands/deslop-reborn.md: three dead crates/*/Cargo.toml globs (family layout added a level), ls crates/ -> family-aware listing, v1-only consumer logic retired, per-crate --features integration phrasing. - .claude/rules/type-placement.md: crates/*/src globs matched nothing; recipes re-pointed and numbers re-measured 2026-08 (3,495 structs/enums, 385 traits, fan-in host_api 53 / common 20 / turns 12). - .claude/rules/skills.md: paths trigger pointed at a nonexistent bundled_skills.rs (rule never fired); SKILLS_REGEX_ACTIVATION_ENABLED / SKILLS_MAX_TOKENS env vars are read by nothing -> documented the real config-file setting and DEFAULT_MAX_SKILL_CONTEXT_TOKENS. - .claude/rules/testing.md, ironclaw-reborn-testing skill, CONTRIBUTING.md, .github/pull_request_template.md, testing-playbook, deslop: the workspace-root `integration` feature is empty with zero consumers - all "cargo test --features integration" guidance re-pointed to crate-level suites. - .claude/skills/reborn-extension-surfaces: four pre-colocation assets/ paths, CapabilitySurfaceKind home, conformance-suite move to ironclaw_extension_contracts, ingestion test move to the registry crate, gate-banned migration exemplar replaced with the live behavioral pin, [mcp] instead-of claim softened (nearai-mcp pins a static [[tools]]). - .claude/skills/ironclaw-reborn-orientation: turn_runner labels, prompt-crate list re-derived (turns/first_party_extension_ports out; host_api, loop_contracts, assistant in), consumer-grep glob fixed. - .claude/skills/reborn-feature + docs/reborn/how-to-port-channel-to-reborn.md: ProductSurface/ProductView/descriptors/caller types live in ironclaw_product_contracts; recipes re-pointed. - CLAUDE.md: dead root --features integration line replaced; project tree redrawn with the ten families; trait homes corrected; ProviderId -> VendorId; CapabilitySurfaceKind + ChannelAdapter homes; [channel.config] -> [channel.connection]/[admin_configuration]; v1 Job State Machine section deleted (no such machine in Reborn); prompt-crates recipe fixed; MCP server name; LLM backend list re-derived from LlmBackendKind. - docs/extensions/building-a-tool.md: product-adapter crates row -> channel surface model; package registration -> PACKAGES collector in ironclaw_extension_support (available_extensions.rs is being dissolved); hosted-MCP policy home -> ironclaw_extension_host/src/mcp.rs; dead v1 bullets dropped. - docs/internal/mutation-audit.md: runnable command blocks re-pointed (family paths; ironclaw_dispatcher example replaced - crate deleted in WS0). - docs/reborn/harness/e2e.md: dispatcher row -> the capabilities dispatch contract suites. docs/reborn/contracts/host-api.md: three ironclaw_dispatcher mentions -> capabilities dispatch module. standard-operations.md: renamed crate + arch-test package name. - scripts: mutation-audit.sh usage header, check-hermetic-env.sh env_helpers pointer, check-generic-without-concrete.sh mirror pointer, telegram_smoke/README regression step (target deleted with v1 in #6375). - .env.example: dead SKILLS_REGEX_ACTIVATION_ENABLED entry -> config-file doc. - docs/qa/telegram-coverage-map.md: nine not-automated reasons re-worded to the crate-level integration tier. Records (dated notes, no rewrites): ADR 0003/0004 path notes (evidence pinned to their measured SHA), FEATURE_PARITY state-migration paragraph marked historical with a git-show recovery pointer, engine-v2 parity record's "coexist on main" claim corrected with a historical note, subagent-spawn legacy scope re-tensed. Pre-family path reproduction count: 73 -> 70 files; every remaining file is a dated record (docs/plans, docs/superpowers, ADRs, audits, CHANGELOG history, historical-marked train docs) or a deliberate past-tense mention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(ws12): tick row 7 — the fresh-agent placement probe passed on the final tree All three placements correct with high confidence, each naming the trait, the tests, and the tempting wrong place it rejected. The probe doubled as a docs audit and independently hit four defects, three of which the stacked guidance PR fixes — it succeeded despite them. WS12 is now 7/7. The restructure is complete. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(target-arch): the product→loop_host recount was wrong on the day it was written Eight importing files across four seams, not seven across three — the fourth being a skill-activation-observer seam (projection.rs, projection/live_progress.rs) this bullet never named, which §6.4.7's own same-day note already implied. Surfaced by the plan-conformance audit. The recount history is 3→5→6→7→8, wrong at four of five attempts. That retires the prose count as a method: the sever slice should land an inventory ratchet before or with the move, not another number. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7263): CodeRabbit round-1 triage — 4 code fixes (2 sabotage-proven, 2 red-first) + 6 doc-truth corrections Code, each verified red-first or by sabotage matrix: - sealed-mint ratchet: per-name sighting floor for TEST_SEAM_MINT_FNS (closed path #12). Proven: renaming test_verified_for_tenant away plus one extra legitimate sibling mention passed the old aggregate floor (silent disarm) and fails the new per-name floor naming the constructor; suite 23/23 after revert. (CodeRabbit's claimed baseline ">2 mentions today" is wrong — each name has exactly one kept sighting — but the doc/enforcement mismatch and at-threshold fragility were real.) - trace credit: non-finite novelty_score/duplicate_score are treated as absent before clamping (clamp preserves NaN, which poisoned online_score and credit_points_estimate); NaN cases added to the #7144 regression test, red first. - trace submission: a 2xx whose body stream dies mid-read now maps through request_failed (network telemetry kind, true I/O cause) instead of collapsing to an empty body that the #7144 strict parse misreported as response_invalid/Submission; truncated-body regression test, red first. - Postgres contract suites (event store + assistant ledger): isolated-DB names now carry a creation epoch and the once-per-binary sweep is age-gated (1h), closing the cross-process window where a sibling's fresh zero-backend database (between CREATE DATABASE and first connection) was sweepable; legacy pid-scheme leftovers still collect immediately. Proven on live Postgres 16: planted stale name swept, planted fresh name survives, 13/13 x2 and 20/20 x2 with zero leftovers. Docs (target-architecture truth pass): - PROPOSAL section 9: the WS6 rename sweep (#7152) had rewritten the source column of the 12 renamed rows to their post-rename names, turning their rename dispositions into no-ops (rows 13/14/28/30/49/51/59/61/64/66/67/70); pre-restructure names restored with a dated footnote. - PROPOSAL:69: removed the superseded 3->5->6->7 recount sentence (the corrected 3->5->6->7->8 passage subsumes it). - PROPOSAL row 34: ToolPermissionOverrideStorePort deletion marked landed (2026-08-05 WS8, matching section 6.5.3; zero workspace hits). - CHECKLIST:631: dated note recording that the WS12 F3 gsuite join landed in this batch (scenario_uninstalled_tool_call_denied_until_active.rs asserts the seeded google token on the gmail.googleapis.com wire; suite run green). - CHECKLIST:632: dated note spending F4 (the audit's 19 was correct at its SHA; the ratchet file now holds 23 tests, re-counted at lines 552/597). - ws12-gauntlet-report P6 heading: first of TWO real failures (one class), matching P8 and the report's own summary. - ws12-mapping-audit rows 49/137: dated D-S closure notes (await-edge store half = journal projection already; resolver retained loop-tier; no shed owed) so the backlog register no longer lists it as in-flight. Not fixed, with evidence: the span-helper macros gate suggestion (info_span!(target = ...) is a hard compile error, E0425 — no silent trap), the webui tracing-subscriber workspace-dep suggestion (no [workspace.dependencies] entry exists; suggestion would not build; 8 siblings use the identical direct shape), and the mapping-audit regeneration (the audit is accurate at its pinned SHA; the in-batch F1 fix is recorded in its dated coordinator note). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7263): CodeRabbit round-2 — rejection-body read keeps its cause; 200 {} is not a submission acknowledgement; lanes.md family dep rule matches measured Cargo.tomls - submission.rs non-2xx path: a failed rejection-body read no longer collapses to an empty detail via .unwrap_or_default() (banned by .claude/rules/error-handling.md); the read error folds into the http_rejection detail so the received status keeps driving the 401/403 auth-retry and the Credential/HttpRejection telemetry split. Regression: submit_preserves_rejection_body_read_failure_cause_with_status. - TraceSubmissionReceipt.status: serde default removed — it fabricated status "submitted" from a proxy's 200 {} (the #7144 synthesis, resurfacing through the wire type's defaults), after which the flush caller recorded Submitted and deleted the only retryable queued copy. The acknowledgement is the server naming what happened to the submission — every workspace fixture sends status and callers persist it unconditionally as server_status — so a status-less 2xx body now fails the strict receipt parse as response_invalid. Regression: submit_rejects_success_response_without_explicit_server_status (covers 200 {} and a status-less non-empty object). - docs(lanes.md): the family Dependency-direction rule no longer claims every lane takes the extension-surface vocabulary crate — measured across crates/lanes/*/Cargo.toml: mcp + sandbox hold ironclaw_extension_contracts under [dependencies], wasm only under [dev-dependencies]; dated ✎ cross-references the ironclaw_wasm entry's 2026-08-05 correction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * review(7263): CodeRabbit round-3 — shared Postgres test provisioner (the "new dep edge" premise measured false), entrypoint self-test armed (sabotage-proven), six doc self-contradictions reconciled Code: - ironclaw_filesystem gains a `postgres_isolation` test-support module — the single home of the per-test isolated-database scaffolding (once-per-binary age-gated stale sweep, epoch-in-name convention, DROP WITH (FORCE) cleanup), parameterised by suite/env-var/prefix/unreachable-policy. Zero new production edges: event_store already normal-deps filesystem, filesystem already owns tokio-postgres, and the dev-dep+feature pattern is the one 17 crates already use. The event-store and product-workflow-ledger suites migrate onto it; both Postgres legs proven live against postgres:16 (12 tests, zero leftover databases). The fabric original keeps its older variant with the differences documented at its IsolatedDatabase. - ironclaw_event_store drops the duplicate tokio-postgres dev-dep (the normal dep already reaches tests). - test-reborn-docker-entrypoint.sh: the missing-argv check now exits the command-substitution subshell instead of incrementing a counter the parent never sees — red-proven (a migrate-but-never-exec entrypoint passed with 7 FAIL lines printed), green after the fix both sabotaged and restored. - trace_commons submission test additionally pins !auth_rejection() for the 503 rejection (the structural assert the API affords; the prescribed payload asserts are refuted — status is private and source is None by design, with the message derived from the structured status in the same constructor). Docs (each reconciled to one canonical statement, measured): - kernel.md: lease ownership decided from code — authorization stores, matches, and expires leases (CapabilityLeaseStore + port + expiry all live there); approvals constructs and issues into that store. The round-1 re-homing of the spliced sentence into approvals was wrong and is corrected in the dated repair note. - app.md: "nothing depends on app" scoped to the three app-layer crates; ironclaw_config's consumers restated by dependency kind (normal: composition, cli, operator, extension_host; dev-only: extension_manager, root integration-tests package). - lanes.md: the mediated-services sentence now states the family law as layer-ladder + injected authority; the no-secrets/network/filesystem-dep claim is scoped to ironclaw_wasm, matching the file's own corrections. - CHECKLIST 429/430: the one open traces clause is named (ScopedFilesystem adoption); the stale "other two" count corrected against the F3a strike. - PROPOSAL:69 + CHECKLIST:72: the project-create route repointed — first_party_extension_ports dissolved into loop_host::skill_activation (WS8, §9 row 55) — still unattempted. - PROPOSAL §9 rows 57/62 synced to §6.8.4 (telegram: dependency-set equality with Slack's four contract-tier crates) and §6.9.4 (webui -> assistant is a charter-permanent edge, §12.11 D-B). - PLAN top summary records Wave 6's design question as resolved (D-S, 2026-08-05). - deploy-reborn-cli-docker.md: the two migration paragraphs unified on the entrypoint's actual behavior — only enabled = false beside signing_secret_env/bot_token_env is migrated; every other retired-key shape fails startup with the migration pointer. - composition-budget.toml: the stale "2398 bp, a true ratchet" header replaced with the WS0-floor truth the baselines test asserts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: move the guidance convention into this PR so its citations resolve families/lanes.md and families/events.md cite docs/reborn/guidance-conventions.md when superseding their 'every crate ships both an AGENTS.md and a CLAUDE.md' requirement, but the file was only on the stacked guidance branch — a forward reference that dangles if this PR merges alone. The convention is the rule those notes invoke, so it belongs with them. Caught by the CodeRabbit round-3 pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): give the hoisted postgres provisioner its safety rationales The round-3 hoist moved test provisioning into a production src/ path, so check_no_panics flagged its four panic/expect sites and reddened Code Style via fast-checks. The gate is right to flag them: it deliberately does NOT exempt #[cfg(feature = "test-support")] modules, because a cargo feature is not a privilege boundary in this workspace (PROPOSAL 12.1a proved exactly that) — so a test-support module still compiles into a build where any sibling enables the feature. Suppressed with the gate's documented inline rationale, which must trail the statement rather than precede it. The panics themselves stay: a configured but unusable Postgres must fail the suite loudly rather than skip it, which is the inert-guard rule the isolation fix exists to serve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ci): classify the three planner-unknown paths this PR touches The Reborn PR test planner fails closed on any unclassified path and raises on the FIRST failure in sorted order, so CI only ever showed .github/pull_request_template.md. Classifying that unmasked two more paths in this PR's own diff: scripts/mutation-audit.sh and scripts/telegram_smoke/README.md. All three are classified; the fail-closed arm is untouched: * .github/pull_request_template.md -> IGNORED_PREFIXES, beside its exact sibling .github/ISSUE_TEMPLATE/ (both GitHub UI templates; classify-test-scope.sh already pairs them in its docs-only arm). * scripts/mutation-audit.sh -> PR_STATIC_CONTROL_PATHS, beside its self-test scripts/test-mutation-audit.sh; both run only in nightly-deep-ci.yml's mutation-frontier job. * scripts/telegram_smoke/ -> QA_HARNESS_PREFIXES; a live, by-hand release smoke harness referenced by no workflow, same class as scripts/reborn_qa_matrix/. Each entry is pinned red-first in test_reborn_pr_test_plan.py (entry commented out, new assertion fails with the exact production error, entry restored, green): a new PR-template test with paired accept-AND-select-nothing assertions plus unknown-.github/-sibling refusal probes, and the two existing class tests extended. Planner self-test: 65 tests OK. The planner CLI over this PR's full 209-path diff now exits 0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
29 KiB
Ironclaw TDD Playbook for Engineers
This playbook explains which tests to write, where they belong, and how to develop a feature or bug fix test-first without adding unnecessary process.
It is written for engineers who are new to Ironclaw. Repository and crate-local
guidance remains authoritative when it is more specific. Start with
.claude/rules/testing.md; for cross-layer Reborn tests, also read
tests/integration/CLAUDE.md.
Core rule
Start with what the user should experience, write a test for that behavior, and implement the smallest amount of code needed to make the test pass.
Do not write every kind of test for every change. Select tests based on the behavior and its risks.
The four test types
1. Unit or contract test
A unit or contract test proves one local rule or one crate-owned public contract.
Use it for:
- pure logic and validation
- typed errors and state transitions
- isolated bug fixes
- public API or policy tables owned by one crate
Examples include rejecting an invalid cron expression or denying an operation when the caller lacks permission.
Run the owning crate's tests:
cargo test -p OWNING_CRATE
2. Hermetic feature test
A hermetic feature test proves a complete Ironclaw behavior without calling a real model or external service. The result should be repeatable and require no API keys.
The Reborn integration harness scripts the vendor model response while keeping the real product workflow, scheduler, agent loop, LLM decorator chain, permissions, capabilities, and persistence in the path.
Use this for most new or changed production-wired Reborn behavior.
cargo test --test reborn_integration_SCENARIO
3. Surface test
A surface test proves behavior through the external surface affected by the change. Choose only the relevant kind:
- Recorded model fixture: the model must choose a particular tool or send particular arguments.
- Browser E2E: the user can see or interact with the behavior in WebUI.
- Backend or runtime integration: the behavior depends on PostgreSQL, libSQL, Docker, WASM, MCP, or another runtime.
Many local changes do not need a surface test.
4. Live canary
A live canary runs a small scenario against a real model or provider. It finds model drift, provider API changes, credential failures, and unexpected prompt behavior.
Live canaries are supplemental. They must not be the only test for a feature, and every pull request should not have to wait for one.
How to choose a test
Ask these questions:
- Can the user see or click the changed behavior in WebUI? Add a browser E2E test.
- Does the model need to choose a particular tool or arguments? Add a recorded fixture and replay test.
- Does the behavior cross Ironclaw components or perform a side effect? Add a hermetic Reborn integration test.
- Is the change only local logic? Add a unit or contract test.
- Does success depend on the current behavior of a real model or provider? Add or run a live canary after deterministic tests pass.
A change can match more than one question. For example, a WebUI approval feature may need both a hermetic integration test and a browser test.
Step-by-step workflow
Step 1: Describe the user behavior
Write one short Given/When/Then example. Describe an observable result, not an internal function call.
Given a user who must approve file writes, when Ironclaw tries to write a file, then the turn waits for approval, the file does not exist before approval, and the file exists after approval.
Step 2: Identify the risks
Mark the risks that apply:
- model behavior
- browser behavior
- side effect
- persistence
- security or permissions
- external provider
- cross-component behavior
These risks determine which test types are needed.
Step 3: Choose the highest-level deterministic test
For most production-wired Reborn behavior, start with a hermetic Reborn integration test. A live canary can be the first scenario designed, but it should not be the first or only automated test.
Step 4: Make the test fail
Run the test before implementing the feature. Confirm it fails because the behavior is missing, not because the test contains a typo or broken setup.
For a bug fix, the test must reproduce the original bug.
Step 5: Write the smallest fix
Implement only enough code to make the test pass. Avoid unrelated refactoring and speculative features.
Step 6: Add important edge cases
Add smaller tests when they improve diagnosis or protect an important rule:
- invalid input
- missing permission
- cancellation
- duplicate requests
- partial failure
- wrong user or tenant
- persistence after restart
Do not repeat the same happy path at every test level.
Step 7: Run tests from fast to slow
Run:
- unit and contract tests
- hermetic Reborn integration tests
- recorded fixture replay, browser E2E, or backend/runtime integration
- live canary
This keeps the development loop fast while preserving outside-in coverage.
How local tests map to CI
The test types above describe what evidence to add. CI lanes describe when that evidence runs. Keep those decisions separate: a browser test does not become a live canary because it runs nightly, and a deterministic provider journey remains hermetic even when the deep lane repeats it in more orders.
| Lane | Purpose | What engineers should expect |
|---|---|---|
| Pull request feedback | Fast, scoped signal on the proposed change | Relevant deterministic subsets, evidence gates, and changed-code checks may run before review. |
| Merge queue | Production gate on the merged result | Queue-covered deterministic checks run in the same shape required before main; path scope is computed by workflow jobs rather than trigger filters. |
Push to main |
Confirm the queue result and publish shared evidence | Repeats queue-covered checks, runs the documented post-merge-only Windows, benchmark, and legacy snapshot checks, warms caches, and publishes reports. |
| Deep scheduled | Exercise expensive breadth | Higher property-test counts, reversed and isolated journey order, mutation audits, browser shards, stress, soak, and live drift run outside the ordinary merge gate. |
| Release artifact | Verify what will actually ship | Smoke the exact built archive or binary rather than treating a development build as release evidence. |
The authoritative workflow contract, required-check names, and scheduling
details live in .github/workflows/README.md.
Its "Known accepted gaps" section names deterministic and informational checks
that are deliberately not merge-gating. Re-derive the current lanes before
changing CI:
rg -n "pull_request:|merge_group:|push:|schedule:|workflow_dispatch:" \
.github/workflows
A green pull request does not imply that scheduled, live, backend, or release tiers ran. Name the tiers actually exercised in the pull request test card.
Where tests belong
| Behavior | Location | Notes |
|---|---|---|
| Private helper or pure local rule | crates/<owning_crate>/src/ in #[cfg(test)] mod tests |
Keep the test next to the implementation. |
| Public crate contract | crates/<owning_crate>/tests/<behavior>_contract.rs |
Test through the crate's public API. |
| Whole Reborn turn or cross-component behavior | tests/integration/ |
Use the scripted-model harness and assert at a meaningful seam. |
| Model tool choice or request shape | tests/fixtures/llm_traces/reborn_qa/ and tests/reborn_qa_recorded_behavior.rs |
Commit only scrubbed fixtures. |
| WebUI behavior | tests/e2e/scenarios/test_<behavior>.py |
Use the Reborn v2 fixtures for WebChat v2. |
| Database or runtime behavior | Owning crate's feature-gated integration suite, or an existing root integration suite | Cover supported production backends. |
| Real model or provider drift | scripts/reborn_webui_v2_live_qa/ |
Reuse the current live-QA lane when possible. |
Codebase examples
Unit or crate contract
Example:
crates/product/ironclaw_webui/tests/webui_v2_descriptors_contract.rs
This contract locks the declared WebChat v2 route surface, including method, path, authentication, body limit, rate limit, CORS, audit class, and allowed effect path.
The test follows this shape:
#[test]
fn route_table_has_exactly_the_expected_routes() {
let routes = webui_v2_routes();
let expected = expected_table();
assert_eq!(
routes.len(),
expected.len(),
"expected {} WebChat v2 routes, found {}",
expected.len(),
routes.len()
);
}
Put a similar test inline when it covers a private helper. Put it under the
owning crate's tests/ directory when it protects a public contract or
caller-facing behavior.
cargo test -p ironclaw_webui
Hermetic Reborn integration
Simple example: tests/integration/greeting.rs
This proves that a synthetic inbound message travels through product workflow, scheduling, the agent loop, the real LLM decorator chain, and persisted thread history. Only the vendor model response is scripted.
let harness = RebornIntegrationHarness::test_default()
.script([RebornScriptedReply::text("Hello! How can I help?")])
.build()
.await
.expect("harness builds");
harness
.submit_turn("hi there")
.await
.expect("turn completes");
harness
.assert_reply_contains("Hello! How can I help?")
.await
.expect("reply finalized in thread history");
For a distinct scenario, add tests/integration/<scenario>.rs and register the
flat test binary in the root Cargo.toml as reborn_integration_<scenario>.
If a scenario shares expensive setup with an existing group, add it to the
group directory and include the module from that group's main.rs instead of
creating another harness.
cargo test --test reborn_integration_greeting
Caller-path side effect
Example:
tests/integration/group_approvals/scenario_gate_then_approve.rs
This drives the real approval and resume path and asserts that the approved
file write actually persisted. It does not stop at Completed status or a mock
call count.
let (run_id, gate_ref) = h
.submit_turn_until_blocked("write the approval file")
.await?;
h.approve_gate(run_id, &gate_ref).await?;
h.wait_for_status(run_id, TurnStatus::Completed).await?;
h.assert_workspace_file_contains("approved.txt", "approved write")
.await?;
Use this caller-path shape whenever a helper controls persistence, egress, dispatch, approval, secrets, or another side effect.
cargo test --test reborn_group_approvals
Recorded model fixture
Examples:
- fixture:
tests/fixtures/llm_traces/reborn_qa/web_status_check.json - contract and replay:
tests/reborn_qa_recorded_behavior.rs
The contract proves that the recorded model response selected builtin.http
with the expected target:
let trace = load_qa_trace(WEB_STATUS_CHECK.fixture);
assert_tool_called_with(&trace, "builtin.http", &["api.github.com"]);
Add tool-choice and key-argument assertions to
tests/reborn_qa_recorded_behavior.rs. Add a replay assertion when the trace
should create or modify durable state. Live recorders stay ignored; contract
and replay tests run hermetically in CI.
scripts/ci/check-reborn-qa-fixtures.sh
cargo test --test reborn_qa_recorded_behavior
Browser E2E
Example:
tests/e2e/scenarios/test_reborn_webui_v2_smoke.py
This starts the standalone Reborn server and proves that an authenticated user reaches the chat shell while an anonymous user reaches the login screen.
async def test_reborn_v2_serves_shell_and_gates_auth(
reborn_v2_server, reborn_v2_browser
):
authed_ctx = await reborn_v2_browser.new_context(
viewport={"width": 1280, "height": 720}
)
authed_page = await authed_ctx.new_page()
await authed_page.goto(
f"{reborn_v2_server}/?token={REBORN_V2_AUTH_TOKEN}"
)
await expect(authed_page.locator(SEL_V2["chat_composer"])).to_be_visible()
For Reborn WebChat v2, use the reborn_v2_* fixtures and SEL_V2 selectors.
If the scenario must be part of the Reborn coverage gate, add its file or
pytest node ID to
tests/e2e/reborn_coverage_tests.txt.
The merged Reborn LCOV report enforces three complementary ratchets:
- the aggregate floor preserves the stronger 85.11% post-process-journal baseline (which supersedes the historical 80.81% floor);
- critical production crates have percentage and covered-line floors in
tests/integration/coverage-floor.toml; - added, instrumentable production code must meet the original committed 90%
changed-line floor. Changed-branch coverage is still required in LCOV and
reported for review, but has no universal percentage floor. The denominator
comes from the pull-request diff intersected with LLVM
DAandBRDArecords, not from a hand-maintained file list. Missing production coverage and changed files with no measured instrumented lines remain hard failures.
Changed-code exemptions live in
tests/integration/changed-coverage-exemptions.toml.
They must name exact line numbers and/or branch-line numbers plus an owner,
reason, issue, and future review date. Globs and whole-file exemptions are not
accepted. Test modules and test-support paths are excluded mechanically; a
production source file missing entirely from LCOV fails rather than
disappearing from the denominator. Added lines in renamed production files
remain in scope.
Branch LCOV exports use --skip-functions because changed-code gates consume
line and branch records only. They pin nightly-2025-11-01 (LLVM 21.1.3):
newer bundled LLVM versions have a confirmed
getInstantiationGroups branch-mapping crash
for async generic Rust code. The coverage-only compiler therefore passes
--ignore-rust-version and bootstraps only the three library features used by
current dependencies that stabilized after Rust 1.93: array_windows,
debug_closure_helpers, and slice_as_array. Normal build, clippy, and test
jobs remain the authoritative current-MSRV checks. Dedicated cache keys
prevent mixing objects from the two compilers. The branch-export self-test
fails if the safe toolchain, exact compatibility envelope, MSRV override, or
record filter is removed, if discovery is empty, or if an expected
workflow/script path goes stale. The blocking workflow also runs the changed
coverage and LCOV-merge sabotage suites before accepting the real report.
The standalone shipping ironclaw binary is separately built under
cargo llvm-cov, driven by the production-composed Python E2E manifest, and
terminated through its graceful shutdown path. The Code Coverage workflow
clears profiles emitted by the instrumented prebuild before starting pytest,
then fails if the E2E run produces no new profiles. It uploads its branch-aware
LCOV as the reborn-shipping-binary-e2e-coverage artifact before sending it to
Codecov.
cd tests/e2e
pytest scenarios/test_reborn_webui_v2_smoke.py
Database or runtime integration
Example:
crates/loop/ironclaw_hooks/tests/parity_matrix.rs
This is the right shape when multiple backends must implement the same behavioral contract. Production-facing persistence behavior should cover both libSQL and PostgreSQL unless the owning contract explicitly says otherwise.
Prefer the owning crate's tests/ directory for one storage or runtime
contract. Use an existing feature-gated crate suite when behavior crosses
composition layers.
Never silently return from a test because Docker or PostgreSQL is missing. Use documented feature gates or a loud, explicit opt-out.
cargo test -p <owning-crate> --features integration # e.g. -p ironclaw_hooks; the workspace-root `integration` feature is empty
Live canary
Examples:
- implementation:
scripts/reborn_webui_v2_live_qa/run_live_qa.py - workflow:
.github/workflows/live-canary.yml
The existing qa_3b_endpoint_status_live_chat case asks Ironclaw to check
whether near.ai returns HTTP 200 and verifies that the current status is
reported.
Add or extend the case under scripts/reborn_webui_v2_live_qa/. Reuse the
current reborn-webui-v2-live-qa lane when possible. Change the workflow only
when a case needs new shard, secret, schedule, or lane wiring.
Authorized maintainers can run one case from a pull request:
/canary cases=qa_3b_endpoint_status_live_chat
Keep mutations isolated and reversible, and scrub uploaded artifacts.
Worked example
Feature: a user asks Ironclaw to create a routine that checks a website every hour.
Possible coverage:
- Unit test: reject an invalid schedule.
- Hermetic integration: script the routine-creation tool call and verify the routine is persisted with the correct schedule.
- Recorded fixture: verify a real model chooses the routine tool with the expected URL and schedule.
- Live canary: ask the current production model to create the routine and verify it succeeds.
A browser test is unnecessary unless the feature changes how routines appear or behave in WebUI.
Pull request test card
The pull request template includes this test card. Complete every field before requesting review; do not remove the section.
### Test card
User behavior:
Risk areas:
- [ ] Model behavior
- [ ] Browser
- [ ] Side effect
- [ ] Persistence
- [ ] Security or permissions
- [ ] External provider
- [ ] Cross-component behavior
Tests added or updated:
- Unit or contract:
- Reborn integration:
- Recorded fixture:
- Browser E2E:
- Backend or runtime:
- Live canary:
What the tests prove:
Commands run:
For every unused field, write Not applicable: <reason>. If an expected test
layer is omitted, explain why in one sentence.
Before creating a new test file
- Search for an existing test that already drives the same caller or workflow.
- Read the owning crate's
AGENTS.md,CLAUDE.md,CONTRACT.md, orREADME.md. - Read
tests/integration/CLAUDE.mdbefore changing the Reborn integration harness. - Write or update the test first and confirm it fails for the expected reason.
- Assert an observable outcome, not only
Completedstatus or a mock call count. - Run the narrowest test during development, then expand based on risk.
- Add
cargo test -p ironclaw_architecture_testswhen dependency or ownership edges change. - Use
bash scripts/reborn-e2e-rust.shwhen a Reborn contract or whole-path behavior changes.
Rules to remember
- Most production changes need a unit or contract test plus a hermetic feature test.
- Every reproducible bug fix needs a regression test.
- Test real outcomes: files written, records persisted, events emitted, requests captured, or permissions enforced.
- Test through the real caller when permissions or side effects are involved.
- Use recorded fixtures only when model behavior matters.
- Use browser tests only when browser behavior matters.
- Live canaries supplement deterministic tests; they never replace them.
- Extend an existing test when it already covers the same workflow.
Provider capability coverage is counted per outcome, not per capability
A capability with one passing happy-path case is not covered. The gate in
tests/e2e/scenarios/test_provider_capability_inventory.py counts
capability x outcome class and enforces two rules, both derived from the
shipped manifests rather than a hand-maintained list:
| Operation kind | Required evidence |
|---|---|
external_write |
A ProviderOperationCase with provider-side readback, an integration_evidence entry, or a journey_evidence entry naming the exact test and its readback assertion helper. |
| read | A ProviderOperationCase with outcome_class = "success" and one with outcome_class = "empty". |
The read/write split comes from each tool's effects in
crates/extensions/packages/*/manifest.toml
(external_write), so shipping a new tool classifies it automatically.
A harvested tool-call name is not evidence for a write. A recorded model
response naming slack__send_message proves the model chose the tool. It says
nothing about whether the provider committed the effect, so it cannot stand in
for a readback. That distinction is the whole point of the rule — before it
existed, eight write capabilities were classified tested on the strength of a
recorded tool-call name, and auditing them found two (google-sheets.write_values
and google-sheets.rename_sheet) whose only journey issues them alongside
append_values against one spreadsheet, so nothing isolates either.
outcome_class covers only the per-operation semantic outcomes. Status and
transport failures belong to the fault profiles below — do not duplicate them
here.
Anything not yet meeting the rules goes in coverage_backlog in
tests/e2e/fixtures/provider_capability_coverage.toml with an owner, reason,
issue, and review condition. It is a ratchet: the gate fails when a backlog
entry names a capability that has since been covered, so entries must be
deleted as work lands.
Provider fault profiles
Use tests/e2e/provider_fault_proxy.py when a provider operation must cross
the real Reborn extension and network path while Emulate retains authoritative
provider state. The proxy supplies reusable HTTP, malformed-response, timeout,
connection-reset, and lost-acknowledgement profiles. Its ledger records only
request metadata, body digests, and credential fingerprints.
Apply profiles by operation equivalence class instead of multiplying every provider operation by every failure. A representative read, idempotent write, and non-idempotent write must assert the model-visible result, proxy attempt count, whether the provider received the request, and direct provider readback. A lost-acknowledgement test must prove the provider committed while the runtime did not report success, and must prove that no blind duplicate request occurred.
Keep missing credentials, credential refresh, and account-scope behavior at their existing auth/runtime seams when a provider proxy cannot create the condition faithfully. Fault state must be reset independently from provider state after every case.
Whole-path journey inventory
Use tests/e2e/journey_types.py and tests/e2e/journey_cases.py to register
representative compositions across ingress, execution, provider state, and
delivery. A JourneyCase is evidence metadata, not a workflow DSL: execution
logic, provider setup, and readback remain in their owning test modules.
Each case names isolated provider worlds, ingress, execution lane, delivery target, observable assertions, and an exact Pytest or Cargo declaration. Provider journeys also bind their recorded trace, replay facts, and scheduled live evidence. Product journeys may bind exact delivery addresses and browser evidence when the cited test proves those claims. Do not populate an optional field from intent: name it only when the referenced test asserts it.
test_journey_coverage.py verifies that the evidence still exists and is
executable. It also derives inbound and outbound channel surfaces from shipped
first-party manifests, so adding a production channel without representative
journey evidence fails CI.
Prefer one representative whole-path case per supported ingress and delivery mechanism. Do not multiply every provider operation by every ingress or move provider-specific assertions into the generic registry.
Run the registry gate directly after adding or changing a journey:
cd tests/e2e
pytest scenarios/test_journey_coverage.py -q
Generated lifecycle and interaction coverage
tests/e2e/state_machine_coverage.py projects journeys, provider operations,
provider faults, and focused Reborn integration tests onto the lifecycle
dimensions that matter across boundaries. It is an evidence inventory, not a
second runtime state machine and not a Cartesian-product test generator.
Update it when production adds or changes a supported:
- ingress, authentication, policy, operation, provider-outcome, lifecycle, or delivery class;
- trigger, retry, cancellation, duplicate, restart, or concurrent-submit sequence;
- terminal-stability, at-most-once-effect, actor-isolation, truthful-uncertainty, or no-orphan-resource invariant;
- high-risk interaction between two of those dimensions.
Prefer projecting an existing JourneyCase, ProviderOperationCase, or
provider fault. Add a focused row only when no existing registry owns the
evidence, and cite the exact executable Pytest or Cargo declaration. Add a
required pair only for an interaction whose combined behavior carries more
risk than either dimension alone.
Re-derive the supported dimensions, sequences, invariants, and selected pairs, then run the fail-loud gate:
rg -n \
"SUPPORTED_DIMENSIONS|REQUIRED_EQUIVALENCE_PAIRS|SequenceClass|StateMachineInvariant" \
tests/e2e/state_machine_coverage.py
cd tests/e2e
pytest scenarios/test_state_machine_coverage.py -q
Promoting failures into regression tests
When a production, live-canary, or QA failure is reproducible, promote it to the lowest deterministic seam that proves the broken rule:
- State the user-visible failure and the expected outcome.
- Remove credentials, personal data, provider identifiers, and irrelevant transcript content before committing any fixture.
- Reproduce the failure with a unit/contract, Reborn integration, recorded fixture, provider operation, journey, or browser test.
- Confirm the regression test fails for the original reason.
- Apply the fix and confirm the same test passes.
- Keep a live canary only when real model or provider drift remains a distinct risk.
The commit hook and
regression-test-check.yml
require test changes for conventionally named fixes and selected high-risk
paths. [skip-regression-check] and the matching label are review-visible
exceptions for genuinely infeasible cases, not substitutes for a reproducible
test. Explain the missing deterministic seam and compensating evidence in the
pull request test card.
Mutation audits test the assertions
Coverage proves that code ran; it does not prove that a test would detect the wrong result. Use a mutation audit when a critical invariant, escaped defect, or suspiciously broad coverage needs assertion-strength evidence.
Start with one file or package, triage every viable survivor, and verify a
real-gap fix against both unmodified and sabotaged code:
./scripts/mutation-audit.sh -p OWNING_CRATE path/to/production.rs
./scripts/mutation-verify-fix.sh -p OWNING_CRATE \
'copy the exact mutant string from the triage queue'
Do not optimize a workspace mutation score. Equivalent mutants and unclear
product contracts are explicit outcomes, and broad mutation work belongs in
the scheduled frontier. Follow
docs/internal/mutation-audit.md for environment
isolation, verdicts, acceptance criteria, and the fail-loud self-test.
Product-surface coverage report
The Reborn E2E lane publishes product-surface-coverage-<sha> as JSON and
Markdown. tests/e2e/product_surface_coverage.py joins the production-derived
capability inventory, ProviderOperationCase, JourneyCase, representative
fault cases, and the existing owned backlog. Do not add a second hand-maintained
capability or journey list to reporting code.
The five evidence axes are contract, journey, faults, browser, and
live. Empty optional cells are reported honestly. Missing production
classifications or a tested capability with no executable evidence fail the
lane; owned gaps, waivers, and live-only rows remain prominent but do not
silently become passing evidence. A harvested live-QA fixture is not current
live evidence unless a stable live result artifact binds back to its typed row.
Scheduled live cells name the exact workflow, job, case id, and result artifact
and remain scheduled until a consumer inspects that result.
Generate the same bird's-eye view locally instead of maintaining a separate capability or journey spreadsheet:
cd tests/e2e
python product_surface_coverage.py \
--json ../../artifacts/product-surface-coverage/matrix.json \
--markdown ../../artifacts/product-surface-coverage/matrix.md
Open artifacts/product-surface-coverage/matrix.md for the human-readable
matrix. In GitHub Actions, download the
product-surface-coverage-<source-commit> artifact from the Reborn E2E run.
External dashboards or Notion pages may summarize or link this report, but the
typed registries and generated matrix remain authoritative.