From a73deb211e7864c00edc86ce641373b132f66187 Mon Sep 17 00:00:00 2001 From: Affaan Mustafa Date: Fri, 7 Aug 2026 14:55:22 -0400 Subject: [PATCH] =?UTF-8?q?docs:=20formalize=20It=C3=B4=20inference=20serv?= =?UTF-8?q?ing=20contract=20(#2708)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/design/ecc-ito-compute-integration.md | 28 +++++ skills/ito-inference/SKILL.md | 134 +++++++++++++++------ tests/ci/ito-inference-skill.test.js | 130 ++++++++++++++++++++ 3 files changed, 255 insertions(+), 37 deletions(-) create mode 100644 tests/ci/ito-inference-skill.test.js diff --git a/docs/design/ecc-ito-compute-integration.md b/docs/design/ecc-ito-compute-integration.md index a21c9bdfd..c346428a9 100644 --- a/docs/design/ecc-ito-compute-integration.md +++ b/docs/design/ecc-ito-compute-integration.md @@ -110,6 +110,34 @@ adapter; the ECC bridge does not expose its paper fixture mode. Managed inference remains unavailable. ECC does not claim that Itô created a model endpoint, deployed a workload, reserved capacity, or moved funds. +### Inference-serving contract + +`skills/ito-inference` is the only canonical serving skill; `ito-serve` is +trigger language, not a second installed skill. The current ECC bridge has no +`serve` verb and rejects it before resolving or spawning the canonical client. +The canonical runtime documents `inference` only as an unsupported compatibility +probe, and MCP remains limited to auth, find, and status. Serving requests +therefore stop before login. + +A future `serve` operation is not releasable until it verifies a completed +booking and fresh serving eligibility, accepts an immutable reviewed manifest, +requires a short-lived single-use confirmation bound to account, action, +manifest digest, and maximum cost, and atomically reserves a caller-provided +idempotency key. CLI arguments carry only an opaque non-authorizing confirmation +reference; bearer confirmation is resolved and consumed server-side. + +Manifest handling must canonicalize the path, reject symlinks, open a regular +file without following links, validate ownership/permissions and bounded size, +and hash bytes from the opened descriptor. The digest must match the value bound +into confirmation before mutation, preventing path-swap and digest-mismatch +attacks. Authentication alone is never workload authority. + +The same canonical client must expose structured, tenant-scoped status, logs, +metrics, cancel, and cleanup with bounded timeouts and revocation-aware errors. +After an ambiguous transport failure, callers reconcile by idempotency key +before retrying. ECC must never replace that control plane with root SSH, local +serving scripts, browser automation, or an unreviewed purchase endpoint. + ## Skill and install shape `skills/ito-compute/SKILL.md` is an opt-in workflow installed through: diff --git a/skills/ito-inference/SKILL.md b/skills/ito-inference/SKILL.md index f2448256d..4a95b6c36 100644 --- a/skills/ito-inference/SKILL.md +++ b/skills/ito-inference/SKILL.md @@ -1,59 +1,119 @@ --- name: ito-inference -description: Serve a model on a completed Itô compute booking through the canonical Itô backend. Use after ito-compute has booked GPU nodes and the user wants an OpenAI-compatible endpoint on that metal. Chains off a booking record; ECC implements no serving stack of its own. +description: Inspect the availability of model serving on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed serving manifest. Use after ito-compute has booked GPU nodes and the user asks for an OpenAI-compatible endpoint, ito-serve, hosted Kimi, or self-hosted open-weights inference. ECC implements no serving stack of its own. metadata: origin: ECC + status: scaffold + aliases: ito-serve, hosted-open-weights --- # Itô Inference -Serve a model on rented Itô metal by delegating to the canonical Itô compute -backend (Layer 0.2). ECC does not implement a parallel serving stack, launch -adapter, or inference server, and does no browser automation. This skill chains -off a **completed booking** produced by `ito-compute`; it never books, reserves, -or spends. +`ito-inference` is the sole canonical ECC skill for inference serving on Itô +compute. Requests naming `ito-serve` route here; do not create or install a +second `ito-serve` skill. ECC never SSHes to nodes, downloads weights, launches +an engine, or exposes an endpoint; it never books, reserves, or spends. -## Prerequisite +## Current production boundary -A completed booking from the `ito-compute` skill: booking id, node IPs, SSH -access, GPU SKU, node count, and fabric, already recorded in harness memory. -Without a booking record, stop — this skill does not provision. +Managed serving is unavailable today. The ECC bridge exposes only `login`, +`auth`, `find`, `status`, and explicitly gated `evals`. It has no `serve` verb. +The canonical runtime documents `inference` only as an unsupported compatibility +probe; ECC does not invoke or depend on it. The MCP surface exposes only auth, +find, and status. The locally enforceable guarantee is that ECC rejects `serve` +before resolving or spawning the credential-bearing canonical client. -## Delegation +Therefore stop before authentication or any command invocation. Report the +missing capability and return to the originating agent. Never substitute a +local runner, SSH helper, browser workflow, purchase endpoint, or any untracked +local `ito-serve` draft. -ECC calls the canonical backend through the `ecc ito` bridge; it never -re-implements serving. Authenticate once with `ecc ito login` (device -authorization; no key in arguments, files, logs, or chat), exactly as -`ito-compute` documents. +## Required entitlement + +When serving is implemented, its first gate is a server-verified completed +booking. Harness memory, an RFQ, a quote, node IPs, or SSH access are not proof +of entitlement. The backend must return fresh serving eligibility bound to the +authenticated account, booking, GPU topology, region, fabric, term, and model +policy. Expired, revoked, mismatched, incomplete, or already-released bookings +fail closed before confirmation. + +## Future CLI and API contract + +The intended command name is `serve`; `inference` may remain only as an +explicitly deprecated compatibility alias after the production contract lands. +The future handoff must be equivalent to: ```sh ecc ito serve \ - --booking \ - --model \ - [--quantization ] \ - [--ttft-ms ] [--tpot-ms ] + --booking \ + --manifest \ + --confirmation-ref \ + --idempotency-key \ + --json ``` -The `--ttft-ms` / `--tpot-ms` SLO is optional; supplying it turns on -disaggregated prefill/decode, which is off by default. +The reviewed manifest must identify the model revision, engine and version, +quantization, tensor/pipeline topology, endpoint exposure policy, artifact +checksums, storage ceiling, runtime limits, optional TTFT/TPOT objectives, and +maximum incremental cost. No raw API key, SSH key, node password, or bearer +token belongs in arguments, manifests, logs, MCP results, or chat. -## What the backend does (Layer 0.2) +The client must canonicalize the manifest path, reject symlinks, open a regular +file without following links, require appropriate ownership and restrictive +permissions, enforce a bounded size, and hash bytes from the opened descriptor. +That digest must exactly equal the digest bound into confirmation before any +workload mutation. A path swap, digest mismatch, oversized file, or mutable +unsafe file fails closed. -The desk backend, not ECC, runs the stages, and this skill only reports them: +The canonical API—not ECC—must own workload creation and return structured JSON +with `ok`, `live_api_contacted`, `notice`, and either `data` or `error`. Serving +data must include stable booking, workload, manifest, and idempotency IDs plus a +state enum; it must not claim an endpoint is live until health and model checks +pass. Errors must include a stable code and safe message without secrets. -1. Fabric gate — never launch on unverified metal. Blocks below 80% of - fabric-expected bus bandwidth; advisory between 80% and 92%; fails loud on - silent NCCL socket fallback. -2. Weights download and shard to the serving layout (desk-side sharded cache - keyed by model, quantization, TP degree). -3. Topology plan (AIConfigurator): TP inside the NVLink domain, PP across nodes; - engine flags emitted as a reviewable file before launch. -4. Launch (vLLM, Dynamo when disaggregating) under systemd, warmup, SLO canary, - and registration of the endpoint URL and config to Graphiti memory. +## Confirmation and execution gates -## Unavailable today +Before workload creation, require all of the following: -The serving operation is not yet wired: the canonical CLI's `inference` verb and -the desk `serve-on-booking` backend are scaffolds. Until they land, this skill -reports the missing capability and stops. Never substitute a local runner or a -purchase endpoint. +1. Fresh entitlement and serving eligibility from the canonical backend. +2. A reviewable immutable manifest and deterministic digest. +3. A separate single-use confirmation bound to account, action, manifest, and + cost, with a short expiry and replay protection. CLI arguments carry only an + opaque, non-authorizing confirmation reference; the server resolves and + consumes the bearer capability out of band. +4. A caller-supplied idempotency key reserved atomically with the workload. +5. Server-side fabric, capacity, model-policy, storage, and cost validation. + +Authentication is identity, not workload authority. A login, API key, quote, +or completed booking never substitutes for the serving confirmation. Inspection +and plan generation must not create a workload. Cancel and cleanup are separate +mutations with their own scoped confirmation and idempotency boundaries. + +## Lifecycle and recovery + +The production surface is incomplete until the same canonical client exposes +tenant-scoped status, logs, metrics, cancel, and cleanup operations. Every +operation needs bounded connect and overall timeouts, revocation-aware errors, +and structured output. After an ambiguous transport failure, query status by +the idempotency key before retrying; never create a second workload merely +because the first response was lost. A revoked credential stops polling and +returns control to the originating agent without starting login automatically. + +Only report `ready` after endpoint health, model identity, and canary inference +all pass. Report intermediate and terminal failure states honestly. Cleanup must +be observable and must not release or modify the underlying booking unless that +separate economic action was explicitly authorized. + +## Proposed backend stages + +These stages describe the future backend, not code that exists in ECC: + +1. Verify entitlement, topology, fabric, and cost gates. +2. Fetch checksum-pinned weights into backend-managed storage. +3. Emit and validate a reviewable topology/engine plan. +4. Launch through the provider control plane, never direct root SSH from ECC. +5. Warm up, test health and model identity, run an SLO canary, then register the + endpoint and redacted configuration. + +Until every gate and lifecycle operation above exists in the canonical runtime, +this skill remains a fail-closed availability check and documentation handoff. diff --git a/tests/ci/ito-inference-skill.test.js b/tests/ci/ito-inference-skill.test.js new file mode 100644 index 000000000..bbc02e920 --- /dev/null +++ b/tests/ci/ito-inference-skill.test.js @@ -0,0 +1,130 @@ +/** + * Contract tests for the installable, fail-closed Itô inference handoff. + */ + +const assert = require("assert"); +const fs = require("fs"); +const os = require("os"); +const path = require("path"); +const { spawnSync } = require("child_process"); + +const REPO_ROOT = path.join(__dirname, "..", ".."); + +function read(relativePath) { + return fs.readFileSync(path.join(REPO_ROOT, relativePath), "utf8"); +} + +function readJson(relativePath) { + return JSON.parse(read(relativePath)); +} + +function test(name, fn) { + try { + fn(); + console.log(` ✓ ${name}`); + return true; + } catch (error) { + console.log(` ✗ ${name}`); + console.error(` ${error.message}`); + return false; + } +} + +console.log("\n=== Testing Itô inference skill lifecycle ===\n"); + +const results = [ + test("uses the canonical serving trigger and fails closed while unavailable", () => { + const skill = read("skills/ito-inference/SKILL.md"); + assert.match(skill, /^name: ito-inference$/m); + assert.match(skill, /self-host|serve a model|OpenAI-compatible endpoint/i); + assert.match(skill, /requests naming .*ito-serve/i); + assert.match(skill, /completed booking/i); + assert.match(skill, /never books, reserves,\s+or spends/i); + assert.match(skill, /serving is unavailable today/i); + assert.match(skill, /report the\s+missing capability and return/i); + assert.match(skill, /stop before authentication/i); + assert.match(skill, /no `serve` verb/i); + assert.match(skill, /`inference`.*unsupported compatibility\s+probe/i); + assert.match(skill, /never substitute a\s+local runner, SSH helper, browser workflow, purchase endpoint/i); + assert.doesNotMatch(skill, /ssh\s+root@|serve-status\.sh/i); + for (const gate of [ + /server-verified completed\s+booking/i, + /fresh serving eligibility/i, + /single-use confirmation/i, + /account, action, manifest, and\s+cost/i, + /idempotency/i, + /status, logs, metrics, cancel, and cleanup/i, + /structured JSON/i, + /ambiguous transport/i, + /reject symlinks/i, + /without following links/i, + /hash bytes from the opened descriptor/i, + /digest must exactly equal/i, + ]) assert.match(skill, gate); + assert.match(skill, /--confirmation-ref /i); + assert.doesNotMatch(skill, /--confirmation-token|--api-key|--access-token/i); + }), + test("keeps unsupported serving outside the executable bridge", () => { + const bridge = read("scripts/ito.js"); + assert.match(bridge, /SUPPORTED_COMMANDS[^\n]+login[^\n]+auth[^\n]+find[^\n]+status[^\n]+evals/); + assert.doesNotMatch(bridge, /SUPPORTED_COMMANDS[^\n]+serve/); + assert.match(bridge, /Unsupported Itô command/); + + const fixtureRoot = fs.mkdtempSync(path.join(os.tmpdir(), "ecc-ito-serve-reject-")); + try { + const canonicalDir = path.join(fixtureRoot, "cli", "ito-compute-cli", "dist", "bin"); + fs.mkdirSync(canonicalDir, { recursive: true }); + const marker = path.join(fixtureRoot, "spawned"); + const executable = path.join(canonicalDir, "ito.js"); + fs.writeFileSync(executable, `require("fs").writeFileSync(${JSON.stringify(marker)}, "spawned");\n`); + const result = spawnSync(process.execPath, [ + path.join(REPO_ROOT, "scripts", "ecc.js"), "ito", "serve", + "--booking", "booking_test", "--model", "model_test", + ], { + encoding: "utf8", + env: { ...process.env, ECC_ITO_CLI_EXECUTABLE: executable }, + }); + assert.notStrictEqual(result.status, 0); + assert.match(result.stderr, /Unsupported Itô command "serve"/); + assert.ok(!fs.existsSync(marker), "unsupported serve spawned the canonical child"); + } finally { + fs.rmSync(fixtureRoot, { recursive: true, force: true }); + } + }), + test("ships canonical inference through the existing opt-in compute module", () => { + const modules = readJson("manifests/install-modules.json").modules; + const module = modules.find((candidate) => candidate.id === "ito-compute"); + assert.ok(module, "ito-compute install module is missing"); + assert.deepStrictEqual(module.paths, [ + "skills/ito-compute", + "skills/ito-inference", + "skills/ito-training", + ]); + assert.deepStrictEqual(module.dependencies, ["platform-configs"]); + assert.strictEqual(module.defaultInstall, false); + assert.strictEqual(module.stability, "beta"); + + const components = readJson("manifests/install-components.json").components; + assert.deepStrictEqual( + components.find((candidate) => candidate.id === "capability:ito-compute"), + { + id: "capability:ito-compute", + family: "capability", + description: "Authenticated Itô GPU inventory, RFQ, status, and explicitly gated node-qualification workflows through the separately installed canonical CLI.", + modules: ["ito-compute"], + } + ); + + const profiles = readJson("manifests/install-profiles.json").profiles; + assert.ok(profiles.full.modules.includes("ito-compute")); + + const packageFiles = readJson("package.json").files; + assert.ok(packageFiles.includes("skills/ito-inference/")); + assert.ok(packageFiles.includes("skills/ito-training/")); + }), +]; + +const failed = results.filter((passed) => !passed).length; +console.log(`\nPassed: ${results.length - failed}`); +console.log(`Failed: ${failed}`); +process.exit(failed > 0 ? 1 : 0);