Files
AI不止语 43a46f7912 refactor(skills): fork 增量改为外挂式并强制声明,保证上游各节不被改动
按维护者决定:保留翻译 skill 内的 fork 增量,但必须做到「不影响上游部分」。
原来的做法做不到这一点 —— 增量是**插在上游步骤序列中间**的。

## 问题

executing-plans 把「处理常见异常」插成了步骤 3,于是上游的
Step 3: Complete Development 被挤成了我们的「步骤 4」。这就不是「不影响上游」,
而是改了上游的结构编号;Remember 也被多加了一条,从上游的 6 条变成 7 条。

## 改法:外挂式

- 「处理常见异常」移出步骤序列,改为独立小节,挂在「何时停下来求助」之后
  (它本就是那一节里「缺少依赖、测试失败、指令不清」三种情形的展开)
- 步骤编号恢复为 1/2/3,与上游逐条对应
- Remember 恢复为上游的 6 条;被多加的「每个任务单独提交」并入外挂节
- 节内首行显式标注:本节是 superpowers-zh 的增量内容,上游没有

using-superpowers 的「中国特色技能路由」同样加上该标注。

现在 14 个翻译 skill 里,上游各节全部逐节对应;多出的 2 节都带标注。

## 让纪律可执行:audit 新增 3c-bis

光靠 README 声明会漂移。新增检查:翻译 skill 的标题数必须等于
「上游标题数 + 本文件里带标注的增量节数」。标注与内容同处一文件,不会各自漂移。
未标注就多出章节 = 隐性分叉,下次同步会被误当成漏译 —— 直接 FAIL。

已双向验证:给 brainstorming 偷偷加一节会被拦下并给出可操作提示,
加上标注后放行。

## 修掉一个让守卫从未生效的 bug(我自己写的)

3c-bis 第一次测试没拦住。查出原因:`grep -c` 匹配到 0 个时输出 "0" 但
**退出码为 1**,写成 `$(grep -c ... || echo 0)` 会拼出 "0\n0",后续整数比较
直接报错、检查静默失效。改用 `; true` 只吞退出码。

全仓扫同一模式,发现测试辅助里也有 3 处(上游同样有这个 bug):
- tests/claude-code/test-helpers.sh:77 assert_count 的实际值
- test-subagent-driven-development-integration.sh 的 task_count / todo_count

这个 bug 的后果是断言**永远无法正常失败** —— 模式没匹配到时比较报错而非判定失败。
三处一并修掉并加注释说明原因。(tests/ 不进 npm 包,仅开发使用。)

修复后 audit 由 154 升到 166 pass —— 因为 3c-bis 现在真的对 14 个 skill 都跑了。

## README

简繁对比表新增一行「翻译 skill 内的增量」,写明仅 2 处、都带标注、
上游各节不被改动、audit 会强制未标注的增量报错。

验证:audit.sh 166 pass / 0 warn / 0 fail、verify-release.sh 90 pass / 0 fail
2026-08-08 14:26:16 +08:00
..

Claude Code Skills Tests

Automated tests for superpowers skills using Claude Code CLI.

Overview

This test suite verifies that skills are loaded correctly and Claude follows them as expected. Tests invoke Claude Code in headless mode (claude -p) and verify the behavior.

Requirements

  • Claude Code CLI installed and in PATH (claude --version should work)
  • Local superpowers plugin installed (see main README for installation)

Running Tests

./run-skill-tests.sh

Run integration tests (slow, 10-30 minutes):

./run-skill-tests.sh --integration

Run specific test:

./run-skill-tests.sh --test test-subagent-driven-development.sh

Run with verbose output:

./run-skill-tests.sh --verbose

Set custom timeout:

./run-skill-tests.sh --timeout 1800  # 30 minutes for integration tests

Test Structure

test-helpers.sh

Common functions for skills testing:

  • run_claude "prompt" [timeout] - Run Claude with prompt
  • assert_contains output pattern name - Verify pattern exists
  • assert_not_contains output pattern name - Verify pattern absent
  • assert_count output pattern count name - Verify exact count
  • assert_order output pattern_a pattern_b name - Verify order
  • create_test_project - Create temp test directory
  • create_test_plan project_dir - Create sample plan file

Test Files

Each test file:

  1. Sources test-helpers.sh
  2. Runs Claude Code with specific prompts
  3. Verifies expected behavior using assertions
  4. Returns 0 on success, non-zero on failure

Example Test

#!/usr/bin/env bash
set -euo pipefail

SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
source "$SCRIPT_DIR/test-helpers.sh"

echo "=== Test: My Skill ==="

# Ask Claude about the skill
output=$(run_claude "What does the my-skill skill do?" 30)

# Verify response
assert_contains "$output" "expected behavior" "Skill describes behavior"

echo "=== All tests passed ==="

Current Tests

Fast Tests (run by default)

test-subagent-driven-development.sh

Tests skill content and requirements (~2 minutes):

  • Skill loading and accessibility
  • Workflow ordering (spec compliance before code quality)
  • Self-review requirements documented
  • Plan reading efficiency documented
  • Spec compliance reviewer skepticism documented
  • Review loops documented
  • Task context provision documented

Integration Tests (use --integration flag)

test-subagent-driven-development-integration.sh

Full workflow execution test (~10-30 minutes):

  • Creates real test project with Node.js setup
  • Creates implementation plan with 2 tasks
  • Executes plan using subagent-driven-development
  • Verifies actual behaviors:
    • Plan read once at start (not per task)
    • Full task text provided in subagent prompts
    • Subagents perform self-review before reporting
    • Spec compliance review happens before code quality
    • Spec reviewer reads code independently
    • Working implementation is produced
    • Tests pass
    • Proper git commits created

What it tests:

  • The workflow actually works end-to-end
  • Our improvements are actually applied
  • Subagents follow the skill correctly
  • Final code is functional and tested

Adding New Tests

  1. Create new test file: test-<skill-name>.sh
  2. Source test-helpers.sh
  3. Write tests using run_claude and assertions
  4. Add to test list in run-skill-tests.sh
  5. Make executable: chmod +x test-<skill-name>.sh

Timeout Considerations

  • Default timeout: 5 minutes per test
  • Claude Code may take time to respond
  • Adjust with --timeout if needed
  • Tests should be focused to avoid long runs

Debugging Failed Tests

With --verbose, you'll see full Claude output:

./run-skill-tests.sh --verbose --test test-subagent-driven-development.sh

Without verbose, only failures show output.

CI/CD Integration

To run in CI:

# Run with explicit timeout for CI environments
./run-skill-tests.sh --timeout 900

# Exit code 0 = success, non-zero = failure

Notes

  • Tests verify skill instructions, not full execution
  • Full workflow tests would be very slow
  • Focus on verifying key skill requirements
  • Tests should be deterministic
  • Avoid testing implementation details