kilo loop stabilization/mobile
This commit is contained in:
parent
1eb9ec22c2
commit
e1239f3d93
102
README.md
102
README.md
|
|
@ -284,17 +284,28 @@ Two layers:
|
|||
system/home paths, …) into the child `kilo run` via `KILO_CONFIG_CONTENT`, so
|
||||
`--auto` still denies them. Skipped when the project already sets a scalar
|
||||
`permission.bash`.
|
||||
2. **Runtime** — every `tool_use` is scanned; a match aborts the session
|
||||
immediately and stops the loop with status `stopped-guard`.
|
||||
2. **Runtime** — every `tool_use` is scanned; a match **aborts that session
|
||||
immediately** (damage control) and, by default, the run **continues with the
|
||||
next iteration**: the still-incomplete stage is retried and the safety note is
|
||||
fed back into the next prompt (`## Loop safety feedback`), so the agent can
|
||||
reach the same result a safer way. In shared-context mode the aborted session
|
||||
is replaced by a fresh one with a synthesized handoff.
|
||||
|
||||
`rm -rf` is **path-aware** in both layers: recursive deletes whose targets are
|
||||
all strictly inside a temp directory (`os.tmpdir()`, `/tmp`, `/var/tmp`) are
|
||||
allowed, so routine scratch cleanup (`rm -rf /tmp/<run>`) no longer blocks a run.
|
||||
Deleting the temp root itself, a relative/project path, or any non-temp absolute
|
||||
path is still blocked. Mixed commands (one temp target plus one unsafe target)
|
||||
are blocked. Pre-execution denies only the catastrophic recursive targets
|
||||
(`~`, `$HOME`, `/home/*`, `/etc*`, `/usr*`, `/opt*`, `/root*`, `/boot*`, `/srv*`,
|
||||
`.git/`), while the runtime scanner remains the catch-all for everything else.
|
||||
Use `--guard-stop` to restore the old behaviour (stop the whole run on the first
|
||||
hit, status `stopped-guard`) and `--max-guard-blocks <n>` (default 3) to cap how
|
||||
many *consecutive* guard-blocked iterations are tolerated before the loop stops.
|
||||
|
||||
`rm -rf` is **path-aware and `cd`-aware** in both layers: recursive deletes whose
|
||||
targets are all strictly inside a temp directory (`os.tmpdir()`, `/tmp`,
|
||||
`/var/tmp`) are allowed, so routine scratch cleanup no longer blocks a run —
|
||||
including the common agent pattern `cd /tmp && rm -rf <scratch>`, where the
|
||||
relative target is resolved against the temp cwd. Deleting the temp root itself,
|
||||
a relative path with no known temp cwd, or any non-temp absolute path is still
|
||||
blocked; `cd /repo && rm -rf src` stays blocked. Mixed commands (one temp target
|
||||
plus one unsafe target) are blocked. Pre-execution denies only the catastrophic
|
||||
recursive targets (`~`, `$HOME`, `/home/*`, `/etc*`, `/usr*`, `/opt*`, `/root*`,
|
||||
`/boot*`, `/srv*`, `.git/`), while the runtime scanner remains the catch-all for
|
||||
everything else.
|
||||
|
||||
### Checkpoints and revert
|
||||
|
||||
|
|
@ -328,11 +339,35 @@ retried instead of failing and losing the turn. Detection covers
|
|||
per attempt up to `--retry-max-delay` (default 60000 ms).
|
||||
- `--max-retries <n>` (default 3); `--no-retry` disables it entirely.
|
||||
- Timeouts (`--max-iteration-minutes`), a command-guard hit, or a user Stop are
|
||||
**not** retried.
|
||||
**not** retried as transient errors. A guard hit instead aborts the session and
|
||||
the loop continues with the next iteration (see `--guard-stop` /
|
||||
`--max-guard-blocks`).
|
||||
- The dashboard marks the iteration `retrying` (`auto-retry ×N`) and the report
|
||||
records the attempt count and the last transient reason.
|
||||
- Optionally notify on retries with `--notify …` (event `model-retry`).
|
||||
|
||||
### Resuming a stopped run
|
||||
|
||||
When a run does not reach `done` (guard, error, stall, budget, abort), the CLI
|
||||
prints the exact command to continue it:
|
||||
|
||||
```
|
||||
Resume this run: kilo-loop --goal <goal> --resume <runId>
|
||||
```
|
||||
|
||||
`--resume <runId>` reloads that run's saved state: stages it already reported as
|
||||
`done` are skipped, `lastPercent` is carried over, and — when the previous run was
|
||||
interrupted by the guard — the safety note is injected into the first iteration
|
||||
again. `--resume last` (or `latest`) resolves to the most recent saved run, so the
|
||||
everyday recovery command is simply:
|
||||
|
||||
```bash
|
||||
kilo-loop --goal .kilo/plans/<plan>.md --resume last
|
||||
```
|
||||
|
||||
Saved runs live under `.kilocode-loop/runs/<runId>/`; `kilo-loop report --last`
|
||||
shows the most recent one without starting anything.
|
||||
|
||||
### Secret redaction
|
||||
Every log line, `state.json`, per-iteration report and the run report is scrubbed
|
||||
before it is written: bearer/Basic headers, `sk-`/`ghp_`/`xox`/AWS/JWT tokens,
|
||||
|
|
@ -485,6 +520,42 @@ After a run finishes, **New run** returns to the launcher (the plan list is
|
|||
refreshed with the new completion state). `POST /api/start` refuses to start a
|
||||
second run while one is active.
|
||||
|
||||
### Reaching the dashboard from a phone (same Wi-Fi)
|
||||
|
||||
By default the dashboard binds loopback (`127.0.0.1`), so only this machine can
|
||||
open it. `--lan` binds **every interface** (`0.0.0.0`) and the console prints the
|
||||
machine's real LAN address to open on another device:
|
||||
|
||||
```bash
|
||||
kilo-loop --lan # or set "lan": true in .kilocode-loop/config.json
|
||||
# Dashboard http://192.168.1.106:7999 (Ctrl+Click to open)
|
||||
```
|
||||
|
||||
Then open that URL on the phone (same Wi-Fi/router). A wildcard bind is never
|
||||
shown as a URL; the tool resolves the primary non-virtual interface (docker/vpn
|
||||
bridges are filtered out). An explicit `--host <ip>` still wins, and `--local`
|
||||
forces loopback back on.
|
||||
|
||||
> **Security:** the dashboard has no authentication and exposes control
|
||||
> endpoints (start/stop/pause, revert which deletes untracked files). Only use
|
||||
> `--lan` on a trusted network, and prefer `--local` (or `--host 127.0.0.1`) on
|
||||
> shared/public Wi-Fi.
|
||||
|
||||
## Scaffolding a plan (`new-plan`)
|
||||
|
||||
Plans must contain `- [ ]` checkbox stages to be discovered at all. Generate one
|
||||
from the project's template instead of hand-writing the format:
|
||||
|
||||
```bash
|
||||
kilo-loop new-plan "My plan title" # writes .kilo/plans/<epoch-ms>-<slug>.md
|
||||
kilo-loop new-plan "My plan title" --print # preview without writing
|
||||
```
|
||||
|
||||
The template is `<project>/.kilo/templates/plan.md` (a sensible default is used if
|
||||
the project has none). Optional flags: `--slug <s>`, `--epoch <ms>`, `--dir <d>`,
|
||||
`--project <p>`, `--print`, `--force`. The scaffolder is also available directly
|
||||
at `tools/kilocode-loop/new-plan.mjs`.
|
||||
|
||||
## CLI reference
|
||||
|
||||
```
|
||||
|
|
@ -513,7 +584,9 @@ Verification & safety
|
|||
--verify "<cmd>" loop-run check; a stage is not done until it exits 0
|
||||
--verify-timeout-minutes <n>
|
||||
--checkpoint / --no-checkpoint
|
||||
--guard / --no-guard
|
||||
--guard / --no-guard deny destructive bash; a hit aborts the session
|
||||
--guard-stop stop the run on the first guard hit (default: continue)
|
||||
--max-guard-blocks <n> consecutive guard hits tolerated (default 3)
|
||||
--redact / --no-redact
|
||||
--max-cost <usd>
|
||||
--max-tokens <n>
|
||||
|
|
@ -532,11 +605,14 @@ Verification & safety
|
|||
Dashboard
|
||||
-p, --port <n> default 7999
|
||||
--host <host> default 127.0.0.1
|
||||
--lan bind 0.0.0.0 so a phone on the same Wi-Fi can open it
|
||||
--local force loopback only (overrides a config "lan": true)
|
||||
--keep-open keep the dashboard after the loop
|
||||
|
||||
Utility
|
||||
new-plan "<title>" scaffold .kilo/plans/<epoch-ms>-<slug>.md from .kilo/templates/plan.md
|
||||
--dry-run simulate the agent (no API calls)
|
||||
--resume <runId> reuse a previous run's session + stage progress
|
||||
--resume <runId|last> reuse a previous run's session + stage progress
|
||||
--run-id <id>
|
||||
--quiet suppress the live console stream
|
||||
--no-color
|
||||
|
|
|
|||
|
|
@ -7,7 +7,9 @@ import { RunController } from '../src/controller.mjs'
|
|||
import { startServer, runsBaseDir } from '../src/server.mjs'
|
||||
import { installTerminalControls } from '../src/terminal.mjs'
|
||||
import { printLog } from '../src/console.mjs'
|
||||
import { loadRun } from '../src/state.mjs'
|
||||
import { listRuns, loadRun } from '../src/state.mjs'
|
||||
import { dashboardUrls } from '../src/net.mjs'
|
||||
import { guardFeedbackText } from '../src/guard.mjs'
|
||||
import { loadGoal } from '../src/goal.mjs'
|
||||
import { listGoalSummaries } from '../src/goals.mjs'
|
||||
import { confirmStart, renderPreflight } from '../src/preflight.mjs'
|
||||
|
|
@ -24,8 +26,11 @@ export const EXIT_CODES = {
|
|||
'stopped-guard': 5,
|
||||
}
|
||||
|
||||
function dashboardLink(url, note) {
|
||||
function dashboardLink(url, note, lan = []) {
|
||||
process.stdout.write(`\n ${c.gray('Dashboard')} ${url} ${c.gray('(Ctrl+Click to open)')}\n`)
|
||||
for (const u of lan) {
|
||||
if (u !== url) process.stdout.write(` ${c.gray('On this network (e.g. phone)')} ${c.bold(u)}\n`)
|
||||
}
|
||||
if (note) process.stdout.write(` ${c.gray(note)}\n`)
|
||||
}
|
||||
|
||||
|
|
@ -112,7 +117,7 @@ async function main() {
|
|||
if (config.launcher) {
|
||||
const controller = new RunController(config)
|
||||
const dashboard = await startServer(config, controller)
|
||||
dashboardLink(dashboard.url, 'No goal given — pick a plan in the dashboard.')
|
||||
dashboardLink(dashboard.url, 'No goal given — pick a plan in the dashboard.', dashboard.urls?.lan || [])
|
||||
installSignalHandlers(async () => {
|
||||
if (controller.isActive()) {
|
||||
controller.requestStop()
|
||||
|
|
@ -140,6 +145,15 @@ async function main() {
|
|||
// so match by equality or prefix against the goal's stage ids.
|
||||
let prevRun = null
|
||||
const resumedStageIds = new Set()
|
||||
if (config.resume === 'last' || config.resume === 'latest') {
|
||||
const [latest] = listRuns(runsBaseDir(config))
|
||||
if (latest?.runId) {
|
||||
config.resume = latest.runId
|
||||
} else {
|
||||
process.stderr.write(' Warning: --resume last requested but no saved runs were found — starting fresh.\n')
|
||||
config.resume = ''
|
||||
}
|
||||
}
|
||||
if (config.resume) {
|
||||
try {
|
||||
prevRun = loadRun(runsBaseDir(config), config.resume)
|
||||
|
|
@ -157,7 +171,7 @@ async function main() {
|
|||
}
|
||||
}
|
||||
|
||||
process.stdout.write(renderPreflight(config, goal, `http://${config.host}:${config.port}`))
|
||||
process.stdout.write(renderPreflight(config, goal, dashboardUrls(config).primary))
|
||||
if (!(await confirmStart(config))) {
|
||||
process.stdout.write(`\n ${c.yellow('Cancelled')} — nothing was started. Edit the goal or config, then run again.\n`)
|
||||
return
|
||||
|
|
@ -173,6 +187,13 @@ async function main() {
|
|||
for (const id of resumedStageIds) orchestrator.completedStages.add(id)
|
||||
const last = (prevRun.iterations || []).slice(-1)[0]
|
||||
if (last?.percent != null) orchestrator.lastPercent = last.percent
|
||||
// If the previous run was interrupted by the safety guard, carry that note
|
||||
// into this run's first iteration so the agent does not retry the command.
|
||||
const blocked = [...(prevRun.iterations || [])].reverse().find((it) => it.guardViolation)
|
||||
if (blocked) {
|
||||
orchestrator.guardFeedback = guardFeedbackText(blocked.guardViolation)
|
||||
orchestrator.log('warn', `Previous run hit the safety guard (${blocked.guardViolation.id}) — the first iteration carries the safety feedback.`)
|
||||
}
|
||||
orchestrator.log('system', `Resumed run ${config.resume} (session ${prevRun.sessionID || '—'}, last ${orchestrator.lastPercent}%).`)
|
||||
}
|
||||
|
||||
|
|
@ -181,7 +202,7 @@ async function main() {
|
|||
const dashboard = await startServer(config, controller)
|
||||
const controls = installTerminalControls(orchestrator)
|
||||
|
||||
dashboardLink(dashboard.url)
|
||||
dashboardLink(dashboard.url, null, dashboard.urls?.lan || [])
|
||||
|
||||
let exitCode = 0
|
||||
const shutdown = async (code) => {
|
||||
|
|
@ -203,6 +224,14 @@ async function main() {
|
|||
try {
|
||||
const final = await runPromise
|
||||
exitCode = EXIT_CODES[final.status] ?? 1
|
||||
// Continuation aid: a run that did not reach `done` can be resumed with the
|
||||
// same goal, carrying its stage progress and (if any) safety feedback.
|
||||
if (final.status !== 'done' && final.runId) {
|
||||
const resume = config.goal
|
||||
? `kilo-loop --goal ${config.goal} --resume ${final.runId}`
|
||||
: `kilo-loop --resume ${final.runId} (pass the same --goal / --goal-text)`
|
||||
process.stdout.write(`\n ${c.yellow('Resume this run')}: ${resume}\n`)
|
||||
}
|
||||
} catch (err) {
|
||||
process.stderr.write(`\n Fatal: ${err?.stack || err}\n`)
|
||||
exitCode = 1
|
||||
|
|
|
|||
10
kilo-loop
10
kilo-loop
|
|
@ -4,6 +4,7 @@
|
|||
# kilo-loop # open the dashboard launcher (pick a plan in the UI)
|
||||
# kilo-loop --goal .kilocode-loop/goal.json -n 5 # run a specific goal directly
|
||||
# kilo-loop --list-goals # print discovered plans + state, then exit
|
||||
# kilo-loop new-plan "My plan title" # scaffold a kilo-loop-compatible plan, then exit
|
||||
# kilo-loop --dry-run -n 3 # offline rehearsal
|
||||
#
|
||||
# Installs node_modules on first use, then runs the loop with the repo root as
|
||||
|
|
@ -12,6 +13,15 @@ set -euo pipefail
|
|||
|
||||
APP_DIR="$(cd "$(dirname "$(readlink -f "$0")")" && pwd)"
|
||||
|
||||
# `kilo-loop new-plan "<title> [flags]"` — deterministic plan scaffolder. Creates
|
||||
# `.kilo/plans/<epoch-ms>-<slug>.md` from `.kilo/templates/plan.md` so new plans
|
||||
# always carry the `- [ ]` stages kilo-loop needs. Handled here (before the loop
|
||||
# runner) so it needs no goal and never starts the dashboard.
|
||||
if [ "${1:-}" = "new-plan" ]; then
|
||||
shift
|
||||
exec node "$APP_DIR/new-plan.mjs" "$@"
|
||||
fi
|
||||
|
||||
# Project = git root of the current directory (submodules resolve to themselves).
|
||||
ROOT="$(git -C "$PWD" rev-parse --show-toplevel 2>/dev/null || printf '%s' "$PWD")"
|
||||
|
||||
|
|
|
|||
179
new-plan.mjs
Normal file
179
new-plan.mjs
Normal file
|
|
@ -0,0 +1,179 @@
|
|||
#!/usr/bin/env node
|
||||
/**
|
||||
* Deterministic plan scaffolder for this repo.
|
||||
*
|
||||
* Creates `.kilo/plans/<epoch-ms>-<slug>.md` from `.kilo/templates/plan.md` so
|
||||
* every new plan is kilo-loop compatible by construction (contains `- [ ]`
|
||||
* stages). kilo-loop only discovers markdown plans with checkbox stages.
|
||||
*
|
||||
* Usage:
|
||||
* node tools/kilocode-loop/new-plan.mjs "My plan title"
|
||||
* node tools/kilocode-loop/new-plan.mjs "My plan title" --slug my-plan
|
||||
* node tools/kilocode-loop/new-plan.mjs "My plan title" --print
|
||||
* node tools/kilocode-loop/new-plan.mjs "My plan title" --force
|
||||
*
|
||||
* Options:
|
||||
* --slug <s> filename slug (default: slugified title)
|
||||
* --epoch <ms> timestamp for the filename (default: now); useful for tests
|
||||
* --dir <d> target directory (default: .kilo/plans)
|
||||
* --project <p> project root (default: git/cwd discovery upward)
|
||||
* --print print the plan to stdout instead of writing a file
|
||||
* --force overwrite an existing file
|
||||
* -h, --help show this help
|
||||
*/
|
||||
import fs from 'node:fs'
|
||||
import path from 'node:path'
|
||||
|
||||
const DEFAULT_DIR = path.join('.kilo', 'plans')
|
||||
const TEMPLATE = path.join('.kilo', 'templates', 'plan.md')
|
||||
const CHECKBOX_RE = /^\s*[-*]\s+\[[ xX]\]/m
|
||||
|
||||
/**
|
||||
* Built-in fallback template, used when the project has no
|
||||
* `.kilo/templates/plan.md`. Guarantees a kilo-loop-compatible plan everywhere.
|
||||
*/
|
||||
const DEFAULT_TEMPLATE = `# {{TITLE}}
|
||||
|
||||
> **Status:** draft.
|
||||
|
||||
## Stages (kilo-loop)
|
||||
|
||||
<!-- One checkbox per deliverable. A plan is discovered only when it contains
|
||||
\`- [ ]\` items. Directives: \`:: verify: <cmd>\` and \`:: acceptance: <text>\`.
|
||||
Mark \`[x]\` only after the stage is actually verified. -->
|
||||
|
||||
- [ ] Stage 1 — <outcome> :: acceptance: <how to confirm>
|
||||
- [ ] Stage 2 — <outcome> :: acceptance: <how to confirm>
|
||||
|
||||
---
|
||||
|
||||
## 0. Goal
|
||||
|
||||
<!-- What we build and why; scope boundaries. -->
|
||||
|
||||
## 1. Current state (verified facts)
|
||||
|
||||
<!-- Facts with \`file:line\`, no assumptions. -->
|
||||
|
||||
## 2. Decisions
|
||||
|
||||
| # | Decision | Choice |
|
||||
|---|----------|--------|
|
||||
| D1 | | |
|
||||
|
||||
## 3. Data model / contracts
|
||||
|
||||
## 4. Backend API
|
||||
|
||||
## 5. Frontend
|
||||
|
||||
## 6. Verification
|
||||
|
||||
<!-- Exact commands/URLs and expected evidence. -->
|
||||
|
||||
## 7. Risks
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| | |
|
||||
`
|
||||
|
||||
function usage() {
|
||||
const txt = fs.readFileSync(new URL(import.meta.url), 'utf8')
|
||||
const block = txt.match(/\/\*\*([\s\S]*?)\*\//)
|
||||
if (block) console.log(block[1].replace(/^\s*\* ?/gm, '').trim())
|
||||
}
|
||||
|
||||
function fail(msg) {
|
||||
console.error(`new-plan: ${msg}`)
|
||||
process.exit(1)
|
||||
}
|
||||
|
||||
function slugify(s) {
|
||||
return s
|
||||
.toLowerCase()
|
||||
.normalize('NFKD')
|
||||
.replace(/[^\p{L}\p{N}]+/gu, '-')
|
||||
.replace(/^-+|-+$/g, '')
|
||||
.slice(0, 80) || 'plan'
|
||||
}
|
||||
|
||||
/** Walk up from `start` until a directory containing `.kilo` (or `.git`) is found. */
|
||||
function findProject(start) {
|
||||
let dir = path.resolve(start)
|
||||
for (;;) {
|
||||
if (fs.existsSync(path.join(dir, '.kilo')) || fs.existsSync(path.join(dir, '.git'))) return dir
|
||||
const parent = path.dirname(dir)
|
||||
if (parent === dir) return path.resolve(start)
|
||||
dir = parent
|
||||
}
|
||||
}
|
||||
|
||||
function parseArgs(argv) {
|
||||
const out = { title: [], slug: '', dir: DEFAULT_DIR, project: '', print: false, force: false, epoch: 0 }
|
||||
for (let i = 0; i < argv.length; i++) {
|
||||
const a = argv[i]
|
||||
if (a === '--slug') out.slug = argv[++i] ?? ''
|
||||
else if (a === '--dir') out.dir = argv[++i] ?? DEFAULT_DIR
|
||||
else if (a === '--project') out.project = argv[++i] ?? ''
|
||||
else if (a === '--epoch') out.epoch = Number(argv[++i]) || 0
|
||||
else if (a === '--print') out.print = true
|
||||
else if (a === '--force') out.force = true
|
||||
else if (a === '-h' || a === '--help') {
|
||||
usage()
|
||||
process.exit(0)
|
||||
} else out.title.push(a)
|
||||
}
|
||||
out.title = out.title.join(' ').trim()
|
||||
return out
|
||||
}
|
||||
|
||||
const args = parseArgs(process.argv.slice(2))
|
||||
if (!args.title) {
|
||||
usage()
|
||||
fail('a plan title is required')
|
||||
}
|
||||
|
||||
const project = args.project ? path.resolve(args.project) : findProject(process.cwd())
|
||||
const templatePath = path.join(project, TEMPLATE)
|
||||
|
||||
let body
|
||||
if (fs.existsSync(templatePath)) {
|
||||
body = fs.readFileSync(templatePath, 'utf8')
|
||||
} else {
|
||||
console.error(`new-plan: no template at ${rel(templatePath)} — using the built-in default`)
|
||||
body = DEFAULT_TEMPLATE
|
||||
}
|
||||
if (!body.includes('{{TITLE}}')) fail(`template is missing the {{TITLE}} placeholder`)
|
||||
|
||||
const slug = slugify(args.slug || args.title)
|
||||
const epoch = args.epoch || Date.now()
|
||||
const fileName = `${epoch}-${slug}.md`
|
||||
const targetDir = path.resolve(project, args.dir)
|
||||
const targetPath = path.join(targetDir, fileName)
|
||||
|
||||
body = body
|
||||
.replaceAll('{{TITLE}}', args.title)
|
||||
.replaceAll('{{SLUG}}', slug)
|
||||
.replaceAll('{{EPOCH}}', String(epoch))
|
||||
.replaceAll('{{DATE}}', new Date(epoch).toISOString())
|
||||
|
||||
// Defensive: keep plans discoverable even if the template is edited badly.
|
||||
if (!CHECKBOX_RE.test(body)) {
|
||||
body += `\n## Stages (kilo-loop)\n\n- [ ] Complete the plan :: acceptance: goal achieved\n`
|
||||
console.error('new-plan: template had no checkbox stages — injected a default Stages section')
|
||||
}
|
||||
|
||||
if (args.print) {
|
||||
process.stdout.write(body)
|
||||
process.exit(0)
|
||||
}
|
||||
|
||||
fs.mkdirSync(targetDir, { recursive: true })
|
||||
if (fs.existsSync(targetPath) && !args.force) fail(`already exists: ${rel(targetPath)} (use --force)`)
|
||||
fs.writeFileSync(targetPath, body)
|
||||
console.log(rel(targetPath))
|
||||
|
||||
function rel(abs) {
|
||||
return path.relative(project, abs).split(path.sep).join('/')
|
||||
}
|
||||
|
|
@ -68,14 +68,14 @@ body { display: flex; flex-direction: column; }
|
|||
rgba(88, 166, 255, 0.13) 66%,
|
||||
rgba(88, 166, 255, 0) 100%);
|
||||
background-size: 240% 100%;
|
||||
animation: topbar-sheen 2.8s linear infinite;
|
||||
animation: topbar-sheen 4.6s linear infinite;
|
||||
}
|
||||
.topbar--running::after { opacity: 1; background: rgba(88, 166, 255, 0.22); }
|
||||
.topbar--running::before {
|
||||
opacity: 1;
|
||||
background: linear-gradient(90deg, rgba(88, 166, 255, 0) 0%, var(--accent) 22%, #7ee787 56%, var(--green) 80%, rgba(63, 185, 80, 0) 100%);
|
||||
box-shadow: 0 0 12px rgba(88, 166, 255, 0.8), 0 0 5px rgba(126, 231, 135, 0.95);
|
||||
animation: topbar-sweep 1.2s cubic-bezier(0.45, 0, 0.35, 1) infinite;
|
||||
animation: topbar-sweep 2.8s cubic-bezier(0.45, 0, 0.35, 1) infinite;
|
||||
}
|
||||
|
||||
/* waiting for operator input — pulsing purple */
|
||||
|
|
|
|||
|
|
@ -22,6 +22,9 @@ export const DEFAULTS = {
|
|||
confirm: true,
|
||||
port: 7999,
|
||||
host: '127.0.0.1',
|
||||
// Bind every interface (0.0.0.0) so another device on the same network — e.g.
|
||||
// a phone over Wi-Fi — can open the dashboard. `--host`/`--local` override it.
|
||||
lan: false,
|
||||
runId: '',
|
||||
resume: '',
|
||||
dryRun: false,
|
||||
|
|
@ -44,9 +47,14 @@ export const DEFAULTS = {
|
|||
maxStaleIterations: 0,
|
||||
// Git checkpoint before every iteration, revertible from the dashboard/CLI.
|
||||
checkpoint: true,
|
||||
// Deny destructive bash commands (git commit/push, rm -rf, …) and abort on
|
||||
// detection.
|
||||
// Deny destructive bash commands (git commit/push, rm -rf, …). On a runtime
|
||||
// hit the session is aborted; by default the loop then continues with the next
|
||||
// iteration (safety feedback is fed back in). Set guardStop to stop the run.
|
||||
guard: true,
|
||||
guardStop: false,
|
||||
// With guardStop off, stop the run anyway after this many consecutive
|
||||
// guard-blocked iterations so a stuck agent cannot spin forever (invalid → 3).
|
||||
maxGuardBlocks: 3,
|
||||
// Scrub secrets out of logs, state and reports before they are persisted.
|
||||
redact: true,
|
||||
// Extra reviewer pass after a stage passes verification.
|
||||
|
|
@ -93,6 +101,7 @@ const STRING_FLAGS = new Set([
|
|||
'max-cost',
|
||||
'max-tokens',
|
||||
'max-stale-iterations',
|
||||
'max-guard-blocks',
|
||||
'max-retries',
|
||||
'retry-delay',
|
||||
'retry-max-delay',
|
||||
|
|
@ -107,12 +116,15 @@ const BOOL_FLAGS = new Set([
|
|||
'shared-context',
|
||||
'confirm',
|
||||
'yes',
|
||||
'lan',
|
||||
'local',
|
||||
'dry-run',
|
||||
'keep-open',
|
||||
'quiet',
|
||||
'color',
|
||||
'checkpoint',
|
||||
'guard',
|
||||
'guard-stop',
|
||||
'redact',
|
||||
'review',
|
||||
'sync-goal',
|
||||
|
|
@ -144,6 +156,7 @@ const INT_FLAGS = new Set([
|
|||
'verify-timeout-minutes',
|
||||
'max-tokens',
|
||||
'max-stale-iterations',
|
||||
'max-guard-blocks',
|
||||
'auto-fresh-tokens',
|
||||
'max-retries',
|
||||
'retry-delay',
|
||||
|
|
@ -263,6 +276,17 @@ export function loadConfig(argv = process.argv.slice(2)) {
|
|||
if (merged.yes) merged.confirm = false
|
||||
merged.confirm = merged.confirm !== false
|
||||
|
||||
// Dashboard binding. Default is loopback; `--lan` (or `"lan": true` in the
|
||||
// project config) binds every interface so a phone on the same Wi-Fi can open
|
||||
// it. An explicit `--host` always wins, and `--local` forces loopback back on.
|
||||
if (parsed.local) {
|
||||
merged.lan = false
|
||||
merged.host = '127.0.0.1'
|
||||
} else if (parsed.host === undefined && merged.lan) {
|
||||
merged.host = '0.0.0.0'
|
||||
}
|
||||
merged.lan = Boolean(merged.lan)
|
||||
|
||||
if (!['always', 'on-question', 'off'].includes(merged.hitl)) {
|
||||
throw new Error(`--hitl must be "always", "on-question" or "off", got "${merged.hitl}"`)
|
||||
}
|
||||
|
|
@ -279,7 +303,7 @@ export function loadConfig(argv = process.argv.slice(2)) {
|
|||
}
|
||||
|
||||
// Clamp numeric limits that may come from a config file (CLI already validates).
|
||||
for (const key of ['maxCost', 'maxTokens', 'maxStaleIterations', 'autoFreshTokens', 'verifyTimeoutMinutes']) {
|
||||
for (const key of ['maxCost', 'maxTokens', 'maxStaleIterations', 'maxGuardBlocks', 'autoFreshTokens', 'verifyTimeoutMinutes']) {
|
||||
const n = Number(merged[key])
|
||||
if (!Number.isFinite(n) || n < 0) merged[key] = DEFAULTS[key]
|
||||
}
|
||||
|
|
@ -321,7 +345,9 @@ Verification & safety:
|
|||
--verify "<cmd>" Command the loop runs itself; a stage is not 'done' until it exits 0
|
||||
--verify-timeout-minutes <n> Timeout for the verify command (default ${DEFAULTS.verifyTimeoutMinutes})
|
||||
--checkpoint / --no-checkpoint Git checkpoint before each iteration (default on; revertible)
|
||||
--guard / --no-guard Deny destructive bash (git commit/push, rm -rf, …) and abort on detection
|
||||
--guard / --no-guard Deny destructive bash (git commit/push, rm -rf, …); a hit aborts the session
|
||||
--guard-stop Stop the whole run on the first guard hit (default: continue to the next iteration)
|
||||
--max-guard-blocks <n> With guard-stop off, stop after N consecutive guard hits (default ${DEFAULTS.maxGuardBlocks})
|
||||
--redact / --no-redact Scrub secrets from logs/state/reports before persisting
|
||||
--max-cost <usd> Stop the loop when total cost reaches this (0 = none)
|
||||
--max-tokens <n> Stop the loop when total tokens reach this (0 = none)
|
||||
|
|
@ -341,6 +367,8 @@ Verification & safety:
|
|||
Dashboard:
|
||||
--port, -p <n> Express dashboard port (default ${DEFAULTS.port})
|
||||
--host <host> Bind host (default ${DEFAULTS.host})
|
||||
--lan Bind every interface (0.0.0.0) so a phone on the same Wi-Fi can open it
|
||||
--local Force loopback only (default; overrides a config "lan": true)
|
||||
--keep-open Keep the dashboard running after the loop ends
|
||||
|
||||
Startup:
|
||||
|
|
@ -349,7 +377,7 @@ Startup:
|
|||
|
||||
Utility:
|
||||
--dry-run Simulate the agent (no API calls) — exercises the whole loop/UI
|
||||
--resume <runId> Resume/continue a previously saved run's session
|
||||
--resume <runId|last> Resume a saved run's session + stage progress ("last" = most recent run)
|
||||
--run-id <id> Explicit run id (default: timestamp)
|
||||
--quiet Suppress the live console stream (reports still print)
|
||||
--no-color Disable ANSI colours
|
||||
|
|
|
|||
|
|
@ -416,6 +416,11 @@ ${
|
|||
? `\n## Verification feedback (from the loop)\nThe last verify run FAILED; fix the cause, do not work around the check:\n\n\`\`\`\n${ctx.verifyFeedback}\n\`\`\`\n`
|
||||
: ''
|
||||
}
|
||||
${
|
||||
ctx.guardFeedback
|
||||
? `\n## Loop safety feedback (from the loop)\nAn earlier iteration was interrupted because a command tripped the unattended-loop safety guard. Do NOT repeat that command; achieve the same result another way (a temp path, a different tool, or a non-destructive approach).\n\n\`\`\`\n${ctx.guardFeedback}\n\`\`\`\n`
|
||||
: ''
|
||||
}
|
||||
|
||||
## Progress from previous iterations
|
||||
${ctx.progressSummary || 'This is the first iteration.'}
|
||||
|
|
@ -497,6 +502,7 @@ export function buildIterationPrompt({
|
|||
handoff = null,
|
||||
filesChanged = [],
|
||||
verifyFeedback = '',
|
||||
guardFeedback = '',
|
||||
}) {
|
||||
const p = loopPaths(config.project)
|
||||
const freshMode = !sharedContext
|
||||
|
|
@ -543,6 +549,7 @@ export function buildIterationPrompt({
|
|||
stageCount: goal.stages?.length || 1,
|
||||
verifyCommand,
|
||||
verifyFeedback,
|
||||
guardFeedback,
|
||||
progressSummary,
|
||||
answersText,
|
||||
progressPath: p.progress,
|
||||
|
|
|
|||
112
src/guard.mjs
112
src/guard.mjs
|
|
@ -46,54 +46,104 @@ function isInside(abs, root) {
|
|||
return abs.startsWith(prefix)
|
||||
}
|
||||
|
||||
function isSafeRmTarget(target, roots) {
|
||||
const cleaned = String(target).replace(/^['"]|['"]$/g, '')
|
||||
if (!cleaned || cleaned === '-' ) return false
|
||||
function stripQuotes(value) {
|
||||
const s = String(value).trim()
|
||||
if (s.length >= 2 && ((s[0] === '"' && s.at(-1) === '"') || (s[0] === "'" && s.at(-1) === "'"))) return s.slice(1, -1)
|
||||
return s.replace(/^['"]|['"]$/g, '')
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve a `cd` target to an absolute path, or null when it cannot be trusted
|
||||
* (expansion/glob, `cd -`, or a relative path with no known cwd).
|
||||
*/
|
||||
function resolveCdTarget(target, cwd) {
|
||||
const cleaned = stripQuotes(target)
|
||||
if (!cleaned || cleaned === '-') return null
|
||||
if (/[$*?`~]/.test(cleaned)) return null
|
||||
if (path.isAbsolute(cleaned)) return path.resolve(cleaned)
|
||||
return cwd ? path.resolve(cwd, cleaned) : null
|
||||
}
|
||||
|
||||
function isSafeRmTarget(target, roots, cwd = null) {
|
||||
const cleaned = stripQuotes(target)
|
||||
if (!cleaned || cleaned === '-') return false
|
||||
// Unresolvable expansions or globs are treated as unsafe (conservative).
|
||||
if (/[$*?`~]/.test(cleaned)) return false
|
||||
if (!path.isAbsolute(cleaned)) return false
|
||||
const abs = path.resolve(cleaned)
|
||||
// A relative target is only trustworthy when a preceding `cd` (in the same
|
||||
// command) established a known working directory — e.g. `cd /tmp && rm -rf x`.
|
||||
let abs
|
||||
if (path.isAbsolute(cleaned)) abs = path.resolve(cleaned)
|
||||
else if (cwd) abs = path.resolve(cwd, cleaned)
|
||||
else return false
|
||||
if (abs === path.parse(abs).root) return false
|
||||
return roots.some((root) => isInside(abs, root))
|
||||
}
|
||||
|
||||
/** Command separators that start a new shell statement (and may change cwd). */
|
||||
const SEGMENT_SPLIT = /(?:&&|\|\||;|\n|\|)/
|
||||
|
||||
/**
|
||||
* True when every `rm` in the command that uses recursive+force flags targets
|
||||
* only paths strictly inside a temporary directory. Deleting scratch dirs like
|
||||
* `/tmp/<run>` is routine; deleting project or system paths is not.
|
||||
* True when every recursive+force `rm` targets only paths strictly inside a
|
||||
* temporary directory. Deleting scratch dirs like `/tmp/<run>` is routine;
|
||||
* deleting project or system paths is not.
|
||||
*
|
||||
* The check is `cd`-aware: `cd /tmp && rm -rf scratch` resolves the relative
|
||||
* target against the cwd set earlier in the same command, so routine scratch
|
||||
* cleanup is not mistaken for a project delete. `cd /repo && rm -rf src` stays
|
||||
* blocked because `/repo` is not a temp root.
|
||||
*/
|
||||
export function isSafeRecursiveRm(command, roots = safeRmRoots()) {
|
||||
const text = String(command || '')
|
||||
const invocations = text.match(/(?:^|[;&|(]\s*)rm\s+[^\n;&|)]*/g)
|
||||
if (!invocations) return false
|
||||
let cwd = null
|
||||
let checked = 0
|
||||
for (const raw of invocations) {
|
||||
const tokens = raw
|
||||
.replace(/^[;&|(]\s*/, '')
|
||||
.split(/\s+/)
|
||||
.filter(Boolean)
|
||||
const args = tokens.slice(1)
|
||||
let hasR = false
|
||||
let hasF = false
|
||||
const targets = []
|
||||
for (const arg of args) {
|
||||
if (arg === '--') continue
|
||||
if (arg.startsWith('-')) {
|
||||
if (/[rR]/.test(arg)) hasR = true
|
||||
if (/[fF]/.test(arg)) hasF = true
|
||||
continue
|
||||
}
|
||||
targets.push(arg)
|
||||
for (const rawSegment of text.split(SEGMENT_SPLIT)) {
|
||||
const segment = rawSegment.trim().replace(/^[()\s]+/, '')
|
||||
if (!segment) continue
|
||||
|
||||
const cd = segment.match(/^cd\s+(?:--\s+)?(.+)$/)
|
||||
if (cd) {
|
||||
cwd = resolveCdTarget(cd[1], cwd)
|
||||
continue
|
||||
}
|
||||
|
||||
const invocations = segment.match(/(?:^|[;&|(]\s*)rm\s+[^\n;&|)]*/g)
|
||||
if (!invocations) continue
|
||||
for (const raw of invocations) {
|
||||
const tokens = raw
|
||||
.replace(/^[;&|(]\s*/, '')
|
||||
.split(/\s+/)
|
||||
.filter(Boolean)
|
||||
const args = tokens.slice(1)
|
||||
let hasR = false
|
||||
let hasF = false
|
||||
const targets = []
|
||||
for (const arg of args) {
|
||||
if (arg === '--') continue
|
||||
if (arg.startsWith('-')) {
|
||||
if (/[rR]/.test(arg)) hasR = true
|
||||
if (/[fF]/.test(arg)) hasF = true
|
||||
continue
|
||||
}
|
||||
targets.push(arg)
|
||||
}
|
||||
if (!hasR || !hasF) continue
|
||||
checked += 1
|
||||
if (!targets.length) return false
|
||||
if (!targets.every((t) => isSafeRmTarget(t, roots, cwd))) return false
|
||||
}
|
||||
if (!hasR || !hasF) continue
|
||||
checked += 1
|
||||
if (!targets.length) return false
|
||||
if (!targets.every((t) => isSafeRmTarget(t, roots))) return false
|
||||
}
|
||||
// No recursive-force rm at all → this rule does not apply.
|
||||
return checked > 0
|
||||
}
|
||||
|
||||
/** Human-readable feedback describing a guard hit, for the next iteration prompt. */
|
||||
export function guardFeedbackText(violation) {
|
||||
if (!violation) return ''
|
||||
const { id, reason, command } = violation
|
||||
return `The safety guard blocked a command in a previous iteration (\`${id}\`: ${reason}).\nDo NOT repeat it.` +
|
||||
(command ? `\n blocked command: ${command}` : '')
|
||||
}
|
||||
|
||||
/**
|
||||
* Return the first matching rule for a shell command, or null.
|
||||
* Recursive `rm` is skipped when every target is a temp directory.
|
||||
|
|
|
|||
60
src/net.mjs
Normal file
60
src/net.mjs
Normal file
|
|
@ -0,0 +1,60 @@
|
|||
import os from 'node:os'
|
||||
|
||||
/**
|
||||
* Network helpers for the dashboard.
|
||||
*
|
||||
* `--lan` binds the dashboard to every interface (`0.0.0.0`) so it is reachable
|
||||
* from another device on the same network (e.g. a phone over Wi-Fi). Binding to
|
||||
* a wildcard address is not usable as a URL, so the CLI prints the machine's
|
||||
* real LAN address instead of `0.0.0.0`.
|
||||
*/
|
||||
|
||||
/** True for addresses that mean "all interfaces". */
|
||||
export function isWildcardHost(host) {
|
||||
const h = String(host || '').trim().toLowerCase()
|
||||
return h === '' || h === '0.0.0.0' || h === '::' || h === '[::]' || h === '*'
|
||||
}
|
||||
|
||||
/** True for addresses reachable only from the same machine. */
|
||||
export function isLoopbackHost(host) {
|
||||
const h = String(host || '').trim().toLowerCase()
|
||||
return h === '127.0.0.1' || h === 'localhost' || h === '::1' || h === '[::1]'
|
||||
}
|
||||
|
||||
/**
|
||||
* Virtual interfaces (docker/vpn/bridges) are excluded: a phone on the same
|
||||
* Wi-Fi cannot reach them, and they only add confusing URLs to the output.
|
||||
*/
|
||||
const VIRTUAL_IFACE = /^(docker|br-|virbr|veth|vmnet|vbox|tun|tap|wg|zt|tailscale|lo\b|utun|awdl|llw|anpi|bridge|ham)/i
|
||||
|
||||
/**
|
||||
* IPv4 addresses of physical, non-internal interfaces, best-effort. Order
|
||||
* follows the OS (the primary interface usually comes first). Empty when the
|
||||
* machine is offline.
|
||||
*/
|
||||
export function lanAddresses() {
|
||||
const out = []
|
||||
for (const [name, infos] of Object.entries(os.networkInterfaces() || {})) {
|
||||
if (VIRTUAL_IFACE.test(name)) continue
|
||||
for (const info of infos || []) {
|
||||
if (info && info.family === 'IPv4' && !info.internal) out.push(info.address)
|
||||
}
|
||||
}
|
||||
return [...new Set(out)]
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the dashboard URLs for a config. `primary` is what the console should
|
||||
* link to; `lan` lists the other-device URLs (empty when bound to loopback).
|
||||
*/
|
||||
export function dashboardUrls(config) {
|
||||
const port = config.port
|
||||
const host = config.host
|
||||
const local = `http://127.0.0.1:${port}`
|
||||
if (isWildcardHost(host)) {
|
||||
const lan = lanAddresses().map((ip) => `http://${ip}:${port}`)
|
||||
return { local, lan, primary: lan[0] || local, wildcard: true }
|
||||
}
|
||||
const url = `http://${host}:${port}`
|
||||
return { local: isLoopbackHost(host) ? url : local, lan: isLoopbackHost(host) ? [] : [url], primary: url, wildcard: false }
|
||||
}
|
||||
|
|
@ -34,7 +34,8 @@ import { printSection } from './console.mjs'
|
|||
import { redact, redactJson } from './redact.mjs'
|
||||
import { runVerify, summariseVerify, verifyCommandFor } from './verify.mjs'
|
||||
import { backoffDelay, transientReason } from './retry.mjs'
|
||||
import { scanToolUse } from './guard.mjs'
|
||||
import { guardFeedbackText, scanToolUse } from './guard.mjs'
|
||||
import { dashboardUrls } from './net.mjs'
|
||||
import {
|
||||
createCheckpoint,
|
||||
isGitRepo,
|
||||
|
|
@ -74,6 +75,10 @@ export class Orchestrator {
|
|||
this.verifyFailures = new Map()
|
||||
this.stopReason = null
|
||||
this.guardViolation = null
|
||||
// Safety feedback carried into the next iteration after a guard hit, and a
|
||||
// counter so a persistently blocked agent cannot spin the loop forever.
|
||||
this.guardFeedback = ''
|
||||
this.consecutiveGuardBlocks = 0
|
||||
this.onReport = null
|
||||
}
|
||||
|
||||
|
|
@ -83,7 +88,7 @@ export class Orchestrator {
|
|||
|
||||
/** Fire-and-forget notification; never blocks or throws into the loop. */
|
||||
_notify(event, payload = {}) {
|
||||
const p = { runId: this.state.runId, goal: this.state.state.goal?.title, url: `http://${this.config.host}:${this.config.port}`, ...payload }
|
||||
const p = { runId: this.state.runId, goal: this.state.state.goal?.title, url: dashboardUrls(this.config).primary, ...payload }
|
||||
return sendNotify(this.config, event, p, { onLog: (m) => this.log('warn', m) }).catch(() => {})
|
||||
}
|
||||
|
||||
|
|
@ -225,7 +230,11 @@ export class Orchestrator {
|
|||
'system',
|
||||
`Run ${state.runId} · goal "${goal.title}" · ${config.iterations} iteration(s) · agent ${config.agent} · ${config.sharedContext ? 'shared context' : 'fresh sessions (handoff)'}${config.dryRun ? ' · DRY-RUN' : ''}`,
|
||||
)
|
||||
this.log('system', `Dashboard: http://${config.host}:${config.port} (Ctrl+Click to open)`)
|
||||
{
|
||||
const urls = dashboardUrls(config)
|
||||
this.log('system', `Dashboard: ${urls.primary} (Ctrl+Click to open)`)
|
||||
for (const u of urls.lan) if (u !== urls.primary) this.log('system', `On this network (e.g. phone): ${u}`)
|
||||
}
|
||||
this.log('system', `Project: ${config.project} · CLI runtime: ${resolveRuntime(config)}`)
|
||||
this.log(
|
||||
'system',
|
||||
|
|
@ -259,12 +268,36 @@ export class Orchestrator {
|
|||
const report = await this._runIteration(goal, i, config.iterations)
|
||||
if (!report) break
|
||||
if (report.status === 'failed') failed = true
|
||||
if (report.guardViolation) {
|
||||
this.stopReason = 'guard'
|
||||
break
|
||||
}
|
||||
|
||||
this._trackStall(report)
|
||||
// A guard hit aborts the destructive session — not the whole run. The loop
|
||||
// continues with the same (still incomplete) stage and hands the safety
|
||||
// feedback to the next iteration so the agent can reach the same result a
|
||||
// safer way. `--guard-stop` restores the old hard-stop; `--max-guard-blocks`
|
||||
// (default 3) stops a persistently blocked agent from spinning.
|
||||
if (report.guardViolation) {
|
||||
this.consecutiveGuardBlocks += 1
|
||||
const limit = Math.max(1, Number(config.maxGuardBlocks) || 3)
|
||||
if (config.guardStop || this.consecutiveGuardBlocks >= limit) {
|
||||
this.stopReason = 'guard'
|
||||
break
|
||||
}
|
||||
this.guardFeedback = guardFeedbackText(report.guardViolation)
|
||||
// A guard hit aborts the session mid-turn, so continuing that session in
|
||||
// shared mode is unreliable: switch to a fresh session with a synthesized
|
||||
// handoff so the next iteration starts clean.
|
||||
if (config.sharedContext) {
|
||||
this._synthesizeHandoff(readProgress(config), 'guard hit aborted the session')
|
||||
this.setSharedContext(false)
|
||||
}
|
||||
this.log(
|
||||
'warn',
|
||||
`Guard blocked \`${report.guardViolation.id}\` in iteration ${i} — session aborted, run continues with safety feedback (${this.consecutiveGuardBlocks}/${limit}; use --guard-stop to stop instead).`,
|
||||
)
|
||||
} else {
|
||||
this.consecutiveGuardBlocks = 0
|
||||
this.guardFeedback = ''
|
||||
this._trackStall(report)
|
||||
}
|
||||
|
||||
const shouldAsk = config.hitl === 'always' || (config.hitl === 'on-question' && report.needsInput)
|
||||
if (shouldAsk && !this._stopped && !state.state.abortRequested) {
|
||||
|
|
@ -406,6 +439,7 @@ export class Orchestrator {
|
|||
handoff,
|
||||
filesChanged: [...this.filesSeen],
|
||||
verifyFeedback,
|
||||
guardFeedback: this.guardFeedback || '',
|
||||
})
|
||||
|
||||
const iterDir = path.join(state.dir, 'iterations')
|
||||
|
|
@ -843,7 +877,9 @@ export class Orchestrator {
|
|||
rec.guardViolation = violation
|
||||
this.guardViolation = violation
|
||||
this.log('error', `Command guard: BLOCKED ${hit.id} — ${hit.reason}\n command: ${hit.command}`, { iteration })
|
||||
this.state.update({ abortRequested: true, guardViolation: violation })
|
||||
// Abort the current session (damage control) but not the whole run:
|
||||
// the loop decides what to do with the hit (continue or --guard-stop).
|
||||
this.state.update({ guardViolation: violation })
|
||||
this._notify('guard', { iteration, detail: `${hit.id}: ${hit.reason}`, verify: hit.command })
|
||||
try {
|
||||
this._currentAbort?.abort()
|
||||
|
|
|
|||
|
|
@ -3,6 +3,7 @@ import readline from 'node:readline'
|
|||
import { c } from './ansi.mjs'
|
||||
import { box } from './reports.mjs'
|
||||
import { resolveRuntime } from './kilocode.mjs'
|
||||
import { dashboardUrls } from './net.mjs'
|
||||
|
||||
function indentBlock(text, pad = ' ') {
|
||||
return String(text || '')
|
||||
|
|
@ -23,6 +24,7 @@ export function renderPreflight(config, goal, url) {
|
|||
if (config.maxCost) limits.push(`cost ≤ $${config.maxCost}`)
|
||||
if (config.maxTokens) limits.push(`tokens ≤ ${config.maxTokens}`)
|
||||
if (config.maxStaleIterations) limits.push(`stale ≤ ${config.maxStaleIterations}`)
|
||||
const otherUrls = dashboardUrls(config).lan.filter((u) => u !== url)
|
||||
const rows = [
|
||||
['Goal', c.bold(goal.title || '—')],
|
||||
['Source', goal.source && goal.source !== 'inline' ? goal.source : '(inline --goal-text)'],
|
||||
|
|
@ -39,6 +41,7 @@ export function renderPreflight(config, goal, url) {
|
|||
['Reviewer', config.review ? `on (${config.reviewAgent})` : 'off'],
|
||||
['Notify', config.notify ? 'configured' : 'off'],
|
||||
['Dashboard', url],
|
||||
...(otherUrls.length ? [['On this network', otherUrls.join(' ')]] : []),
|
||||
['Mode', config.dryRun ? c.yellow('DRY-RUN (no API calls)') : 'live'],
|
||||
]
|
||||
|
||||
|
|
|
|||
|
|
@ -4,6 +4,7 @@ import { fileURLToPath } from 'node:url'
|
|||
import express from 'express'
|
||||
|
||||
import { listRuns, loadRun, readRunLog } from './state.mjs'
|
||||
import { dashboardUrls } from './net.mjs'
|
||||
import { listCheckpoints, checkpointDir } from './checkpoint.mjs'
|
||||
import { RunController } from './controller.mjs'
|
||||
import {
|
||||
|
|
@ -288,11 +289,14 @@ export function startServer(config, target) {
|
|||
server.close(() => r())
|
||||
server.closeAllConnections?.()
|
||||
})
|
||||
const urls = dashboardUrls(config)
|
||||
resolve({
|
||||
app,
|
||||
server,
|
||||
controller,
|
||||
url: `http://${config.host}:${config.port}`,
|
||||
// A wildcard bind (0.0.0.0) is not a usable URL; report the LAN address.
|
||||
url: urls.primary,
|
||||
urls,
|
||||
close,
|
||||
})
|
||||
})
|
||||
|
|
|
|||
|
|
@ -15,6 +15,8 @@ test('parseArgs handles the new safety/verification flags', () => {
|
|||
'5',
|
||||
'--auto-fresh-tokens',
|
||||
'90000',
|
||||
'--max-guard-blocks',
|
||||
'2',
|
||||
'--no-checkpoint',
|
||||
'--no-guard',
|
||||
'--no-redact',
|
||||
|
|
@ -23,6 +25,7 @@ test('parseArgs handles the new safety/verification flags', () => {
|
|||
assert.equal(out.maxCost, 1.25)
|
||||
assert.equal(out.maxTokens, 500000)
|
||||
assert.equal(out.maxStaleIterations, 3)
|
||||
assert.equal(out.maxGuardBlocks, 2)
|
||||
assert.equal(out.verifyTimeoutMinutes, 5)
|
||||
assert.equal(out.autoFreshTokens, 90000)
|
||||
assert.equal(out.checkpoint, false)
|
||||
|
|
@ -31,6 +34,21 @@ test('parseArgs handles the new safety/verification flags', () => {
|
|||
assert.equal(out.syncGoal, false)
|
||||
})
|
||||
|
||||
test('parseArgs handles --guard-stop', () => {
|
||||
assert.equal(parseArgs(['--guard-stop']).guardStop, true)
|
||||
assert.equal(parseArgs(['--no-guard-stop']).guardStop, false)
|
||||
})
|
||||
|
||||
test('--lan binds all interfaces; --local and --host override it', () => {
|
||||
assert.equal(loadConfig(['--goal-text', 'x']).host, '127.0.0.1')
|
||||
assert.equal(loadConfig(['--goal-text', 'x']).lan, false)
|
||||
assert.equal(loadConfig(['--goal-text', 'x', '--lan']).host, '0.0.0.0')
|
||||
assert.equal(loadConfig(['--goal-text', 'x', '--lan']).lan, true)
|
||||
assert.equal(loadConfig(['--goal-text', 'x', '--lan', '--local']).host, '127.0.0.1')
|
||||
assert.equal(loadConfig(['--goal-text', 'x', '--lan', '--local']).lan, false)
|
||||
assert.equal(loadConfig(['--goal-text', 'x', '--lan', '--host', '192.168.1.9']).host, '192.168.1.9')
|
||||
})
|
||||
|
||||
test('parseArgs rejects negative and non-numeric budgets', () => {
|
||||
assert.throws(() => parseArgs(['--max-cost', '-1']), /Invalid number/)
|
||||
assert.throws(() => parseArgs(['--max-tokens', 'abc']), /Invalid number/)
|
||||
|
|
@ -44,6 +62,8 @@ test('loadConfig exposes safe defaults for the new options', () => {
|
|||
assert.equal(config.maxStaleIterations, 0)
|
||||
assert.equal(config.checkpoint, true)
|
||||
assert.equal(config.guard, true)
|
||||
assert.equal(config.guardStop, false)
|
||||
assert.equal(config.maxGuardBlocks, 3)
|
||||
assert.equal(config.redact, true)
|
||||
assert.equal(config.review, false)
|
||||
assert.equal(config.autoFreshTokens, 0)
|
||||
|
|
|
|||
|
|
@ -30,6 +30,31 @@ test('scanCommand allows recursive deletes that only touch temp dirs', () => {
|
|||
assert.equal(isSafeRecursiveRm('rm -rf /tmp/kilo/tf-test'), true)
|
||||
})
|
||||
|
||||
test('scanCommand allows relative recursive deletes after cd into a temp dir', () => {
|
||||
const cmds = [
|
||||
'cd /tmp && rm -rf tpl-onlineshop',
|
||||
'cd /tmp/kilo && rm -rf ./scratch',
|
||||
'cd /tmp && rm -rf a b && git clone ssh://x/a b',
|
||||
'mkdir -p /tmp/x && cd /tmp/x && rm -rf build',
|
||||
'cd /var/tmp && rm -rf cache',
|
||||
]
|
||||
for (const cmd of cmds) assert.equal(scanCommand(cmd), null, `expected allowed: ${cmd}`)
|
||||
})
|
||||
|
||||
test('scanCommand blocks relative recursive deletes after cd into a non-temp dir', () => {
|
||||
const cmds = [
|
||||
'cd /repo && rm -rf src',
|
||||
'cd /home/joe/work && rm -rf node_modules',
|
||||
'cd /tmp && rm -rf /home/joe/work', // an unsafe absolute target wins
|
||||
'cd /tmp && rm -rf ..', // escapes the temp root
|
||||
]
|
||||
for (const cmd of cmds) {
|
||||
const hit = scanCommand(cmd)
|
||||
assert.ok(hit, `expected blocked: ${cmd}`)
|
||||
assert.equal(hit.id, 'rm-recursive-force')
|
||||
}
|
||||
})
|
||||
|
||||
test('scanCommand still blocks recursive deletes outside temp dirs', () => {
|
||||
const cmds = [
|
||||
'rm -rf build',
|
||||
|
|
|
|||
35
test/net.test.mjs
Normal file
35
test/net.test.mjs
Normal file
|
|
@ -0,0 +1,35 @@
|
|||
import { test } from 'node:test'
|
||||
import assert from 'node:assert/strict'
|
||||
|
||||
import { isWildcardHost, isLoopbackHost, dashboardUrls, lanAddresses } from '../src/net.mjs'
|
||||
|
||||
test('host classification', () => {
|
||||
assert.equal(isWildcardHost('0.0.0.0'), true)
|
||||
assert.equal(isWildcardHost('::'), true)
|
||||
assert.equal(isWildcardHost('127.0.0.1'), false)
|
||||
assert.equal(isLoopbackHost('127.0.0.1'), true)
|
||||
assert.equal(isLoopbackHost('localhost'), true)
|
||||
assert.equal(isLoopbackHost('::1'), true)
|
||||
assert.equal(isLoopbackHost('192.168.1.5'), false)
|
||||
})
|
||||
|
||||
test('dashboardUrls on loopback has no LAN URLs', () => {
|
||||
const u = dashboardUrls({ host: '127.0.0.1', port: 7999 })
|
||||
assert.equal(u.primary, 'http://127.0.0.1:7999')
|
||||
assert.deepEqual(u.lan, [])
|
||||
assert.equal(u.wildcard, false)
|
||||
})
|
||||
|
||||
test('dashboardUrls on a wildcard host resolves to a usable LAN address', () => {
|
||||
const u = dashboardUrls({ host: '0.0.0.0', port: 7999 })
|
||||
assert.equal(u.wildcard, true)
|
||||
assert.ok(u.primary.startsWith('http://'))
|
||||
assert.ok(!u.primary.includes('0.0.0.0'), `primary must not be the wildcard address: ${u.primary}`)
|
||||
assert.deepEqual(u.lan, lanAddresses().map((ip) => `http://${ip}:7999`))
|
||||
})
|
||||
|
||||
test('dashboardUrls on a specific non-loopback host exposes it', () => {
|
||||
const u = dashboardUrls({ host: '10.0.0.5', port: 8000 })
|
||||
assert.equal(u.primary, 'http://10.0.0.5:8000')
|
||||
assert.deepEqual(u.lan, ['http://10.0.0.5:8000'])
|
||||
})
|
||||
70
test/new-plan.test.mjs
Normal file
70
test/new-plan.test.mjs
Normal file
|
|
@ -0,0 +1,70 @@
|
|||
import { test } from 'node:test'
|
||||
import assert from 'node:assert/strict'
|
||||
import fs from 'node:fs'
|
||||
import os from 'node:os'
|
||||
import path from 'node:path'
|
||||
import { execFileSync } from 'node:child_process'
|
||||
import { fileURLToPath } from 'node:url'
|
||||
|
||||
import { discoverGoalFiles } from '../src/goals.mjs'
|
||||
|
||||
const SCRIPT = fileURLToPath(new URL('../new-plan.mjs', import.meta.url))
|
||||
|
||||
function tmpProject() {
|
||||
return fs.mkdtempSync(path.join(os.tmpdir(), 'kilo-new-plan-'))
|
||||
}
|
||||
|
||||
function run(args, project) {
|
||||
return execFileSync(process.execPath, [SCRIPT, ...args, '--project', project], { encoding: 'utf8' })
|
||||
}
|
||||
|
||||
test('new-plan writes a discoverable plan from the project template', () => {
|
||||
const project = tmpProject()
|
||||
fs.mkdirSync(path.join(project, '.kilo', 'templates'), { recursive: true })
|
||||
fs.writeFileSync(
|
||||
path.join(project, '.kilo', 'templates', 'plan.md'),
|
||||
'# {{TITLE}}\n\n## Stages (kilo-loop)\n\n- [ ] Ship it :: verify: echo ok :: acceptance: shipped\n',
|
||||
)
|
||||
|
||||
const out = run(['My Feature'], project).trim()
|
||||
const abs = path.join(project, out)
|
||||
assert.ok(fs.existsSync(abs), `plan file exists: ${out}`)
|
||||
assert.match(path.basename(out), /^\d+-my-feature\.md$/)
|
||||
|
||||
const body = fs.readFileSync(abs, 'utf8')
|
||||
assert.match(body, /^# My Feature$/m)
|
||||
assert.match(body, /- \[ \] Ship it/)
|
||||
|
||||
// The whole point: kilo-loop must discover it.
|
||||
assert.ok(discoverGoalFiles(project).some((f) => path.resolve(f) === path.resolve(abs)))
|
||||
})
|
||||
|
||||
test('new-plan falls back to a built-in template when the project has none', () => {
|
||||
const project = tmpProject()
|
||||
const out = run(['Standalone'], project).trim()
|
||||
const body = fs.readFileSync(path.join(project, out), 'utf8')
|
||||
assert.match(body, /^# Standalone$/m)
|
||||
assert.match(body, /## Stages \(kilo-loop\)/)
|
||||
assert.match(body, /- \[ \] Stage 1/)
|
||||
})
|
||||
|
||||
test('new-plan --print does not write a file', () => {
|
||||
const project = tmpProject()
|
||||
const out = run(['Preview only', '--print'], project)
|
||||
assert.match(out, /^# Preview only$/m)
|
||||
assert.equal(fs.existsSync(path.join(project, '.kilo', 'plans')), false)
|
||||
})
|
||||
|
||||
test('new-plan refuses to overwrite without --force', () => {
|
||||
const project = tmpProject()
|
||||
const first = run(['Same', '--slug', 'same', '--epoch', '1700000000000'], project).trim()
|
||||
assert.equal(first, '.kilo/plans/1700000000000-same.md')
|
||||
assert.throws(() => run(['Same', '--slug', 'same', '--epoch', '1700000000000'], project), /already exists/)
|
||||
// --force makes it succeed again.
|
||||
assert.doesNotThrow(() =>
|
||||
execFileSync(process.execPath, [
|
||||
SCRIPT, 'Same', '--slug', 'same', '--epoch', '1700000000000', '--force', '--project', project,
|
||||
]),
|
||||
)
|
||||
assert.ok(fs.existsSync(path.join(project, first)))
|
||||
})
|
||||
116
test/orchestrator-guard.test.mjs
Normal file
116
test/orchestrator-guard.test.mjs
Normal file
|
|
@ -0,0 +1,116 @@
|
|||
import { test } from 'node:test'
|
||||
import assert from 'node:assert/strict'
|
||||
import fs from 'node:fs'
|
||||
import os from 'node:os'
|
||||
import path from 'node:path'
|
||||
|
||||
import { Orchestrator } from '../src/orchestrator.mjs'
|
||||
|
||||
const sink = { write: () => true }
|
||||
|
||||
function makeProject() {
|
||||
const project = fs.mkdtempSync(path.join(os.tmpdir(), 'kilo-guard-orch-'))
|
||||
fs.mkdirSync(path.join(project, '.kilocode-loop'), { recursive: true })
|
||||
fs.writeFileSync(path.join(project, 'goal.json'), JSON.stringify({ title: 'Guard goal', stages: [{ id: 's1', title: 'Do work' }] }))
|
||||
return project
|
||||
}
|
||||
|
||||
function makeConfig(project, overrides = {}) {
|
||||
return {
|
||||
project,
|
||||
goal: 'goal.json',
|
||||
goalText: '',
|
||||
iterations: 3,
|
||||
agent: 'code-design',
|
||||
model: '',
|
||||
variant: '',
|
||||
auto: true,
|
||||
sharedContext: false,
|
||||
hitl: 'off',
|
||||
confirm: false,
|
||||
port: 7997,
|
||||
host: '127.0.0.1',
|
||||
dryRun: true,
|
||||
keepOpen: false,
|
||||
quiet: true,
|
||||
color: false,
|
||||
promptExtra: '',
|
||||
maxIterationMinutes: 0,
|
||||
contextWarnTokens: 150000,
|
||||
verify: '',
|
||||
verifyTimeoutMinutes: 30,
|
||||
maxCost: 0,
|
||||
maxTokens: 0,
|
||||
maxStaleIterations: 0,
|
||||
checkpoint: false,
|
||||
guard: true,
|
||||
guardStop: false,
|
||||
maxGuardBlocks: 3,
|
||||
redact: true,
|
||||
review: false,
|
||||
reviewAgent: '',
|
||||
autoFreshTokens: 0,
|
||||
notify: '',
|
||||
syncGoal: true,
|
||||
...overrides,
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Runner that emits a destructive bash tool_use only on the requested iteration,
|
||||
* so the loop's runtime guard aborts that session but (by default) keeps going.
|
||||
*/
|
||||
function guardOnceRunner(guardIterations = [1]) {
|
||||
return async ({ prompt, onEvent, signal }) => {
|
||||
const iteration = Number(prompt.match(/iteration (\d+)\//)?.[1] || 1)
|
||||
if (guardIterations.includes(iteration)) {
|
||||
onEvent({ type: 'tool_use', part: { tool: 'bash', state: { input: { command: 'cd /repo && rm -rf src' } } } })
|
||||
}
|
||||
return {
|
||||
sessionID: `ses_iter_${iteration}`,
|
||||
exitCode: signal.aborted ? 130 : 0,
|
||||
error: null,
|
||||
texts: [],
|
||||
reasoning: [],
|
||||
toolCalls: [],
|
||||
stderr: '',
|
||||
durationMs: 5,
|
||||
interrupted: signal.aborted,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
test('a guard hit aborts the session but the run continues to the next iteration', async () => {
|
||||
const project = makeProject()
|
||||
const final = await new Orchestrator(makeConfig(project, { runner: guardOnceRunner([1]), iterations: 2, runId: 'guard-continue' }), { out: sink }).run()
|
||||
|
||||
assert.notEqual(final.status, 'stopped-guard')
|
||||
assert.equal(final.stopReason, null)
|
||||
assert.equal(final.iterations.length, 2)
|
||||
assert.equal(final.iterations[0].guardViolation.id, 'rm-recursive-force')
|
||||
// The safety feedback must reach the next iteration's prompt.
|
||||
const prompt2 = fs.readFileSync(path.join(project, '.kilocode-loop', 'runs', 'guard-continue', 'iterations', '02-prompt.md'), 'utf8')
|
||||
assert.match(prompt2, /Loop safety feedback/)
|
||||
assert.match(prompt2, /rm-recursive-force/)
|
||||
})
|
||||
|
||||
test('--guard-stop restores the hard stop on the first hit', async () => {
|
||||
const project = makeProject()
|
||||
const final = await new Orchestrator(makeConfig(project, { runner: guardOnceRunner(), guardStop: true, runId: 'guard-stop' }), { out: sink }).run()
|
||||
|
||||
assert.equal(final.status, 'stopped-guard')
|
||||
assert.equal(final.stopReason, 'guard')
|
||||
assert.equal(final.iterations.length, 1)
|
||||
})
|
||||
|
||||
test('a persistently blocked agent stops after maxGuardBlocks iterations', async () => {
|
||||
const project = makeProject()
|
||||
const final = await new Orchestrator(
|
||||
makeConfig(project, { runner: guardOnceRunner([1, 2, 3]), maxGuardBlocks: 2, iterations: 5, runId: 'guard-limit' }),
|
||||
{ out: sink },
|
||||
).run()
|
||||
|
||||
assert.equal(final.status, 'stopped-guard')
|
||||
assert.equal(final.stopReason, 'guard')
|
||||
assert.equal(final.iterations.length, 2)
|
||||
})
|
||||
|
|
@ -135,7 +135,7 @@ test('retries stop after maxRetries and the iteration fails', async () => {
|
|||
assert.equal(final.iterations[0].status, 'failed')
|
||||
})
|
||||
|
||||
test('a guard violation is never retried', async () => {
|
||||
test('a guard violation is never retried as a transient error', async () => {
|
||||
const project = makeProject([{ id: 's1', title: 'One' }])
|
||||
let calls = 0
|
||||
const runner = async ({ onEvent }) => {
|
||||
|
|
@ -143,7 +143,9 @@ test('a guard violation is never retried', async () => {
|
|||
onEvent?.({ type: 'tool_use', part: { tool: 'bash', state: { input: { command: 'git commit -m x' } } } })
|
||||
return failResult('ses_x', 'Connection reset by server')
|
||||
}
|
||||
const final = await new Orchestrator(makeConfig(project, { runner }), { out: sink }).run()
|
||||
// --guard-stop: the hit stops the run immediately (no transient retry); with
|
||||
// the default the loop would instead continue to the next iteration.
|
||||
const final = await new Orchestrator(makeConfig(project, { runner, guardStop: true }), { out: sink }).run()
|
||||
assert.equal(calls, 1)
|
||||
assert.equal(final.stopReason, 'guard')
|
||||
})
|
||||
|
|
|
|||
Loading…
Reference in a new issue