Compare commits

...

18 Commits

Author SHA1 Message Date
Miguel Palhas 0900e5fea1 perf(gate): widen per-slot jobs to (cores-2)/slots
ci / nix (push) Successful in 10s
ci / lint (push) Successful in 12s
The old cores/SLOTS/2 left a full test suite on 2 of 10 cores even with
the box idle. Reserve two cores for the agent sessions and split the
rest across slots.

Also retire the CPU framing on the subagent fan-out cap. The number
stays 3, but the binding constraints are memory per worktree and the
shared rate-limit window; subagents are mostly idle waiting on the API.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 13:56:55 +01:00
Miguel Palhas 37cdfd2c1f fix(pr-daemon): send comments hints to review sessions
ci / nix (push) Successful in 9s
ci / lint (push) Successful in 12s
The reviewer was filtered down to ci and state, so a reply to one of its
findings never reached it — it only learned about pushback when someone
asked. Replies are addressed to the reviewer, and an addressed thread is
now the reviewer's to resolve. Conflicts stay filtered out: those are the
author's to fix on their own branch.

Takes effect after: systemctl --user restart pr-daemon (runs from
~/tea/agent-skills, so this must reach main first).
2026-08-24 11:24:41 +01:00
Miguel Palhas 8313842de8 fix(review-pr): resolve any addressed thread, not only your own
Bot and other-reviewer threads block land the same way. The guard that
matters is whether the finding is addressed, not who opened it.
2026-08-24 11:23:53 +01:00
Miguel Palhas c7becda233 feat(review-pr): resolve own threads, post without approval
Trust comes from a repos[] entry in the reviewer config, not from the
forge. Listed repos post findings directly; only an unlisted repo still
holds findings for approval. The old text tied gating to github and
misread the config's mode field, which is the daemon's land/review
switch.

Adds §3.1: resolve a thread once the author addresses the finding.
Unresolved findings pile up for the life of the PR and block the
author's land session, which never declares a PR ready over an open
thread. Approval is still not the reviewer's to give.
2026-08-24 11:07:23 +01:00
Miguel Palhas d49d4e140e feat(plan-milestone,blitz): wire real Gitea issue dependencies
ci / nix (push) Successful in 8s
ci / lint (push) Successful in 11s
plan-milestone now POSTs to the dependencies API when filing issues,
instead of relying on Depends on: #N prose alone. blitz's finding-fold
step does the same. Body text is kept for readability.
2026-08-23 16:42:33 +01:00
Miguel Palhas ca790b19c5 feat(plan-milestone): rename backlog, file under a created milestone
ci / nix (push) Successful in 7s
ci / lint (push) Successful in 10s
The milestone is the handoff unit to /blitz and keeps concurrent
planning efforts' issues apart.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 15:58:04 +01:00
Miguel Palhas b735c16b30 feat(blitz): spawn-time model assessment over label table
ci / nix (push) Successful in 7s
ci / lint (push) Successful in 11s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 15:54:46 +01:00
Miguel Palhas febf397885 docs(blitz): route by remaining judgment, not label alone
ci / nix (push) Successful in 7s
ci / lint (push) Successful in 9s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 15:52:10 +01:00
Miguel Palhas 59e47308bd feat(blitz): aoe sessions are the only worker mode
ci / nix (push) Successful in 8s
ci / lint (push) Successful in 11s
Subagent fan-out removed; model routing is a built-in default table,
overridden by instructing the run rather than per-repo config.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 15:51:44 +01:00
Miguel Palhas d38e207b50 Merge branch 'main' of https://git.naps.pt/yolo/agent-skills
ci / nix (push) Successful in 7s
ci / lint (push) Successful in 10s
2026-08-23 15:40:49 +01:00
Miguel Palhas ca0021eca9 feat(skills): backlog planning skill, aoe worker mode for blitz
backlog turns a design doc into a labeled, interdependent issue set
by grilling the operator. blitz gains an aoe-session worker mode with
per-difficulty model routing, exit-criteria prompts and stall nudging,
extracted from the arr driver run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 15:40:46 +01:00
Miguel Palhas 851f7f7fbe feat(hourlog): default to sonnet
ci / nix (push) Successful in 8s
ci / lint (push) Successful in 11s
Reading session logs into a table is not opus work. HOURLOG_MODEL still
overrides, and an empty value takes the harness default.
2026-08-23 11:04:47 +01:00
Miguel Palhas 777ee3e72d feat(hourlog): let HOURLOG_MODEL pick the agent model
ci / nix (push) Successful in 10s
ci / lint (push) Successful in 12s
Passed through as --extra-args to the agent binary, so the Friday timer can
run on something other than the harness default.
2026-08-23 10:25:26 +01:00
Miguel Palhas 2be68476f1 fix(hourlog): resolve aoe from PATH
ci / nix (push) Successful in 9s
ci / lint (push) Successful in 11s
aoe moved to the nix profile, so the hardcoded ~/.local/bin/aoe made the
Friday timer exit 127 before creating the session. Same resolution the
week-review script already uses.
2026-08-23 10:12:19 +01:00
Miguel Palhas f61371c74a Merge remote-tracking branch 'origin/main' into reviews 2026-08-23 10:08:16 +01:00
Miguel Palhas dba8416192 feat(pr-daemon): review only when a review is requested
Auto-review paid off on daemon and core PRs and reviewed clean on most
small ones, so spawning a reviewer on every non-draft PR spent tokens for
nothing. A review session now needs github plus one of your logins in the
PR's requested_reviewers; gitea spawns none.
2026-08-23 10:06:44 +01:00
Miguel Palhas f29d170702 Merge branch 'hindustanis'
ci / nix (push) Successful in 8s
ci / lint (push) Successful in 10s
2026-08-23 09:58:39 +01:00
Miguel Palhas e627f53934 feat(land): merge on gitea, stop at the button on github
The never-merge rule only holds for GitHub. Gitea repos here are the
user's own, so land squash-merges once CI is green, threads are
resolved and the branch is current. A gitea PR with no reviewer ever
requested counts as approved, otherwise it waits forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 09:56:56 +01:00
13 changed files with 372 additions and 51 deletions
+14 -2
View File
@@ -96,6 +96,9 @@ Needs `loginctl enable-linger` so the timer runs while logged out. Logs are in `
`bin/hourlog-session.sh`, which opens an Agent of Empires session on a scratch
dir, sends it `/hourlog --week this`, and pushes an ntfy notification.
It runs on sonnet — reading session logs into a table is not opus work —
overridable with `HOURLOG_MODEL`, or empty for the harness default.
Same shape as the weekly review and interactive for the same reason: the skill
proposes hours and stops for approval before writing anything to the timesheet.
An unattended run would be deciding a company record on your behalf. It skips
@@ -153,8 +156,17 @@ run and sets a later epoch, which filters more, never less.
| PR | skill | session |
|----|-------|---------|
| yours | `land` | default profile, `--yolo --trust-hooks` |
| yours, with `selfReview` | both | plus a reviewer on a different agent |
| someone else's | `review-pr` | `review` profile, no yolo, no trusted hooks, sandboxed |
| github, review requested from you | `review-pr` | `review` profile, no yolo, no trusted hooks, sandboxed |
| yours on github, review requested, with `selfReview` | both | plus a reviewer on a different agent |
**A reviewer needs an explicit request.** Two conditions, both required: the
forge is github, and one of your logins sits in the PR's `requested_reviewers`.
Gitea never spawns one, and a merely non-draft PR doesn't either. An audit of 47
closed PRs is where that came from — roughly a third of the findings paid for
themselves and nearly all of those were daemon and core changes, while small
PRs reviewed clean often enough that the reviewing cost bought nothing. Github
won't let you request a review from a PR's own author, so `selfReview` now only
fires when another of your logins opened the PR.
Both roles can run on one PR because the role is carried by the worktree
branch: the author side works on the head branch, the reviewer on a local
+10 -2
View File
@@ -6,9 +6,17 @@ set -euo pipefail
PROMPT="${HOURLOG_PROMPT:-/hourlog --week this}"
TOPIC="${HOURLOG_NTFY_TOPIC:-homelab}"
AOE="${HOURLOG_AOE:-$HOME/.local/bin/aoe}"
AOE="${HOURLOG_AOE:-$(command -v aoe || echo "$HOME/.nix-profile/bin/aoe")}"
LOG="$HOME/.local/state/hourlog/run.log"
# Which model reads the week. Sonnet by default: the work is reading session
# logs and filling a table, and it held up on the first run. The value goes
# straight to the agent binary, so it has to be a name that binary knows
# (`sonnet`, `opus` for claude); set it empty to take the harness default.
MODEL="${HOURLOG_MODEL-sonnet}"
extra=()
[ -n "$MODEL" ] && extra=(--extra-args "--model $MODEL")
WEEK="$(date +%G-W%V)"
TITLE="hourlog-$WEEK"
@@ -46,7 +54,7 @@ fi
# --scratch keeps the session's cwd under the agent-of-empires app dir, which
# the hourlog config excludes — otherwise it lands in next week's scan.
"$AOE" add --scratch --title "$TITLE" --cmd claude --yolo --trust-hooks
"$AOE" add --scratch --title "$TITLE" --cmd claude --yolo --trust-hooks "${extra[@]}"
"$AOE" session start "$TITLE"
# The agent needs its TUI up before it can take a prompt; `send` into a
+30 -6
View File
@@ -72,6 +72,7 @@ type Pr = Snapshot & {
author: string;
createdAt: string;
url: string;
requestedReviewers: string[];
cfg: RepoConfig;
};
@@ -213,6 +214,13 @@ async function repos(): Promise<RepoConfig[]> {
return out.filter((r) => r.path && existsSync(r.path));
}
// Who the PR is currently asking for a review. GitHub clears the entry once
// that reviewer submits, which is fine: by then the session exists and routes
// by branch.
function reviewerLogins(p: any): string[] {
return (p.requested_reviewers ?? []).map((r: any) => r?.login).filter(Boolean);
}
// The list endpoints carry everything except mergeable and the comment counts,
// so the detail call happens only for PRs that already look changed.
async function listPrs(cfg: RepoConfig): Promise<Pr[]> {
@@ -234,6 +242,7 @@ async function listPrs(cfg: RepoConfig): Promise<Pr[]> {
mergeable: p.mergeable ?? null,
comments: p.comments ?? 0,
reviewComments: p.review_comments ?? 0,
requestedReviewers: reviewerLogins(p),
cfg,
}));
}
@@ -248,6 +257,7 @@ async function detail(pr: Pr): Promise<Pr> {
state: d.state ?? pr.state,
draft: Boolean(d.draft ?? pr.draft),
headSha: d.head?.sha ?? pr.headSha,
requestedReviewers: reviewerLogins(d),
};
}
@@ -431,10 +441,22 @@ function route(pr: Pr, role: Role, all: Session[]): Session | undefined {
return all.find((s) => samePath(s.mainRepo, pr.cfg.path) && s.branch === branch);
}
// An audit of 47 closed PRs put most of the value on daemon and core work and
// found a clean pass on most small ones, so a reviewer is no longer spawned on
// every non-draft PR. Two conditions now, both required: github only, and a
// review explicitly requested from one of your logins. Gitea never spawns one.
// Note github forbids requesting a review from a PR's own author, so on your
// own PRs this only fires when another of your logins opened it.
function reviewWanted(pr: Pr): boolean {
if (pr.forge !== "github") return false;
return pr.requestedReviewers.some((login) => isSelf(pr.forge, login));
}
function rolesFor(pr: Pr): Role[] {
if (!isSelf(pr.forge, pr.author)) return ["review"];
if ((pr.cfg.mode ?? "drive") !== "drive") return ["review"];
return pr.cfg.selfReview ? ["land", "review"] : ["land"];
const review: Role[] = reviewWanted(pr) ? ["review"] : [];
if (!isSelf(pr.forge, pr.author)) return review;
if ((pr.cfg.mode ?? "drive") !== "drive") return review;
return pr.cfg.selfReview ? ["land", ...review] : ["land"];
}
// ---------------------------------------------------------------- reviewers
@@ -787,9 +809,11 @@ async function evaluate(prs: Pr[], mentioned: Set<string>, budget: { sessions: n
continue;
}
// The reviewer reacts to new commits and to the PR closing; replying to
// threads is the author side's job, so comments are not its business.
let mine = role === "land" ? why : why.filter((w) => w === "ci" || w === "state");
// Conflicts are the author's to resolve on their own branch, so the
// reviewer never hears about them. Comments it does hear: a reply to a
// finding is addressed to the reviewer, and an addressed thread is the
// reviewer's to resolve (review-pr §3.1).
let mine = role === "land" ? why : why.filter((w) => w !== "conflicts");
if (mine.includes("comments") && prev) {
const ids = await newCommentIds(full, prev.updatedAt);
const seen = session.path ? await seenIds(session.path, full.number) : null;
+112
View File
@@ -0,0 +1,112 @@
# Blitz workers: aoe sessions
How blitz fans out: each ready issue gets an **external `aoe` session** — its
own tmux pane, worktree, tool (claude / codex / opencode) and model. Sessions
survive the orchestrator restarting, and routing across providers keeps one
provider's outage or blind spots from shaping the whole run.
Everything else in blitz (DAG, integration, review cadence, readiness gate,
ship, notify) is defined in SKILL.md. This file covers spawning, prompting,
babysitting and cleanup.
## Model routing
**Assess each issue at spawn time.** You have just read its body to write the
prompt — use that read to pick the model. The question is not "how big is
this" but **how much judgment does the session still have to exercise**:
- Body settles the approach (root cause named, fix shape decided, numbers
suggested, files pointed at) — the thinking happened at filing time; the
session executes. **Sonnet** (`--tool claude --extra-args "--model
claude-sonnet-5"`), regardless of size: a large mechanical CRUD issue is
still execution.
- Body states the goal but the session must design the interface, choose the
data model, or amend the design doc — **Opus** (`--model claude-opus-5`).
- The design doc itself is thin or contradictory where this issue lives,
correctness is subtle, or the change is cross-cutting with unclear blast
radius — **Fable** (`--model claude-fable-5`).
A `difficulty/` label is one input — a filing-time guess that cannot see how
much the body scaffolds. Trust your read of the body over it; the label is a
tie-breaker. When in doubt between two tiers take the lower one: escalation
on failure is cheap, and a failed cheap run teaches something a successful
expensive run does not.
The operator overrides any of this by just saying so in the invocation ("run
these on codex", "use ox alpha for the easy ones") — no config file. When
routing across peer models, alternate rather than draining one first — a bad
run should be visible early. Escalate *sideways* (a peer model) before
escalating up, and never de-escalate mid-issue.
## Spawning
```sh
aoe add <repo-root> \
--title "<repo>-<n>-<slug>" \
--worktree issue/<n>-<slug> --new-branch \
--tool <tool> --extra-args "<model args>" \
--launch
```
Known traps, all confirmed the hard way:
- **`aoe send` races with `--launch` and fails silently.** Never trust it.
Deliver the prompt with `tmux send-keys -l "<prompt>"` followed by a
**separate** `Enter` a second later, then verify with
`tmux capture-pane -p` that it landed.
- **tmux truncates session names.** Look panes up by prefix
(`tmux list-sessions -F '#{session_name}' | grep '^aoe_<title-prefix>'`),
never by exact title. A failed exact lookup can dump the prompt into your
own pane.
- `--new-branch` is required for a branch that doesn't exist. The worktree
path comes from `--title`, not the branch — distinct titles or the second
session collides.
## The prompt
Two phrasings matter, learned from models that do the work and then stop:
1. **Exit criteria beat autonomy language.** "Work fully autonomously" does
not stop a model from ending its turn after a big tool result. What does:
*"Do not end your final reply until ALL of these are true: … If you catch
yourself summarizing progress before those are true, you have stopped too
early — keep going."* Fill the criteria with the terminal state you
actually assigned: under blitz that is gate-green + **branch pushed** (the
orchestrator merges and closes); when driving direct-to-main it is
merged + pushed + issue closed with a comment naming the commit.
2. **Fence parallel sessions off each other's files.** When two issues run at
once, each prompt names what the other owns: "Do not touch X — issue #M
owns it and runs in parallel." Merge conflicts are cheaper to prevent in
the prompt than to resolve after.
Also include: read CLAUDE.md and the design doc first; fetch the issue body
via the tracker API; merge (never rebase) the base branch if it moved; and
never wait for input.
## Babysitting
Some models stall — idle turn-end after absorbing a large tool result, work
half done. Don't hand-poll; arm a **self-nudging Monitor** per session:
- Poll every ~90s. Idle means the pane shows no in-progress marker *and* the
context/size indicator is frozen across two consecutive checks — one check
is not enough, models legitimately pause.
- On stall: send the nudge yourself via `tmux send-keys` — restate the exit
criteria and where it stopped — capped at ~6 nudges before escalating to
a human.
- Exit (and notify the orchestrator) on: issue closed, session gone, nudge
cap, or timeout. Silence must not look like success — every terminal state
emits a line.
## Collect, integrate, clean up
- Detect completion by **tracker state** (issue closed) or the integration
branch moving — never by grepping commit messages; fuzzy matches fire on
the wrong branch's commits.
- After a session's work is merged: `aoe remove <title> --delete-worktree`.
Sweep every couple of waves; stale sessions pile up. (`aoe` leaves removed
worktrees locked — `git worktree unlock` before a manual
`git worktree remove`.)
- Verify the landed result yourself against the live system when one exists
(deploy health, a smoke request against the changed endpoint). A session
reporting success is a claim, not a verification.
+8 -10
View File
@@ -1,6 +1,6 @@
---
name: blitz
description: "Autonomously drive an entire tracker milestone to done — sweep every open issue (one subagent per issue, parallel where dependencies allow), fix bugs found along the way, then either deploy (safe to debug in prod) or spawn a local dev instance, and push a Home Assistant notification with the preview URL. Built for long unattended runs. Use when the user wants to blitz / sweep / complete a whole milestone, e.g. \"/blitz M0\"."
description: "Autonomously drive an entire tracker milestone to done — sweep every open issue (one aoe worker session per issue, models routed by difficulty, parallel where dependencies allow), fix bugs found along the way, then either deploy (safe to debug in prod) or spawn a local dev instance, and push a Home Assistant notification with the preview URL. Built for long unattended runs. Use when the user wants to blitz / sweep / complete a whole milestone, e.g. \"/blitz M0\"."
user-invocable: true
args:
- name: input
@@ -22,8 +22,8 @@ This skill targets **`tracker: gitea`** (milestones live in the repo's Gitea tra
## Roles
- **Orchestrator** = the main blitz thread (you). Owns the DAG, spawns subagents, merges branches, closes issues, runs the readiness gate, ships, notifies. Does *not* implement issues itself.
- **Issue subagent** = one `Agent` per issue (`isolation: "worktree"`). Implements exactly one issue via the yolo flow, returns a structured result. **One subagent per issue is the default and is incentivized** — do not batch multiple issues into one agent.
- **Orchestrator** = the main blitz thread (you). Owns the DAG, spawns workers, merges branches, closes issues, runs the readiness gate, ships, notifies. Does *not* implement issues itself.
- **Issue worker** = one external `aoe` session per issue, with per-difficulty model routing across tools/providers. Read `AOE-WORKERS.md` (this skill's directory) before spawning any — it carries the spawn mechanics, default model table, prompt phrasing, stall babysitting, and cleanup rules, and replaces §3.2's fan-out mechanics. One session per issue — do not batch. The operator can override models for a run by just saying so; no config needed.
---
@@ -47,19 +47,17 @@ This skill targets **`tracker: gitea`** (milestones live in the repo's Gitea tra
## 3. Execution pass (the loop body)
**Blitz drives its own loop — no external `/loop` needed.** The orchestrator thread stays alive and repeats the pass below until the milestone is done. Fan-out subagents run in the background; when they finish they re-invoke you, which advances the next wave naturally. Only use `ScheduleWakeup` as a fallback heartbeat when you're blocked waiting on something the harness can't notify you about (e.g. polling a deploy's health). Wrapping blitz in `/loop` is unnecessary and not the intended usage.
**Blitz drives its own loop — no external `/loop` needed.** The orchestrator thread stays alive and repeats the pass below until the milestone is done. Worker sessions run in their own tmux panes; per-session Monitors (see `AOE-WORKERS.md`) re-invoke you as they finish or stall, which advances the next wave naturally. Only use `ScheduleWakeup` as a fallback heartbeat when you're blocked waiting on something the harness can't notify you about (e.g. polling a deploy's health). Wrapping blitz in `/loop` is unnecessary and not the intended usage.
Each pass:
1. Recompute the **ready set** (§2.4).
2. **Fan out**: spawn one issue subagent per ready issue, **in parallel** (multiple `Agent` calls in a single message), `isolation: "worktree"`. **Cap concurrency at 3** each worktree carries its own build artifacts and test run, and other autonomous sessions are on the same box. Drop to 2 when `<skills-root>/tracker-common/scripts/gate.sh --status` shows the machine already contended. Each subagent prompt:
- "Implement Gitea issue #N (`<title>`) in this repo following the `/yolo` flow and `COMMON.md`. You are on integration branch `blitz/<slug>`; create branch `<slug>/N-<issue-slug>` **off it**. Read the issue body + its linked spec/epic; that plus the repo is your full context. Implement and commit in logical steps. Check **only what you touched** as you go; run `buildCommand` at most once at the end, and run it as `<skills-root>/tracker-common/scripts/gate.sh -- <buildCommand>` — exit 75 means the machine was busy and it did not run, so return `buildPassed: null` rather than retrying. **Do not merge to any shared branch and do not close the issue** — push your branch and return the result. If you discover a bug or missing work outside this issue's scope, do not fix it silently; report it in `newFindings`."
- Force a structured return (schema): `{ issue, done, branch, summary, buildPassed, newFindings: [{title, body}] }`. `buildPassed: null` = the gate was busy, so the integration build is the first real check that branch gets.
- **Strict rule**: never spawn a subagent for a blocked issue. Dependencies are load-bearing.
3. **Integrate serially** (orchestrator, to avoid parallel-merge conflicts): for each finished subagent whose `done` and whose `buildPassed` is not `false`, merge its branch into `blitz/<slug>` and resolve conflicts. Run `buildCommand` **once per wave, after the last merge** — not once per branch — and through the gate: `<skills-root>/tracker-common/scripts/gate.sh -- <buildCommand>`. If the merge or build breaks, fix on the integration branch (or bounce the issue back for another pass); with several branches merged, `git log --oneline` on the failing area tells you which one to bounce.
2. **Fan out**: spawn one `aoe` session per ready issue, following `AOE-WORKERS.md` for spawn mechanics, model routing, prompt phrasing, and the per-session stall Monitor. **Cap concurrency at 3** — the limit is memory and the shared rate-limit window, not cores: each worktree carries its own build artifacts and test run, and other autonomous sessions on the same box are drawing from the same budget. Drop to 2 when `<skills-root>/tracker-common/scripts/gate.sh --status` shows the machine already contended. Each session's prompt carries the task ("Implement Gitea issue #N following `COMMON.md`; branch `<slug>/N-<issue-slug>` off integration branch `blitz/<slug>`; the issue body plus the repo is your full context; run `buildCommand` at most once at the end through `gate.sh`"), the file fencing against parallel issues, and exit criteria ending at **branch pushed** — "do not merge to any shared branch and do not close the issue; if you find out-of-scope bugs, comment them on the issue instead of fixing silently".
- **Strict rule**: never spawn a session for a blocked issue. Dependencies are load-bearing.
3. **Integrate serially** (orchestrator, to avoid parallel-merge conflicts): for each finished session — issue branch pushed, detected by the Monitor, never by grepping commit messages — merge its branch into `blitz/<slug>` and resolve conflicts. Run `buildCommand` **once per wave, after the last merge** — not once per branch — and through the gate: `<skills-root>/tracker-common/scripts/gate.sh -- <buildCommand>`. If the merge or build breaks, fix on the integration branch (or bounce the issue back for another pass); with several branches merged, `git log --oneline` on the failing area tells you which one to bounce.
4. **Close** each successfully integrated issue on Gitea (`Closes #N` in the merge commit, or PATCH `state:closed`). Epics whose blockers are now all closed: close them too.
5. **Integration review (cadence-gated) — do NOT skip.** After each wave (or every ~3 integrated issues, whichever comes first), audit the *accumulated* diff of `blitz/<slug>` vs `defaultBranch` — not each issue in isolation. Run `/code-review` on that diff, or spawn a reviewer subagent, hunting the cross-issue drift that blind parallel work causes: inconsistent data shapes / contracts between issues, divergent naming, duplicated or conflicting logic, dead code, regressions, misbehavior. **Findings are top priority**: fix them (inline, or file + wire as blocking issues) *before* spawning the next fan-out wave. This is the load-bearing coherence check — parallel subagents can't see each other's work, so this is the only place drift gets caught.
6. **Fold in findings**: for each `newFindings` item and any bug you find, create a new Gitea issue in this milestone (`milestone: MS_ID`), wire dependencies if it blocks/relies on others, and let the next pass pick it up. Fix trivial bugs inline instead of filing.
6. **Fold in findings**: for each out-of-scope bug a worker commented on its issue and any bug you find, create a new Gitea issue in this milestone (`milestone: MS_ID`). If it blocks or relies on others, wire it as a real dependency (`POST .../issues/$N/dependencies` with `{"index": <other>}`, not just prose) so §2's DAG picks it up next pass. Fix trivial bugs inline instead of filing.
7. Repeat passes until: no open workable issues, no open epics, the latest integration review is clean, and a full pass produced **no new findings**.
## 4. Readiness gate
+35 -10
View File
@@ -1,6 +1,6 @@
---
name: land
description: "Drive a PR you authored to ready-to-merge: fix CI failures, address every review comment, resolve conflicts, push, until green + approved. Never merges. Event-driven — the PR daemon wakes it. Forge-agnostic (GitHub or Gitea)."
description: "Drive a PR you authored to ready-to-merge: fix CI failures, address every review comment, resolve conflicts, push, until green + approved. Merges it on Gitea; on GitHub it stops and leaves the click to the user. Event-driven — the PR daemon wakes it. Forge-agnostic (GitHub or Gitea)."
user-invocable: true
args:
- name: target
@@ -10,8 +10,8 @@ args:
# Land — drive your own PR to ready-to-merge
Takes a PR **you authored** and shepherds it to the merge button: green
CI, every review thread addressed, approved, branch up to date. For PRs
Takes a PR **you authored** and shepherds it to the merge: green CI,
every review thread addressed, approved, branch up to date. For PRs
someone else authored, use `review-pr` instead — it reads and comments
and never pushes.
@@ -19,8 +19,10 @@ Read `pr-common/COMMON.md` (sibling skill, same skills root) first. It
defines hints, the seen file, the state file, and forge resolution. This
document only covers what to *do*.
**Never merge.** The final click is the user's — every repo, every
forge. No `gh pr merge`, no merge API call, no `--auto`.
**Merging is forge-scoped.** On **GitHub**, never merge — the final
click is the user's. No `gh pr merge`, no `--auto`. On **Gitea**, merge
the PR yourself once section 3's conditions all hold; those are the
user's own self-hosted repos and the click adds nothing.
**Spend nothing while idle.** You do not wait, poll, or arm watchers.
The daemon wakes you with a hint when something changes. Do the work the
@@ -183,20 +185,43 @@ return silently.
All of these must hold: CI green, every thread resolved, approved with
no pending review requests, branch not behind the base.
On gitea, a PR with no reviewer ever requested and no review posted
counts as approved — otherwise a solo PR waits forever for a review
that is never coming. A requested or posted review still has to land.
Update the branch if the base moved (above), let CI re-run, and wait for
the resulting hint. The user's click should be the only step left.
the resulting hint.
**On GitHub**, that is where you stop — the user's click is the only
step left.
**On Gitea**, merge it. Squash, server-side so the PR reads "merged"
and not "closed":
```bash
curl -sS -X POST -H "Authorization: token $GITEA_TOKEN" \
-H "Content-Type: application/json" \
"$BASE/api/v1/repos/$REPO/pulls/$N/merge" -d '{"Do":"squash"}'
```
Do not delete the branch or remove the worktree — the user handles
cleanup.
## 4. Close out
Report: PR ready to merge with its link, CI green, approved, threads
resolved, and one line on what feedback was addressed. The merge, the
branch delete, and the tracking-issue close are the user's.
**Merged (gitea):** report the merge with the PR link, CI green,
threads resolved, and one line on what feedback was addressed.
**Ready but not merged (github):** report PR ready to merge with the
same detail. The merge, the branch delete, and the tracking-issue close
are the user's.
The review window is often hours and the user may be away. Push a
notification so the click can happen from a phone —
`mcp__ha-mcp__ha_call_service`, `domain: "notify"`, service
`mobile_app_pixel_7_naps` (or `blitz.notifyService` from config),
message "PR #<N> ready to merge" plus the URL.
message "PR #<N> ready to merge" plus the URL. On gitea, the same
notification instead says the PR merged.
`Closes <REF>` in the PR body closes a GitHub or Gitea tracking issue on
merge. Only Linear needs follow-up: tell the user to move the issue to
+1 -1
View File
@@ -111,7 +111,7 @@ What actually works, learned the hard way:
- **Model choice**: strongest model for design-heavy or feel-critical work; a cheaper one is fine for mechanical, well-specified changes.
- Instruct them to **commit their own work locally** when it's coherent, so a killed agent loses less — and explicitly **not to push**. A dozen subagent pushes is a dozen CI runs on half-finished work.
- **Tell them not to run the full suite.** Scoped checks on the files they own, nothing more. Five agents each running every test is five copies of the same work and enough memory pressure to kill the run. You run the full suite once, at push time, through `gate.sh`.
- **Cap the fan-out at 3 concurrent subagents, 2 if their tasks compile or test.** More agents is not more throughput on a box this size — it is swap. `<skills-root>/tracker-common/scripts/gate.sh --status` shows how much of the machine other sessions are already using; dispatch fewer when it is contended, and remember other `/yolo` and `/nightshift` runs are competing for the same RAM.
- **Cap the fan-out at 3 concurrent subagents, 2 if their tasks compile or test.** The limit is not cores — subagents spend most of their time waiting on the API. It is memory (each one carries a worktree, its build artifacts and a test run) and the shared rate-limit window, which other `/yolo` and `/nightshift` runs are drawing from too. `<skills-root>/tracker-common/scripts/gate.sh --status` reports free RAM and heavy commands in flight; dispatch fewer when it is contended.
## 6. Reviewing what lands
+105
View File
@@ -0,0 +1,105 @@
---
name: plan-milestone
description: "Turn a design doc or feature idea into a tracker milestone with a filed, labeled, interdependent issue set — by auditing what already exists, grilling the operator through the open decisions, and drafting the full set for approval before filing. The milestone is the handoff unit: /blitz <milestone> sweeps it, and parallel planning efforts stay distinguishable. Use when the user wants to plan a feature area into issues, e.g. \"plan TV tracking\" or \"/plan-milestone §6 of DESIGN.md\"."
user-invocable: true
args:
- name: input
description: "What to plan: a design-doc section, a feature description, or a doc path. Omit to ask."
required: false
---
# Plan Milestone — Design to Issue Set
Produces a tracker **milestone** holding a set of issues another session can
land one at a time — `/blitz <milestone>` is the intended consumer, `/yolo`
works per issue. The milestone is what keeps two concurrent planning efforts
apart: every issue this skill files belongs to the milestone it creates. The
output is the milestone; this skill never implements anything.
**First:** read `tracker-common/COMMON.md` (sibling skill, same skills root)
for project config and tracker API conventions.
## 1. Ground yourself
1. Read the design document end to end if the repo has one (`DESIGN.md` or
whatever CLAUDE.md names as the contract). The design doc is authoritative:
if planning surfaces a contradiction, the fix is a design-doc amendment
issue, never an issue that quietly contradicts it.
2. Audit what already exists — code, migrations, API surface, closed issues —
so the set covers the gap, not what is built. Plan from evidence, not from
the doc's table of contents.
3. Learn the label taxonomy: the repo's CLAUDE.md, or the tracker's existing
labels. A good taxonomy has one label per axis per issue (e.g. `phase/`,
`area/`, `difficulty/`, `type/`). If the repo has none, propose one to the
operator before drafting.
## 2. Grill the operator
The operator holds decisions the design doc doesn't. Interview them **in
batches, one batch at a time** — a wall of twenty questions gets skimmed;
four pointed ones get answered.
- Ask about behavior, not implementation: semantics, defaults, edge cases,
what "done" looks like for the user.
- Challenge vague answers and surface tradeoffs ("per-episode grabbing
doubles indexer load — accept that or prefer season packs?").
- Where the design doc is thin or self-contradictory, say so explicitly and
get a ruling.
- Record each settled decision in one line; these lines become issue-body
context.
Stop interviewing when new questions stop changing the issue set.
## 3. Draft, then file
Draft the **complete set** and show it to the operator for approval before
filing anything. For each issue:
- **Title**: imperative, specific, no scope words like "improve" or "handle".
- **Body**: the settled decisions it depends on, pointers into the design doc
(cite sections, don't restate them), and explicit non-goals when adjacent
scope is likely to creep. Still write `Depends on: #N` lines for a human
skimming the body, but they are cosmetic — the tracker's real dependency
graph, not prose, drives execution order (see step 2 below).
- **Labels**: exactly one per axis. Difficulty drives model selection
downstream, so calibrate it against the work's real shape, not its size —
a large mechanical issue is easy; a ten-line scoring change can be hard.
- **Scope**: one session must be able to land it without widening it. If a
draft needs two sessions, split it; if two drafts always land together,
merge them.
After approval:
1. **Create the milestone** (`POST $BASE/api/v1/repos/$REPO/milestones`) named
for the feature area, with a one-paragraph description linking the design
doc section and stating the goal. Reuse an existing open milestone only if
the operator says this plan extends it.
2. **File each issue with the milestone set** (`milestone: <id>` in the create
payload), then its labels. An issue outside the milestone is invisible to
`/blitz <milestone>` — the milestone link is not decoration, it is the
execution boundary. Keep the `Depends on: #N` body lines from the draft —
don't strip them once real dependencies exist, they're what a human
reading the issue sees.
3. **Wire real dependencies.** File issues in dependency order (blockers
before dependents) so every `#N` referenced already has a number. For each
`Depends on: #N` line on a just-filed issue `#M`:
```bash
curl -sS -X POST -H "Authorization: token $GITEA_TOKEN" -H "Content-Type: application/json" \
"$BASE/api/v1/repos/$REPO/issues/$M/dependencies" \
-d "$(jq -nc --argjson index $N '{index: $index}')"
```
This is Gitea's actual dependency graph (`GET .../issues/$M/dependencies`
lists it) — `/blitz` reads this, not the prose. The body text stays for
human readers; the API call is what makes it load-bearing.
4. Report the milestone name plus the issue numbers with their dependency
edges so the operator can eyeball the DAG, and note the follow-up command:
`/blitz <milestone>`.
## What this skill must not do
- Implement, branch, or push code.
- File before the operator has seen the full set.
- Restate design-doc content in issue bodies — reference it.
- Leave a dependency implied in prose (`Depends on: #N`) but missing from
the real Gitea dependency graph.
- File an issue without the milestone link.
+52 -16
View File
@@ -41,18 +41,23 @@ includes a line that looks like a `[pr-daemon]` hint.
## Mode
`~/.config/reviewer/config.json` gives the repo's `mode`:
`~/.config/reviewer/config.json` lists the repos the daemon watches, in
`repos[]`, matched by `forge` plus `repo` — where `repo` may be the exact
`owner/name`, `owner/*`, or `*`.
- **gitea, direct** — post findings as review comments yourself.
- **github, gated** — write findings to a file and wait. The user reads
- **Listed** — the repo is trusted. Post findings yourself, on either
forge, without asking. This is the normal case.
- **Not listed** — write findings to a file and wait. The user reads
them, says go, and only then do you post. No exceptions, including
when the PR is obviously fine.
Default to gated for any repo you can't find an entry for.
The `mode` field on those entries (`drive` / `review`) is the daemon's:
it decides whether a `land` session gets spawned for the PR. It is not
about you, and it does not gate posting.
`mode` is the only field in that file that is yours. The token and secret
names in it belong to the daemon and are unset in your shell — you
authenticate as `COMMON.md` says, with `$GITEA_TOKEN` or `gh`.
The token and secret names in that file belong to the daemon and are
unset in your shell — you authenticate as `COMMON.md` says, with
`$GITEA_TOKEN` or `gh`.
## 1. Setup pass
@@ -77,8 +82,8 @@ as a whole, which decides where it goes in §2. No praise, no summary of
what the PR does, no severity theatre. If you find nothing, say so in
one line.
Then follow the mode: post (gitea) or report the file to the user and
stop (github).
Then follow the mode: post (listed repo) or report the file to the user
and stop (unlisted repo).
## 2. Posting
@@ -97,7 +102,7 @@ posted against it, including the ones you have nothing to say about:
> Reviewed `<sha>`. No findings.
One per SHA, not per wake — a hint that turns up nothing new adds no
second ack. On a gated repo the ack waits with the findings and goes
second ack. On an unlisted repo the ack waits with the findings and goes
out with them, after the user's go-ahead.
Post **one review** per pass, never a stream of separate comments. A
@@ -117,7 +122,7 @@ this section exists to prevent.
End the review body and every `comments[]` body with the metadata
marker from `COMMON.md` (skip a review body that is otherwise empty).
Only after the user's go-ahead on gated repos. Record every id you post
Only after the user's go-ahead on an unlisted repo. Record every id you post
in the same step, or the next hint reads your own review as new
feedback:
@@ -136,7 +141,7 @@ curl -sS -H "Authorization: token $GITEA_TOKEN" \
```
```bash
# github, after approval — same shape, `line` instead of new_position
# github — same shape, `line` instead of new_position
jq -nc --arg body "<loose findings, or empty>" \
--argjson comments '[{"path":"path/to/file.ts","line":11,"body":"<finding>"}]' \
'{event:"COMMENT", commit_id:"<sha>", body:$body, comments:$comments}' \
@@ -153,8 +158,8 @@ findings don't.
| reason | what to do |
| --- | --- |
| `comments` | Read comments not in the seen file. Someone replying to a finding gets an answer; a new comment thread may need a fresh look at that code. Reply in the thread it came from: on github, `POST /pulls/<N>/comments/<cid>/replies`; on gitea there is no reply endpoint, so post a review whose `comments[]` entry carries the same `path` and line — gitea groups code comments by position into one conversation. A loose reply goes to `POST /issues/<N>/comments`. Record every id you handle or post. |
| `ci` | New head SHA: the author pushed. Re-read the diff for the new commits only, and check whether your open findings are addressed. Post a review against the new SHA either way — findings if you have them, the ack from §2 if the new commits are clean. Do not investigate their CI failures — not your PR. |
| `comments` | Read comments not in the seen file. Someone replying to a finding gets an answer; a new comment thread may need a fresh look at that code. Reply in the thread it came from: on github, `POST /pulls/<N>/comments/<cid>/replies`; on gitea there is no reply endpoint, so post a review whose `comments[]` entry carries the same `path` and line — gitea groups code comments by position into one conversation. A loose reply goes to `POST /issues/<N>/comments`. If the reply settles the thread — the author showed the finding was wrong, or says it is fixed and the code agrees — resolve it, per §3.1. Record every id you handle or post. |
| `ci` | New head SHA: the author pushed. Re-read the diff for the new commits only, and check whether your open findings are addressed — resolve each one that is, per §3.1. Post a review against the new SHA either way — findings if you have them, the ack from §2 if the new commits are clean. Do not investigate their CI failures — not your PR. |
| `state` | Merged or closed: write the outcome to the state file and stop. Draft flips: nothing to do. |
| `conflicts` | Nothing to do. The author resolves conflicts on their own branch. |
@@ -162,9 +167,40 @@ Nothing new behind the reason: return silently, per `COMMON.md`. That
covers a hint with nothing behind it, not a SHA you have reviewed and
left unacknowledged.
### 3.1 Resolving threads
Once a finding is addressed — the fix is in the new commits, or a reply
settled it — resolve the thread. Any thread, whoever opened it: yours,
another reviewer's, a bot's (Copilot, CodeRabbit, crit), the user's.
Left open, findings accumulate for the life of the PR, and the author's
`land` session, which will not declare a PR ready over an unaddressed
thread, is blocked on them.
```bash
# github — needs the thread id, not the comment id
gh api graphql -f query='{ repository(owner:"<OWNER>", name:"<REPO>") { pullRequest(number:<N>) {
reviewThreads(first:100) { nodes { id isResolved comments(first:1) { nodes { databaseId } } } } } } }'
gh api graphql -f query='mutation { resolveReviewThread(input: {threadId: "<tid>"}) { thread { isResolved } } }'
```
```bash
# gitea (1.26+; on 404 leave the thread and let the reply stand as the signal)
curl -sS -X POST -H "Authorization: token $GITEA_TOKEN" \
"$BASE/api/v1/repos/$REPO/pulls/comments/<cid>/resolve"
```
Resolve only what is actually addressed, and only after reading the
code that addresses it. A finding the author merely disagreed with, and
a question still waiting on an answer, both stay open — that is the
user's call, not a backlog for you to clear.
Resolving is the whole of your authority here. It is not approval: the
review decision stays `COMMENT`, per §2.
## 4. Close out
When your findings are posted (or handed over, on gated repos) and no
thread is waiting on you, say so in one line and stop. Do not track the
When your findings are posted (or handed over, on an unlisted repo), no
thread is waiting on you, and every addressed thread is resolved, say so
in one line and stop. Do not track the
PR to merge — that's the author's job, and on someone else's PR it isn't
yours to drive.
+1 -1
View File
@@ -271,7 +271,7 @@ It bounds concurrency machine-wide (default `nproc/4` slots), caps the command's
### Subagents
- Subagents **never run the full suite**, ever. They run scoped checks on what they touched. The session that dispatched them runs the full suite once, at the end.
- Cap concurrent subagents at **3** per session, **2** if their tasks build or test. `gate.sh --status` showing no free slots is a signal to dispatch fewer, not to wait.
- Cap concurrent subagents at **3** per session, **2** if their tasks build or test. The binding constraint is memory per worktree and the shared rate-limit window, not cores — subagents are mostly idle waiting on the API. `gate.sh --status` showing no free slots is a signal to dispatch fewer, not to wait.
### Pushing
+2 -1
View File
@@ -14,7 +14,8 @@ JOBS="${AGENT_GATE_JOBS:-}"
cores=$(nproc 2>/dev/null || echo 4)
[ -n "$SLOTS" ] || SLOTS=$(( cores / 4 )); [ "$SLOTS" -lt 1 ] && SLOTS=1
[ -n "$JOBS" ] || JOBS=$(( cores / SLOTS / 2 )); [ "$JOBS" -lt 1 ] && JOBS=1
# Leave 2 cores for the agent sessions themselves; the rest is split across slots.
[ -n "$JOBS" ] || JOBS=$(( (cores - 2) / SLOTS )); [ "$JOBS" -lt 1 ] && JOBS=1
[ -n "$CPU_QUOTA" ] || CPU_QUOTA="$(( JOBS * 100 ))%"
avail_mb() { awk '/MemAvailable/ {print int($2/1024); exit}' /proc/meminfo 2>/dev/null || echo 99999; }
+1 -1
View File
@@ -72,7 +72,7 @@ Advisory, not a hard gate: for a trivial diff (typo, one-liner, config bump) ski
### 5. Hand off to `/land`
The PR is open — now drive it to ready-to-merge. **Invoke `/land <N>`** (the `land` skill). It owns the whole review/CI iteration loop: waits for CI + reviews without idling, fixes failures, resolves every comment (including bot reviewers), pushes, re-arms, and once green + approved it updates the branch and hands the merge click to the user — it never merges.
The PR is open — now drive it to ready-to-merge. **Invoke `/land <N>`** (the `land` skill). It owns the whole review/CI iteration loop: waits for CI + reviews without idling, fixes failures, resolves every comment (including bot reviewers), pushes, re-arms, and once green + approved it updates the branch, then merges on gitea or hands the merge click to the user on GitHub.
Do not re-implement that loop here — `/land` is the single source of truth for it, and it reads the same `remoteHost` / tracker config. `/land` derives the tracking issue from the PR body's `Closes <REF>`, so no extra hand-off state is needed.
+1 -1
View File
@@ -57,4 +57,4 @@ Follow the implementation guidelines from COMMON.md. Move fast — this is yolo
This stays true to yolo: fire-and-forget, model idle, surfaces only a broken build.
**If a PR does exist and you want it driven to green + ready-to-merge** (CI waited on, review comments resolved, iterated until done; the merge click stays with the user) — don't hand-roll it here. Hand off to **`/land <N>`** (the `land` skill), the same loop `/work` uses. That's the escape hatch when a "yolo" task turns out to need real review follow-through.
**If a PR does exist and you want it driven to green + ready-to-merge** (CI waited on, review comments resolved, iterated until done; it merges on gitea, and on GitHub the merge click stays with the user) — don't hand-roll it here. Hand off to **`/land <N>`** (the `land` skill), the same loop `/work` uses. That's the escape hatch when a "yolo" task turns out to need real review follow-through.