GitHub-style runner: research and author the design plan in docs/ (no implementation until reviewed) #288

Closed
opened 2026-09-10 13:56:40 +00:00 by crueber · 3 comments
Owner

What's requested

Produce a researched design plan for a GitHub-Actions-style runner system in walhub: builds and jobs that execute as part of walhub, reported into the existing checks/statuses surface. The deliverable of this ticket is documentation only — a plan in docs/. No implementation. Nothing gets built until Chris reviews the plan.

Context (what exists that the plan must build on)

  • Statuses, not jobs, today. internal/checks implements the commit-status model: external CI POSTs per-commit results (POST …/api/checks/statuses/{sha}, wct_ CI-token auth), the CAS'd checks index feeds the PR merge gating and the /checks UI. There is no execution anywhere — walhub reports what external runners did.
  • Precedent for plan-first features. The mirror feature (#240) followed exactly this shape: researched plan doc → user review → implementation tickets. Its doc is docs/features/11_mirror.md, with numbered decision sections (R1 (a)/(b) markers), status header linking the Forgejo issue, and "plan revision is normative on conflict."
  • Doc conventions to follow: feature spec in docs/features/NN_name.md (next number is 12); architecture/design detail in docs/go/ (the 17-numbered series); dependency budget is LAW (AGENTS.md law 1: backend = chi, BurntSushi/toml, x/net only — the plan must work within this or explicitly propose an amendment with rationale); Make is the only task runner; go test ≥95% per-package gate; everything persists in the object store (no external DB).
  • Relevant neighbors: the task system (internal/wal/tasks.go — narrated long work with (repo,kind) single-flight), the WAL engine, bundles/maintenance scheduling (internal/maintain — an existing loop goroutine pattern), egress rules (internal/egress — SSRF posture for outbound runner traffic), and SSH transport (internal/sshd).

The plan must research and decide (numbered decision sections, mirror-doc style)

  1. Runner model. Self-hosted pull-based agents (GitHub Actions runner shape: poll/claim job, execute, report) vs server-executed jobs (walhub spawns directly) vs hybrid. Consideration: walhub instances are disposable; runners likely run elsewhere — what's the trust/coordinator split? Include the job claim/lease/heartbeat protocol and how it maps onto the object store + SSE (no external queue — or argue for one as a law-1 amendment).
  2. Job definition. Workflow file format and location (.walhub/*.yml in-repo? store-side config?), trigger model (push, PR, manual, schedule — note the mirror-schedule machinery in internal/mirror as prior art), matrix/strategy scope for v1.
  3. Execution environment. Container-per-job vs bare subprocess; image resolution/registry auth; what the runner does and does not get access to (clone tokens — reuse the wct_ shape?; secrets storage — the object-store-only constraint makes secret management a real design problem, don't hand-wave it).
  4. Integration surfaces. Checks/statuses reporting (extend internal/checks ReportInput with a richer check model — annotations, logs, steps — or version the API); log streaming (SSE envelope per 07 §9.3); artifacts (store layout, retention, size caps); UI (repo Checks tab extension, run/job detail pages).
  5. Security posture. Egress rules for runners, isolation boundaries, secret redaction (the scrubURL/scrubError precedent), abuse caps (max concurrent, max runtime, max bytes — the import [import] bounds section is the pattern), and what happens on the multi-instance placement model.
  6. Rollout slices. A staged path: e.g. R1 = manual single-job runs server-side, R2 = pull-based external runners, R3 = schedules/matrices — each slice independently shippable, mapped to acceptance criteria.
  7. Explicit non-goals for v1 and open questions surfaced for Chris's review.

Constraints for the author

  • Research before writing: survey how Gitea/Forgejo (act/act_runner), Woodpecker CI, and GitHub Actions structure their runner protocols — walhub already runs on Forgejo + Woodpecker, so the author can cite lived experience. Cite what's borrowed vs invented.
  • The plan doc must respect the dependency law or explicitly draft the amendment text for user approval (a container/runtime dependency for job execution is the most likely candidate — call it out early, don't bury it).
  • Write the doc as if it will be normative: numbered decisions with rationale, wire sketches, store key layouts, and a testing-strategy section per 15_testing.md.
  • Do not create implementation tickets, touch internal/, or start any code. Stop at the reviewed plan.

Acceptance criteria

  • docs/features/12_runner.md (or docs/go/18_runner.md if the author judges it architecture-not-feature) contains the full researched plan with numbered decision sections, prior-art citations, and rollout slices.
  • Every law-1 (dependency) conflict is drafted as an explicit amendment for user decision, not silently assumed.
  • The plan covers: runner model + protocol, job format + triggers, execution/isolation, secrets, checks integration, logs, artifacts, security caps, UI surfaces, testing strategy, and non-goals.
  • The doc ends with an explicit "awaiting review — do not implement" marker naming the reviewer (Chris).
  • No code changes anywhere in the repo.
## What's requested Produce a **researched design plan** for a GitHub-Actions-style **runner system** in walhub: builds and jobs that execute as part of walhub, reported into the existing checks/statuses surface. **The deliverable of this ticket is documentation only — a plan in `docs/`. No implementation. Nothing gets built until Chris reviews the plan.** ## Context (what exists that the plan must build on) - **Statuses, not jobs, today.** `internal/checks` implements the commit-status model: external CI `POST`s per-commit results (`POST …/api/checks/statuses/{sha}`, `wct_` CI-token auth), the CAS'd checks index feeds the PR merge gating and the `/checks` UI. There is no execution anywhere — walhub reports what *external* runners did. - **Precedent for plan-first features.** The mirror feature (#240) followed exactly this shape: researched plan doc → user review → implementation tickets. Its doc is `docs/features/11_mirror.md`, with numbered decision sections (R1 (a)/(b) markers), status header linking the Forgejo issue, and "plan revision is normative on conflict." - **Doc conventions to follow:** feature spec in `docs/features/NN_name.md` (next number is 12); architecture/design detail in `docs/go/` (the 17-numbered series); dependency budget is LAW (`AGENTS.md` law 1: backend = chi, BurntSushi/toml, x/net only — the plan must work within this or explicitly propose an amendment with rationale); Make is the only task runner; `go test` ≥95% per-package gate; everything persists in the object store (no external DB). - **Relevant neighbors:** the task system (`internal/wal/tasks.go` — narrated long work with `(repo,kind)` single-flight), the WAL engine, bundles/maintenance scheduling (`internal/maintain` — an existing loop goroutine pattern), egress rules (`internal/egress` — SSRF posture for outbound runner traffic), and SSH transport (`internal/sshd`). ## The plan must research and decide (numbered decision sections, mirror-doc style) 1. **Runner model.** Self-hosted pull-based agents (GitHub Actions runner shape: poll/claim job, execute, report) vs server-executed jobs (walhub spawns directly) vs hybrid. Consideration: walhub instances are disposable; runners likely run elsewhere — what's the trust/coordinator split? Include the job claim/lease/heartbeat protocol and how it maps onto the object store + SSE (no external queue — or argue for one as a law-1 amendment). 2. **Job definition.** Workflow file format and location (`.walhub/*.yml` in-repo? store-side config?), trigger model (push, PR, manual, schedule — note the mirror-schedule machinery in `internal/mirror` as prior art), matrix/strategy scope for v1. 3. **Execution environment.** Container-per-job vs bare subprocess; image resolution/registry auth; what the runner does and does not get access to (clone tokens — reuse the `wct_` shape?; secrets storage — the object-store-only constraint makes secret management a real design problem, don't hand-wave it). 4. **Integration surfaces.** Checks/statuses reporting (extend `internal/checks` ReportInput with a richer check model — annotations, logs, steps — or version the API); log streaming (SSE envelope per 07 §9.3); artifacts (store layout, retention, size caps); UI (repo Checks tab extension, run/job detail pages). 5. **Security posture.** Egress rules for runners, isolation boundaries, secret redaction (the `scrubURL`/`scrubError` precedent), abuse caps (max concurrent, max runtime, max bytes — the import `[import]` bounds section is the pattern), and what happens on the multi-instance placement model. 6. **Rollout slices.** A staged path: e.g. R1 = manual single-job runs server-side, R2 = pull-based external runners, R3 = schedules/matrices — each slice independently shippable, mapped to acceptance criteria. 7. **Explicit non-goals** for v1 and open questions surfaced for Chris's review. ## Constraints for the author - **Research before writing:** survey how Gitea/Forgejo (act/act_runner), Woodpecker CI, and GitHub Actions structure their runner protocols — walhub already runs on Forgejo + Woodpecker, so the author can cite lived experience. Cite what's borrowed vs invented. - The plan doc must respect the dependency law or explicitly draft the amendment text for user approval (a container/runtime dependency for job execution is the most likely candidate — call it out early, don't bury it). - Write the doc as if it will be normative: numbered decisions with rationale, wire sketches, store key layouts, and a testing-strategy section per `15_testing.md`. - **Do not** create implementation tickets, touch `internal/`, or start any code. Stop at the reviewed plan. ## Acceptance criteria - [ ] `docs/features/12_runner.md` (or `docs/go/18_runner.md` if the author judges it architecture-not-feature) contains the full researched plan with numbered decision sections, prior-art citations, and rollout slices. - [ ] Every law-1 (dependency) conflict is drafted as an explicit amendment for user decision, not silently assumed. - [ ] The plan covers: runner model + protocol, job format + triggers, execution/isolation, secrets, checks integration, logs, artifacts, security caps, UI surfaces, testing strategy, and non-goals. - [ ] The doc ends with an explicit "awaiting review — do not implement" marker naming the reviewer (Chris). - [ ] No code changes anywhere in the repo.
Author
Owner

Design plan ready for review: PR #304 (branch docs/issue-288) — docs/features/12_runner.md (DRAFT, D1–D8 + R1/R2/R3 slices, open questions Q1–Q6 for Chris). Docs only, no implementation. Please review the plan there; this PR must not merge until the plan is approved.

Design plan ready for review: PR #304 (branch docs/issue-288) — docs/features/12_runner.md (DRAFT, D1–D8 + R1/R2/R3 slices, open questions Q1–Q6 for Chris). Docs only, no implementation. Please review the plan there; this PR must not merge until the plan is approved.
Author
Owner

PR #304 review (DOC-ONLY plan for #288) — findings:

VERIFICATION

  • Diff is docs-only: docs/features/12_runner.md (+495 new) + docs/features/README.md table row (+1). No code, config, CLI, or compose changes (git diff --numstat confirms).
  • Markdown eyeball: sections 0-10 + Decisions section, 3 closed code fences, token table well-formed, status header + closing awaiting-review-do-not-implement marker naming Chris present.
  • No test runs (docs only — noted), no browser (nothing browser-facing; markdown eyeball only — noted), no docker/compose, no system packages. Main worktree untouched, still clean on main.
  • Two small doc fixes pushed to origin/docs/issue-288 as 6f88bdf (scratch worktree, main untouched): header D1-D9 corrected to D1-D8 (only D1-D8 exist); Q6 (YAML-vs-TOML) lived only in section 10, now cross-referenced in the section 9 open-questions list.

LAW COMPLIANCE (of the PROPOSED design)

  • Law 1 (deps): zero amendments proposed; execution deps kept out-of-tree (reference runner not in server budget, Forgejo server/act_runner split cited). Reserve container-socket amendment drafted-not-proposed in section 10 — correct handling. One non-blocking nit for implementation time: D6 claims secretbox is already inside the law-1 budget via the SSH amendment — same module (x/crypto) but a different package (nacl/secretbox vs ssh) than the SSH-server-transport-only parenthetical covers. Suggest a one-line parenthetical amendment (SSH + secretbox) with the implementation change.
  • Law 8 (seams): new internal/actions package, Seam 4 event-sink fan-out, Seam 5 task kind actions-run, additive ReportInput extension per 14 section 14.12 field rule (verified the rule exists), frozen-overwritable-list amendment deferred to the implementation change per 14 section 14.11 rule 2 + law 12. Frozen Principal untouched (wct_ precedent followed). No seam violation proposed.
  • Law 4 (bucket): all coordinator state bucket-native (actions/ family), CAS is the lock, probe-don-t-list, hot-window queue cap. Passes the wipe-every-instance test.
  • Law 6 (round trips): budgets stated (enqueue <=4, poll-hit <=3, poll-miss 1 probe per 5s window) with sim-pinning intent. No gratuitous sequential store reads on hot paths.
  • Law 3 (concurrency): Concurrency subsections present (sections 3 and D6/D7); no repo locks taken; claim = single CAS with 409-retry, sweep requeue with CAS re-read, coordinator-allocated log seqs. Matches the 13_concurrency hazard+avoidance convention.
  • Law 12 (docs with code): Decisions entries D1-D8 + R-slices + no-amendment note, each with rationale, dated #288. Matches the mirror-doc precedent.

PLAN COHERENCE

  • D1 pull-based split: sound (NAT/disposability, three prior arts converged). Trust boundary honest, R1 exception explicit.
  • D2/D3 queue+claim: CAS-claim with loser retries, 60s lease + 30s grace, sweeper-only requeue, no steal of live heartbeats; exactly-once correctly disclaimed for idempotent effects (coordinator seqs, 412-on-resend, last-write-wins checks per doc 05 section 2). Stuck-queue surfacing answers the act_runner lived pain.
  • D4 tokens: art_/wrt_/wjt_ narrow scopes, hash-only storage, single-use reg tokens, machine-users rejected with role-lattice rationale. Mirrors the wct_ shape correctly.
  • D5 workflows: in-repo pinning to trigger sha, strict-parse fail-closed, label-subset scheduling; matrix/schedule deferrals clean.
  • D6 secrets: sealed secretbox envelopes + operator-held key, 503-without-key fail-closed, coordinator-side interpolation, env-only injection, double redaction (runner hints + server scrub backstop). In-flight pinning documented as intended.
  • D7 reporting: additive checks extension (gate/combined view unmodified), immutable log chunks with bounded viewer windows, content-addressed artifact blobs, SSE live-tail on the collab stream. UI/SDK follow the 08 pattern.
  • D8 security: honest non-sandbox statement, [actions] caps in the [import]-bounds pattern, fork-PR secretless-until-approved default, runner egress as runner-operator firewall. Fail-closed throughout.
  • R1/R2/R3 independently shippable with acceptance criteria each; R1 proves the reporting spine before widening trust. Sound ordering.
  • D1-D8 each carry rationale. Q1-Q6 all genuine owner calls, none answerable from the tree (R1 exec model, secrets key mgmt, org-runner scope, .walhub vs .github path, hosting-bill defaults, YAML-vs-TOML). Non-goals sane (marketplace, reusable workflows, OIDC federation, org pools, GPU, cache, annotations, environments, per-branch routing, Win/Mac agent).
  • Placement: docs/features/12_runner.md correct (collaboration layer, next number 12, no docs/go/ number claimed). README row correct; 05 dependency accurate.

MERGE RECOMMENDATION: ready to merge (as DRAFT plan doc; implementation stays gated on #288 approval per the doc's own marker). No structural blockers; remaining items (secretbox parenthetical one-liner, Q-answers) belong to implementation planning, not this doc.

PR #304 review (DOC-ONLY plan for #288) — findings: VERIFICATION - Diff is docs-only: docs/features/12_runner.md (+495 new) + docs/features/README.md table row (+1). No code, config, CLI, or compose changes (git diff --numstat confirms). - Markdown eyeball: sections 0-10 + Decisions section, 3 closed code fences, token table well-formed, status header + closing awaiting-review-do-not-implement marker naming Chris present. - No test runs (docs only — noted), no browser (nothing browser-facing; markdown eyeball only — noted), no docker/compose, no system packages. Main worktree untouched, still clean on main. - Two small doc fixes pushed to origin/docs/issue-288 as 6f88bdf (scratch worktree, main untouched): header D1-D9 corrected to D1-D8 (only D1-D8 exist); Q6 (YAML-vs-TOML) lived only in section 10, now cross-referenced in the section 9 open-questions list. LAW COMPLIANCE (of the PROPOSED design) - Law 1 (deps): zero amendments proposed; execution deps kept out-of-tree (reference runner not in server budget, Forgejo server/act_runner split cited). Reserve container-socket amendment drafted-not-proposed in section 10 — correct handling. One non-blocking nit for implementation time: D6 claims secretbox is already inside the law-1 budget via the SSH amendment — same module (x/crypto) but a different package (nacl/secretbox vs ssh) than the SSH-server-transport-only parenthetical covers. Suggest a one-line parenthetical amendment (SSH + secretbox) with the implementation change. - Law 8 (seams): new internal/actions package, Seam 4 event-sink fan-out, Seam 5 task kind actions-run, additive ReportInput extension per 14 section 14.12 field rule (verified the rule exists), frozen-overwritable-list amendment deferred to the implementation change per 14 section 14.11 rule 2 + law 12. Frozen Principal untouched (wct_ precedent followed). No seam violation proposed. - Law 4 (bucket): all coordinator state bucket-native (actions/ family), CAS is the lock, probe-don-t-list, hot-window queue cap. Passes the wipe-every-instance test. - Law 6 (round trips): budgets stated (enqueue <=4, poll-hit <=3, poll-miss 1 probe per 5s window) with sim-pinning intent. No gratuitous sequential store reads on hot paths. - Law 3 (concurrency): Concurrency subsections present (sections 3 and D6/D7); no repo locks taken; claim = single CAS with 409-retry, sweep requeue with CAS re-read, coordinator-allocated log seqs. Matches the 13_concurrency hazard+avoidance convention. - Law 12 (docs with code): Decisions entries D1-D8 + R-slices + no-amendment note, each with rationale, dated #288. Matches the mirror-doc precedent. PLAN COHERENCE - D1 pull-based split: sound (NAT/disposability, three prior arts converged). Trust boundary honest, R1 exception explicit. - D2/D3 queue+claim: CAS-claim with loser retries, 60s lease + 30s grace, sweeper-only requeue, no steal of live heartbeats; exactly-once correctly disclaimed for idempotent effects (coordinator seqs, 412-on-resend, last-write-wins checks per doc 05 section 2). Stuck-queue surfacing answers the act_runner lived pain. - D4 tokens: art_/wrt_/wjt_ narrow scopes, hash-only storage, single-use reg tokens, machine-users rejected with role-lattice rationale. Mirrors the wct_ shape correctly. - D5 workflows: in-repo pinning to trigger sha, strict-parse fail-closed, label-subset scheduling; matrix/schedule deferrals clean. - D6 secrets: sealed secretbox envelopes + operator-held key, 503-without-key fail-closed, coordinator-side interpolation, env-only injection, double redaction (runner hints + server scrub backstop). In-flight pinning documented as intended. - D7 reporting: additive checks extension (gate/combined view unmodified), immutable log chunks with bounded viewer windows, content-addressed artifact blobs, SSE live-tail on the collab stream. UI/SDK follow the 08 pattern. - D8 security: honest non-sandbox statement, [actions] caps in the [import]-bounds pattern, fork-PR secretless-until-approved default, runner egress as runner-operator firewall. Fail-closed throughout. - R1/R2/R3 independently shippable with acceptance criteria each; R1 proves the reporting spine before widening trust. Sound ordering. - D1-D8 each carry rationale. Q1-Q6 all genuine owner calls, none answerable from the tree (R1 exec model, secrets key mgmt, org-runner scope, .walhub vs .github path, hosting-bill defaults, YAML-vs-TOML). Non-goals sane (marketplace, reusable workflows, OIDC federation, org pools, GPU, cache, annotations, environments, per-branch routing, Win/Mac agent). - Placement: docs/features/12_runner.md correct (collaboration layer, next number 12, no docs/go/ number claimed). README row correct; 05 dependency accurate. MERGE RECOMMENDATION: ready to merge (as DRAFT plan doc; implementation stays gated on #288 approval per the doc's own marker). No structural blockers; remaining items (secretbox parenthetical one-liner, Q-answers) belong to implementation planning, not this doc.
Author
Owner

Plan landed in PR #304 (review clean; implementation gated on plan approval per the doc marker). Closing the planning ticket — implementation to follow on approval.

Plan landed in PR #304 (review clean; implementation gated on plan approval per the doc marker). Closing the planning ticket — implementation to follow on approval.
crueber added this to the v1 milestone 2026-09-10 22:20:53 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
crueber/walhub#288
No description provided.