Self-heal process for missing repos and missing data #209

Closed
opened 2026-09-08 19:13:09 +00:00 by crueber · 6 comments
Owner

Full plan in comments below (planning subagent output). Summary: (a) empty repos (manifest, no refs — like acme/waveB) stop being treated as errors: guided empty state, zero toasts, zero tasks; (b) refs-present-but-objects-missing: detect via fsck, classify, degrade gracefully, guide the admin; re-fetch only from explicitly configured upstream; give-up rule with stalled flag. No resurrection promises.

Full plan in comments below (planning subagent output). Summary: (a) empty repos (manifest, no refs — like acme/waveB) stop being treated as errors: guided empty state, zero toasts, zero tasks; (b) refs-present-but-objects-missing: detect via fsck, classify, degrade gracefully, guide the admin; re-fetch only from explicitly configured upstream; give-up rule with stalled flag. No resurrection promises.
Author
Owner

TICKET 1 — Self-heal process for missing repos / missing data

0. Problem statement (verified live 2026-09-06)

acme/waveB has a manifest with head:null, branches:0, tags:0 — a repo that
exists (manifest present, summary 200) but has zero refs. Tree / resolve /
commits all 404 (ErrNotFound: "unborn HEAD" — internal/api/bind_wal.go:181,
"HEAD" — :187). The UI surfaces these 404s as error-tray toasts
(web/src/lib/data.js:start() line 121: every fetch failure → reportError)
instead of guiding the user. There is no working "create empty repo" flow
(see Ticket 2), so an empty repo is a dead end: it looks broken, not new.

Two distinct states are being conflated and must be handled separately:

  • (a) Empty repo (manifest present, no refs). This is a legitimate state,
    not damage. Every repo passes through it: Registry.createSlow
    (internal/wal/registry.go:224-263) does PutCreate of
    Manifest{HeadSeq:0, MinSeq:0, Revision:1} with no refs, then inits the
    local bare repo. "Heal" here means stop treating it as an error —
    guided empty state, never toasts.
  • (b) Refs present but objects missing. This is damage (or a cold
    cache). Recovery Buzzword-Bingo ("self-heal") must not promise resurrection:
    with no upstream copy, missing objects are unrecoverable, full stop. The
    plan is: detect (fsck) → classify → degrade gracefully → guide the admin,
    with re-fetch only where a byte source actually exists.

1. How waveB came to exist (repo creation today — all paths)

Determined from code, not guessed. Every path below ends in the same
empty-manifest state, which is why "empty" is normal:

  1. Explicit create endpoint (the "no working flow" of the ticket — the
    endpoint EXISTS, the UI flow does not).

    PUT /{o}/{r} or PUT …/api (docs/go/06_server_http.md:207,
    internal/api/summary.go:53-85 repoPut): gate AuthWrite
    (require_write), ?object_format=sha1|sha256 (400 on bogus,
    internal/api/gaps5_test.go:381), Create → 201 {owner,name,full_name},
    exists → 409 "repository already exists". CLI twin:
    walhub repo create <REPO> [--object-format] (cmd/walhub/main.go:64,
    docs/go/11_config_cli.md:384). Both produce exactly the waveB state
    (manifest, HeadSeq 0, no refs). There is no UI page that calls it —
    no "New repository" button, no repo.create() in web/sdk/src/repo.js
    (that file has refs/tree/blob/commits/commit/overview only), no route in
    web/src/pages/. So: creatable by API/CLI, undiscoverable in UI.
  2. Auto-create on push (the default path; first-run default true —
    internal/config/firstrun.go:18, docs/go/06_server_http.md:248).
    Two interception points, both require_write-gated BEFORE creation:
    • GET …/info/refs?service=git-receive-pack on unknown repo + AutoCreate
      → advertises an empty ref list instead of 404, creating the repo
      (internal/server/smart.go:143-157). The push that follows fills it.
    • POST …/git-receive-pack → engine.Repo(ctx, id, create=true, …)
      (smart.go:419-428); not-found with auto-create off → 404.
      An empty repo survives from this path iff the push never landed (client
      aborted after info/refs, push rejected by policy, pack too large —
      smart.go:414-418 413, push-pipeline refusal e.g. managed refs
      internal/git/managed.go). waveB is plausibly one of these residues.
  3. Fork provisioning residue. web/src/lib/data.js:62-71 documents it:
    a fork writes repos/<o>/<r>/fork.json before the child manifest
    exists
    , so listings name a child whose manifest-gated reads 404. The
    inverse (manifest without refs) arises if the fork target was created
    (manifest Create won the CAS) but no refs were ever pushed/copied.
  4. Import path (docs/features/10_git_import.md): POST /api/v1/repos/imports
    → repo-import task → target manifest created, then content fetched. A
    failed/never-run import leaves manifest-without-refs.
  5. Test seed. internal/api/*_test.go, cmd/walhub/*_test.go create
    demo/empty-style repos routinely; a dev-server pointed at a test store
    (or a copied data dir) inherits them.

Net: waveB needs no exotic explanation — empty-manifest is the normal
post-create state
, and at least four production paths strand it when the
first push never arrives. Ticket 2 (placeholder) turns this from residue
into a first-class state.

2. What "heal" can and cannot mean (hard-nosed)

State Signal Heal = Explicitly NOT
(a) manifest, HeadSeq 0, no refs summary 200, head:null, branches/tags 0 UI guidance + push docs; zero toasts, zero tasks any background job; nothing is broken
(b1) refs advertised, serving copy cold/missing objects git fsck fails locally but bucket packs intact re-materialize from bucket (local cache rebuild — always safe, bucket is truth, AGENTS law 4) not data loss; no admin needed
(b2) bucket packs missing, upstream.git configured fsck missing[] non-empty existing repair unit: fetch 500-oid batches via §7.9 helper, publish repair pack, set repaired_seq (internal/maintain/repair.go, plan.go:263-268) only works WITH upstream
(b3) bucket packs missing, NO upstream same signal, Upstream.Git == "" → repair predicate false, damage sits forever detection + admin guidance + graceful degradation (this ticket's new work) resurrection; fabricating objects; rewriting refs to hide loss

The existing machinery already covers (b1) implicitly (sync replays from the
bucket; local state is a cache — registry.go openSlow step 4) and (b2)
(the repair unit). The gap this ticket closes: (a) UX + (b3) visibility.

3. Design

3.1 State classification (one probe, no LIST — AGENTS law 4/6)

Add a pure classifier on data the summary path already fetches (no new hot-path
round trips; summary today = 1 manifest GET + local ref read):

RepoHealth = Empty        // manifest ok, HeadSeq==0, refs==0
           | Healthy      // refs resolve, objects present
           | Degraded     // refs present, some objects missing (fsck-backed)
           | Missing      // manifest absent (404 — not this ticket's subject;
                          //  handled by tolerateMissing + "not found" shell)
  • Empty is decided inline in the summary handler from the manifest + ref
    counts it already holds. Cost: 0 new requests (law 6 safe).
  • Degraded is decided ONLY from the cached fsck.pb report
    (internal/maintain/util.go:getFsckReport — one conditional GET, off the
    hot path; fsck unit 6 runs on its existing interval predicate
    plan.go:186-189, never inline on a request). No request goroutine ever
    runs git fsck (law 3 / 14 §14.11 rule 5).
  • New API field, additive per 14 §14.12: summaryBody.health: "empty" | "healthy" | "degraded" (+ optional missing_total when degraded). Old
    clients ignore it. head:null stays the wire signal for unborn (frozen —
    summary.go:16 "the one sanctioned null").

3.2 API behavior per state (no new endpoints except one admin op)

  • GET …/api (summary): always 200 for existing repos incl. empty (already
    true). Adds health. SWR + ETag(head sha / "empty") unchanged.
  • resolve / tree / blob / commits / commit on an empty repo:
    stay 404 (frozen wire behavior; the SDK and bind_wal_test.go:440-444
    pin "unborn HEAD" → ErrNotFound), BUT the 404 body gains a stable
    machine-readable marker, e.g. plain-text prefix empty repository: (plain-text
    errors are the frozen convention — features README P-conventions / 07 §2).
    Rationale: the UI must distinguish "empty, guide me" from "broken, toast
    me" without parsing prose; a prefix is additive and greppable.
  • GET …/overview (WAL health JSON, no-store): include the fsck summary
    (missing_total, repaired_seq, last audit at) when fsck.pb exists —
    it already loads it for the snapshot (maintain.go:285-297); surfacing is
    a read-only projection. This is the admin's machine interface.
  • New: POST …/api/ops/repair-check (require_write; joins existing
    (repo,kind) task semantics, §9.4)? Prefer reusing the existing fsck
    op over a new endpoint
    : the ops surface already exposes maintenance
    units (GET …/ops lists OpSpec; POST …/ops/{op} starts/joins).
    Specify: POST …/api/ops/fsck triggers an out-of-schedule audit as the
    existing fsck kind (Seam 5, see 3.3), SSE-attachable. Only add
    repair-check if the ops table cannot address fsck on demand — decide
    at implementation time, note in Decisions.

3.3 Task design (Seam 5 — 14_extensibility.md §14.7)

No new periodic task for empty repos. Empty is not damage; a sweeper
would burn LIST/GET budget against human-rate state (law 6) and violate
"no LIST on a hot path" for zero benefit. On-demand + existing cadence only:

  1. On-demand audit = existing KindFsck (internal/maintain/units.go:32).
    Trigger: POST …/ops/fsck (joins in-flight same (repo,kind) per the
    frozen join semantics — a second click attaches, never duplicates).
    Bounds: one unit per repo per pass (loop discipline, 14 §14.7
    Concurrency); git fsck --connectivity-only --no-dangling exact argv
    (fsck.go:55-60); missing list bounded at fsckMissingBound (100k,
    units.go:65) with unbounded missing_total. Report overwrite to
    fsck.pb (frozen overwritable family — no spec change needed).
    Cost per run: local subprocess + 1 Overwrite PUT. Gives up: never — an
    audit always completes with a report; repair is what gives up (below).
  2. Repair stays the existing KindRepair unit, predicate UNCHANGED
    (plan.go:265-268: RepairedSeq==0 && (Missing||Total) && Upstream.Git != ""). No new kind, no new lease (repair is lease-free by
    design — cheap + idempotent, repair.go:16-18). Bounds: 500-oid batches
    (repairBatch), publish via ordinary CAS ladder, repaired_seq=head
    disarms re-fire (repair.go:52-53). When upstream is absent the unit
    simply never fires — that is the (b3) case, handled by 3.4, not by
    forcing a fetch from nothing.
  3. Fork/network-peer re-fetch ("if any exist"): the ONLY sanctioned
    extension, and it is configuration, not magic. If Upstream.Git points
    at a peer that has the objects (a fork parent, a second walhub via the
    follow path — follow.go), the existing repair unit already fetches from
    it. What the plan adds: document that upstream.git MAY target a
    fork-network sibling, and the repair-check response SHOULD name the
    configured upstream (or "none configured") so the admin knows where a
    repair would fetch from. No cross-repo object snooping, no implicit
    peer discovery (that would be LIST-by-another-name and a privacy hole
    across private repos — P6 require_read applies to every read).
  4. Give-up rule (normative): repair attempts are bounded by the existing
    pass structure (one unit per repo per pass; publish failure keeps
    repaired_seq==0 → next pass retries — wave4b_test.go:512). After N
    consecutive error outcomes (suggest N=5, config maintenance.repair_retries,
    default 5), the unit stops retrying and the fsck.pb consumer surfaces
    repair_stalled:true + last error in overview. The data stays as-is;
    the admin guidance (3.4) takes over. A stalled repair must never block
    checkpoint/bundles/compaction (priority order already guarantees repair
    is #2 and skippable via Skip).

3.4 (b3) admin guidance + graceful degradation (the actual new UX)

When health==degraded with no repair path (no upstream / stalled):

  • Repo shell (Repo.jsx): amber (not red) banner under the header:
    "Some objects are missing (fsck: N missing). Reads may fail; pushes of
    new refs still work. Details in Settings → WAL." Links to the WAL page.
    Never a toast; never blocks navigation.
  • Tree/Blob/Commits pages: a fetch failure whose error carries the
    degraded marker renders an inline notice ("this object is missing from
    the store — see WAL health") instead of reportError. Mechanism: extend
    the tolerateMissing pattern (data.js:77-82) with a tolerateDegraded
    wrapper keyed on the new 404 prefix / the summary health the shell
    already holds (no extra fetch — the shell's summary entry is shared via
    context).
  • Settings → WAL tab: new "Object health" section rendering the
    overview fsck projection: missing_total, bounded sample of missing
    oids, repaired_seq, last audit time, configured upstream (or "none —
    set upstream.git to enable repair"), stalled flag + last error, and the
    exact CLI to re-run the audit (POST …/ops/fsck, plus the walhub
    Seam-7 twin if added). This is the admin repair surface — read + trigger,
    no new mutation semantics.
  • Degradation guarantee (state explicitly): pushes of NEW objects/refs
    keep working (publish path never consults fsck); reads of missing objects
    404 with the marker; reads of present objects are unaffected. Document
    that deletion of the repo (DELETE …/api, admin, summary.go:87-101)
    remains available as the last resort, with the fork/GC semantics of
    01_identity_permissions.md §5.1 (children unaffected).

3.5 (a) guided empty state (closes the waveB complaint)

  • Repo shell: Repo.jsx:496 already renders <span class="pill">empty</span>
    on head:null — keep, and add an EmptyRepoGuide block on the Code tab
    when summary.health=="empty" (equivalently !head && !branches && !tags):
    clone URL (reuse CloneMenu URL builders — server clone_url verbatim,
    issue #124 rule), git remote add origin <url> + git push -u origin main
    (copy buttons, same copyText helper), plus protocol toggle parity.
    No recipes fetch needed; static commands only.
  • Tree/Commits/Blob pages: when the shell knows the repo is empty, do
    NOT issue the doomed resolve/commits?n=1 fetch at all (the "by-design
    empty-repo commits?n=1 fetch" noted in docs/go/12_web_ui.md:517 becomes
    a suppressed fetch: selArgs()-style null means "do not fetch" —
    Tree.jsx:103-105 already establishes the pattern). Redirect
    /:owner/:repo/tree/* etc. to the Code-tab empty guide when empty.
    This kills the toasts at the source instead of filtering them downstream.
  • Data layer: useResolved short-circuits when the cached summary for
    the repo is known-empty (read the repo:{full} cache entry; no new fetch).
    Fallback (summary not yet loaded): the 404-with-empty repository: prefix
    maps to silent + guide, never reportError. reportError itself is
    untouched (frozen tray behavior for real errors).

4. How waveB specifically gets diagnosed under this plan

  1. GET /acme/waveB/api → 200, health:"empty" → shell shows empty guide
    (no toasts). Verdict: never broken, just unborn.
  2. If instead refs existed with missing objects: overview shows
    missing_total, upstream "none", repair never fires → amber banner +
    WAL-tab guidance ("set upstream.git or restore from a clone via push of
    the missing refs").
  3. Forensics for "how did it get empty": GET …/api/wal / log segments
    show zero PUSH entries; meta/import.json / fork.json presence
    distinguishes import/fork residue. No new endpoint — existing WAL reads.

5. EVIDENCE / perf notes

  • Empty path adds zero store round trips: health:"empty" derives from
    manifest+refs the summary already loads (law 6: warm-refs budget intact).
    The suppressed commits?n=1/resolve fetches remove 1–2 requests per
    empty-repo page load — measurable improvement, assert in sim/Tier-2 e2e:
    "empty repo Code tab issues 0 manifest-gated reads beyond summary".
  • overview fsck projection: +0 requests when fsck.pb absent (probe miss
    is the existing snapshot load); +0 when present (already loaded per pass —
    projection only). No hot-path change: overview is no-store admin page.
  • On-demand ops/fsck: bounded 1 subprocess + 1 PUT, (repo,kind)
    single-flight dedupes stampedes; NOT on any push/fetch path.
  • No LIST added anywhere (classification by probe; repair batches by exact
    oid fetch — FetchObjectsAsPack, units.go:159). Round-trip budgets
    (push ≤5, warm refs 1) untouched — record "no change" entries in
    docs/EVIDENCE.md with the harness (internal/devtools/) per AGENTS.md.
  • Coverage: new classifier + projection need table-driven httptest
    (≥95% pkg gate, -race); UI: node --test for the empty-guide render
    predicate + suppressed-fetch logic; real-browser check (AGENTS ladder #8)
    of /acme/waveB, a degraded fixture, and /setup console-clean.

6. Acceptance criteria

  1. acme/waveB (or fixture-identical empty repo) loads /, /tree/*,
    /commits with zero toasts, an "empty" pill, and a push-guide block
    showing the verbatim server clone_url + push -u origin main.
  2. resolve/tree/commits on empty repos still 404 for API clients (frozen),
    with the empty repository: marker prefix.
  3. Summary carries additive health (empty|healthy|degraded); no existing
    field changes; discovery/SDK updated (repo.get() typedef).
  4. Degraded fixture (refs + fsck.pb with missing, no upstream): amber
    banner, inline (not toast) read errors, WAL tab shows missing_total +
    "no upstream configured" + re-audit trigger; pushes of new refs succeed.
  5. Upstream-configured fixture: existing repair unit fires and heals
    (already covered by maintain_test.go:83+ — regression, not new).
  6. make cover ≥95% holds for touched packages; make sim budgets pass
    unchanged; real Chromium drive clean console (except the suppressed —
    i.e. absent — empty fetch).
  7. Doc updates in the same change (AGENTS law 12): docs/go/07_api.md
    (health field + 404 marker), docs/go/10_maintenance.md (give-up rule +
    repair_stalled), docs/go/12_web_ui.md (empty guide + banner),
    Decisions appends.

7. Risks / non-goals

  • Scope creep into resurrection: explicitly out. Any proposal to
    "reconstruct" commits from commit render caches, reflogs, or peer
    guessing is rejected — caches are not truth (law 4) and guessing breaks
    hash integrity. Detection + guidance only.
  • Peer-fetch privacy: re-fetch ONLY from explicitly configured
    upstream.git. Never auto-discover "peers" via listings — cross-repo
    reads bypass P6 visibility reasoning.
  • Toast-filter overreach: the fix suppresses fetches / classifies
    emptiness; it must NOT globally mute 404s (real "repo deleted under you"
    flows — #200 tolerateMissing → "not found" shell — depend on them).
  • fsck cost on huge repos: audit is subprocess + full local copy
    (fullCopy gate, plan.go:287); on-demand trigger needs the same
    TryAcquire per-repo semaphore discipline as git handlers (503 +
    Retry-After when busy), never a blocking wait.
  • No new overwritable bucket family (uses fsck.pb, already frozen) —
    keeps the 14 §14.11 frozen-list untouched. If repair_stalled needs
    durability beyond the in-memory snapshot, it rides inside fsck.pb as an
    additive optional field (14 §14.12 field rule), not a new key.

8. Work breakdown (implementer waves)

  1. Backend: classifier + health field + 404 marker + overview fsck
    projection (+ httptest). 2. Ops: on-demand fsck trigger wiring (if not
    already addressable) + stall rule + config key. 3. Web: empty guide,
    fetch suppression, degraded banner + WAL section (+ node tests). 4. Docs
    • EVIDENCE entries + browser proof. Each wave shippable alone.
# TICKET 1 — Self-heal process for missing repos / missing data ## 0. Problem statement (verified live 2026-09-06) `acme/waveB` has a manifest with `head:null, branches:0, tags:0` — a repo that exists (manifest present, summary 200) but has zero refs. Tree / resolve / commits all 404 (`ErrNotFound`: "unborn HEAD" — `internal/api/bind_wal.go:181`, "HEAD" — `:187`). The UI surfaces these 404s as error-tray toasts (`web/src/lib/data.js:start()` line 121: every fetch failure → `reportError`) instead of guiding the user. There is no working "create empty repo" flow (see Ticket 2), so an empty repo is a dead end: it looks broken, not new. Two distinct states are being conflated and must be handled separately: - **(a) Empty repo (manifest present, no refs).** This is a *legitimate* state, not damage. Every repo passes through it: `Registry.createSlow` (`internal/wal/registry.go:224-263`) does `PutCreate` of `Manifest{HeadSeq:0, MinSeq:0, Revision:1}` with no refs, then inits the local bare repo. "Heal" here means **stop treating it as an error** — guided empty state, never toasts. - **(b) Refs present but objects missing.** This is *damage* (or a cold cache). Recovery Buzzword-Bingo ("self-heal") must not promise resurrection: with no upstream copy, missing objects are unrecoverable, full stop. The plan is: **detect (fsck) → classify → degrade gracefully → guide the admin**, with re-fetch only where a byte source actually exists. ## 1. How `waveB` came to exist (repo creation today — all paths) Determined from code, not guessed. Every path below ends in the same empty-manifest state, which is why "empty" is normal: 1. **Explicit create endpoint (the "no working flow" of the ticket — the endpoint EXISTS, the UI flow does not).** `PUT /{o}/{r}` or `PUT …/api` (`docs/go/06_server_http.md:207`, `internal/api/summary.go:53-85 repoPut`): gate `AuthWrite` (`require_write`), `?object_format=sha1|sha256` (400 on bogus, `internal/api/gaps5_test.go:381`), `Create` → 201 `{owner,name,full_name}`, exists → **409** "repository already exists". CLI twin: `walhub repo create <REPO> [--object-format]` (`cmd/walhub/main.go:64`, `docs/go/11_config_cli.md:384`). Both produce exactly the waveB state (manifest, HeadSeq 0, no refs). There is **no UI page that calls it** — no "New repository" button, no `repo.create()` in `web/sdk/src/repo.js` (that file has refs/tree/blob/commits/commit/overview only), no route in `web/src/pages/`. So: creatable by API/CLI, undiscoverable in UI. 2. **Auto-create on push** (the default path; first-run default `true` — `internal/config/firstrun.go:18`, `docs/go/06_server_http.md:248`). Two interception points, both `require_write`-gated BEFORE creation: - `GET …/info/refs?service=git-receive-pack` on unknown repo + AutoCreate → advertises an **empty ref list** instead of 404, creating the repo (`internal/server/smart.go:143-157`). The push that follows fills it. - `POST …/git-receive-pack` → `engine.Repo(ctx, id, create=true, …)` (`smart.go:419-428`); not-found with auto-create off → 404. An empty repo survives from this path iff the push never landed (client aborted after `info/refs`, push rejected by policy, pack too large — `smart.go:414-418` 413, push-pipeline refusal e.g. managed refs `internal/git/managed.go`). waveB is plausibly one of these residues. 3. **Fork provisioning residue.** `web/src/lib/data.js:62-71` documents it: a fork writes `repos/<o>/<r>/fork.json` **before the child manifest exists**, so listings name a child whose manifest-gated reads 404. The inverse (manifest without refs) arises if the fork target was created (manifest `Create` won the CAS) but no refs were ever pushed/copied. 4. **Import path** (`docs/features/10_git_import.md`): `POST /api/v1/repos/imports` → `repo-import` task → target manifest created, then content fetched. A failed/never-run import leaves manifest-without-refs. 5. **Test seed.** `internal/api/*_test.go`, `cmd/walhub/*_test.go` create `demo/empty`-style repos routinely; a dev-server pointed at a test store (or a copied data dir) inherits them. Net: waveB needs no exotic explanation — **empty-manifest is the normal post-create state**, and at least four production paths strand it when the first push never arrives. Ticket 2 (placeholder) turns this from residue into a first-class state. ## 2. What "heal" can and cannot mean (hard-nosed) | State | Signal | Heal = | Explicitly NOT | |---|---|---|---| | (a) manifest, HeadSeq 0, no refs | summary 200, `head:null`, branches/tags 0 | UI guidance + push docs; zero toasts, zero tasks | any background job; nothing is broken | | (b1) refs advertised, serving copy cold/missing objects | `git fsck` fails locally but bucket packs intact | re-materialize from bucket (local cache rebuild — always safe, bucket is truth, AGENTS law 4) | not data loss; no admin needed | | (b2) bucket packs missing, `upstream.git` configured | fsck `missing[]` non-empty | existing `repair` unit: fetch 500-oid batches via §7.9 helper, publish repair pack, set `repaired_seq` (`internal/maintain/repair.go`, `plan.go:263-268`) | only works WITH upstream | | (b3) bucket packs missing, NO upstream | same signal, `Upstream.Git == ""` → repair predicate false, damage sits forever | **detection + admin guidance + graceful degradation** (this ticket's new work) | resurrection; fabricating objects; rewriting refs to hide loss | The existing machinery already covers (b1) implicitly (sync replays from the bucket; local state is a cache — `registry.go openSlow` step 4) and (b2) (the repair unit). The gap this ticket closes: **(a) UX + (b3) visibility**. ## 3. Design ### 3.1 State classification (one probe, no LIST — AGENTS law 4/6) Add a pure classifier on data the summary path already fetches (no new hot-path round trips; summary today = 1 manifest GET + local ref read): ``` RepoHealth = Empty // manifest ok, HeadSeq==0, refs==0 | Healthy // refs resolve, objects present | Degraded // refs present, some objects missing (fsck-backed) | Missing // manifest absent (404 — not this ticket's subject; // handled by tolerateMissing + "not found" shell) ``` - `Empty` is decided inline in the summary handler from the manifest + ref counts it already holds. Cost: 0 new requests (law 6 safe). - `Degraded` is decided ONLY from the cached `fsck.pb` report (`internal/maintain/util.go:getFsckReport` — one conditional GET, off the hot path; fsck unit 6 runs on its existing interval predicate `plan.go:186-189`, never inline on a request). No request goroutine ever runs `git fsck` (law 3 / 14 §14.11 rule 5). - New API field, additive per 14 §14.12: `summaryBody.health: "empty" | "healthy" | "degraded"` (+ optional `missing_total` when degraded). Old clients ignore it. `head:null` stays the wire signal for unborn (frozen — `summary.go:16` "the one sanctioned null"). ### 3.2 API behavior per state (no new endpoints except one admin op) - `GET …/api` (summary): always 200 for existing repos incl. empty (already true). Adds `health`. SWR + ETag(head sha / "empty") unchanged. - `resolve` / `tree` / `blob` / `commits` / `commit` on an **empty** repo: stay 404 (frozen wire behavior; the SDK and `bind_wal_test.go:440-444` pin "unborn HEAD" → ErrNotFound), BUT the 404 body gains a stable machine-readable marker, e.g. plain-text prefix `empty repository: ` (plain-text errors are the frozen convention — features README P-conventions / 07 §2). Rationale: the UI must distinguish "empty, guide me" from "broken, toast me" without parsing prose; a prefix is additive and greppable. - `GET …/overview` (WAL health JSON, no-store): include the fsck summary (`missing_total`, `repaired_seq`, last audit `at`) when `fsck.pb` exists — it already loads it for the snapshot (`maintain.go:285-297`); surfacing is a read-only projection. This is the admin's machine interface. - New: `POST …/api/ops/repair-check` (require_write; joins existing `(repo,kind)` task semantics, §9.4)? **Prefer reusing the existing fsck op over a new endpoint**: the ops surface already exposes maintenance units (`GET …/ops` lists `OpSpec`; `POST …/ops/{op}` starts/joins). Specify: `POST …/api/ops/fsck` triggers an out-of-schedule audit as the existing `fsck` kind (Seam 5, see 3.3), SSE-attachable. Only add `repair-check` if the ops table cannot address `fsck` on demand — decide at implementation time, note in Decisions. ### 3.3 Task design (Seam 5 — `14_extensibility.md §14.7`) **No new periodic task for empty repos.** Empty is not damage; a sweeper would burn LIST/GET budget against human-rate state (law 6) and violate "no LIST on a hot path" for zero benefit. On-demand + existing cadence only: 1. **On-demand audit = existing `KindFsck` (`internal/maintain/units.go:32`).** Trigger: `POST …/ops/fsck` (joins in-flight same `(repo,kind)` per the frozen join semantics — a second click attaches, never duplicates). Bounds: one unit per repo per pass (loop discipline, 14 §14.7 Concurrency); `git fsck --connectivity-only --no-dangling` exact argv (`fsck.go:55-60`); missing list bounded at `fsckMissingBound` (100k, `units.go:65`) with unbounded `missing_total`. Report overwrite to `fsck.pb` (frozen overwritable family — no spec change needed). Cost per run: local subprocess + 1 Overwrite PUT. Gives up: never — an audit always completes with a report; *repair* is what gives up (below). 2. **Repair stays the existing `KindRepair` unit, predicate UNCHANGED** (`plan.go:265-268`: `RepairedSeq==0 && (Missing||Total) && Upstream.Git != ""`). No new kind, no new lease (repair is lease-free by design — cheap + idempotent, `repair.go:16-18`). Bounds: 500-oid batches (`repairBatch`), publish via ordinary CAS ladder, `repaired_seq=head` disarms re-fire (`repair.go:52-53`). When upstream is absent the unit simply never fires — that is the (b3) case, handled by 3.4, not by forcing a fetch from nothing. 3. **Fork/network-peer re-fetch ("if any exist"):** the ONLY sanctioned extension, and it is *configuration*, not magic. If `Upstream.Git` points at a peer that has the objects (a fork parent, a second walhub via the follow path — `follow.go`), the existing repair unit already fetches from it. What the plan adds: document that `upstream.git` MAY target a fork-network sibling, and the `repair-check` response SHOULD name the configured upstream (or "none configured") so the admin knows *where* a repair would fetch from. No cross-repo object snooping, no implicit peer discovery (that would be LIST-by-another-name and a privacy hole across private repos — P6 `require_read` applies to every read). 4. **Give-up rule (normative):** repair attempts are bounded by the existing pass structure (one unit per repo per pass; publish failure keeps `repaired_seq==0` → next pass retries — `wave4b_test.go:512`). After N consecutive error outcomes (suggest N=5, config `maintenance.repair_retries`, default 5), the unit stops retrying and the `fsck.pb` consumer surfaces `repair_stalled:true` + last error in `overview`. The data stays as-is; the admin guidance (3.4) takes over. A stalled repair must never block checkpoint/bundles/compaction (priority order already guarantees repair is #2 and skippable via `Skip`). ### 3.4 (b3) admin guidance + graceful degradation (the actual new UX) When `health==degraded` with no repair path (no upstream / stalled): - **Repo shell** (`Repo.jsx`): amber (not red) banner under the header: "Some objects are missing (fsck: N missing). Reads may fail; pushes of new refs still work. Details in Settings → WAL." Links to the WAL page. Never a toast; never blocks navigation. - **Tree/Blob/Commits pages**: a fetch failure whose error carries the degraded marker renders an inline notice ("this object is missing from the store — see WAL health") instead of `reportError`. Mechanism: extend the `tolerateMissing` pattern (`data.js:77-82`) with a `tolerateDegraded` wrapper keyed on the new 404 prefix / the summary `health` the shell already holds (no extra fetch — the shell's summary entry is shared via context). - **Settings → WAL tab**: new "Object health" section rendering the `overview` fsck projection: `missing_total`, bounded sample of missing oids, `repaired_seq`, last audit time, configured upstream (or "none — set `upstream.git` to enable repair"), stalled flag + last error, and the exact CLI to re-run the audit (`POST …/ops/fsck`, plus the `walhub` Seam-7 twin if added). This is the admin repair surface — read + trigger, no new mutation semantics. - **Degradation guarantee** (state explicitly): pushes of NEW objects/refs keep working (publish path never consults fsck); reads of missing objects 404 with the marker; reads of present objects are unaffected. Document that deletion of the repo (`DELETE …/api`, admin, `summary.go:87-101`) remains available as the last resort, with the fork/GC semantics of `01_identity_permissions.md §5.1` (children unaffected). ### 3.5 (a) guided empty state (closes the waveB complaint) - **Repo shell**: `Repo.jsx:496` already renders `<span class="pill">empty</span>` on `head:null` — keep, and add an `EmptyRepoGuide` block on the Code tab when `summary.health=="empty"` (equivalently `!head && !branches && !tags`): clone URL (reuse `CloneMenu` URL builders — server `clone_url` verbatim, issue #124 rule), `git remote add origin <url>` + `git push -u origin main` (copy buttons, same `copyText` helper), plus protocol toggle parity. No recipes fetch needed; static commands only. - **Tree/Commits/Blob pages**: when the shell knows the repo is empty, do NOT issue the doomed `resolve`/`commits?n=1` fetch at all (the "by-design empty-repo `commits?n=1` fetch" noted in `docs/go/12_web_ui.md:517` becomes a *suppressed* fetch: `selArgs()`-style null means "do not fetch" — `Tree.jsx:103-105` already establishes the pattern). Redirect `/:owner/:repo/tree/*` etc. to the Code-tab empty guide when empty. This kills the toasts at the source instead of filtering them downstream. - **Data layer**: `useResolved` short-circuits when the cached summary for the repo is known-empty (read the `repo:{full}` cache entry; no new fetch). Fallback (summary not yet loaded): the 404-with-`empty repository:` prefix maps to silent + guide, never `reportError`. `reportError` itself is untouched (frozen tray behavior for real errors). ## 4. How waveB specifically gets diagnosed under this plan 1. `GET /acme/waveB/api` → 200, `health:"empty"` → shell shows empty guide (no toasts). Verdict: never broken, just unborn. 2. If instead refs existed with missing objects: `overview` shows `missing_total`, upstream "none", repair never fires → amber banner + WAL-tab guidance ("set upstream.git or restore from a clone via push of the missing refs"). 3. Forensics for "how did it get empty": `GET …/api/wal` / log segments show zero PUSH entries; `meta/import.json` / `fork.json` presence distinguishes import/fork residue. No new endpoint — existing WAL reads. ## 5. EVIDENCE / perf notes - Empty path adds **zero** store round trips: `health:"empty"` derives from manifest+refs the summary already loads (law 6: warm-refs budget intact). The suppressed `commits?n=1`/resolve fetches *remove* 1–2 requests per empty-repo page load — measurable improvement, assert in sim/Tier-2 e2e: "empty repo Code tab issues 0 manifest-gated reads beyond summary". - `overview` fsck projection: +0 requests when `fsck.pb` absent (probe miss is the existing snapshot load); +0 when present (already loaded per pass — projection only). No hot-path change: `overview` is no-store admin page. - On-demand `ops/fsck`: bounded 1 subprocess + 1 PUT, `(repo,kind)` single-flight dedupes stampedes; NOT on any push/fetch path. - No LIST added anywhere (classification by probe; repair batches by exact oid fetch — `FetchObjectsAsPack`, `units.go:159`). Round-trip budgets (push ≤5, warm refs 1) untouched — record "no change" entries in `docs/EVIDENCE.md` with the harness (`internal/devtools/`) per AGENTS.md. - Coverage: new classifier + projection need table-driven httptest (≥95% pkg gate, `-race`); UI: `node --test` for the empty-guide render predicate + suppressed-fetch logic; real-browser check (AGENTS ladder #8) of `/acme/waveB`, a degraded fixture, and `/setup` console-clean. ## 6. Acceptance criteria 1. `acme/waveB` (or fixture-identical empty repo) loads `/`, `/tree/*`, `/commits` with **zero toasts**, an "empty" pill, and a push-guide block showing the verbatim server `clone_url` + `push -u origin main`. 2. `resolve/tree/commits` on empty repos still 404 for API clients (frozen), with the `empty repository:` marker prefix. 3. Summary carries additive `health` (`empty|healthy|degraded`); no existing field changes; discovery/SDK updated (`repo.get()` typedef). 4. Degraded fixture (refs + `fsck.pb` with missing, no upstream): amber banner, inline (not toast) read errors, WAL tab shows missing_total + "no upstream configured" + re-audit trigger; pushes of new refs succeed. 5. Upstream-configured fixture: existing repair unit fires and heals (already covered by `maintain_test.go:83+` — regression, not new). 6. `make cover` ≥95% holds for touched packages; `make sim` budgets pass unchanged; real Chromium drive clean console (except the suppressed — i.e. absent — empty fetch). 7. Doc updates in the same change (AGENTS law 12): `docs/go/07_api.md` (health field + 404 marker), `docs/go/10_maintenance.md` (give-up rule + `repair_stalled`), `docs/go/12_web_ui.md` (empty guide + banner), Decisions appends. ## 7. Risks / non-goals - **Scope creep into resurrection**: explicitly out. Any proposal to "reconstruct" commits from `commit` render caches, reflogs, or peer guessing is rejected — caches are not truth (law 4) and guessing breaks hash integrity. Detection + guidance only. - **Peer-fetch privacy**: re-fetch ONLY from explicitly configured `upstream.git`. Never auto-discover "peers" via listings — cross-repo reads bypass P6 visibility reasoning. - **Toast-filter overreach**: the fix suppresses fetches / classifies emptiness; it must NOT globally mute 404s (real "repo deleted under you" flows — #200 `tolerateMissing` → "not found" shell — depend on them). - **fsck cost on huge repos**: audit is subprocess + full local copy (`fullCopy` gate, `plan.go:287`); on-demand trigger needs the same `TryAcquire` per-repo semaphore discipline as git handlers (503 + Retry-After when busy), never a blocking wait. - **No new overwritable bucket family** (uses `fsck.pb`, already frozen) — keeps the 14 §14.11 frozen-list untouched. If `repair_stalled` needs durability beyond the in-memory snapshot, it rides inside `fsck.pb` as an additive optional field (14 §14.12 field rule), not a new key. ## 8. Work breakdown (implementer waves) 1. Backend: classifier + `health` field + 404 marker + `overview` fsck projection (+ httptest). 2. Ops: on-demand `fsck` trigger wiring (if not already addressable) + stall rule + config key. 3. Web: empty guide, fetch suppression, degraded banner + WAL section (+ node tests). 4. Docs + EVIDENCE entries + browser proof. Each wave shippable alone.
Author
Owner

REVIEW — Ticket 1 self-heal plan (against CODE, not docs)

Verdict: PROCEED-WITH-FIXES (no redesign needed; 3 blocking scoping/correctness fixes, all cuttable).

VERIFIED CORRECT (checked in tree):

  • "unborn HEAD" / "HEAD" 404 sites: internal/api/bind_wal.go:181,187. Confirmed.
  • Empty-manifest post-create state: internal/wal/registry.go:224-263 (HeadSeq 0, Revision 1, no refs). Confirmed.
  • Auto-create interception points: internal/server/smart.go:143-157 (info/refs advert) and :419-428 (receive-pack Repo(create)). Confirmed, both require_write-gated upstream.
  • PUT 201/409 + ?object_format 400: internal/api/summary.go:53-85; pinned by gaps5_test.go:381-388. Confirmed.
  • Repair predicate + batch + bound: plan.go:265-268 (RepairedSeq==0 && (Missing||Total) && Upstream.Git!=""), units.go:62 (500), units.go:65 (100k). Confirmed.
  • fsck argv: maintain/fsck.go:60 fsck --connectivity-only --no-dangling. Confirmed.
  • POST ops/fsck already works: opsTable has fsck (ops.go:13), opStart is AuthWrite (ops.go:62-65), gaps_test.go:360 proves POST, Wal.jsx already drives ops.run. So: NO new repair-check endpoint — close that option in Decisions, reuse wins.
  • Frontend claims: Repo.jsx:496 empty pill, data.js:77 tolerateMissing (404->missing, status-based via SDK errors.js notFound), Tree.jsx:105 selArgs-null suppression pattern, useResolved always fetches today (data.js:326 — no empty short-circuit). All confirmed.
  • pushPipeline shared by HTTP (smart.go:439) and SSH (bind_ssh.go:181): degradation guarantee ("pushes of new refs keep working, fsck never consulted") is structurally sound.

BLOCKING:

  • [B1] Overview fsck-projection cost is miscounted. The plan claims overview "already loads fsck.pb for the snapshot (maintain.go:285-297); surfacing is a read-only projection, +0 requests". Wrong path: that load is maintain.Snapshot (maintainer loop), NOT the serving path. bind_wal.go Overview (425-463) builds from the manifest only and never touches fsck.pb. Serving missing_total/repaired_seq/last-audit costs +1 conditional GET probe per overview call. Acceptable (no-store admin page, off hot path, law 6 safe) — but §5 must state +1, not +0.
  • [B2] repair_stalled durability needs a real schema decision. fsck.pb is protobuf under law 5 (append-only, golden fixtures). "Rides inside fsck.pb as an additive optional field, no spec change" still requires a new field NUMBER, proto edit, fixture regen + round-trip test, and a Decisions entry (+02_storage_protobuf.md note). Alternatively derive stalled-ness read-side (report age > X with upstream set and RepairedSeq==0) with ZERO new state. Pick one; don't hand-wave it.
  • [B3] Give-up rule (N=5, maintenance.repair_retries) contradicts tested behavior: wave4b_test.go:512 pins "publish failure keeps repaired_seq at 0 so the next pass retries", and the plan.go predicate has no counter. A stop-after-N changes the predicate and needs: counter durability (in-memory resets on restart = restart-dependent behavior; durable = another fsck.pb field -> B2), config spec (11_config_cli.md + setup schema + env overlay — missing from the §6 doc list), and test updates. RECOMMENDATION for v1: cut the give-up, keep retry-forever, surface derived-stalled read-side per B2-alt. Detection+guidance (the ticket's actual gap) works either way.

SHOULD-FIX:

  • [S1] Two health vocabularies will confuse: new summary health: empty|healthy|degraded vs existing overview Health.status: ok|degraded|error (api/env.go:206-212, sdk types.js:33). Scope them explicitly in 07_api.md (summary-health = repo state; overview-Health = dashboard). Also fix the ETag claim: empty repos get etag "" today (summary.go:46-49), not '"empty"' — define (don't misdescribe) ETag with the new field.
  • [S2] Pin the 404 empty repository: prefix precisely: exact predicate (HeadSeq==0 && refs==0 at handler time), enumerate touched error sites (bind_wal.go:181,187,230,317,336,349,370,388,392), audit per-handler cost (most hold the snapshot; name any that gains a GET), and assert damage-404s (refs>0, objects missing) NEVER carry the prefix or the UI misclassifies degraded as empty.
  • [S3] Frontend cache read needs a mechanism: data.js's cache Map is module-private, so useResolved cannot "read the repo:{full} cache entry" without a new exported peek (or via useRepo context). Name it. Spec Commits.jsx n=1 suppression symmetrically with Tree.
  • [S4] Law 3: add an explicit ### Concurrency subsection (hazard + avoidance). Material exists (§7 risks: TryAcquire/503+Retry-After per smart.go:134, never fsck on a request goroutine, (repo,kind) join) — promote it to normative form.
  • [S5] Doc list: add 11_config_cli.md (any surviving config key), 02_storage_protobuf.md (any proto field); fix the "§9.4" cite (ops join semantics = 07_api.md §12.2 / ops.go).

NITS:

  • §4 forensics cites "GET …/api/wal" — no such route in routes.go; name the real endpoint (overview / push-history / tasks).
  • "§3 table + PUT row" style cites are fine; keep the EVIDENCE "no change" budget entries (sim must assert empty-Code-tab issues 0 manifest-gated reads beyond summary).

CROSS-PLAN (joint with #210): #209 health-from-refs composes cleanly with stale-marker-on-real IFF the joint summary shape is defined — see #210 review B1 (blocking there, tracked here as dependency). Order: #209 health first, #210 adds the placeholder projection onto it. Do not let #209's classifier branch on any marker it doesn't own yet.

REVIEW — Ticket 1 self-heal plan (against CODE, not docs) Verdict: PROCEED-WITH-FIXES (no redesign needed; 3 blocking scoping/correctness fixes, all cuttable). VERIFIED CORRECT (checked in tree): - "unborn HEAD" / "HEAD" 404 sites: internal/api/bind_wal.go:181,187. Confirmed. - Empty-manifest post-create state: internal/wal/registry.go:224-263 (HeadSeq 0, Revision 1, no refs). Confirmed. - Auto-create interception points: internal/server/smart.go:143-157 (info/refs advert) and :419-428 (receive-pack Repo(create)). Confirmed, both require_write-gated upstream. - PUT 201/409 + ?object_format 400: internal/api/summary.go:53-85; pinned by gaps5_test.go:381-388. Confirmed. - Repair predicate + batch + bound: plan.go:265-268 (RepairedSeq==0 && (Missing||Total) && Upstream.Git!=""), units.go:62 (500), units.go:65 (100k). Confirmed. - fsck argv: maintain/fsck.go:60 `fsck --connectivity-only --no-dangling`. Confirmed. - POST ops/fsck already works: opsTable has `fsck` (ops.go:13), opStart is AuthWrite (ops.go:62-65), gaps_test.go:360 proves POST, Wal.jsx already drives ops.run. So: NO new repair-check endpoint — close that option in Decisions, reuse wins. - Frontend claims: Repo.jsx:496 empty pill, data.js:77 tolerateMissing (404->missing, status-based via SDK errors.js notFound), Tree.jsx:105 selArgs-null suppression pattern, useResolved always fetches today (data.js:326 — no empty short-circuit). All confirmed. - pushPipeline shared by HTTP (smart.go:439) and SSH (bind_ssh.go:181): degradation guarantee ("pushes of new refs keep working, fsck never consulted") is structurally sound. BLOCKING: - [B1] Overview fsck-projection cost is miscounted. The plan claims overview "already loads fsck.pb for the snapshot (maintain.go:285-297); surfacing is a read-only projection, +0 requests". Wrong path: that load is maintain.Snapshot (maintainer loop), NOT the serving path. bind_wal.go Overview (425-463) builds from the manifest only and never touches fsck.pb. Serving missing_total/repaired_seq/last-audit costs +1 conditional GET probe per overview call. Acceptable (no-store admin page, off hot path, law 6 safe) — but §5 must state +1, not +0. - [B2] `repair_stalled` durability needs a real schema decision. fsck.pb is protobuf under law 5 (append-only, golden fixtures). "Rides inside fsck.pb as an additive optional field, no spec change" still requires a new field NUMBER, proto edit, fixture regen + round-trip test, and a Decisions entry (+02_storage_protobuf.md note). Alternatively derive stalled-ness read-side (report age > X with upstream set and RepairedSeq==0) with ZERO new state. Pick one; don't hand-wave it. - [B3] Give-up rule (N=5, `maintenance.repair_retries`) contradicts tested behavior: wave4b_test.go:512 pins "publish failure keeps repaired_seq at 0 so the next pass retries", and the plan.go predicate has no counter. A stop-after-N changes the predicate and needs: counter durability (in-memory resets on restart = restart-dependent behavior; durable = another fsck.pb field -> B2), config spec (11_config_cli.md + setup schema + env overlay — missing from the §6 doc list), and test updates. RECOMMENDATION for v1: cut the give-up, keep retry-forever, surface derived-stalled read-side per B2-alt. Detection+guidance (the ticket's actual gap) works either way. SHOULD-FIX: - [S1] Two health vocabularies will confuse: new summary `health: empty|healthy|degraded` vs existing overview `Health.status: ok|degraded|error` (api/env.go:206-212, sdk types.js:33). Scope them explicitly in 07_api.md (summary-health = repo state; overview-Health = dashboard). Also fix the ETag claim: empty repos get etag "" today (summary.go:46-49), not '"empty"' — define (don't misdescribe) ETag with the new field. - [S2] Pin the 404 `empty repository:` prefix precisely: exact predicate (HeadSeq==0 && refs==0 at handler time), enumerate touched error sites (bind_wal.go:181,187,230,317,336,349,370,388,392), audit per-handler cost (most hold the snapshot; name any that gains a GET), and assert damage-404s (refs>0, objects missing) NEVER carry the prefix or the UI misclassifies degraded as empty. - [S3] Frontend cache read needs a mechanism: data.js's cache Map is module-private, so useResolved cannot "read the repo:{full} cache entry" without a new exported peek (or via useRepo context). Name it. Spec Commits.jsx n=1 suppression symmetrically with Tree. - [S4] Law 3: add an explicit `### Concurrency` subsection (hazard + avoidance). Material exists (§7 risks: TryAcquire/503+Retry-After per smart.go:134, never fsck on a request goroutine, (repo,kind) join) — promote it to normative form. - [S5] Doc list: add 11_config_cli.md (any surviving config key), 02_storage_protobuf.md (any proto field); fix the "§9.4" cite (ops join semantics = 07_api.md §12.2 / ops.go). NITS: - §4 forensics cites "GET …/api/wal" — no such route in routes.go; name the real endpoint (overview / push-history / tasks). - "§3 table + PUT row" style cites are fine; keep the EVIDENCE "no change" budget entries (sim must assert empty-Code-tab issues 0 manifest-gated reads beyond summary). CROSS-PLAN (joint with #210): #209 health-from-refs composes cleanly with stale-marker-on-real IFF the joint summary shape is defined — see #210 review B1 (blocking there, tracked here as dependency). Order: #209 health first, #210 adds the placeholder projection onto it. Do not let #209's classifier branch on any marker it doesn't own yet.
Author
Owner

Plan revision R1 (review findings — R1 wins on conflict)

Blocking resolutions (normative)

  • B1 — overview fsck projection costs +1 conditional GET per overview call (no-store admin page, off hot path). §5 corrected (not +0).
  • B2 — repair_stalled: derive read-side (report age > X + upstream set + RepairedSeq==0) — zero new protobuf state, no fixture regen.
  • B3 — give-up rule CUT for v1. No stop-after-N, no counter, no config key. Detection + guidance works without it; wave4b_test.go:512 retry-forever stays.

Joint with #210

  • Summary health field + placeholder projection defined in the #210 R1 (single shared shape); #209 implements health (empty|healthy|degraded) first.

Should-fix adoptions

Single ok|degraded|error overview vocabulary vs summary empty|healthy|degraded scoped in 07_api.md; ETag empty→"" as today; 404-marker predicate pinned + all 9 bind_wal.go sites costed; frontend cache peek exported; ### Concurrency subsection added; POST …/ops/fsck confirmed already addressable (no new endpoint); repair-check option closed.

# Plan revision R1 (review findings — R1 wins on conflict) ## Blocking resolutions (normative) - **B1 — overview fsck projection costs +1 conditional GET** per overview call (no-store admin page, off hot path). §5 corrected (not +0). - **B2 — `repair_stalled`: derive read-side** (report age > X + upstream set + `RepairedSeq==0`) — zero new protobuf state, no fixture regen. - **B3 — give-up rule CUT for v1.** No stop-after-N, no counter, no config key. Detection + guidance works without it; `wave4b_test.go:512` retry-forever stays. ## Joint with #210 - Summary `health` field + placeholder projection defined in the #210 R1 (single shared shape); #209 implements `health` (`empty|healthy|degraded`) first. ## Should-fix adoptions Single `ok|degraded|error` overview vocabulary vs summary `empty|healthy|degraded` scoped in `07_api.md`; ETag empty→`""` as today; 404-marker predicate pinned + all 9 `bind_wal.go` sites costed; frontend cache peek exported; `### Concurrency` subsection added; `POST …/ops/fsck` confirmed already addressable (no new endpoint); `repair-check` option closed.
Author
Owner

waves 1-3 implemented in PR #216 (branch feat/issue-209): #216 — health on summary, empty-repo 404 marker at all 9 sites, overview fsck projection, EmptyRepoGuide + suppression + degraded banner/WAL section, docs + EVIDENCE E13. R1 resolutions honored (B1 +1 stated, B2 derived stall, B3 give-up cut). Do NOT merge yet — awaiting review.

waves 1-3 implemented in PR #216 (branch feat/issue-209): https://git.packden.us/crueber/walhub/pulls/216 — health on summary, empty-repo 404 marker at all 9 sites, overview fsck projection, EmptyRepoGuide + suppression + degraded banner/WAL section, docs + EVIDENCE E13. R1 resolutions honored (B1 +1 stated, B2 derived stall, B3 give-up cut). Do NOT merge yet — awaiting review.
Author
Owner

Review: PR #216 (feat/issue-209, commit 31d00b3) — self-heal

Spec checked: plan (comment 1801) + code-review (1808, PROCEED-WITH-FIXES) + R1 (1812, normative: B1 +1 stated, B2 derived stall, B3 give-up cut, joint shape to #210). Verified in scratch worktree at the exact PR commit; main worktree untouched (still clean).

R1 rulings — all honored

  • B1: overview fsck projection costs +1 GET — stated in 07_api §12.1, 10_maintenance §9.3/Decisions, EVIDENCE E13, and pinned by TestSummaryOverviewRoundTrips (overview ± report = exactly 1 probe). Correct scope: no-store admin page, off law-6 paths.
  • B2: repair_stalled derived read-side (fsckProjection: upstream set + RepairedSeq==0 + missing signal + report older than one fsck_interval; nil timestamp never stalls). Zero protobuf change — internal/store/proto untouched, law 5 clean.
  • B3: give-up CUT — no counter, no config key, no predicate change (plan.go:265-268 untouched). Retry-forever stands.
  • Joint (#210): no placeholder/stale-marker fields anywhere in the diff; classifier branches only on refs + fsck.pb. #209-first ordering respected.
  • Should-fix: vocabularies scoped in 07 §9.1 (repo-state vs dashboard); ETag ""-when-unborn as-today + ~degraded suffix; 9/9 bind_wal.go sites enumerated in 07 §9.10 with per-handler cost; peekCached exported (S3); ### Concurrency in health.go + 10_maintenance §9.3; POST …/ops/fsck reuse confirmed (no repair-check), closed in Decisions.

Review checklist (per-item verdicts)

  • Health trips: empty folds from in-hand snapshot + in-memory ManifestSnapshot (+0); the fsck.pb probe is behind if health != RepoHealthEmpty (summary.go:40) — branch verified. Non-empty summary +1, openly stated in E13.
  • ETag: ~degraded flip busts SWR (pinned by TestSummaryDegraded304: stale sha → 200, suffixed → 304); unborn "" unchanged (wire test pins absent ETag).
  • 404 marker: 9/9 sites (Resolve×3, Tree×1, Blob×2, Commits×1, Commit×2), exact predicate (Head==nil && Branches==0 && Tags==0 && (nil || HeadSeq==0)), fail-closed on any manifest/snapshot error. Damage-404s keep identical format strings (verified in diff) + TestEmptyMarkerSites pins frozen prefixes and asserts the marker never appears on damage.
  • No new endpoints: zero route changes; WAL re-audit drives existing ops.run("fsck"). No new bucket families, no LIST (grep clean on new code), git argv untouched.
  • useResolved: Step-0 suppression only on known-empty summary; resolve-catch maps only marker 404s; sha-step maps only 404 + known-degraded — real errors still reach the tray. tolerateDegraded requires known-degraded, no global mute (#200 shell intact).
  • Banner/WAL card read existing data (summary signal / overview projection); the only POST is the sanctioned existing fsck op.
  • E13 honest: counting decorator over the memory store through the production probeFsck → GetBytes path; verdict explicitly scoped ("budgets never call here"); empty +0 / non-empty +1 / overview +1 all asserted.
  • Deps: go.mod/go.sum/package.json/pnpm-lock diff empty. Docs + Decisions appended (law 12).

Verification (scratch worktree, no browser per instructions)

  • gofmt -l clean, go vet ./internal/api/... clean.
  • go test -race -count=1 ./internal/api/... → ok; coverage 95.3% (≥95% gate holds; only internal/api touched on the Go side).
  • node --test web/test/unit/*.test.js → 410/410 pass (incl. new empty-degraded.test.js 10/10). Note: scratch worktree had no web/node_modules; ran via a temporary symlink to main's install, removed afterward. Suite takes ~4 min.
  • Sim/e2e not run (no internal/wal or publish/sync changes; E13 covers the cost claim). No browser drive (noted as instructed).

Nits (non-blocking, left for author — no push)

  1. "Conditional GET" wording (overview.go, health.go, health209_test.go, 07 §9.1/§12.1, 10 §9.3): the probe passes GetOptions{} (no IfNoneMatch), i.e. a plain exact-key GET. Costed correctly as +1; R1 used the same term, so this is terminology only.
  2. SummaryData.MissingTotal (env.go:394-396) is never populated in non-test code (wire body carries its own). Dead internal field — drop it, or keep if #210's placeholder projection intends to ride the view struct.
  3. fsckHasMissing includes Problems>0 (fsck stderr non-missing lines), so a problems-only old report + upstream derives repair_stalled even though the repair predicate (plan.go:266) only fires on Missing/MissingTotal. Self-consistent with summary health (same helper); flag if the admin copy should distinguish corruption from missing.
  4. Cold-race note: on first load the summary may not be cached when useResolved runs, so one resolve may still fire and map silently via the marker fallback (zero toasts either way — covered by test "empty-prefix fallback…summary not loaded"). "Zero doomed requests" holds once the shell entry exists.

No code pushed (nothing found that warrants a branch push; nits are author judgment + possible #210 interaction).

MERGE RECOMMENDATION: ready to merge

## Review: PR #216 (feat/issue-209, commit 31d00b3) — self-heal Spec checked: plan (comment 1801) + code-review (1808, PROCEED-WITH-FIXES) + R1 (1812, normative: B1 +1 stated, B2 derived stall, B3 give-up cut, joint shape to #210). Verified in scratch worktree at the exact PR commit; main worktree untouched (still clean). ### R1 rulings — all honored - **B1**: overview fsck projection costs +1 GET — stated in 07_api §12.1, 10_maintenance §9.3/Decisions, EVIDENCE E13, and pinned by `TestSummaryOverviewRoundTrips` (overview ± report = exactly 1 probe). Correct scope: no-store admin page, off law-6 paths. - **B2**: `repair_stalled` derived read-side (`fsckProjection`: upstream set + `RepairedSeq==0` + missing signal + report older than one `fsck_interval`; nil timestamp never stalls). Zero protobuf change — `internal/store/proto` untouched, law 5 clean. - **B3**: give-up CUT — no counter, no config key, no predicate change (`plan.go:265-268` untouched). Retry-forever stands. - **Joint (#210)**: no placeholder/stale-marker fields anywhere in the diff; classifier branches only on refs + `fsck.pb`. #209-first ordering respected. - **Should-fix**: vocabularies scoped in 07 §9.1 (repo-state vs dashboard); ETag `""`-when-unborn as-today + `~degraded` suffix; 9/9 `bind_wal.go` sites enumerated in 07 §9.10 with per-handler cost; `peekCached` exported (S3); `### Concurrency` in `health.go` + 10_maintenance §9.3; `POST …/ops/fsck` reuse confirmed (no `repair-check`), closed in Decisions. ### Review checklist (per-item verdicts) - Health trips: empty folds from in-hand snapshot + in-memory `ManifestSnapshot` (+0); the `fsck.pb` probe is behind `if health != RepoHealthEmpty` (`summary.go:40`) — branch verified. Non-empty summary +1, openly stated in E13. - ETag: `~degraded` flip busts SWR (pinned by `TestSummaryDegraded304`: stale sha → 200, suffixed → 304); unborn `""` unchanged (wire test pins absent ETag). - 404 marker: 9/9 sites (Resolve×3, Tree×1, Blob×2, Commits×1, Commit×2), exact predicate (`Head==nil && Branches==0 && Tags==0 && (nil || HeadSeq==0)`), fail-closed on any manifest/snapshot error. Damage-404s keep identical format strings (verified in diff) + `TestEmptyMarkerSites` pins frozen prefixes and asserts the marker never appears on damage. - No new endpoints: zero route changes; WAL re-audit drives existing `ops.run("fsck")`. No new bucket families, no LIST (grep clean on new code), git argv untouched. - `useResolved`: Step-0 suppression only on known-empty summary; resolve-catch maps only marker 404s; sha-step maps only 404 + known-degraded — real errors still reach the tray. `tolerateDegraded` requires known-degraded, no global mute (#200 shell intact). - Banner/WAL card read existing data (summary signal / overview projection); the only POST is the sanctioned existing fsck op. - E13 honest: counting decorator over the memory store through the production `probeFsck → GetBytes` path; verdict explicitly scoped ("budgets never call here"); empty +0 / non-empty +1 / overview +1 all asserted. - Deps: `go.mod/go.sum/package.json/pnpm-lock` diff empty. Docs + Decisions appended (law 12). ### Verification (scratch worktree, no browser per instructions) - `gofmt -l` clean, `go vet ./internal/api/...` clean. - `go test -race -count=1 ./internal/api/...` → ok; coverage **95.3%** (≥95% gate holds; only `internal/api` touched on the Go side). - `node --test web/test/unit/*.test.js` → **410/410 pass** (incl. new `empty-degraded.test.js` 10/10). Note: scratch worktree had no `web/node_modules`; ran via a temporary symlink to main's install, removed afterward. Suite takes ~4 min. - Sim/e2e not run (no `internal/wal` or publish/sync changes; E13 covers the cost claim). No browser drive (noted as instructed). ### Nits (non-blocking, left for author — no push) 1. "Conditional GET" wording (`overview.go`, `health.go`, `health209_test.go`, 07 §9.1/§12.1, 10 §9.3): the probe passes `GetOptions{}` (no `IfNoneMatch`), i.e. a plain exact-key GET. Costed correctly as +1; R1 used the same term, so this is terminology only. 2. `SummaryData.MissingTotal` (`env.go:394-396`) is never populated in non-test code (wire body carries its own). Dead internal field — drop it, or keep if #210's placeholder projection intends to ride the view struct. 3. `fsckHasMissing` includes `Problems>0` (fsck stderr non-missing lines), so a problems-only old report + upstream derives `repair_stalled` even though the repair predicate (`plan.go:266`) only fires on Missing/MissingTotal. Self-consistent with summary `health` (same helper); flag if the admin copy should distinguish corruption from missing. 4. Cold-race note: on first load the summary may not be cached when `useResolved` runs, so one resolve may still fire and map silently via the marker fallback (zero toasts either way — covered by test "empty-prefix fallback…summary not loaded"). "Zero doomed requests" holds once the shell entry exists. No code pushed (nothing found that warrants a branch push; nits are author judgment + possible #210 interaction). ### MERGE RECOMMENDATION: ready to merge
Author
Owner

Implemented in PR #216 (review: all R1 rulings verified, probes costed, masking impossible; 95.3% + 410/410), merged. Closing.

Implemented in PR #216 (review: all R1 rulings verified, probes costed, masking impossible; 95.3% + 410/410), merged. Closing.
crueber added this to the v1 milestone 2026-09-10 22:27:13 +00:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
crueber/walhub#209
No description provided.