Mirror anon/walhub: object-level requests hang to 504 (refs fine) — fix the materialization wedge and make mirror health self-healing #320

Closed
opened 2026-09-11 11:42:49 +00:00 by crueber · 3 comments
Owner

What's wrong

The mirrored repo anon/walhub is broken for content browsing: every object-level request hangs until the proxy gives up (504 Gateway Time-out), while ref-level requests answer instantly. The repo is unusable as a mirror — the whole point of mirroring walhub is to browse its code, and the Code tab can't load a single file.

Reproduced live (2026-09-11, against hub.packden.us)

Request Result
GET /anon/walhub/api (summary) 200 — head: main @ 15a5960…, branches: 37, health: "healthy", mirror.last_result: "ok", due: true
GET /anon/walhub/api/refs 200
GET /anon/walhub/api/refs/branches 200
GET /anon/walhub/api/commits?n=1 hang → client timeout / 504 (tried repeatedly, > 90 s)
GET /anon/walhub/api/blob/main/README.md hang → 504 Gateway Time-out (openresty)
GET /anon/walhub/api/tree/main hang

Note the contradictions the UI surfaces from this: the summary says health: "healthy", mirror.last_result: "ok", branches: 37 — yet nothing object-level works. The repo looks fine in every listing and is dead everywhere it matters.

Root cause analysis (code evidence)

  • Ref-level routes (summary/refs) read the manifest snapshot only — no objects needed. Object-level routes call localView(ctx, id, SyncServe) (internal/api/bind_wal.go:377/403/437/465/672 — tree, blob, commits, commit) which triggers engine materialization of the serving copy (engine.Sync, syncLevelToWal). That materialization is what hangs for this repo.
  • The mirror sync (internal/mirror/sync.go:245-247) does a full clone --mirror → ingestPacks → PublishRefs, so the bucket has packs; the wedge is in the on-demand serve-level materialization path (engine.Sync → locks → fetch/materialize against the cache dir) for a repo with 37 refs and an upstream-sized pack set.
  • The hourly schedule (due: true with last_synced_at 06:00Z) means syncs keep succeeding and keep landing refs — each sync potentially invalidates/extents the serving copy again, so the wedge isn't a one-time bad state; it plausibly re-arms itself.
  • Whatever the precise internal wedge (lock contention with the mirror sync's own engine writes, a pack materialization loop, cache-dir state, or an unbounded retry), the observable contract failure is: an unhealthy repo reports health: "healthy" and mirror.last_result: "ok" and never surfaces the failure anywhere, and there is no automatic recovery.

What's needed

  1. Fix the wedge so object-level requests on this mirror work (or fail fast with an honest error, not a proxy 504). Investigate: engine Sync/serve materialization for mirror-published repos — lock interaction between the hourly sync task and concurrent SyncServe requests; per-request timeouts on materialization; whether the mirror's ingested pack set is being re-materialized from scratch on each request (cache-dir invalidation churn).
  2. Self-healing (the standing requirement):
    • The mirror sync loop should verify serve-readiness as part of its success criteria — last_result: "ok" must mean "refs synced AND objects servable" (e.g. a post-sync probe: one blob fetch at head). A sync that publishes refs the instance can't serve is not a successful sync.
    • When object-level requests fail/timeout, mark the repo degraded (the mechanism exists — DegradedNotice, health projection #209) instead of reporting healthy, and surface the state in the mirror projection (last_result should say the truth).
    • A bounded self-heal: on degraded detection, the maintainer/mirror loop re-materializes the serving copy (tear down cache-dir state for the repo, re-materialize, re-probe) with retry backoff — the #209 self-heal machinery is the precedent.
  3. Guardrails so this can't silently recur: the 504-through-proxy shape means requests have no server-side deadline — materialization paths need bounded timeouts that answer 503/degraded instead of hanging. And the summary's health field must reflect object-level servability, not just ref/manifest health.

Acceptance criteria

  • anon/walhub (or a comparable mirror) serves blob/tree/commits at head without proxy timeouts.
  • A mirror sync's last_result only reports ok when a post-sync servability probe passes; a wedged/un-servable state reports honestly (last_result/health degraded).
  • Object-level materialization has a bounded server-side timeout; failure answers 503 + degraded state, never an indefinite hang (no proxy 504s).
  • Self-heal: a degraded mirror is re-materialized automatically (bounded retries/backoff) and flips back to healthy when servable — no operator action required.
  • The wedge's root cause is identified and stated in the PR (lock interaction vs cache-dir state vs unbounded work), with a regression test for the specific failure mode.
  • Health/summary/mirror-projection fields agree with reality in both healthy and degraded states (listing pages must not advertise a repo that can't serve).
## What's wrong The mirrored repo [`anon/walhub`](https://hub.packden.us/anon/walhub) is broken for content browsing: **every object-level request hangs until the proxy gives up (504 Gateway Time-out)**, while ref-level requests answer instantly. The repo is unusable as a mirror — the whole point of mirroring walhub is to browse its code, and the Code tab can't load a single file. ## Reproduced live (2026-09-11, against hub.packden.us) | Request | Result | |---|---| | `GET /anon/walhub/api` (summary) | **200** — `head: main @ 15a5960…`, `branches: 37`, `health: "healthy"`, `mirror.last_result: "ok"`, `due: true` | | `GET /anon/walhub/api/refs` | **200** | | `GET /anon/walhub/api/refs/branches` | **200** | | `GET /anon/walhub/api/commits?n=1` | **hang → client timeout / 504** (tried repeatedly, > 90 s) | | `GET /anon/walhub/api/blob/main/README.md` | **hang → 504 Gateway Time-out (openresty)** | | `GET /anon/walhub/api/tree/main` | **hang** | Note the contradictions the UI surfaces from this: the summary says `health: "healthy"`, `mirror.last_result: "ok"`, `branches: 37` — yet nothing object-level works. The repo *looks* fine in every listing and is dead everywhere it matters. ## Root cause analysis (code evidence) - Ref-level routes (summary/refs) read the manifest snapshot only — no objects needed. Object-level routes call `localView(ctx, id, SyncServe)` (`internal/api/bind_wal.go:377/403/437/465/672` — tree, blob, commits, commit) which triggers engine materialization of the serving copy (`engine.Sync`, syncLevelToWal). **That materialization is what hangs** for this repo. - The mirror sync (`internal/mirror/sync.go:245-247`) does a full `clone --mirror` → `ingestPacks` → `PublishRefs`, so the *bucket* has packs; the wedge is in the on-demand serve-level materialization path (`engine.Sync` → locks → fetch/materialize against the cache dir) for a repo with 37 refs and an upstream-sized pack set. - The hourly schedule (`due: true` with `last_synced_at 06:00Z`) means syncs keep succeeding and keep landing refs — each sync potentially invalidates/extents the serving copy again, so the wedge isn't a one-time bad state; it plausibly re-arms itself. - Whatever the precise internal wedge (lock contention with the mirror sync's own engine writes, a pack materialization loop, cache-dir state, or an unbounded retry), the observable contract failure is: **an unhealthy repo reports `health: "healthy"` and `mirror.last_result: "ok"` and never surfaces the failure anywhere**, and there is no automatic recovery. ## What's needed 1. **Fix the wedge** so object-level requests on this mirror work (or fail fast with an honest error, not a proxy 504). Investigate: engine Sync/serve materialization for mirror-published repos — lock interaction between the hourly sync task and concurrent `SyncServe` requests; per-request timeouts on materialization; whether the mirror's ingested pack set is being re-materialized from scratch on each request (cache-dir invalidation churn). 2. **Self-healing (the standing requirement):** - The mirror sync loop should verify serve-readiness as part of its success criteria — `last_result: "ok"` must mean "refs synced AND objects servable" (e.g. a post-sync probe: one `blob` fetch at head). A sync that publishes refs the instance can't serve is not a successful sync. - When object-level requests fail/timeout, mark the repo `degraded` (the mechanism exists — `DegradedNotice`, health projection #209) instead of reporting `healthy`, and surface the state in the mirror projection (`last_result` should say the truth). - A bounded self-heal: on degraded detection, the maintainer/mirror loop re-materializes the serving copy (tear down cache-dir state for the repo, re-materialize, re-probe) with retry backoff — the #209 self-heal machinery is the precedent. 3. **Guardrails so this can't silently recur:** the 504-through-proxy shape means requests have no server-side deadline — materialization paths need bounded timeouts that answer 503/degraded instead of hanging. And the summary's `health` field must reflect object-level servability, not just ref/manifest health. ## Acceptance criteria - [ ] `anon/walhub` (or a comparable mirror) serves blob/tree/commits at head without proxy timeouts. - [ ] A mirror sync's `last_result` only reports `ok` when a post-sync servability probe passes; a wedged/un-servable state reports honestly (`last_result`/health degraded). - [ ] Object-level materialization has a bounded server-side timeout; failure answers 503 + degraded state, never an indefinite hang (no proxy 504s). - [ ] Self-heal: a degraded mirror is re-materialized automatically (bounded retries/backoff) and flips back to healthy when servable — no operator action required. - [ ] The wedge's root cause is identified and stated in the PR (lock interaction vs cache-dir state vs unbounded work), with a regression test for the specific failure mode. - [ ] Health/summary/mirror-projection fields agree with reality in both healthy and degraded states (listing pages must not advertise a repo that can't serve).
crueber added this to the v1 milestone 2026-09-11 11:42:49 +00:00
Author
Owner

Fix is up: PR #329 (branch fix/issue-320) — please review, do not merge yet (per task instructions the merge is yours).

Root cause: unbounded work, not cache-dir churn. Sync(LevelServe) serializes on packMu plus the (repo,materialize) single-flight whose body runs on the registry-lifetime ctx with no deadline at any layer, so one stalled materialize wedges every later object-level request while refs (no packMu) stay instant — exactly the observed shape. The hourly sync kept reporting ok/healthy because success never probed servability.

What the PR does: bounds the serve wait (45s to 503 plus degraded) and the materialize body (10m cap); gates last_result ok on a post-publish Sync plus cat-file -e probe (no-op syncs exempt); records the sticky serve-health sidecar and surfaces it in health plus mirror.degraded_reason (ETag-covered); self-heals degraded mirrors via the mirror-heal task with backoff (no teardown, mirror doc untouched).

Verification: race-clean wal/api/mirror, cover gate at or above 95 on all touched packages, contract plus e2e plus full node suite green; hung-store regression tests pin the wedge (bounded timeout, refs-instant divergence, resume plus clear). Two environment notes: TestUIAssetConcepts fails on stale web/dist here (pre-existing, no pnpm to rebuild), and there is no sim tier in this tree (no TestSim scenarios or make target — round-trip harness plus contract plus e2e run instead). Live instance untouched; reproduced locally only.

Fix is up: PR #329 (branch fix/issue-320) — please review, do not merge yet (per task instructions the merge is yours). Root cause: unbounded work, not cache-dir churn. Sync(LevelServe) serializes on packMu plus the (repo,materialize) single-flight whose body runs on the registry-lifetime ctx with no deadline at any layer, so one stalled materialize wedges every later object-level request while refs (no packMu) stay instant — exactly the observed shape. The hourly sync kept reporting ok/healthy because success never probed servability. What the PR does: bounds the serve wait (45s to 503 plus degraded) and the materialize body (10m cap); gates last_result ok on a post-publish Sync plus cat-file -e probe (no-op syncs exempt); records the sticky serve-health sidecar and surfaces it in health plus mirror.degraded_reason (ETag-covered); self-heals degraded mirrors via the mirror-heal task with backoff (no teardown, mirror doc untouched). Verification: race-clean wal/api/mirror, cover gate at or above 95 on all touched packages, contract plus e2e plus full node suite green; hung-store regression tests pin the wedge (bounded timeout, refs-instant divergence, resume plus clear). Two environment notes: TestUIAssetConcepts fails on stale web/dist here (pre-existing, no pnpm to rebuild), and there is no sim tier in this tree (no TestSim scenarios or make target — round-trip harness plus contract plus e2e run instead). Live instance untouched; reproduced locally only.
Author
Owner

Review: PR #329 (fix/issue-320) — verified, ready to merge

Reviewed the full diff (28 files, +2430/-77) in a scratch worktree; all verification below is on commit 0da5fcd. No fixes pushed — nothing blocking found. Two non-blocking observations at the end.

Root cause — convincing

The unbounded-work claim checks out against the code: serve-level Sync joins the (repo,materialize) single-flight body running on the registry-lifetime task ctx with no deadline at any layer (server.request_timeout is indeed config-only, never enforced — confirmed), while refs-level syncs never touch packMu (internal/wal/handle.go). One stalled bulk fetch wedges every later object request via packMu queue + single-flight fan-in. Correctly NOT blamed on lock order (syncMu→packMu→rw already correct) or cache-dir churn (present-check resume sound; no teardown added — agreed, deletion would not address the mechanism).

Timeouts — sane, fail path honest

  • serve_sync_timeout 45s < typical proxy 60s: server answers 503 itself. Retry-After: 15 is centralized in writePlain (internal/api/env.go:670) so ALL object routes (tree/blob/commits/commit, pinned by TestServeTimeoutAnswers503) get it. WalErrTimeout falls through notFoundOr to the generic 503 — never 404, never hang.
  • serve_materialize_timeout 10m caps the detached body; hung store GET cannot outlive it; resume at file granularity + tmp sweep at body start. Non-positive config falls back to compiled defaults — no unbounded wait configurable by accident. Both values are config (not hardcoded) with setup UI round-trip test.
  • Client disconnect marks nothing (ctx.Err() branch) — correct.

Lock discipline — clean

  • Mark/clear store calls run AFTER packMu release on detached 10s-bounded ctx (handle.go Sync; servehealth.go) — no lock held across store calls. Lock order unchanged. serveDegraded flag gates success-clears to instances that actually failed (zero steady-state round trips). Overwrite-always marker: contention is repeated identical verdicts, no CAS needed — and correctly NOT on the manifest CAS path (sideband only, never the ACK point).

Probe argv — pinned (law 2)

git --git-dir=

cat-file -e added to 04_git.md §12 + argv list, Runner.ProbeObject honors GitTimeout. -e (existence, no output) is the cheapest proof. Verdict-not-retry on non-zero exit: correct.

Heal loop — bounded, can't fight sync

  • Backoff 15m×2^(n-1)/24h-capped via shared BackoffDelay ladder; rate-bounded, never count-capped. Never takes sync lease, never touches mirror doc ( ConsecutiveFailures/LastResult stay the sync's story), never tears down cache-dir. Due mirrors skip heal (their sync carries its own probe); non-due marked mirrors heal one minute out. No-op syncs never probe (prior verdict stands, stays 11 ops) — no cold-cache flap.
  • Sticky-until-reproven with no TTL (fail closed); cleared only on proven servability (demand success gated by flag, probe pass, heal pass). Unparseable marker reads degraded; transport-error reads unmarked (fail open, next serve re-proves) — reasonable, documented.

ETag / budgets / hot path

  • mirrorHash covers degraded_reason; summary ~degraded suffix; bust-cache test pins no-304-across-flip. Healthy non-mirror summary +2 exact-key GETs (was +1, same cost class, off law-6 paths — stated in-code and in 07_api.md). First sync 24 ops (was 22: probe GET + blind clear-DELETE). No LIST anywhere (exact-key probes; 404s free). Mirror hook fills degraded_reason with +0 extra summary GETs.

Verification (scratch worktree, removed afterward)

  • go test -race: wal, mirror, api, config, store(+fault,proto), cmd/walhub — all clean. Coverage: wal 95.4 / api 95.5 / mirror 96.8 / config 95.7 / store 95.0 — gate holds.
  • Contract suite (memory+filesystem): pass. gofmt/vet/build: clean.
  • server pkg: 11 failures, all ui-shell-missing/dist-absent — reproduced IDENTICALLY on unmodified base 42ebfa0 in a no-dist worktree (pre-existing/environmental, as the PR discloses). New TestSetupServeTimeoutKeysRoundTrip passes. TestVerifyTokenWireNegatives flaked once under full-suite load, passes -count=3 in isolation; touches no PR file (auth untouched).
  • e2e: only TestE2E_InvalidConfigSetupOnly, ui-shell cause, identical on base.
  • node --test: setup-form+setup-advanced 67/67 pass (covers the 2-line FIELDS change); 9 files fail on missing solid-js (no node_modules in scratch, no installs per scope) — identical 9 on base.
  • Main worktree left clean (only pre-existing untracked .opencode/).

Non-blocking observations (no action required before merge)

  1. packPhaseError checks pctx.Err()==DeadlineExceeded before ctx.Err(): a parent ctx with its own earlier deadline would misclassify as serve-timeout + mark degraded. Unreachable today (no caller sets a parent deadline on these paths — request_timeout unenforced, loop/task ctxs deadline-free). Consider a deadline-comparison guard only if parent deadlines ever appear.
  2. mirror.View.NextSyncAt gained omitempty (alongside additive degraded_reason). Pre-1.0 so allowed; flagging since it is a (tiny) wire-shape change beyond pure addition.

Acceptance criteria mapping

  1. Fail-fast half proven (503+degraded in 45s, refs instant); live anon/walhub serving to be confirmed post-deploy (no live contact per PR, correct call). 2. Probe-gated last_result: code + tests. 3. Bounded materialization + 503/degraded: code + tests. 4. Self-heal with backoff, no operator: code + healDue/probe/heal-round tests. 5. Root cause stated (PR + 05 §5.2/§5.3 + 11_mirror §3): unbounded work, with hung-store regression harness. 6. Agreement (summary/mirror/ETag): table + bust-cache tests.

MERGE RECOMMENDATION: ready to merge.

## Review: PR #329 (fix/issue-320) — verified, ready to merge Reviewed the full diff (28 files, +2430/-77) in a scratch worktree; all verification below is on commit 0da5fcd. No fixes pushed — nothing blocking found. Two non-blocking observations at the end. ### Root cause — convincing The unbounded-work claim checks out against the code: serve-level Sync joins the (repo,materialize) single-flight body running on the registry-lifetime task ctx with no deadline at any layer (server.request_timeout is indeed config-only, never enforced — confirmed), while refs-level syncs never touch packMu (internal/wal/handle.go). One stalled bulk fetch wedges every later object request via packMu queue + single-flight fan-in. Correctly NOT blamed on lock order (syncMu→packMu→rw already correct) or cache-dir churn (present-check resume sound; no teardown added — agreed, deletion would not address the mechanism). ### Timeouts — sane, fail path honest - serve_sync_timeout 45s < typical proxy 60s: server answers 503 itself. Retry-After: 15 is centralized in writePlain (internal/api/env.go:670) so ALL object routes (tree/blob/commits/commit, pinned by TestServeTimeoutAnswers503) get it. WalErrTimeout falls through notFoundOr to the generic 503 — never 404, never hang. - serve_materialize_timeout 10m caps the detached body; hung store GET cannot outlive it; resume at file granularity + tmp sweep at body start. Non-positive config falls back to compiled defaults — no unbounded wait configurable by accident. Both values are config (not hardcoded) with setup UI round-trip test. - Client disconnect marks nothing (ctx.Err() branch) — correct. ### Lock discipline — clean - Mark/clear store calls run AFTER packMu release on detached 10s-bounded ctx (handle.go Sync; servehealth.go) — no lock held across store calls. Lock order unchanged. serveDegraded flag gates success-clears to instances that actually failed (zero steady-state round trips). Overwrite-always marker: contention is repeated identical verdicts, no CAS needed — and correctly NOT on the manifest CAS path (sideband only, never the ACK point). ### Probe argv — pinned (law 2) git --git-dir=<dir> cat-file -e <oid> added to 04_git.md §12 + argv list, Runner.ProbeObject honors GitTimeout. -e (existence, no output) is the cheapest proof. Verdict-not-retry on non-zero exit: correct. ### Heal loop — bounded, can't fight sync - Backoff 15m×2^(n-1)/24h-capped via shared BackoffDelay ladder; rate-bounded, never count-capped. Never takes sync lease, never touches mirror doc ( ConsecutiveFailures/LastResult stay the sync's story), never tears down cache-dir. Due mirrors skip heal (their sync carries its own probe); non-due marked mirrors heal one minute out. No-op syncs never probe (prior verdict stands, stays 11 ops) — no cold-cache flap. - Sticky-until-reproven with no TTL (fail closed); cleared only on proven servability (demand success gated by flag, probe pass, heal pass). Unparseable marker reads degraded; transport-error reads unmarked (fail open, next serve re-proves) — reasonable, documented. ### ETag / budgets / hot path - mirrorHash covers degraded_reason; summary ~degraded suffix; bust-cache test pins no-304-across-flip. Healthy non-mirror summary +2 exact-key GETs (was +1, same cost class, off law-6 paths — stated in-code and in 07_api.md). First sync 24 ops (was 22: probe GET + blind clear-DELETE). No LIST anywhere (exact-key probes; 404s free). Mirror hook fills degraded_reason with +0 extra summary GETs. ### Verification (scratch worktree, removed afterward) - go test -race: wal, mirror, api, config, store(+fault,proto), cmd/walhub — all clean. Coverage: wal 95.4 / api 95.5 / mirror 96.8 / config 95.7 / store 95.0 — gate holds. - Contract suite (memory+filesystem): pass. gofmt/vet/build: clean. - server pkg: 11 failures, all ui-shell-missing/dist-absent — reproduced IDENTICALLY on unmodified base 42ebfa0 in a no-dist worktree (pre-existing/environmental, as the PR discloses). New TestSetupServeTimeoutKeysRoundTrip passes. TestVerifyTokenWireNegatives flaked once under full-suite load, passes -count=3 in isolation; touches no PR file (auth untouched). - e2e: only TestE2E_InvalidConfigSetupOnly, ui-shell cause, identical on base. - node --test: setup-form+setup-advanced 67/67 pass (covers the 2-line FIELDS change); 9 files fail on missing solid-js (no node_modules in scratch, no installs per scope) — identical 9 on base. - Main worktree left clean (only pre-existing untracked .opencode/). ### Non-blocking observations (no action required before merge) 1. packPhaseError checks pctx.Err()==DeadlineExceeded before ctx.Err(): a parent ctx with its own earlier deadline would misclassify as serve-timeout + mark degraded. Unreachable today (no caller sets a parent deadline on these paths — request_timeout unenforced, loop/task ctxs deadline-free). Consider a deadline-comparison guard only if parent deadlines ever appear. 2. mirror.View.NextSyncAt gained omitempty (alongside additive degraded_reason). Pre-1.0 so allowed; flagging since it is a (tiny) wire-shape change beyond pure addition. ### Acceptance criteria mapping 1. Fail-fast half proven (503+degraded in 45s, refs instant); live anon/walhub serving to be confirmed post-deploy (no live contact per PR, correct call). 2. Probe-gated last_result: code + tests. 3. Bounded materialization + 503/degraded: code + tests. 4. Self-heal with backoff, no operator: code + healDue/probe/heal-round tests. 5. Root cause stated (PR + 05 §5.2/§5.3 + 11_mirror §3): unbounded work, with hung-store regression harness. 6. Agreement (summary/mirror/ETag): table + bust-cache tests. **MERGE RECOMMENDATION: ready to merge.**
Author
Owner

Fixed by PR #329 (review clean — root cause, timeouts, lock discipline, probe, heal, ETag all verified), merged. Closing.

Fixed by PR #329 (review clean — root cause, timeouts, lock discipline, probe, heal, ETag all verified), merged. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
crueber/walhub#320
No description provided.