Fix #320: mirror wedge + self-healing #329

Merged
crueber merged 1 commit from fix/issue-320 into main 2026-09-11 15:22:44 +00:00
Owner

Root cause (stated per acceptance criterion 5)

UNBOUNDED WORK — not cache-dir churn, not a lock-order violation. RepoHandle.Sync(LevelServe) holds packMu across the (repo,materialize) single-flight join/leader-wait, and that task body runs on the registry-lifetime ctx with no deadline at any layer: no handler/request timeout (server.request_timeout is config-only, never enforced anywhere), no packMu/join bound, no store-GET timeout. So one stalled materialize (slow bulk fetch of the upstream-sized pack set after new mirror packs land, or a hung store GET) wedges every later object-level request — each either queues on packMu or joins the same orphaned task with a deadline-free request ctx — indefinitely, while LevelRefs (summary/refs: manifest only, never touch packMu) stays instant. The hourly mirror sync (clone to ingest to PublishRefs) never consults serve-readiness, so last_result ok plus health healthy keep reporting. Lock order was already correct (syncMu to packMu to rw, TryLock-only writers); the present-check resume is sound, so no cache teardown was added.

What was built

  1. Bounded serve materialization (05 section 5.2): server.serve_sync_timeout (45s, under proxy 60s) bounds the pack-phase wait to WalErrTimeout and HTTP 503 + Retry-After 15; server.serve_materialize_timeout (10m) caps the detached body so a hung store GET cannot outlive it — next serve/heal starts fresh and resumes at file granularity (tmp orphans swept at body start). Client disconnect marks nothing. Non-positive config falls back to compiled defaults (no unbounded wait to configure by accident).
  2. Serve-health sidecar meta/serve-health.json (05 section 5.2.1): sticky-until-reproven (no TTL — fail closed), overwrite-always, best-effort sideband. Marked by the engine on pack-phase failure and by sync probe/heal outcomes; cleared only on proven servability.
  3. Probe-gated mirror sync (11 section 3): post-publish Sync(LevelServe) plus git cat-file -e at the new head with a patient background budget. Pass means ok plus marker cleared; fail means failed servability probe verdict plus backoff plus marker (refs stay published). No-op fires never probe (prior verdict stands; avoids cold-cache flap and keeps no-op at 11 ops).
  4. Self-heal (11 section 3.1): new mirror-heal task kind; the loop's non-due branch fires it when the marker's backoff elapses (rate-bounded, never count-capped). Re-drives Sync plus probe, no teardown, never touches the mirror doc. Pass deletes the marker (healthy again); fail refreshes attempts plus one, anchored.
  5. Agreement (07 section 9.1): summary health degraded from fsck (short-circuit) or serve-health — direct probe for non-mirrors, the hook's degraded_reason for mirrors (plus zero extra GETs); mirrorHash covers it so the flip busts SWR.

Tests / verification

  • New: internal/wal/servetimeout_test.go (wait bound, refs-instant-during-wedge, joiner herd, resume plus clear, cancel-marks-nothing, body cap incl. terminal-record proof, stress, zero-config defaults), servehealth_test.go, internal/mirror/heal_test.go (probe shapes incl. hung-store loop plus recovery, gating incl. no-op scope, healDue table, heal pass/fail/round integration, ProbeObject argv), internal/api/serve320_test.go (health table incl. empty-skip, ETag-bust, 503 plus Retry-After on all object routes), cmd/walhub/mirror320_test.go, setup round-trip test.
  • go test -race clean on wal/api/mirror; cover gate: wal 95.3 / api 95.5 / mirror 96.8 / config 95.7 / store 95.0 / server 98.5 (all at or above 95). contract suite, internal/e2e (45s), full node --test (606/606) green. gofmt and go vet clean.
  • Budget notes (documented in-code plus docs): healthy non-mirror summary plus 2 exact-key GETs (was plus 1 — same R1-B1 cost class, off law-6 paths); first sync 24 ops (was 22: probe GET plus blind clear-DELETE); no-op sync unchanged at 11.

Deviations / environment notes

  • TestUIAssetConcepts (server) fails here: the scratch worktree has no built web/dist concepts and no pnpm to rebuild; main's dist is identically stale — pre-existing/environmental, untouched by this change.
  • Sim tier N/A: no TestSim scenarios and no make sim target exist in this tree (docs describe them); equivalents run: round-trip harness plus contract plus e2e.
  • No live-instance contact: reproduced locally only (hung-store harness); anon/walhub itself untouched. No browser (server-side only; no web behavior change — two setup-FIELDS entries, integrity-tested).
  • No teardown in heal (deliberate, documented): deletion would be data-loss-adjacent without addressing the diagnosed mechanism.
## Root cause (stated per acceptance criterion 5) UNBOUNDED WORK — not cache-dir churn, not a lock-order violation. RepoHandle.Sync(LevelServe) holds packMu across the (repo,materialize) single-flight join/leader-wait, and that task body runs on the registry-lifetime ctx with no deadline at any layer: no handler/request timeout (server.request_timeout is config-only, never enforced anywhere), no packMu/join bound, no store-GET timeout. So one stalled materialize (slow bulk fetch of the upstream-sized pack set after new mirror packs land, or a hung store GET) wedges every later object-level request — each either queues on packMu or joins the same orphaned task with a deadline-free request ctx — indefinitely, while LevelRefs (summary/refs: manifest only, never touch packMu) stays instant. The hourly mirror sync (clone to ingest to PublishRefs) never consults serve-readiness, so last_result ok plus health healthy keep reporting. Lock order was already correct (syncMu to packMu to rw, TryLock-only writers); the present-check resume is sound, so no cache teardown was added. ## What was built 1. Bounded serve materialization (05 section 5.2): server.serve_sync_timeout (45s, under proxy 60s) bounds the pack-phase wait to WalErrTimeout and HTTP 503 + Retry-After 15; server.serve_materialize_timeout (10m) caps the detached body so a hung store GET cannot outlive it — next serve/heal starts fresh and resumes at file granularity (tmp orphans swept at body start). Client disconnect marks nothing. Non-positive config falls back to compiled defaults (no unbounded wait to configure by accident). 2. Serve-health sidecar meta/serve-health.json (05 section 5.2.1): sticky-until-reproven (no TTL — fail closed), overwrite-always, best-effort sideband. Marked by the engine on pack-phase failure and by sync probe/heal outcomes; cleared only on proven servability. 3. Probe-gated mirror sync (11 section 3): post-publish Sync(LevelServe) plus git cat-file -e at the new head with a patient background budget. Pass means ok plus marker cleared; fail means failed servability probe verdict plus backoff plus marker (refs stay published). No-op fires never probe (prior verdict stands; avoids cold-cache flap and keeps no-op at 11 ops). 4. Self-heal (11 section 3.1): new mirror-heal task kind; the loop's non-due branch fires it when the marker's backoff elapses (rate-bounded, never count-capped). Re-drives Sync plus probe, no teardown, never touches the mirror doc. Pass deletes the marker (healthy again); fail refreshes attempts plus one, anchored. 5. Agreement (07 section 9.1): summary health degraded from fsck (short-circuit) or serve-health — direct probe for non-mirrors, the hook's degraded_reason for mirrors (plus zero extra GETs); mirrorHash covers it so the flip busts SWR. ## Tests / verification - New: internal/wal/servetimeout_test.go (wait bound, refs-instant-during-wedge, joiner herd, resume plus clear, cancel-marks-nothing, body cap incl. terminal-record proof, stress, zero-config defaults), servehealth_test.go, internal/mirror/heal_test.go (probe shapes incl. hung-store loop plus recovery, gating incl. no-op scope, healDue table, heal pass/fail/round integration, ProbeObject argv), internal/api/serve320_test.go (health table incl. empty-skip, ETag-bust, 503 plus Retry-After on all object routes), cmd/walhub/mirror320_test.go, setup round-trip test. - go test -race clean on wal/api/mirror; cover gate: wal 95.3 / api 95.5 / mirror 96.8 / config 95.7 / store 95.0 / server 98.5 (all at or above 95). contract suite, internal/e2e (45s), full node --test (606/606) green. gofmt and go vet clean. - Budget notes (documented in-code plus docs): healthy non-mirror summary plus 2 exact-key GETs (was plus 1 — same R1-B1 cost class, off law-6 paths); first sync 24 ops (was 22: probe GET plus blind clear-DELETE); no-op sync unchanged at 11. ## Deviations / environment notes - TestUIAssetConcepts (server) fails here: the scratch worktree has no built web/dist concepts and no pnpm to rebuild; main's dist is identically stale — pre-existing/environmental, untouched by this change. - Sim tier N/A: no TestSim scenarios and no make sim target exist in this tree (docs describe them); equivalents run: round-trip harness plus contract plus e2e. - No live-instance contact: reproduced locally only (hung-store harness); anon/walhub itself untouched. No browser (server-side only; no web behavior change — two setup-FIELDS entries, integrity-tested). - No teardown in heal (deliberate, documented): deletion would be data-loss-adjacent without addressing the diagnosed mechanism.
Root cause (unbounded work, not cache-dir churn): Sync(LevelServe) held
packMu across the (repo,materialize) single-flight whose body runs on the
registry-lifetime ctx with no deadline at any layer (no handler timeout,
no packMu/single-flight bound, no store GET timeout; server.request_timeout
is config-only and unenforced). One stalled materialize wedged every later
object-level request while LevelRefs (no packMu) stayed instant, and the
hourly sync kept reporting last_result ok + healthy without ever probing
servability.

05 §5.2/§5.2.1, 07 §9.1, features/11 §3/§3.1/§5/§7, 04 §12, 11 config table;
Decisions appended in each. New knobs server.serve_sync_timeout (45s wait)
+ server.serve_materialize_timeout (10m body cap), non-positive → defaults.
Post-publish servability probe (Sync + cat-file -e) gates last_result ok;
serve-health sidecar (sticky-until-reproven) drives health degraded +
mirror.degraded_reason (ETag-covered); mirror-heal task re-materializes
with backoff and flips back on pass. No-op syncs never probe; heals never
tear down cache state or touch the mirror doc.
Sign in to join this conversation.
No description provided.