Mirror anon/walhub: object-level requests hang to 504 (refs fine) — fix the materialization wedge and make mirror health self-healing #320
Labels
No labels
actions
bug
cli
duplicate
enhancement
fork
forum
git storage
help wanted
insights
invalid
issues
moderation
oidc
ownership transfer
packages
pr/merge protection rules
projects
pull requests
question
releases
sponsorships
tags
webhooks
wiki
wontfix
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
crueber/walhub#320
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
What's wrong
The mirrored repo
anon/walhubis broken for content browsing: every object-level request hangs until the proxy gives up (504 Gateway Time-out), while ref-level requests answer instantly. The repo is unusable as a mirror — the whole point of mirroring walhub is to browse its code, and the Code tab can't load a single file.Reproduced live (2026-09-11, against hub.packden.us)
GET /anon/walhub/api(summary)head: main @ 15a5960…,branches: 37,health: "healthy",mirror.last_result: "ok",due: trueGET /anon/walhub/api/refsGET /anon/walhub/api/refs/branchesGET /anon/walhub/api/commits?n=1GET /anon/walhub/api/blob/main/README.mdGET /anon/walhub/api/tree/mainNote the contradictions the UI surfaces from this: the summary says
health: "healthy",mirror.last_result: "ok",branches: 37— yet nothing object-level works. The repo looks fine in every listing and is dead everywhere it matters.Root cause analysis (code evidence)
localView(ctx, id, SyncServe)(internal/api/bind_wal.go:377/403/437/465/672— tree, blob, commits, commit) which triggers engine materialization of the serving copy (engine.Sync, syncLevelToWal). That materialization is what hangs for this repo.internal/mirror/sync.go:245-247) does a fullclone --mirror→ingestPacks→PublishRefs, so the bucket has packs; the wedge is in the on-demand serve-level materialization path (engine.Sync→ locks → fetch/materialize against the cache dir) for a repo with 37 refs and an upstream-sized pack set.due: truewithlast_synced_at 06:00Z) means syncs keep succeeding and keep landing refs — each sync potentially invalidates/extents the serving copy again, so the wedge isn't a one-time bad state; it plausibly re-arms itself.health: "healthy"andmirror.last_result: "ok"and never surfaces the failure anywhere, and there is no automatic recovery.What's needed
SyncServerequests; per-request timeouts on materialization; whether the mirror's ingested pack set is being re-materialized from scratch on each request (cache-dir invalidation churn).last_result: "ok"must mean "refs synced AND objects servable" (e.g. a post-sync probe: oneblobfetch at head). A sync that publishes refs the instance can't serve is not a successful sync.degraded(the mechanism exists —DegradedNotice, health projection #209) instead of reportinghealthy, and surface the state in the mirror projection (last_resultshould say the truth).healthfield must reflect object-level servability, not just ref/manifest health.Acceptance criteria
anon/walhub(or a comparable mirror) serves blob/tree/commits at head without proxy timeouts.last_resultonly reportsokwhen a post-sync servability probe passes; a wedged/un-servable state reports honestly (last_result/health degraded).Fix is up: PR #329 (branch fix/issue-320) — please review, do not merge yet (per task instructions the merge is yours).
Root cause: unbounded work, not cache-dir churn. Sync(LevelServe) serializes on packMu plus the (repo,materialize) single-flight whose body runs on the registry-lifetime ctx with no deadline at any layer, so one stalled materialize wedges every later object-level request while refs (no packMu) stay instant — exactly the observed shape. The hourly sync kept reporting ok/healthy because success never probed servability.
What the PR does: bounds the serve wait (45s to 503 plus degraded) and the materialize body (10m cap); gates last_result ok on a post-publish Sync plus cat-file -e probe (no-op syncs exempt); records the sticky serve-health sidecar and surfaces it in health plus mirror.degraded_reason (ETag-covered); self-heals degraded mirrors via the mirror-heal task with backoff (no teardown, mirror doc untouched).
Verification: race-clean wal/api/mirror, cover gate at or above 95 on all touched packages, contract plus e2e plus full node suite green; hung-store regression tests pin the wedge (bounded timeout, refs-instant divergence, resume plus clear). Two environment notes: TestUIAssetConcepts fails on stale web/dist here (pre-existing, no pnpm to rebuild), and there is no sim tier in this tree (no TestSim scenarios or make target — round-trip harness plus contract plus e2e run instead). Live instance untouched; reproduced locally only.
Review: PR #329 (fix/issue-320) — verified, ready to merge
Reviewed the full diff (28 files, +2430/-77) in a scratch worktree; all verification below is on commit
0da5fcd. No fixes pushed — nothing blocking found. Two non-blocking observations at the end.Root cause — convincing
The unbounded-work claim checks out against the code: serve-level Sync joins the (repo,materialize) single-flight body running on the registry-lifetime task ctx with no deadline at any layer (server.request_timeout is indeed config-only, never enforced — confirmed), while refs-level syncs never touch packMu (internal/wal/handle.go). One stalled bulk fetch wedges every later object request via packMu queue + single-flight fan-in. Correctly NOT blamed on lock order (syncMu→packMu→rw already correct) or cache-dir churn (present-check resume sound; no teardown added — agreed, deletion would not address the mechanism).
Timeouts — sane, fail path honest
Lock discipline — clean
Probe argv — pinned (law 2)
git --git-dir=
cat-file -e added to 04_git.md §12 + argv list, Runner.ProbeObject honors GitTimeout. -e (existence, no output) is the cheapest proof. Verdict-not-retry on non-zero exit: correct.Heal loop — bounded, can't fight sync
ETag / budgets / hot path
Verification (scratch worktree, removed afterward)
42ebfa0in a no-dist worktree (pre-existing/environmental, as the PR discloses). New TestSetupServeTimeoutKeysRoundTrip passes. TestVerifyTokenWireNegatives flaked once under full-suite load, passes -count=3 in isolation; touches no PR file (auth untouched).Non-blocking observations (no action required before merge)
Acceptance criteria mapping
MERGE RECOMMENDATION: ready to merge.
Fixed by PR #329 (review clean — root cause, timeouts, lock discipline, probe, heal, ETag all verified), merged. Closing.