Visibility still bounces public/private after #381+#391 — continued investigation #394

Closed
opened 2026-09-12 16:47:17 +00:00 by crueber · 5 comments
Owner

Follow-up to #381 (summary off SWR) and #391 (save hardening): the visibility value still flips between public and private across refreshes/saves.

What is already ruled out (with evidence in #381/#391):

  • Browser HTTP-cache stale-serve on the summary (now ccMutable + ETag economics proven), detailed route (#384), and owner profile (#385).
  • Store staleness on memory/filesystem backends (handler-level repro: fresh PUTs stick, two instances converge).
  • Silent save failures staying silent (all failure paths now reseed + loud; fresh CAS version + one 409 retry).

Remaining hypotheses to verify in order:

  1. The reporting instance is not running a build containing #381/#391 (check the deployed version/commit first — cheapest).
  2. S3-class backend conditional-GET semantics (03_store_backends.md documents the contract; it was never verified against the deployed backend class — rustfs/MinIO/S3 or a caching proxy in front).
  3. Multi-instance LRU divergence (PUT on one instance, GET on another before conditional revalidation).
  4. Two-writer version races (Access tab full-doc PUT vs Settings visibility-only PUT interleaving).
  5. A second read path (any component reading visibility through a WAL-replayed/manifest-cached surface rather than the direct doc — #391 item 4 was never fully closed).

Acceptance: root cause with evidence stated; bounce gone across 10+ save/refresh cycles on the reporting setup; regression test for the specific mode.

Follow-up to #381 (summary off SWR) and #391 (save hardening): the visibility value still flips between public and private across refreshes/saves. What is already ruled out (with evidence in #381/#391): - Browser HTTP-cache stale-serve on the summary (now ccMutable + ETag economics proven), detailed route (#384), and owner profile (#385). - Store staleness on memory/filesystem backends (handler-level repro: fresh PUTs stick, two instances converge). - Silent save failures staying silent (all failure paths now reseed + loud; fresh CAS version + one 409 retry). Remaining hypotheses to verify in order: 1. The reporting instance is not running a build containing #381/#391 (check the deployed version/commit first — cheapest). 2. S3-class backend conditional-GET semantics (03_store_backends.md documents the contract; it was never verified against the deployed backend class — rustfs/MinIO/S3 or a caching proxy in front). 3. Multi-instance LRU divergence (PUT on one instance, GET on another before conditional revalidation). 4. Two-writer version races (Access tab full-doc PUT vs Settings visibility-only PUT interleaving). 5. A second read path (any component reading visibility through a WAL-replayed/manifest-cached surface rather than the direct doc — #391 item 4 was never fully closed). Acceptance: root cause with evidence stated; bounce gone across 10+ save/refresh cycles on the reporting setup; regression test for the specific mode.
Author
Owner

Investigation update: root cause NOT provable from this workspace — no invented fix, no PR. Ruled out with evidence: (1) repro harness driving the real routeAccess handler against memory, filesystem, AND rustfs S3 rig — alternating writers, 8-way concurrent writers, cross-instance readers, conditional-GET, CAS enforcement — zero bounces, zero stale reads, versions strictly +1; (2) code audit: every visibility read funnels through one path (badge ← summary repoVisibility, selects ← GET access, listings ← fillVisibilityFlags → RepoVisibility, gates ← CheckRead → GetAccess; single wiring in cmd/walhub/collab.go:120) — no second read path; ETags cover visibility everywhere; stale full-doc PUTs 409, never silently clobber. Prime suspect, unverifiable from here: the reporting instance runs a stale build — anon-probed / → 401, /setup → 403, /healthz says version dev, Server header is a container ID not a commit; timeline fits pre-#391 code or a cached pre-#391 JS bundle (which keeps the old silent-failure save path until hard refresh). Needed: instance identity + build (walhub version / image digest), one repro cycle with the #391 access PUT/GET server log lines, and the backend class if S3-but-not-rustfs. If it still bounces on ≥ d174bf6 with a fresh bundle, look outside the repo (proxy/CDN 304 synthesis, multi-replica) — the harness re-creates in minutes.

Investigation update: root cause NOT provable from this workspace — no invented fix, no PR. Ruled out with evidence: (1) repro harness driving the real routeAccess handler against memory, filesystem, AND rustfs S3 rig — alternating writers, 8-way concurrent writers, cross-instance readers, conditional-GET, CAS enforcement — zero bounces, zero stale reads, versions strictly +1; (2) code audit: every visibility read funnels through one path (badge ← summary repoVisibility, selects ← GET access, listings ← fillVisibilityFlags → RepoVisibility, gates ← CheckRead → GetAccess; single wiring in cmd/walhub/collab.go:120) — no second read path; ETags cover visibility everywhere; stale full-doc PUTs 409, never silently clobber. Prime suspect, unverifiable from here: the reporting instance runs a stale build — anon-probed / → 401, /setup → 403, /healthz says version dev, Server header is a container ID not a commit; timeline fits pre-#391 code or a cached pre-#391 JS bundle (which keeps the old silent-failure save path until hard refresh). Needed: instance identity + build (walhub version / image digest), one repro cycle with the #391 access PUT/GET server log lines, and the backend class if S3-but-not-rustfs. If it still bounces on ≥ d174bf6 with a fresh bundle, look outside the repo (proxy/CDN 304 synthesis, multi-replica) — the harness re-creates in minutes.
Author
Owner

Live-container findings (5fc9fe4569df, serves hub.packden.us): (1) stored access.json is coherent (v7 private, updated 17:18:16Z); all 5 PUTs 200; every GET in the window is state-coherent (200s carry current, 304s agree) — no stale response exists server-side. (2) User bundle index-CLpaNq0N.js loaded 16:46:27Z from the current image — bundle is current, pre-#391 code ruled out for this tab. (3) ETag logic verified sound (version-keyed, exact match; post-PUT GETs correctly 200). (4) No render cache, no second read path, single backend container (no split-brain; the stray host walhub process is gone). Secondary defect filed separately: client fetch storm (~5/sec bursts across all endpoints). Next step is a coordinated repro: one save+refresh cycle with a timestamp, correlated against the server log to see exactly what was sent vs displayed.

Live-container findings (5fc9fe4569df, serves hub.packden.us): (1) stored access.json is coherent (v7 private, updated 17:18:16Z); all 5 PUTs 200; every GET in the window is state-coherent (200s carry current, 304s agree) — no stale response exists server-side. (2) User bundle index-CLpaNq0N.js loaded 16:46:27Z from the current image — bundle is current, pre-#391 code ruled out for this tab. (3) ETag logic verified sound (version-keyed, exact match; post-PUT GETs correctly 200). (4) No render cache, no second read path, single backend container (no split-brain; the stray host walhub process is gone). Secondary defect filed separately: client fetch storm (~5/sec bursts across all endpoints). Next step is a coordinated repro: one save+refresh cycle with a timestamp, correlated against the server log to see exactly what was sent vs displayed.
Author
Owner

Fix ready for review: PR #399 (#399, branch fix/issue-394). Reseed-when-clean in both selects + headless helper with node cover (9/9 new, 777/777 suite green, vite build green). No backend change, no new deps.

Fix ready for review: PR #399 (https://git.packden.us/crueber/walhub/pulls/399, branch fix/issue-394). Reseed-when-clean in both selects + headless helper with node cover (9/9 new, 777/777 suite green, vite build green). No backend change, no new deps.
Author
Owner

Review of PR #399 (fix/issue-394, reseed-when-clean) — verified in scratch worktree at 9c1290a. No browser (per brief: node tests + reasoning; noted explicitly).

(1) State machine — CORRECT. web/src/lib/visibilityReseed.js:78-83: null next → unchanged; unseeded → seed both; clean → reseed both; dirty → same reference back. Settings.jsx:122-126 effect maps signals through reseed and sets only on field change (settles, no loop); save success rebases on echo (Settings.jsx:170-172), failure reseeds truth then rebases (Settings.jsx:186-190) — settles clean, no stuck-dirty (only exception: truth null leaves form untouched, correct since no truth arrived).

(2) Screenshot scenario — RESOLVES. Committed test pins v6/public → v7/private reseed to private while clean (visibility-reseed.test.js:44-53); ad-hoc node drive of the helper through seed-public → v7/private → dirty → v8-lands-no-clobber → rebase-clean all passed. Effect fires on every shared-entry delivery while clean, so the select flips with no user action.

(3) Loading-vs-public — CORRECT. docVisibility null/undefined → null (visibilityReseed.js:45-48); unseeded select renders value='' disabled + 'loading current visibility…' line (Settings.jsx:224-236, Access.jsx:172-183); save disabled while unseeded (Settings.jsx:242); saveVisibility guards isVisibility (Settings.jsx:151). Committed test pins loading-is-not-dirty and null-delivery no-op (test:26-35).

(4) Access.jsx symmetric fix — CORRECT, no behavior lost. Old render-time seed() (seeded once when !getBase) is subsumed by the createEffect (Access.jsx:154-158): first load (!getBase → reset) plus follow-while-clean. Baseline rebases ONLY while clean — dirty keeps original CAS version so concurrent saves still 409 into load() (Access.jsx:88-106) instead of last-writer-winning. Dirty marker added (Access.jsx:266-268). One nuance, not a defect: clean-form poll redelivery calls reset() unconditionally, but Solid === equality suppresses same-value signal writes, so no flicker.

(5) No clobber race — SAFE. Both effects test clean/dirty synchronously at effect time against current signals; entry-update and keystroke in either order preserve the edit (reseed-then-keystroke ends user-value-on-new-baseline; keystroke-then-entry hits the dirty guard). Committed same-reference pin (test:55-66).

(6) Poll interplay — NO-OPS. Settings sets fire only on actual change (Settings.jsx:124-125); same-truth redelivery asserted content-equal in test:50-52.

(7) No backend change — CONFIRMED. Diff is exactly 5 files: docs/go/12_web_ui.md, web/src/lib/visibilityReseed.js, web/src/pages/Access.jsx, web/src/pages/Settings.jsx, web/test/unit/visibility-reseed.test.js. Zero .go files.

(8) No new deps; docs accurate — CONFIRMED. No package.json/pnpm-lock/go.mod changes; helper imports nothing. 12_web_ui.md decision entry matches the code (blank-disabled-while-unseeded, warn-line marker, ?? public only on present docs, dirty-keeps-CAS 409 path, seed() deletion).

Tests (scratch worktree, node_modules symlinked from main checkout): non-smoke suite 777/777 green (exit 0); visibility-reseed.test.js 9/9; vite build + esbuild SDK bundle green. Full glob run shows 778 pass / 2 fail, both in smoke.test.js which probes live :8080 — that port holds an unrelated instance (healthz 200, left untouched); with WALHUB_TEST_WEB_BASE_URL pointed at an empty port the smoke file skips 3/3 as designed. AGENTS.md laws 1/7/8/12 hold (no new deps, no background work, web-only seam, decision appended).

MERGE RECOMMENDATION: ready to merge.

Review of PR #399 (fix/issue-394, reseed-when-clean) — verified in scratch worktree at 9c1290a. No browser (per brief: node tests + reasoning; noted explicitly). (1) State machine — CORRECT. web/src/lib/visibilityReseed.js:78-83: null next → unchanged; unseeded → seed both; clean → reseed both; dirty → same reference back. Settings.jsx:122-126 effect maps signals through reseed and sets only on field change (settles, no loop); save success rebases on echo (Settings.jsx:170-172), failure reseeds truth then rebases (Settings.jsx:186-190) — settles clean, no stuck-dirty (only exception: truth null leaves form untouched, correct since no truth arrived). (2) Screenshot scenario — RESOLVES. Committed test pins v6/public → v7/private reseed to private while clean (visibility-reseed.test.js:44-53); ad-hoc node drive of the helper through seed-public → v7/private → dirty → v8-lands-no-clobber → rebase-clean all passed. Effect fires on every shared-entry delivery while clean, so the select flips with no user action. (3) Loading-vs-public — CORRECT. docVisibility null/undefined → null (visibilityReseed.js:45-48); unseeded select renders value='' disabled + 'loading current visibility…' line (Settings.jsx:224-236, Access.jsx:172-183); save disabled while unseeded (Settings.jsx:242); saveVisibility guards isVisibility (Settings.jsx:151). Committed test pins loading-is-not-dirty and null-delivery no-op (test:26-35). (4) Access.jsx symmetric fix — CORRECT, no behavior lost. Old render-time seed() (seeded once when !getBase) is subsumed by the createEffect (Access.jsx:154-158): first load (!getBase → reset) plus follow-while-clean. Baseline rebases ONLY while clean — dirty keeps original CAS version so concurrent saves still 409 into load() (Access.jsx:88-106) instead of last-writer-winning. Dirty marker added (Access.jsx:266-268). One nuance, not a defect: clean-form poll redelivery calls reset() unconditionally, but Solid === equality suppresses same-value signal writes, so no flicker. (5) No clobber race — SAFE. Both effects test clean/dirty synchronously at effect time against current signals; entry-update and keystroke in either order preserve the edit (reseed-then-keystroke ends user-value-on-new-baseline; keystroke-then-entry hits the dirty guard). Committed same-reference pin (test:55-66). (6) Poll interplay — NO-OPS. Settings sets fire only on actual change (Settings.jsx:124-125); same-truth redelivery asserted content-equal in test:50-52. (7) No backend change — CONFIRMED. Diff is exactly 5 files: docs/go/12_web_ui.md, web/src/lib/visibilityReseed.js, web/src/pages/Access.jsx, web/src/pages/Settings.jsx, web/test/unit/visibility-reseed.test.js. Zero .go files. (8) No new deps; docs accurate — CONFIRMED. No package.json/pnpm-lock/go.mod changes; helper imports nothing. 12_web_ui.md decision entry matches the code (blank-disabled-while-unseeded, warn-line marker, ?? public only on present docs, dirty-keeps-CAS 409 path, seed() deletion). Tests (scratch worktree, node_modules symlinked from main checkout): non-smoke suite 777/777 green (exit 0); visibility-reseed.test.js 9/9; vite build + esbuild SDK bundle green. Full glob run shows 778 pass / 2 fail, both in smoke.test.js which probes live :8080 — that port holds an unrelated instance (healthz 200, left untouched); with WALHUB_TEST_WEB_BASE_URL pointed at an empty port the smoke file skips 3/3 as designed. AGENTS.md laws 1/7/8/12 hold (no new deps, no background work, web-only seam, decision appended). MERGE RECOMMENDATION: ready to merge.
Author
Owner

Fixed by PR #399 (review clean — all 8 points pass, screenshot scenario pinned by test, no clobber race, poll-safe; 777/777), merged. Closing.

Fixed by PR #399 (review clean — all 8 points pass, screenshot scenario pinned by test, no clobber race, poll-safe; 777/777), merged. Closing.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
crueber/walhub#394
No description provided.