From 5b52673b75734f40821fb84cebba7856218c4208 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Thu, 2 Jul 2026 08:26:35 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20snapshot=20=E2=80=94=202026-07-02=20(?= =?UTF-8?q?Deckard=20trial=E2=86=92revert=20to=20qwopus;=20MTP=20concurren?= =?UTF-8?q?cy=20verdict=20=3D=20not-kept-on-shared-gen;=20worldtree=20#332?= =?UTF-8?q?=20diagnosis=20+=20scoped-view/tunnel=20+=20CI-race=20lesson;?= =?UTF-8?q?=20mtf-dev=20granite=20harness=20spike;=20/books=20ESH=20mount)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- persistent-memory.md | 99 ++++++++++++++++++++++++++++++++++---------- 1 file changed, 77 insertions(+), 22 deletions(-) diff --git a/persistent-memory.md b/persistent-memory.md index 1e83132..d6958d6 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-06-25_ +_Last updated: 2026-07-02_ ## Repo purpose @@ -110,30 +110,37 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-06-25:_ +_As of 2026-07-02:_ -- **Worldtree persona-render config arc COMPLETE (demo + personal).** Pre-synced + - deployed three bind-mount config deltas on corviduo-dev demo+personal: #314 - relational-stance (relational_extractor role + relational_stance block + stance/ dir), - #322 affect/mood gate (affect-render-baseline-allow policy rule + mood_tier_deployment_cap), - #317 valence teardown (REMOVED relational_valence — a **boot-blocking removal**). All - deployed green; the full two-layer persona render (relational + mood) is live on - demo+personal; Worldtree at v0.37.38. Recipe + the additions-vs-removals lesson in - auto-memory `reference_corviduo_dev_emergency_ops`. +- **`gen` model = qwopus (`qwen3.5-122-a10b`); the Deckard trial is CONCLUDED.** Trialed + `robbatt/Qwen3.6-40B-Deckard-NVFP4` behind the `gen` aliases — it won writing quality + decisively but lost on speed (~36 vs qwopus's ~90 tok/s); reverted (git `681eb70`). + Deckard kept STAGED on ana-ml2 as T1's quality benchmark + (`/tank/aimodels/qwen36-40b-deckard-{bf16,nvfp4}`; container `vllm-deckard-40b` + stopped-but-kept). Full arc in auto-memory `reference_gen_qwopus_122b`. -- **nh3-extdev (10.100.50.42) stood up as an althing v0.17.1 mesh PEER — MODEL B.** Full - provision in one session: system zellij + althing (`/usr/local/bin`, both users), - CLAUDE.md, and the mesh (handles `mailman`+`ldp-dev`, ↔nh3-dev). MODEL B (operator's - multi-user choice): receiver runs as the **`althing-svc`** service account + a - group-shared **`/srv/althing`** root (group `althing`; infra-ops+lkraven members; - `ALTHING_ROOT=/srv/althing` via `/etc/profile.d/althing.sh`). Live + verified (two OS - users sharing ONE config/DB). Layout + access facts in auto-memory - `reference_nh3_extdev_althing_mesh`. Open: forseti's multi-user ADR. +- **T1 (mtf-dev qwopus writing-LoRA) IN PROGRESS.** Attention-only LoRA (full-attn q/k/v/o + + the bf16 GDN `in_proj_qkv`/`out_proj`); serve via swappable-LoRA-on-NVFP4 (test QUEUED, + gated on the first T1 adapter) or merge+requant. Trainer harness seam PROVEN via a + granite-8b spike (green 2026-07-02). NVFP4 quant-structure + serve-path facts in + `reference_gen_qwopus_122b`. -- **zellij native web client piloted on nh3-dev** (`zellij-web.service` :8443, native TLS) - ALONGSIDE the ttyd seats (ttyd KEPT as the no-auth fallback). Verdict: wins on - clipboard/auth/TLS/simplicity, loses device-independent fonts; bind left `0.0.0.0` per - operator. auto-memory `reference_zellij_web_seat`. +- **Worldtree #332 scoped-log view + tunnel = STANDING ASSET.** `wt_gateway_logs` view + + `wt_readonly` role on the litellm DB + `wt-db-tunnel` systemd on corviduo-dev, for + worldtree-dev's embed-recall regression-watches. #332 fix verified in prod (15×→1.01× + re-embed). Teardown steps + the IP-pin caveat in auto-memory + `reference_wt_gateway_scoped_log_view`. + +- **`/books` mounted (transient) on nh3-dev** for a books/corpus ingestion: + `10.0.50.50:/mnt/books` (ESH NAS) → `/mnt/books`, NFSv4 `ro,soft` (soft dodges the ESH + D-state hang). NOT fstab — re-mount via `infra-ops@10.100.10.50` if nh3-dev reboots. + +- **granite-4.1-8b bf16 kept** at `irv-ml1:/home/lkraven/granite-4.1-8b-bf16` (17G, + reap-on-request; the T1-harness-spike base). comfyui was borrowed off the A6000 for the + spike + RESTORED healthy. + +- **Recently completed (2026-06-22..25, now in Recent decisions):** Worldtree #314/#322/#317 + persona-render config arc; nh3-extdev althing v0.17.1 mesh peer (Model B); zellij web-seat pilot. - **Backups — recovered + hardened (2026-06-20), STILL OPEN:** rotate the 5 disclosed rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB @@ -164,6 +171,33 @@ _As of 2026-06-25:_ ## Recent decisions +- `[2026-07-02]` **mtf-dev granite harness-spike provisioned on irv-ml1 (comfyui displaced + for the A6000) → ran GREEN.** Freed the A6000 by stopping comfyui (operator-coordinated + with comfy-dev), staged granite-4.1-8b bf16 to `irv-ml1:/home/lkraven`, mtf-dev's TRL + SFT→DPO→eval seam proved end-to-end (DPO genuinely learned, 0.833 acc; anti-slop ~0 = the + expected null on clean-writing granite-instruct); comfyui restored healthy. De-risks T1's + trainer harness ahead of the real qwopus train. + +- `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel + provisioned + fix verified.** The persona-recitation gate re-embedded persona segments + every turn (15× re-embed, 95% cross-turn recurrence, one anchor ×206/min → worldtree-gateway + was ~95% of the embedding backend load); worldtree-dev's cross-turn content-hash cache + dropped it to 1.01× in prod. Built a boundary-safe read-only `wt_gateway_logs` view + + `wt_readonly` role + `wt-db-tunnel` (corviduo-dev) for their regression-watch. auto-memory + `reference_wt_gateway_scoped_log_view`. + +- `[2026-07-01]` **qwopus native MTP speculative-decode tested on `gen` → NOT kept.** qwopus + HAS a full native MTP head (785 tensors, in NVFP4 + bf16). Measured +12% single-stream but + **−15–20% AGGREGATE at moderate concurrency** (N=4: 250→200 tok/s) + it silently drops + `min_p`/`logit_bias` → reverted to clean baseline. The win is banked for T1 (MTP as a + per-deployment option, preserved through requant). `reference_gen_qwopus_122b`. + +- `[2026-07-01]` **Deckard trial → reverted to qwopus (`gen`).** `robbatt/Qwen3.6-40B-Deckard-NVFP4` + won writing "in every way" but at ~36 vs ~90 tok/s (dense-40B vs MoE-~10B-active); + spec-decode rescue ruled out (aeon DFlash image is arm64/DGX-Spark-only; Deckard's + NVFP4/bf16/base-27B all LACK MTP; EAGLE-head training a multi-week non-starter). git + `b63c48b` (repoint) + `681eb70` (revert). Deckard kept staged as T1's writing benchmark. + - `[2026-06-25]` **althing re-architected to the v0.17 lean multi-machine bus; nh3-extdev stood up as a mesh peer (MODEL B).** v0.15.0's lean-bus cut RIPPED moderation/chamber/forseti-daemon/agent-runner/redis-valkey; **v0.17 = per-box @@ -240,6 +274,27 @@ _118 older entries archived to archival-memory.md._ ## Tried and abandoned +- `[2026-07-01]` **A personal-Worldtree CI deploy that fails ~85s in with "not found / + unauthorized" is usually the pull-only-vs-build RACE, not registry-auth.** `deploy-personal.yml` + is PULL-ONLY ("image must already be built by a push to main") but fires on the + `staging/vX` tag push SIMULTANEOUSLY with `deploy.yml`'s main build → it tries to pull the + image ~2.5 min BEFORE the build finishes pushing it → step-4 "Verify image exists" aborts + "not found". Misattributed to registry-auth twice (the earlier cff3328 saga too). DIAGNOSE: + the image tag (12-char short-sha, NOT 7) exists in the registry + the VM login succeeds ⇒ + it's the race. FIX: re-run once the build's done (image now present), OR gate deploy-personal + on `workflow_run: completed`. + +- `[2026-07-01]` **MTP/spec-decode on a SHARED serving model helps single-stream but HURTS + moderate-concurrency aggregate throughput + silently ignores `min_p`/`logit_bias`.** + Measured on qwopus `gen`: N=1 +12%, N=4 −20% aggregate. Don't bolt spec-decode onto the + fleet `gen` for a single-stream win — reserve it for dedicated/interactive deployments. + +- `[2026-07-02]` **irv-ml1 `/worktank` ROOT is root-owned — lkraven can't write there (and + irv-ml1 sudo needs a password non-interactively) → stage model pulls to `/home`.** Also + PIN THE A6000 BY UUID for training runs: nvidia-smi index 1 is the A6000, but native-CUDA + ordering can differ vs docker, and the 3090 (index 0) is usually near-full → land there and + OOM. `CUDA_VISIBLE_DEVICES=GPU-`. + - `[2026-06-25]` **althing "unreachable: — retry later" can MASK an app-level 500.** A multi-hour "intermittent :8087 / firewall flap" hunt was a red herring: raw network was always clean (curl POST to the receiver `:8087` worked; a connect probe = 0