memory: snapshot — 2026-07-02 (Deckard trial→revert to qwopus; MTP concurrency verdict = not-kept-on-shared-gen; worldtree #332 diagnosis + scoped-view/tunnel + CI-race lesson; mtf-dev granite harness spike; /books ESH mount)
This commit is contained in:
+77
-22
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-06-25_
|
||||
_Last updated: 2026-07-02_
|
||||
|
||||
## Repo purpose
|
||||
|
||||
@@ -110,30 +110,37 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-06-25:_
|
||||
_As of 2026-07-02:_
|
||||
|
||||
- **Worldtree persona-render config arc COMPLETE (demo + personal).** Pre-synced +
|
||||
deployed three bind-mount config deltas on corviduo-dev demo+personal: #314
|
||||
relational-stance (relational_extractor role + relational_stance block + stance/ dir),
|
||||
#322 affect/mood gate (affect-render-baseline-allow policy rule + mood_tier_deployment_cap),
|
||||
#317 valence teardown (REMOVED relational_valence — a **boot-blocking removal**). All
|
||||
deployed green; the full two-layer persona render (relational + mood) is live on
|
||||
demo+personal; Worldtree at v0.37.38. Recipe + the additions-vs-removals lesson in
|
||||
auto-memory `reference_corviduo_dev_emergency_ops`.
|
||||
- **`gen` model = qwopus (`qwen3.5-122-a10b`); the Deckard trial is CONCLUDED.** Trialed
|
||||
`robbatt/Qwen3.6-40B-Deckard-NVFP4` behind the `gen` aliases — it won writing quality
|
||||
decisively but lost on speed (~36 vs qwopus's ~90 tok/s); reverted (git `681eb70`).
|
||||
Deckard kept STAGED on ana-ml2 as T1's quality benchmark
|
||||
(`/tank/aimodels/qwen36-40b-deckard-{bf16,nvfp4}`; container `vllm-deckard-40b`
|
||||
stopped-but-kept). Full arc in auto-memory `reference_gen_qwopus_122b`.
|
||||
|
||||
- **nh3-extdev (10.100.50.42) stood up as an althing v0.17.1 mesh PEER — MODEL B.** Full
|
||||
provision in one session: system zellij + althing (`/usr/local/bin`, both users),
|
||||
CLAUDE.md, and the mesh (handles `mailman`+`ldp-dev`, ↔nh3-dev). MODEL B (operator's
|
||||
multi-user choice): receiver runs as the **`althing-svc`** service account + a
|
||||
group-shared **`/srv/althing`** root (group `althing`; infra-ops+lkraven members;
|
||||
`ALTHING_ROOT=/srv/althing` via `/etc/profile.d/althing.sh`). Live + verified (two OS
|
||||
users sharing ONE config/DB). Layout + access facts in auto-memory
|
||||
`reference_nh3_extdev_althing_mesh`. Open: forseti's multi-user ADR.
|
||||
- **T1 (mtf-dev qwopus writing-LoRA) IN PROGRESS.** Attention-only LoRA (full-attn q/k/v/o
|
||||
+ the bf16 GDN `in_proj_qkv`/`out_proj`); serve via swappable-LoRA-on-NVFP4 (test QUEUED,
|
||||
gated on the first T1 adapter) or merge+requant. Trainer harness seam PROVEN via a
|
||||
granite-8b spike (green 2026-07-02). NVFP4 quant-structure + serve-path facts in
|
||||
`reference_gen_qwopus_122b`.
|
||||
|
||||
- **zellij native web client piloted on nh3-dev** (`zellij-web.service` :8443, native TLS)
|
||||
ALONGSIDE the ttyd seats (ttyd KEPT as the no-auth fallback). Verdict: wins on
|
||||
clipboard/auth/TLS/simplicity, loses device-independent fonts; bind left `0.0.0.0` per
|
||||
operator. auto-memory `reference_zellij_web_seat`.
|
||||
- **Worldtree #332 scoped-log view + tunnel = STANDING ASSET.** `wt_gateway_logs` view +
|
||||
`wt_readonly` role on the litellm DB + `wt-db-tunnel` systemd on corviduo-dev, for
|
||||
worldtree-dev's embed-recall regression-watches. #332 fix verified in prod (15×→1.01×
|
||||
re-embed). Teardown steps + the IP-pin caveat in auto-memory
|
||||
`reference_wt_gateway_scoped_log_view`.
|
||||
|
||||
- **`/books` mounted (transient) on nh3-dev** for a books/corpus ingestion:
|
||||
`10.0.50.50:/mnt/books` (ESH NAS) → `/mnt/books`, NFSv4 `ro,soft` (soft dodges the ESH
|
||||
D-state hang). NOT fstab — re-mount via `infra-ops@10.100.10.50` if nh3-dev reboots.
|
||||
|
||||
- **granite-4.1-8b bf16 kept** at `irv-ml1:/home/lkraven/granite-4.1-8b-bf16` (17G,
|
||||
reap-on-request; the T1-harness-spike base). comfyui was borrowed off the A6000 for the
|
||||
spike + RESTORED healthy.
|
||||
|
||||
- **Recently completed (2026-06-22..25, now in Recent decisions):** Worldtree #314/#322/#317
|
||||
persona-render config arc; nh3-extdev althing v0.17.1 mesh peer (Model B); zellij web-seat pilot.
|
||||
|
||||
- **Backups — recovered + hardened (2026-06-20), STILL OPEN:** rotate the 5 disclosed
|
||||
rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB
|
||||
@@ -164,6 +171,33 @@ _As of 2026-06-25:_
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-07-02]` **mtf-dev granite harness-spike provisioned on irv-ml1 (comfyui displaced
|
||||
for the A6000) → ran GREEN.** Freed the A6000 by stopping comfyui (operator-coordinated
|
||||
with comfy-dev), staged granite-4.1-8b bf16 to `irv-ml1:/home/lkraven`, mtf-dev's TRL
|
||||
SFT→DPO→eval seam proved end-to-end (DPO genuinely learned, 0.833 acc; anti-slop ~0 = the
|
||||
expected null on clean-writing granite-instruct); comfyui restored healthy. De-risks T1's
|
||||
trainer harness ahead of the real qwopus train.
|
||||
|
||||
- `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel
|
||||
provisioned + fix verified.** The persona-recitation gate re-embedded persona segments
|
||||
every turn (15× re-embed, 95% cross-turn recurrence, one anchor ×206/min → worldtree-gateway
|
||||
was ~95% of the embedding backend load); worldtree-dev's cross-turn content-hash cache
|
||||
dropped it to 1.01× in prod. Built a boundary-safe read-only `wt_gateway_logs` view +
|
||||
`wt_readonly` role + `wt-db-tunnel` (corviduo-dev) for their regression-watch. auto-memory
|
||||
`reference_wt_gateway_scoped_log_view`.
|
||||
|
||||
- `[2026-07-01]` **qwopus native MTP speculative-decode tested on `gen` → NOT kept.** qwopus
|
||||
HAS a full native MTP head (785 tensors, in NVFP4 + bf16). Measured +12% single-stream but
|
||||
**−15–20% AGGREGATE at moderate concurrency** (N=4: 250→200 tok/s) + it silently drops
|
||||
`min_p`/`logit_bias` → reverted to clean baseline. The win is banked for T1 (MTP as a
|
||||
per-deployment option, preserved through requant). `reference_gen_qwopus_122b`.
|
||||
|
||||
- `[2026-07-01]` **Deckard trial → reverted to qwopus (`gen`).** `robbatt/Qwen3.6-40B-Deckard-NVFP4`
|
||||
won writing "in every way" but at ~36 vs ~90 tok/s (dense-40B vs MoE-~10B-active);
|
||||
spec-decode rescue ruled out (aeon DFlash image is arm64/DGX-Spark-only; Deckard's
|
||||
NVFP4/bf16/base-27B all LACK MTP; EAGLE-head training a multi-week non-starter). git
|
||||
`b63c48b` (repoint) + `681eb70` (revert). Deckard kept staged as T1's writing benchmark.
|
||||
|
||||
- `[2026-06-25]` **althing re-architected to the v0.17 lean multi-machine bus; nh3-extdev
|
||||
stood up as a mesh peer (MODEL B).** v0.15.0's lean-bus cut RIPPED
|
||||
moderation/chamber/forseti-daemon/agent-runner/redis-valkey; **v0.17 = per-box
|
||||
@@ -240,6 +274,27 @@ _118 older entries archived to archival-memory.md._
|
||||
|
||||
## Tried and abandoned
|
||||
|
||||
- `[2026-07-01]` **A personal-Worldtree CI deploy that fails ~85s in with "not found /
|
||||
unauthorized" is usually the pull-only-vs-build RACE, not registry-auth.** `deploy-personal.yml`
|
||||
is PULL-ONLY ("image must already be built by a push to main") but fires on the
|
||||
`staging/vX` tag push SIMULTANEOUSLY with `deploy.yml`'s main build → it tries to pull the
|
||||
image ~2.5 min BEFORE the build finishes pushing it → step-4 "Verify image exists" aborts
|
||||
"not found". Misattributed to registry-auth twice (the earlier cff3328 saga too). DIAGNOSE:
|
||||
the image tag (12-char short-sha, NOT 7) exists in the registry + the VM login succeeds ⇒
|
||||
it's the race. FIX: re-run once the build's done (image now present), OR gate deploy-personal
|
||||
on `workflow_run: completed`.
|
||||
|
||||
- `[2026-07-01]` **MTP/spec-decode on a SHARED serving model helps single-stream but HURTS
|
||||
moderate-concurrency aggregate throughput + silently ignores `min_p`/`logit_bias`.**
|
||||
Measured on qwopus `gen`: N=1 +12%, N=4 −20% aggregate. Don't bolt spec-decode onto the
|
||||
fleet `gen` for a single-stream win — reserve it for dedicated/interactive deployments.
|
||||
|
||||
- `[2026-07-02]` **irv-ml1 `/worktank` ROOT is root-owned — lkraven can't write there (and
|
||||
irv-ml1 sudo needs a password non-interactively) → stage model pulls to `/home`.** Also
|
||||
PIN THE A6000 BY UUID for training runs: nvidia-smi index 1 is the A6000, but native-CUDA
|
||||
ordering can differ vs docker, and the 3090 (index 0) is usually near-full → land there and
|
||||
OOM. `CUDA_VISIBLE_DEVICES=GPU-<uuid>`.
|
||||
|
||||
- `[2026-06-25]` **althing "unreachable: <machine> — retry later" can MASK an app-level
|
||||
500.** A multi-hour "intermittent :8087 / firewall flap" hunt was a red herring: raw
|
||||
network was always clean (curl POST to the receiver `:8087` worked; a connect probe = 0
|
||||
|
||||
Reference in New Issue
Block a user